Real-time automated monitoring and control of ultrafiltration / diafiltration (UF / DF) conditioning and dilution processes
Real-time monitoring and control of UF/DF pools using Raman spectroscopy and machine learning addresses inefficiencies in large-scale biomolecule manufacturing, enabling immediate product release and reducing process deviations.
Patent Information
- Application Number
- JP2025534125
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-12
- Filing Date
- 2023-12-12
- Publication Date
- 2026-01-21
AI Technical Summary
Existing large-scale biomolecule manufacturing processes face inefficiencies in monitoring and controlling product quality attributes of ultrafiltration/diafiltration (UF/DF) conditioned pools, leading to time-consuming and resource-intensive offline quality control methods that hinder real-time process adjustments and product release.
Implementing real-time automated monitoring and control using Raman spectroscopy and machine learning models to predict product quality attributes of UF/DF pools, allowing for immediate release decisions based on in-line measurements.
Enables rapid, efficient monitoring and control of UF/DF pool quality attributes, facilitating immediate product release and reducing process deviations by leveraging Raman spectroscopy and machine learning for real-time quality assessment.
Smart Images

Figure 2026502096000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 432,033, filed December 12, 2023, the entire contents of which are incorporated herein by reference and relied upon. [Technical Field]
[0002] Field This description is generally directed to real-time (e.g., online, in-line, etc.) automated monitoring and control of ultrafiltration / diafiltration (UF / DF) conditioning or dilution processes. More particularly, this description provides methods and systems for real-time automated monitoring and control of product quality attributes of a UF / DF conditioning or dilution process downstream of a UF / DF process using machine learning models. The outcome of such methods and systems can be prediction and / or measurement of quality attributes at the end of (target) UF / DF conditioning for immediate release of a drug substance in place of manual quality control testing. [Background technology]
[0003] background Large-scale manufacturing of biomolecules can be a highly involved process and typically involves robust quality control mechanisms, such as sampling of biomolecules or cell culture media for offline measurement of various product quality attributes of the manufacturing process. For example, monoclonal antibodies purified from cell culture media or product may be received in a UF / DF collection vessel for conditioning, and samples of the monoclonal antibodies may be taken from the collection vessel for offline evaluation of product quality attributes. However, such quality control mechanisms can be time- and resource-consuming. Therefore, there is a need for a technology that facilitates rapid, in-line monitoring and control of product quality attributes of the UF / DF-conditioned pool in the collection vessel. Summary of the Invention
[0004] overview The present disclosure provides techniques that allow product quality to be automatically monitored and controlled in real time, avoiding undesirable process deviations. Furthermore, final measurements of quality attributes at the end of UF / DF conditioning can be used to facilitate immediate release of the product. In various embodiments, according to a first aspect, a method for predicting one or more product quality attributes of a conditioned ultrafiltration / diafiltration (UF / DF) pool in a bioprocess is provided. In various embodiments, the method includes receiving Raman measurements of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of the UF / DF operation and receiving the UF / DF pool for conditioning with a buffer to form the conditioned UF / DF pool. The Raman measurements may be received as in-line measurements. For example, the Raman measurements may be received in real time or near real time. The Raman measurements may be measurements of the conditioned UF / DF pool at a current time, and predicting one or more product quality attributes of the conditioned UF / DF pool may include predicting one or more product quality attributes of the conditioned UF / DF pool at a current time. Furthermore, the method includes predicting one or more product quality attributes of the conditioned UF / DF pool using a machine learning model trained to receive the Raman measurements of the conditioned UF / DF pool as input and generate one or more predicted product quality attributes of the conditioned UF / DF pool as output. The machine learning model may be trained using training data including a plurality of Raman measurements of the conditioned UF / DF pool and corresponding measurements of the one or more product quality attributes of the conditioned UF / DF pool.
[0005] In an embodiment, the one or more product quality attributes include the osmolality of the conditioned UF / DF pool.
[0006] In embodiments, the product of the bioprocess comprises one or more proteins in solution, and the one or more product quality attributes comprises protein concentration of the conditioned UF / DF pool.
[0007] In embodiments, the UF / DF pool is a purified UF / DF pool received in a collection vessel from upstream ultrafiltration and diafiltration processes.
[0008] In embodiments, the conditioned UF / DF pool comprises antibodies.
[0009] In an embodiment, the antibody is a monoclonal antibody (mAb).
[0010] In embodiments, the machine learning model is trained to predict one or more product quality attributes using a training dataset that includes one or more product quality attributes of conditioned UF / DF pools conditioned with different amounts of buffer added thereto and associated Raman measurements of conditioned UF / DF pools conditioned with different amounts of buffer.
[0011] In embodiments, the trained machine learning model is a trained multivariate linear model, a trained neural network, a trained deep learning model, or a trained ensemble model.
[0012] In an embodiment, the trained machine learning model is one of a design of experiments (DOE) partial least squares (PLS) regression model, a trend analysis (TA) partial least squares (PLS) regression model, or a combination thereof.
[0013] In embodiments, the method further includes generating an indication of whether the conditioned UF / DF pool is ready for immediate release based on the one or more predicted product quality characteristics. The indication may be generated by comparing the one or more predicted product quality characteristics to one or more predetermined thresholds for each of the one or more predicted product quality characteristics. In embodiments, predicting the one or more product quality characteristics of the conditioned UF / DF pool includes predicting the one or more product quality characteristics of the conditioned UF / DF pool at a current time. In other embodiments, predicting the one or more product quality characteristics of the conditioned UF / DF pool includes predicting the one or more product quality characteristics of the conditioned UF / DF pool at a future time (e.g., an endpoint of the conditioning process).
[0014] In an embodiment, the steps of receiving Raman measurements and predicting one or more product quality characteristics are repeated one or more times at a predetermined frequency.
[0015] In an embodiment, receiving as input the Raman measurements of the conditioned UF / DF pool includes receiving as input one or more of: image data of a spectrum corresponding to the Raman measurements; a series of peaks in the spectrum; a data series including intensity values for each of a plurality of wavenumbers; and a stride between a minimum and a maximum of a plurality of wavenumbers in the spectrum; or a combination thereof.
[0016] In various embodiments, a method for controlling a UF / DF pool conditioning process is provided according to a second aspect. In various embodiments, the method includes receiving, by a processor, a first Raman measurement of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of the UF / DF operation and receiving the UF / DF pool for conditioning with a buffer to form the conditioned UF / DF pool. The method further includes predicting, by the processor, one or more product quality characteristics of the conditioned UF / DF pool using a trained machine learning model that receives the first Raman measurement as input, the one or more product quality characteristics including a first protein concentration of the conditioned UF / DF pool. The method further includes calculating, by the processor, a weight index of the collection vessel corresponding to a second protein concentration of the conditioned UF / DF pool that differs from the first protein concentration of the conditioned UF / DF pool. The method further includes outputting, by the processor, instructions to add or stop adding buffer to the collection container based on the calculated weight indicator of the collection container.
[0017] In an embodiment, outputting, by the processor, an instruction to add or stop adding buffer to the collection container includes the processor sending a buffer pump signal to a buffer pump operably coupled to the collection container, the buffer pump signal being configured to instruct the buffer pump to add or stop adding buffer to the collection container.
[0018] In embodiments, the method includes calculating, by a processor, a weight index of a collection container corresponding to a first protein concentration of the conditioned UF / DF, and determining, by the processor, an amount of buffer to add to the collection container based on the weight index of the collection container corresponding to the first protein concentration and the weight index of the collection container corresponding to the second protein concentration. For example, the amount of buffer may depend on the difference between the weight index of the collection container corresponding to the first protein concentration and the weight index of the collection container corresponding to the second protein concentration.
[0019] In an embodiment, the method further includes determining, by the processor, a flow rate for adding the determined amount of buffer to the collection container based on the weight index of the collection container corresponding to the first protein concentration and the weight index of the collection container corresponding to the second protein concentration.
[0020] In embodiments, determining the flow rate of the determined amount of buffer to add to the collection vessel may include determining whether to increase or decrease the flow rate.
[0021] In embodiments, the instructions to add or stop adding buffer include instructions for a determined amount of buffer to add to the collection vessel. The second protein concentration may be a predetermined concentration, such as, for example, a target protein concentration of the product. Calculating by the processor a weight index of the collection vessel corresponding to the protein concentration of the conditioned UF / DF pool may use a predetermined relationship between the weight index and the protein concentration. The predetermined relationship may be in the form of a mathematical function or a look-up table. The weight index may be a value indicative of the amount (e.g., volume) of product in the collection vessel.
[0022] In embodiments, the instructions to add or stop adding buffer include instructions for a flow rate for adding the determined amount of buffer to the collection vessel.
[0023] In embodiments, the first Raman measurement corresponds to a current time point of the conditioned UF / DF pool. In some aspects, the first predicted protein concentration and / or weight index may correspond to a current time point of the conditioned UF / DF pool.
[0024] In embodiments, the processor may be in wired or wireless communication with the Raman spectrometer. In some cases, the processor may be in remote communication with the Raman spectrometer. In some aspects, the latency between reception and output is about 2 seconds or less.
[0025] In an embodiment, the processor is an edge node, a processor running a virtual machine, a processor of a cloud server, or a processor of a cloud serverless solution.
[0026] In an embodiment, the one or more product quality attributes of the conditioned UF / UD pool further include an osmolality of the conditioned UF / DF pool, and the predicting includes predicting the osmolality of the conditioned UF / DF pool based on analysis of the first Raman measurements.
[0027] In embodiments, the trained machine learning model may be configured to receive the first Raman measurement as an input and generate as an output a first protein concentration of the conditioned UF / DF pool and an osmolality of the conditioned UF / DF pool. The trained machine learning model may include multiple models, including a first trained machine learning model configured to take the first Raman measurement as an input and generate as an output a first protein concentration of the conditioned UF / DF pool, and a second trained machine learning model configured to take the first Raman measurement as an input and generate as an output an osmolality of the conditioned UF / DF pool. The trained machine learning model may include one or more models, each configured to take the first Raman measurement as an input and generate as an output both the first protein concentration of the conditioned UF / DF pool and the osmolality of the conditioned UF / DF pool.
[0028] In embodiments, the method further includes receiving, by the processor and from the Raman spectrometer, a second Raman measurement of the conditioned UF / DF pool after transmission. The method may further include using the processor to predict a third protein concentration of the conditioned UF / DF pool using a trained machine learning model that incorporates the second Raman measurement as input. Additionally, the method may further include comparing, by the processor, the second protein concentration to the third protein concentration to determine the effectiveness of adding buffer to the collection vessel.
[0029] In embodiments, the method does not use any measurements obtained by extracting a sample of the conditioned UF / DF pool from a collection vessel.
[0030] In embodiments, the conditioned UF / DF pool comprises an antibody. For example, the antibody may comprise a monoclonal antibody (mAb).
[0031] In an embodiment, the method may further include, in response to an instruction to add buffer, adding buffer in an amount sufficient to increase the weight index of the collection container to within the calculated weight threshold of the weight index.
[0032] In embodiments, the trained machine learning model is a trained multivariate linear model, a trained neural network, a trained deep learning model, or a trained ensemble model.
[0033] In an embodiment, the trained machine learning model is one of a design of experiments (DOE) partial least squares (PLS) regression model, a trend analysis (TA) partial least squares (PLS) regression model, or a combination thereof.
[0034] In embodiments, the method may further include receiving, by the processor, a second Raman measurement of the conditioned UF / DF pool after output from the Raman spectrometer. The method may further include predicting, by the processor using a trained machine learning model that takes the second Raman measurement as input, a third protein concentration of the conditioned UF / DF pool. The method may further include generating, by the processor, an indication of whether the conditioned UF / DF pool is ready for immediate release based on the predicted third protein concentration.
[0035] In various embodiments, according to a third aspect, a system is provided that includes a UF / DF pool collection system and a communications module. In various embodiments, the UF / DF pool collection system is downstream of the UF / DF operation and configured to receive and store a buffer-conditioned UF / DF pool. The UF / DF pool collection system further includes a Raman spectrometer operably coupled to the collection vessel and configured to perform Raman measurements on the conditioned UF / DF pool. In various embodiments, the communications module is operably coupled to the remote computing platform and the Raman spectrometer and configured to (i) receive Raman measurements from the Raman spectrometer and upload the Raman measurements to the remote computing platform for predicting one or more product quality characteristics by the remote computing platform based on the Raman measurements, and (ii) transmit a signal received from the remote computing platform and related to the one or more product quality characteristics to the UF / DF pool collection system. In various embodiments, a latency between receiving the Raman measurements at the communications module and transmitting the signal to the UF / DF pool collection system satisfies a latency threshold.
[0036] In an embodiment, the system may further include a remote computing platform including a processor configured to receive the Raman measurements from the communication module, predict one or more product quality characteristics, and output a signal related to the one or more product quality characteristics.
[0037] In embodiments, the one or more product quality characteristics include a first protein concentration of the conditioned UF / DF pool. The processor is further configured to calculate a weight index of the collection vessel corresponding to a second protein concentration of the conditioned UF / DF pool that is different from the first protein concentration of the conditioned UF / DF pool. The UF / DF pool collection system may further include a buffer pump operably coupled to the collection vessel. The signal may be a buffer pump signal configured to instruct the buffer pump to add or stop adding buffer to the collection vessel based on the weight index of the collection vessel.
[0038] In an embodiment, the one or more product quality attributes may include the osmolality of the conditioned UF / DF pool.
[0039] In an embodiment, the latency threshold may be approximately two seconds (eg, two seconds).
[0040] In embodiments, the conditioned UF / DF pool comprises an antibody. For example, the antibody may comprise a monoclonal antibody (mAb).
[0041] In various embodiments, a computer-implemented method for providing a tool for predicting product quality attributes and / or controlling an ultrafiltration / diafiltration (UF / DF) pool conditioning process is provided, according to a fourth aspect. The method includes obtaining a training dataset, the training dataset including a plurality of Raman measurements of a conditioned UF / DF pool obtained from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of an ultrafiltration / diafiltration (UF / DF) operation and receiving the UF / DF pool for conditioning with a buffer, and corresponding measurements of one or more product quality attributes of the conditioned UF / DF pool. The method further includes training a machine learning model using the training dataset to predict one or more product quality attributes of the conditioned UF / DF pool using the Raman measurements of the conditioned UF / DF pool as input. The method according to this aspect may have any of the features described with respect to the first and second aspects.
[0042] In various embodiments, according to a fifth aspect, there is provided a system including one or more processors and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods described herein.
[0043] In various embodiments, according to a fifth aspect, there is provided one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods described herein. [Brief explanation of the drawings]
[0044] For a more complete understanding of the principles disclosed herein and their advantages, reference is now made to the following descriptions taken in conjunction with the accompanying drawings.
[0045] [Figure 1] 1 is a schematic workflow for real-time (e.g., online) monitoring and control of a conditioned or diluted ultrafiltration / diafiltration (UF / DF) pool, according to various embodiments.
[0046] [Figure 2] FIG. 1 is a block diagram of a product quality attribute prediction system according to a non-limiting example of the present disclosure.
[0047] [Figure 3] 1 is a flowchart of a machine learning model-based process for predicting product quality attributes of a conditioned or diluted UF / DF pool according to a non-limiting example of the present disclosure.
[0048] [Figure 4] 1 is a flowchart of a machine learning model-based process for controlling a UF / DF pool according to a non-limiting example of the present disclosure.
[0049] [Figure 5] 1 illustrates two categories of experiments and analyses that were conducted to develop a model for real-time monitoring and control of UF / DF conditioning processes and target conditions, according to a non-limiting example of the present disclosure.
[0050] [Figure 6A] 1 shows plots of Raman spectra of conditioned UF / DF pool endpoints with various levels of osmolality in a preliminary scale-up analysis according to a non-limiting example of the present disclosure.
[0051] [Figure 6B]FIG. 6B shows a scatter plot of actual versus predicted osmolality of conditioned UF / DF pool endpoints using the Raman spectra of FIG. 6A as model input, according to a non-limiting example of the present disclosure.
[0052] [Figure 7A] 1 shows plots of Raman spectra of conditioned UF / DF pools in a design of experiments (DOE) analysis with various levels of protein concentration, according to non-limiting examples of the present disclosure.
[0053] [Figure 7B] FIG. 7B shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools in a DOE analysis using the Raman spectra of FIG. 7A as model input, according to a non-limiting example of the present disclosure.
[0054] [Figure 8A] 1 shows plots of Raman spectra of conditioned UF / DF pools in trend analysis with various levels of protein concentration, according to various embodiments.
[0055] [Figure 8B] FIG. 8B shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools in a trend analysis using the Raman spectra of FIG. 8A as model input, according to a non-limiting example of the present disclosure.
[0056] [Figure 9] 8B shows a scatter plot of actual versus predicted osmolality of conditioned UF / DF pools in a trend analysis using the Raman spectra of FIG. 8A as model input, according to various embodiments.
[0057] [Figure 10A]FIG. 10 shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools when predicting endpoint test data from a DOE analysis using a comprehensive model developed by combining trend and DOE training datasets, according to a non-limiting example of the present disclosure.
[0058] [Figure 10B] FIG. 10 shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools when using a comprehensive model to predict end-to-end test data from trend analysis, according to a non-limiting example of the present disclosure.
[0059] [Figure 11A] FIG. 10 shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools in a trend analysis using a multivariate linear regression model, according to a non-limiting example of the present disclosure.
[0060] [Figure 11B] 1 shows a scatter plot of actual versus predicted protein concentrations of conditioned UF / DF pools in a trend analysis using a deep neural network model, according to a non-limiting example of the present disclosure.
[0061] [Figure 12] 8A shows a time series plot of experimental data comparing real-time protein concentration predictions over time using a model (the "process trend model") trained on the trend analysis data shown in FIG. 8A. According to a non-limiting example of the present disclosure, the time series plot of protein concentration predictions is overlaid with measurements of protein concentration of samples taken throughout the experiment at various time points, measured using a protein concentration sensor, as described herein.
[0062] [Figure 13]13 shows a time series plot of experimental data comparing osmolality predictions using Raman spectra collected from the real-time experiment of FIG. 12 with the osmolality of samples taken throughout the experiment and measured using an osmometer as input to a process trend model according to a non-limiting example of the present disclosure. The experiment associated with FIGS. 12 and 13 illustrates the disclosed monitoring of a UF / DF pool performed using the conditions shown in Table 3 (further described below), where the first step corresponds to filling the collection vessel with the UF / DF diluted pool (mAb feed), and each addition step represents a subsequent dilution with conditioning buffer (20 mM sodium acetate, 1057 mM (40%) trehalose, 0.20% w / v polysorbate 20, pH 5.3).
[0063] [Figure 14] FIG. 13 shows a scatter plot of actual (true measurements) versus predicted protein concentrations (averaged over a window of the last five measurements) corresponding to the time series plot of FIG. 12 , according to a non-limiting example of the present disclosure.
[0064] [Figure 15] 14 shows a scatter plot of actual osmolality versus predicted osmolality (averaged over a window of the last five measurements) corresponding to the time series plot of FIG. 13 , according to a non-limiting example of the present disclosure.
[0065] [Figure 16] 10 shows a plot illustrating real-time control of a conditioned UF / DF pool with a target protein concentration of 50 mg / mL in the collection vessel based on real-time adjustment of the weight of the collection vessel using the feed flow rate as a control variable, according to a non-limiting example of the present disclosure.
[0066] [Figure 17]10 shows a plot illustrating real-time control of a conditioned UF / DF pool with a target protein concentration of 45 mg / mL in the collection vessel based on real-time adjustment of the weight of the collection vessel using the feed flow rate as a control variable, according to a non-limiting example of the present disclosure.
[0067] [Figure 18] FIG. 1 is a block diagram of a computer system according to a non-limiting example of the present disclosure.
[0068] [Figure 19] 1 illustrates an exemplary neural network that may be used to implement a deep learning neural network, according to a non-limiting example of the present disclosure.
[0069] It should be understood that the figures are not necessarily drawn to scale, and that objects in the figures are not necessarily drawn to scale relative to each other. The figures are intended to provide clarity and understanding of various embodiments of the devices, systems, and methods disclosed herein. Wherever possible, the same reference numerals will be used throughout the figures to refer to the same or like parts. Furthermore, it should be appreciated that the figures are not intended to limit the scope of the present teachings in any way. DETAILED DESCRIPTION OF THE INVENTION
[0070] Detailed Description I. Overview Antibodies are defensive proteins produced by the immune system in response to the presence of foreign substances called antigens. Antibodies can include monoclonal antibodies (mAbs), polyclonal antibodies (pAbs), antibody fragments (e.g., Fab', Fab, F(ab')2), single-domain antibodies (DABs), Fvs, and single-chain Fvs (scFvs). Monoclonal antibodies (mAbs) are popular for both therapeutic and research purposes. Several improvements in large-scale manufacturing processes have improved the quality of large-scale mAb production. Efficient recovery and purification of mAbs from cell culture media are critical aspects of the production process. At least for clinical antibodies, the purification process must produce mAbs that are safe and reliable for use in human patients. This includes monitoring product quality attributes, including protein characteristics (e.g., concentration) and impurities, including host cell proteins, DNA, viruses, endotoxins, aggregates, concentration, excipients, and other species that may affect patient safety, efficacy, or potency. One or more of these parameters, including process intermediates, may be critical process parameters for unit operation performance. These parameters need to be monitored throughout production, and the mAb product should be tested at various stages of the downstream process to ensure acceptable levels of impurities are present.
[0071] Quality control can be achieved in antibody production processes by analyzing intermediate and formulated drug substance samples. Examples of such samples include samples of ultrafiltration / diafiltration (UF / DF) conditioned pools, which contain drug substance after the purification process in the UF / DF skid or tank of an antibody production system. Sample quality control can be performed using offline methods for each production lot. For example, samples may be removed from the UF / DF tank and subjected to offline testing to measure product quality attributes. In some examples, samples may be monitored and analyzed inline, i.e., at the UF / DF process site. For example, Raman spectroscopy may be utilized to monitor the UF / DF process, as discussed in U.S. Patent Application Publication No. 2020 / 062802, entitled "Use of Raman Spectroscopy in Downstream Purification," which is incorporated herein by reference in its entirety.
[0072] Once a UF / DF batch is purified in the UF / DF process, the UF / DF bulk is transferred to a collection vessel for dilution and conditioning. During the collection process, the UF / DF pool is conditioned with different buffers (e.g., conditioning buffer, dilution buffer, etc.). API quality attributes at the end of the conditioning process are key indicators for determining the release of the API batch; i.e., off-target quality attributes can result in the batch being discarded. The present disclosure provides a solution for automated real-time monitoring and control of the conditioned and / or diluted UF / DF pool in the collection vessel, ensuring real-time, on-target control of quality attributes and avoiding process deviations. Furthermore, because quality attributes measured at the end of the conditioning stage (referred to herein as the "endpoint") are used to determine the release of the API batch, the final measurement of quality attributes at the end of conditioning (referred to herein as the "target condition") obtained from real-time model predictions can be used to make automated release testing decisions, replacing or augmenting existing manual quality control methods. Alternative manual and offline measurements of product quality attributes can be slow, inefficient, and prone to causing significant product release delays.
[0073] The present disclosure provides processes and systems related to measuring product quality attributes of conditioned or diluted UF / DF pools in collection vessels by utilizing Raman spectroscopy as an in-line probe within the collection vessel. Examples of product quality attributes that can be monitored and controlled by the technology disclosed herein include color, clarity / opalescence, physical state, pH, osmolality, polysorbate 20 concentration, one or more protein concentrations of the conditioned UF / DF pool, methionine concentration, N-acetyltryptophan concentration, high molecular weight forms, low molecular weight forms, acidic regions, basic regions, high mannose, afucosylation, glucosylation, oxidation, microbiological purity, bacterial endotoxin, potency, and / or the like of the media and molecules (e.g., mAbs) in the collection vessel. In particular, this disclosure describes techniques for using an in-line Raman spectrometer to acquire Raman spectral data of the conditioned UF / DF pool in a collection vessel, and predicting product quality attributes, such as osmolality and protein concentration, of the conditioned UF / DF pool using a machine learning model that incorporates the Raman measurements as input data and is trained to predict product quality attributes based on an analysis of the input data. Collection operations, including conditioning of the UF / DF pool in the collection vessel, can then be controlled in real time based on the predicted product quality attributes. Furthermore, the predicted quality attributes can replace or augment existing manual, offline quality control testing to enable real-time release testing of drugs, such as mAbs and other molecules.
[0074] II. Definition The present disclosure is not limited to these exemplary embodiments and applications or the manner in which the exemplary embodiments and applications operate or are described herein. Further, the figures may show simplified or partial views, and the dimensions of the elements in the figures may be exaggerated or out of proportion.
[0075] Furthermore, when the terms "on," "mounted," "connected," "coupled," or similar terms are used herein, one element (e.g., a component, material, layer, substrate, etc.) may be "on," "mounted," "connected," or "coupled" to another element, regardless of whether the element is directly on, directly attached to, connected to, or coupled to another element, or whether there are one or more intervening elements between the one and the other element. Furthermore, when a reference is made to a list of elements (e.g., elements a, b, c), such reference is intended to include any one of the listed elements by itself, any combination of fewer than all of the listed elements, and / or all combinations of the listed elements. The section divisions herein are for ease of reference only and do not limit any combination of elements described.
[0076] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those of ordinary skill in the art. Furthermore, unless the context requires otherwise, singular terms shall include the plural and plural terms shall include the singular. Generally, the nomenclature utilized in connection with, and techniques of, chemistry, biochemistry, molecular biology, pharmacology, and toxicology described herein are those well known and commonly used in the art.
[0077] As used herein, "substantially" means sufficient to function for its intended purpose. Thus, the term "substantially" allows for slight, insignificant variations from an absolute or perfect state, dimension, measurement, result, etc., as would be expected by one of ordinary skill in the art, but does not noticeably affect overall performance. For parameters or characteristics that are numerical values or can be expressed as numerical values, "substantially" means within 10%.
[0078] The term "ones" means two or more.
[0079] As used herein, the term "plurality" can be 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.
[0080] As used herein, the term "about" refers to a range of normal error for each readily known value. Reference herein to a value or parameter with "about" includes (and describes) embodiments directed to the value or parameter itself. For example, a statement referring to "about X" includes a statement of "X." In some embodiments, "about" may refer to ±15%, ±10%, ±5%, or ±1%, as would be understood by one of ordinary skill in the art.
[0081] As used herein, the term "set" means one or more. For example, a set of items includes one or more items.
[0082] As used herein, the phrase “at least one of,” when used in conjunction with a list of items, means that different combinations of one or more of the listed items may be used, or that only one of the items in the list may be required. An item may be a specific object, thing, step, action, process, or category. In other words, “at least one of” means that any combination or number of items from the list may be used, but not all of the items in the list may be required. For example, without limitation, “at least one of item A, item B, or item C” means item A, item A and item B, item B, item A, item B, and item C, item B and item C, or item A and item C. In some cases, “at least one of item A, item B, or item C” means, but is not limited to, two items A, one item B, and ten items C, or four items B and seven items C, or some other suitable combination.
[0083] As used herein, "machine learning" may include the practice of using algorithms to analyze data, learn from it, and then make decisions or predictions about something in the world. Machine learning uses algorithms that can learn from data without relying on rule-based programming.
[0084] Embodiments of the present disclosure relate to UF / DF pools containing antibodies. As used herein, the term "antibody" refers to any immunological binding agent, such as IgG, IgM, IgA, IgD, and IgE, and also refers to any antibody-like molecule having an antigen-binding region, including antibody fragments such as Fab', Fab, F(ab')2, single-domain antibodies (DAB), Fv, and scFv (single-chain Fv). In certain embodiments, antibodies may be monoclonal or humanized.
[0085] As used herein, the term "vessel" refers to an apparatus suitable for containing and growing cell cultures, including manufacturing scale. Additionally, the term can also refer to a harvest vessel. A harvest vessel is an apparatus suitable for containing a UF / DF pool for a harvest process in which the UF / DF pool is diluted and conditioned using a dilution buffer and a conditioning buffer, respectively.
[0086] As used herein, the term "conditioning" refers to the addition of a dilution and / or conditioning buffer to a solution or medium, such as a UF / DF pool.
[0087] The term "cell culture," as used herein, refers to the growth of cells in an artificial environment under suitable conditions.
[0088] As used herein, the term "molecule" refers to a substance produced by a cell and may include carbohydrates, lipids, nucleic acids, and proteins. The terms "molecule" and "biomolecule" may be used interchangeably.
[0089] As used herein, "artificial neural network" or "neural network" (NN) may refer to a mathematical model or computational algorithm for training and / or deploying a mathematical model (i.e., for making predictions). A neural network, sometimes referred to as a neural net, may use one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as the input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current value of each parameter set. In various embodiments, a reference to a "neural network" may refer to one or more neural networks.
[0090] Neural networks may process information in two ways: when they are being trained, they are in training mode; and when they put what they have learned into practice, they are in inference (or prediction) mode. Neural networks may learn through a feedback process (e.g., backpropagation) that allows the network to adjust the weight coefficients of individual nodes in intermediate hidden layers (modify the behavior of individual nodes) so that the outputs match those of the training data. In other words, neural networks may receive training data (training examples) and learn how to arrive at the correct output even when presented with a new range or set of inputs. The neural network may include, for example, but not be limited to, at least one of a feedforward neural network (FNN), a recurrent neural network (RNN), a modular neural network (MNN), a convolutional neural network (CNN), a residual neural network (ResNet), an ordinary differential equation neural network (neural-ODE), or another type of neural network.
[0091] As used herein, a "bioprocess" (also referred to herein as a "biomassing process") refers to a process in which biological components, such as cells, parts thereof, such as organelles, or multicellular structures, such as organoids or spheroids, are maintained in a liquid medium in an artificial environment, such as a bioreactor. In embodiments, a bioprocess refers to cell culture. A bioprocess typically results in a product, which may include one or more compounds produced as a result of the activity of the biomass and / or biological components.
[0092] As used herein, "target condition" or "target matrix condition" refers to a state, instant, or stage of a UF / DF pool bioprocess at which a product quality attribute is targeted for monitoring, prediction, or control. The state, instant, or stage may be predefined or fixed. Thus, a "fixed target matrix condition" may refer to a fixed or predefined moment in a bioprocess (e.g., a harvest stage of a UF / DF pool) at which a product quality attribute can be measured or predicted. Such predefined moments include, but are not limited to, the start of the harvest stage (e.g., immediately after or immediately after the UF / DF pool is received in a harvest vessel from a UF / DF tank for harvest), the end of the harvest stage (e.g., after the harvest operation of the conditioned UF / DF pool is completed), or any other moment in a bioprocess.
[0093] As used herein, "endpoint" refers to a point near the end of a bioprocess at which a target outcome of the bioprocess is analyzed to help determine the effectiveness of a product developed via the bioprocess.
[0094] As used herein, "product quality characteristics" refers to the physical and / or chemical properties of a product that includes one or more biomolecules in solution that are the product of a bioprocess.
[0095] As used herein, "UF / DF pool" refers to a solution containing one or more biomolecules that are the product of a bioprocess, which solution has been obtained by applying ultrafiltration and diafiltration to the output of the bioprocess.
[0096] As used herein, "Raman measurements" may refer to Raman spectroscopic data obtained by a Raman spectrometer by detecting inelastic Raman scattering from a solution (e.g., a conditioned UD / DF pool). The Raman measurements may be used as data input for training and / or applying various machine learning models described herein. For example, the Raman measurements may be in the form of spectral image data, a series of peaks in the spectrum, a data series having intensities for each wavenumber with a given stride between the minimum and maximum values in the spectrum, or a combination thereof.
[0097] As used herein, "design of experiments (DOE) analysis" refers to the analysis of multiple measurements of a product quality attribute (e.g., protein concentration, osmolality concentration, etc.) over a predetermined range using target matrix conditions at a predefined moment in a bioprocess (e.g., the end of a UF / DF pool conditioning process). For example, the predetermined range may encompass all product quality attributes (e.g., protein concentrations) expected at a predefined moment (e.g., endpoint) of a bioprocess (e.g., a particular UF / DF process). A design of experiments (DOE) partial least squares (PLS) regression model (also referred to herein as a "DOE PLS model") refers to a PLS regression-based machine learning model trained using training data including multiple measurements of a product quality attribute over a DOE analysis, plus corresponding Raman measurements taken on the conditioned UF / DF pool at that predefined moment.
[0098] As used herein, "trend analysis (TA)" refers to the analysis of the range of measurements of a product quality attribute (e.g., protein concentration, osmolality concentration, etc.) from the beginning to the end of a bioprocess ("end-to-end"). For example, for a UF / DF conditioning process starting with a diluted pool at time t_0 and ending with a conditioned pool at time t_final, a TA PLS model may be trained using training data that includes the range of protein concentrations between the beginning and end of the conditioning phase (e.g., including the initial protein concentration at time t_0, the final protein concentration at time t_final, and the protein concentrations of any additions therebetween of the UF / DF pools under dilution or conditioning). A trend analysis (TA) partial least squares (PLS) regression model (also referred to herein as a "TA PLS model") refers to a PLS regression-based machine learning model that is trained using training data that includes measurements of product quality attributes (e.g., protein concentration, osmolality concentration, etc.) from the beginning to the end ("end-to-end") of the UF / DF conditioning process, along with corresponding Raman measurements made on the end-to-end conditioned UF / DF pool.
[0099] III. Real-time monitoring and control of the conditioned or diluted UF / DF pool in the collection vessel 1 illustrates a schematic workflow 100 for real-time monitoring and control of a conditioned UF / DF pool, according to various embodiments. As described above, the UF / DF pool purified in UF / DF tank 110 may be received in collection vessel 120 for dilution or conditioning with a buffer to generate a final formulation of the product (e.g., a therapeutic drug substance). For example, cell culture media configured to produce a monoclonal antibody may undergo a UF / DF operation or process in UF / DF tank 110 to remove impurities, reduce volume, and generally achieve conditions favorable for therapeutic development and storage. The UF / DF pool from UF / DF tank 110 may then be transferred to collection vessel 120 for dilution and conditioning with a buffer to form a final formulation of the product (e.g., a monoclonal antibody therapeutic).
[0100] Currently, the quality of the drug substance in collection vessel 120 is assessed through several manual steps, including sampling the conditioned or diluted UF / DF pool from collection vessel 120 and performing offline measurements of product quality attributes of the conditioned UF / DF pool. However, such manual sampling and offline measurements may be too slow to allow for real-time corrections to manufacturing processes and procedures as well as to allow for immediate release of product. For example, calculation of a bolus addition to collection vessel 120 based on offline measurements may be too late for the ongoing therapeutic drug production process, and any subsequent manual and / or offline corrective actions may be too late and far from real-time, potentially resulting in process deviations and product release delays.
[0101] In various embodiments, Raman spectrometer 130 can be used to perform in-line Raman spectroscopy measurements of the conditioned UF / DF pool in collection vessel 120 for real-time monitoring and control of the conditioned UF / DF pool. In-line Raman spectroscopy in collection vessel 120 refers to the use of an in-line or in-situ probe 140 in collection vessel 120 to perform Raman spectroscopy measurements. In such cases, probe 140 detects the energy of photons inelastically scattered by analytes or molecules in the conditioned UF / DF pool. The photon energy is associated with multiple vibrational modes of the analytes or molecules in the UF / DF pool, which Raman spectrometer 130 can analyze to generate Raman spectroscopy data corresponding to the conditioned UF / DF pool. In-line Raman spectroscopy is advantageous in that it is a non-invasive and non-destructive technique that allows for automated, real-time monitoring and control of the conditioned UF / DF pool (e.g., by enabling assessment of product quality attributes based on Raman measurements, as discussed in more detail below) without requiring sampling of the conditioned UF / DF pool. In various embodiments, the term "Raman measurements" may refer to Raman spectroscopy data acquired by Raman spectrometer 130 by detecting inelastic Raman scattering from the conditioned UF / DF pool in collection vessel 120. Raman measurements may be used as data input for training and / or applying various machine learning models described herein. For example, Raman measurements may be in the form of image data of a spectrum (e.g., the entire spectrum), a series of peaks in the spectrum, a data series with intensities for each wavenumber with a given stride between the minimum and maximum wavenumbers in the spectrum, or a combination thereof.The use of Raman spectroscopy to characterize therapeutic agents is discussed in "Multi-attribute Raman spectroscopy (MARS) for monitoring product quality attributes in formulated monoclonal antibody therapeutics" by B. Wei et al., mAbs, 14:1, 2007564 (2022), which is incorporated herein by reference in its entirety.
[0102] After Raman measurements of the conditioned UF / DF pool in collection vessel 120 are obtained by Raman spectrometer 130, in various embodiments, Raman spectrometer 130 transmits the Raman measurements to computing platform 150 via a communications module operably coupled to Raman spectrometer 130 and the computing platform (e.g., at a processor of computing platform 150). In some examples, computing platform 150 may be on-site, i.e., at the location where the manufacturing process for producing the therapeutic drug substance is occurring (e.g., where UF / DF tank 110 or collection vessel 120 is located). In some cases, computing platform 150 may be off-site or remote. For example, remote computing platform 150 may be able to access Raman spectrometer 130 via a wireless network. In some cases, the computing platform may be an edge node (e.g., an IoT edge node), an on-premise virtual machine, a cloud server, a cloud serverless solution, or a combination thereof.
[0103] In various embodiments, computing platform 150 processes the Raman measurements of the conditioned UF / DF pool received from Raman spectrometer 130 and hosts a machine learning model that can be trained to predict one or more product quality attributes of the conditioned UF / DF pool. For example, the machine learning model may be a multivariate linear model, a neural network model, a deep learning model, an ensemble model, a regression model, or a combination thereof. Product quality attributes of the conditioned UF / DF pool that can be predicted by the machine learning model include, but are not limited to, the protein concentration, osmolality concentration, etc. of the conditioned UF / DF pool.
[0104] In some examples, a machine learning model is trained to predict product quality characteristics using a training dataset that includes in-line Raman measurements of conditioned UF / DF pools and associated offline or real-time measured product quality characteristics (i.e., of these conditioned UF / DF pools). For example, the training dataset may include product quality characteristics of conditioned UF / DF pools that have been measured using offline methods, such as by applying a measurement instrument (e.g., an osmometer, a protein sensor, etc.) to measure the product quality characteristics on samples of the conditioned UF / DF pools removed from each conditioned UF / DF pool. Additionally, or alternatively, the training dataset may include product quality attributes of the conditioned UF / DF pools measured in-line, for example, via a measurement device (e.g., an osmometer, a protein sensor, etc.) located near, within, or at each conditioned UF / DF pool (e.g., as they are undergoing conditioning), so that samples do not need to be removed from each conditioned UF / DF pool or removed from the conditioning process for measurement of any product quality attributes. The training dataset may further include Raman measurements of each such conditioned UF / DF pool. The training dataset is provided to a machine learning model so that the machine learning model can learn to process the in-line or offline Raman measurements and predict product quality attributes based on the analysis. In some cases, the training dataset may be a design of experiments (DOE) dataset of in-line Raman measurements and associated product quality attributes of the conditioned UF / DF pool obtained to train a machine learning model to predict product quality attributes at target matrix conditions (e.g., the endpoint of the conditioning or recovery phase of the UF / DF pool in the recovery vessel).DOE is a systematic statistical method for analyzing and quantifying the influence and interaction of factors that affect output, and therefore can be applied at the data collection stage to arrive at a valid output. In another example, the training dataset may be a trend analysis dataset (also referred to herein as a "trend analysis dataset") of in-line Raman measurements of a conditioned UF / DF pool and associated product quality attributes obtained to train a machine learning model to predict product quality attributes for a matrix of conditions from start to finish (e.g., from the start to the end of conditioning the UF / DF pool in a collection vessel).
[0105] As described above, Raman spectrometer 130 may transmit Raman measurements of the UF / DF pool in the collection vessel to computing platform 150 using a communications module operably coupled to computing platform 150. Upon receiving the Raman measurements, in various embodiments, computing platform 150 may execute a trained machine learning model hosted thereon to analyze the Raman measurements and predict product quality attributes of the conditioned UF / DF pool based on the analysis. For example, the trained machine learning model may analyze the Raman measurements and predict product quality attributes of the conditioned UF / DF pool, such as, but not limited to, the protein concentration, osmolality concentration, etc. of the conditioned UF / DF pool.
[0106] In various embodiments, monitoring and controlling the conditioned UF / DF pool in collection receptacle 120 includes adding or discontinuing buffer into collection receptacle 120 to adjust product quality attributes of the conditioned UF / DF pool to values indicative of a target product quality. Additionally or alternatively, monitoring and controlling the conditioned UF / DF pool in collection receptacle 120 may further include altering the flow rate of buffer into collection receptacle 120, preventing the next scheduled bolus addition of buffer, temporarily stopping the continuous flow of buffer, triggering the next bolus addition of buffer, etc. That is, for example, the conditioning process of the UF / DF pool in collection receptacle 120 may vary the amount and / or concentration of buffer in collection receptacle 120 over time (e.g., via adding, discontinuing, a changed flow rate, or a changed bolus schedule of buffer) to condition or dilute (or discontinue diluting or conditioning) the UF / DF pool received from UF / DF tank 110. For example, a first product quality characteristic of the conditioned UF / DF pool predicted by the machine learning model can be indicative of the product quality at the endpoint of the conditioning process. Then, as part of the UF / DF conditioning process, a buffer (e.g., a conditioning or dilution buffer) can be added (or, if already in progress, its addition can be terminated) to collection vessel 120 to achieve a second product quality characteristic of the UF / DF pool that is indicative of the target product quality. Alternatively, the first product quality characteristic of the conditioned UF / DF pool predicted by the machine learning model can be indicative of the product quality at a current point in the dilution process. In such an embodiment, in response to determining that the first product quality characteristic predicted at the current point is not the desired or target product quality characteristic, the UF / DF pool may be further conditioned (e.g., via a conditioning or dilution buffer) as part of the UF / DF conditioning process to achieve a second product quality characteristic of the UF / DF pool that is indicative of the target product quality characteristic.
[0107] In such an embodiment, computing platform 150 may determine the amount of buffer that needs to be added to collection vessel 120 for the conditioned UF / DF pool to achieve the second or target product quality characteristic. Alternatively, computing platform 150 may determine whether the continuous addition of buffer should be terminated so that the conditioned UF / DF pool achieves or does not significantly deviate from the second or target product quality characteristic. For example, computing platform 150 may determine a weight indicator of collection vessel 120 that corresponds to the second product quality characteristic (e.g., the weight of collection vessel 120, the height and / or volume of the conditioned UF / DF pool in collection vessel 120, etc.), and thereby calculate the amount of buffer that needs to be added to collection vessel 120 so that collection vessel 120 reaches the determined weight indicator (e.g., or whether the addition of buffer should be terminated).
[0108] As a non-limiting example, the first product quality characteristic predicted by the machine learning model can be a first protein concentration or osmolality of the conditioned UF / DF pool, indicative of product quality at a given time point (e.g., current, start, intermediate, or endpoint) during the conditioning process. Furthermore, the first product quality characteristic (e.g., first protein concentration or osmolality) can be higher or lower than a second respective product quality characteristic (e.g., second respective protein concentration or osmolality) associated with a product having a target product quality. For example, the second product quality characteristic (e.g., second protein concentration or osmolality) can be a second value (desired or target value) of the same characteristic of product quality (e.g., protein concentration, osmolality, etc.). In such a case, the computing platform 150 may determine a weight of the collection container 120 corresponding to the second product quality characteristic (e.g., a second protein concentration or osmolality) based on, for example, a predetermined formula or a table relating product quality characteristics to weight indicators (e.g., in this example, relating protein concentration or osmolality to the weight of the collection container 120). The computing platform 150 may then determine or calculate an amount of buffer that needs to be added to the collection container 120 so that the weight of the collection container 120 increases to the determined weight. For example, the computing platform may have information regarding the weight of the collection container 120 corresponding to the first product quality characteristic (e.g., a first protein concentration or osmolality) (“first weight”). In such a case, the amount of buffer added to the collection container 120 is a weight at least substantially equal to the difference between the determined weight and the first weight. In other words, the added buffer will increase the weight of the collection container 120 from the first weight to the determined weight.
[0109] In such an embodiment, computing platform 150 may output instructions to add or stop adding buffer to the collection receptacle based on the calculated weight index of the collection receptacle. For example, computing platform 150 may generate a buffer pump signal and send the buffer pump signal to buffer supply pump 160, which is configured to pump buffer (e.g., from a buffer reservoir) into collection receptacle 120. The buffer pump signal may be configured to instruct buffer supply pump 160 to add buffer to collection receptacle 120 in an amount that will cause collection receptacle 120 to achieve a weight index corresponding to the second product quality characteristic (alternatively, the buffer pump signal may instruct buffer supply pump 160 to stop adding buffer such that the weight index of collection receptacle 120 does not significantly deviate from the determined weight index).
[0110] In various embodiments, computing platform 150 may be configured to generate an indication based on the predicted product quality (e.g., without further manual quality control processes) that indicates the product is ready for immediate release. For example, a machine learning model may be used to predict product quality characteristics (e.g., protein concentration, osmolality concentration, etc.) of the UF / DF pool in the collection vessel (e.g., during any given time in the bioprocess). Computing platform 150 may then generate an indication based on the predicted product quality that indicates whether a batch of the UF / DF pool in the collection vessel at a given time is ready for immediate release. That is, the indication may indicate that the batch may not require further product quality testing. For example, computing platform 150 may compare the predicted product quality characteristic with a threshold product quality characteristic and generate an indication based on the comparison. For example, when the comparison indicates that the predicted product quality characteristic is within (or outside) a predetermined range of the threshold product quality characteristic, the indication may indicate that the batch is ready (or not ready) for immediate release. In various embodiments, the generation of instructions may be real-time, i.e., instructions may be generated during or at the end of the collection process, thereby reducing delays associated with manual product quality testing.
[0111] In various embodiments, monitoring and control of the conditioning of the UF / DF pool in collection receptacle 120 as shown in workflow 100 may occur in real time or near real time, without the need for manual sampling of the UF / DF pool from collection receptacle 120. That is, the process from the arrival of the UF / DF pool at collection receptacle 120 to the addition of buffer to collection receptacle 120 by buffer supply pump 160 or its completion may occur automatically in real time or near real time, without the need for manual control or manual sampling to determine, for example, product quality attributes of the UF / DF pool in collection receptacle 120. In some examples, the duration between the transmission of the Raman measurements by the Raman spectrometer 130 to the computing platform 150 and the arrival of the buffer pump signal at the buffer supply pump can be between about 0.1 seconds and about 3 seconds, between about 0.5 seconds and about 2.5 seconds, between about 1 second and about 2 seconds, between about 1 second and about 1.5 seconds, between about 1.5 seconds and about 2 seconds, about 1 second, about 2 seconds, and values and subranges therebetween, as well as ranges combining any of the specified lower and upper limits.
[0112] FIG. 2 is a block diagram of a product quality attribute prediction system 200 according to various embodiments. In various embodiments, the quality attribute prediction system 200 may correspond to the computing platform 150 of FIG. 1. For example, the quality attribute prediction system 200 may be an edge node (e.g., an IoT edge node), a virtual machine (e.g., an on-premises virtual machine), a cloud server, a cloud serverless solution, or a combination thereof, remotely communicating with a collection vessel for the conditioned UF / DF pool. The quality attribute prediction system 200 uses a machine learning model 210 to predict product quality attributes of the conditioned UF / DF pool in the collection vessel. Examples of product quality attributes include the osmolality or protein concentration of the conditioned UF / DF pool. In some examples, the UF / DF pool may be received in the collection vessel from a UF / DF tank, and the UF / DF pool may be conditioned with a buffer to support a collection operation to recover the active pharmaceutical ingredient (e.g., mAb) therein. In such a case, in-line Raman measurements 206 of the UF / DF pool may be received at quality attribute prediction system 200 such that a machine learning model 210 trained on a training dataset 208 of Raman measurements and associated product quality attributes can predict product quality attributes 212 of the conditioned UF / DF pool to predict the product quality attributes. Quality attribute prediction system 200 includes a computing platform 202, data storage 214, a set of input devices 216, and a display system 204.
[0113] Computing platform 202 may take a variety of forms. In various embodiments, computing platform 202 includes a single computer (or computer system) or multiple computers that communicate with each other. In other examples, computing platform 202 takes the form of a cloud computing platform. In various embodiments, computing platform 202 may be communicatively coupled to data storage 214, display system 204, set of input devices 216, or a combination thereof. In various embodiments, data storage 214, display system 204, set of input devices 216, or a combination thereof may be considered part of computing platform 202 or may be otherwise integrated with computing platform 202. Thus, in some examples, computing platform 202, data storage 214, set of input devices 216, and display system 204 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together.
[0114] In various embodiments, a Raman spectrometer performs Raman scans of the conditioned UF / DF pool in the collection vessel to generate in-line Raman measurements 206. The in-line Raman measurements 206 can be Raman spectral data of the conditioned UF / DF pool. For example, a Raman spectrometer (e.g., a Raman RXN2 spectrometer from Kaiser Optical Systems, Inc.) can use an in-line probe operably coupled to the collection vessel to perform one or more Raman scans of the conditioned UF / DF pool at a desired rate (e.g., scans every few minutes). The Raman scans can be performed over a wide range of wavenumbers (e.g., from about 100 / cm to about 3000 / cm), and the Raman spectrometer can generate Raman spectroscopic data of the UF / DF pool based on the Raman scans, the Raman spectroscopic data including peaks at those wavenumbers corresponding to Raman shifts (e.g., those wavenumbers at which inelastic scattering of photons by components of the conditioned UF / DF pool occurs).
[0115] In various embodiments, the inline Raman measurements 206 are provided to machine learning model 210 for analysis and use in predicting product quality attributes of the conditioned UF / DF pool. In some cases, additive data associated with the conditioned UF / DF pool or a collection vessel containing the conditioned UF / DF pool may also be provided to machine learning model 210. For example, a pH probe may be operably coupled to the collection vessel and may measure the pH of the conditioned UF / DF pool (e.g., may store such measurements in data storage 214). As another example, a temperature probe may be operably coupled to the collection vessel and may measure the temperature of the conditioned UF / DF pool or collection vessel (e.g., may store such measurements in data storage 214). Thus, additive data associated with the conditioned UF / DF pool or collection vessel may be the measured pH and / or temperature of the conditioned UF / DF pool, and such data may also be provided as input data to machine learning model 210.
[0116] The machine learning model 210 may be, but is not limited to, a neural network, a decision tree, a random forest, a support vector machine, a Bayesian network, a regression model, a multivariate linear model, an ensemble model, etc., or a combination thereof. The neural network may be a deep neural network, a convolutional neural network (CNN), an artificial neural network (ANN), a recurrent neural network (RNN), a modular neural network (MNN), a residual neural network (RNN), an ordinary differential equation neural network (neural-ODE), a squeeze and excitation embedding neural network, MobileNet, etc. The ANN may be a long short-term memory (LSTM) neural network. The regression model may be a gradient boosting machine (GBM) model (e.g., XGBoost). Regression models can also be linear regression models (e.g., including multivariate linear regression models), logistic regression models, polynomial regression models, ridge regression models, least absolute shrinkage and selection operator (lasso) regression models, partial least squares (PLS) regression models, principal component regression models, etc. Lasso regression models are regularized linear regression models that perform regularization and variable or feature selection to improve the predictive accuracy of the model. Lasso regression uses an L1 penalty for both fitting and penalizing model coefficients, and performs variable selection and regularization simultaneously to simultaneously improve both the predictive accuracy and interpretability of the model. PLS methods assume input / observed variables generated by a system or process driven by a small number of latent (e.g., not directly observed or measured) variables and model the relationship between the observed and latent variables. Data input for these models may include, but is not limited to, spectral image data (e.g., corresponding to Raman measurements), a data series with intensities for each wavenumber with a given stride between the minimum and maximum wavenumbers in the spectrum, and a series of peaks in a spectrum.
[0117] In various embodiments, the machine learning model 210 may be trained with a training dataset 208 of Raman measurements and associated product quality attributes, such that the machine learning model 210 can take the in-line Raman measurements 206 as input data and predict the product quality attributes 212 of the UF / DF pool. In some examples, the training dataset 208 may be a dataset of Raman measurements of the UF / DF pool and associated product quality attributes acquired under fixed target matrix conditions. As used herein, a fixed target matrix condition may refer to a predefined moment in the bioprocess (e.g., a harvest phase of the UF / DF pool) at which the product quality attribute can be measured or predicted. The predefined moment may include, but is not limited to, the start of the harvest phase (e.g., immediately or shortly after the UF / DF pool is received into a harvest vessel from a UF / DF tank for harvest), the end of the harvest phase (e.g., after the harvest operation of the conditioned UF / DF pool is completed), any other moment (e.g., duration or time point) of the harvest phase, etc. The Raman measurements for the training dataset may be obtained by in-line or offline Raman spectroscopy performed on a UF / DF pool conditioned at a fixed target matrix condition, and the associated product quality attributes may also be obtained or measured offline or in real time. For example, a sample of the conditioned UF / DF pool may be withdrawn at the fixed target matrix condition (e.g., at the end of the recovery phase), and its protein concentration may be measured offline using a protein concentration sensor (e.g., the SoloVPE System® by C Technologies, Inc.™). As another example, an osmometer may be utilized on the sample to measure the osmolality of the UF / DF pool. In such a case, the Raman measurements and the measured product quality attributes (e.g., measured protein concentration, measured osmolality, etc.) may be combined to form the training dataset 208 for that fixed target matrix condition. Such a training dataset 208, including Raman and product quality attribute measurements for the fixed target matrix condition, may be referred to as a design of experiments (DOE) dataset.In various embodiments, the machine learning model 210 trained using the DOE training dataset may be referred to as a DOE machine learning model.
[0118] In some examples, the training dataset 208 may be a dataset of Raman measurements and associated product quality attributes of the UF / DF pool obtained for a range of recovery phase matrix conditions. This range of recovery phase matrix conditions may correspond to a selected duration of the UF / DF pool recovery phase. For example, the duration may be the entire duration of the recovery phase from the beginning (e.g., the end of receiving or diluting the UF / DF pool in a recovery vessel) to the end (e.g., the end of the recovery or conditioning phase of the conditioned UF / DF pool). That is, the Raman measurements and associated product quality attributes may be obtained for matrix conditions from the beginning to the end of the conditioned UF / DF pool recovery operation. In some examples, the selected duration may be shorter than the entire duration of the recovery operation or conditioning phase. The Raman measurements and associated product quality attributes of the UF / DF pool obtained over a range of recovery phase matrix conditions facilitate analysis and determination of the behavior and trends of the UF / DF pool over the duration of the recovery operation. Thus, a training dataset 208 containing such measurements for a range of recovery stage matrix conditions may be referred to as a trend analysis (TA) training dataset. In various embodiments, a machine learning model 210 trained using a TA training dataset may be referred to as a "TA machine learning model" or a "process trend model."
[0119] In various embodiments, Raman measurements and associated product quality characteristics for a range of recovery stage matrix conditions may be obtained by performing multiple measurements across a desired target matrix condition. For example, different amounts of buffer may be added to the collection vessel of the conditioned UF / DF pool over the duration of the recovery run to reach the desired matrix condition, and Raman measurements and associated product quality characteristics of the UF / DF pool may be obtained after each buffer addition to obtain Raman measurements and associated product quality characteristics for that range of recovery stage matrix conditions (e.g., recovery stage matrix conditions from start to finish). In various examples, Raman measurements may be obtained by in-line Raman spectroscopy performed on the conditioned UF / DF pool. Additionally, associated product quality characteristics may also be obtained or measured offline. For example, samples of the conditioned UF / DF pool may be drawn across a range of recovery stage matrix conditions, and their protein concentrations may be measured offline using a protein concentration sensor (e.g., the SoloVPE System® by C Technologies, Inc.™). As another example, an osmometer may be utilized on the sample to measure the osmolality of the UF / DF pool over a range of recovery stage matrix conditions. In such a case, the Raman measurements may be combined with measured product quality attributes (e.g., measured protein concentration, measured osmolality, etc.) obtained for a range of recovery stage matrix conditions to form a TA training dataset.
[0120] In various embodiments, as described above, the in-line Raman measurements 206 can be Raman spectroscopic measurements of a UF / DF pool obtained by an in-line Raman spectrometer operably coupled to a collection vessel of the UF / DF pool undergoing a collection operation (e.g., including conditioning or dilution with buffer). In such cases, the trained machine learning model 210 and the in-line Raman measurements 206 can be used to determine product quality attributes, such as, but not limited to, protein concentration, osmolality concentration, etc., of the conditioned UF / DF pool in real time or near real time, without necessarily having to manually sample the conditioned UF / DF pool. For example, a technician may want to adjust the amount of buffer being added to the collection vessel in real time to improve the yield of valuable molecules (e.g., mAbs) in the collection vessel (e.g., to increase or decrease product quality attributes), while avoiding product contamination, wasted time and resources, and other deviations that may be associated with manual product sampling and product characteristic determination.
[0121] In such cases, in-line Raman measurements 206 of the conditioned UF / DF pool may be provided to a machine learning model 210, which analyzes the in-line Raman measurements 206 and predicts product quality attributes 212 of the conditioned UF / DF pool. In some examples, the in-line Raman measurements 206 may be obtained for a fixed target matrix condition, and in such examples, the in-line Raman measurements 206 may be provided to a DOE machine learning model for prediction of product quality attributes of the conditioned UF / DF pool. In some examples, the in-line Raman measurements 206 may be obtained for a range of recovery stage matrix conditions (e.g., recovery stage matrix conditions from start to finish), and in such cases, the in-line Raman measurements 206 may be provided to a TA machine learning model. In such cases, in various embodiments, the buffer used for the collection operation of the conditioned UF / DF pool (e.g., the conditioned UF / DF pool of the inline Raman measurements 206) may be the same as the buffer used for the collection operation of the UF / DF pool from which the training data set 208 is obtained.
[0122] FIG. 3 illustrates a machine learning model-based process 300 for predicting one or more product quality attributes of a bioprocess (e.g., UF / DF pool conditioning) according to various embodiments. In various embodiments, process 300 is implemented using a system, such as product quality attribute prediction system 200 of FIG. 2 and / or the processor of computing platform 150 of FIG. 1, to predict product quality attributes of a conditioned UF / DF pool in a collection vessel. For example, a UF / DF pool purified in a UF / DF skid or tank may be received in a collection vessel for a collection operation (which may include conditioning, e.g., via buffer addition (e.g., conditioning buffer, dilution buffer, etc.)), and a technician may be tasked with monitoring and controlling the collection operation. For example, to improve the yield of a collection operation (e.g., recovery of valuable molecules, such as mAbs, in the conditioned UF / DF pool), a technician may want to adjust the product quality attributes of the conditioned UF / DF pool in real time, without necessarily manually removing a sample of the conditioned UF / DF pool. In such cases, a technician may utilize process 300 to obtain a prediction of one or more product quality attributes (e.g., protein concentration, osmolality, etc.) of the conditioned UF / DF pool, which may then be used to determine whether to add or terminate buffer addition to the collection vessel.
[0123] In step 310, the processor receives Raman measurements of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool. The collection vessel (e.g., collection vessel 120) is downstream of the UF / DF operation (e.g., performed in UF / DF tank 110) and receives the buffer-conditioned UF / DF pool. For example, the processor may receive real-time Raman measurements of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel that receives the UF / DF pool for buffer conditioning downstream of an ultrafiltration / diafiltration (UF / DF) operation. In some examples, the UF / DF pool is a purified UF / DF pool received in the collection vessel from an upstream filtration vessel after undergoing purification. In some examples, the conditioned UF / DF pool includes an antibody, such as a monoclonal antibody (mAb).
[0124] In step 320, a processor predicts one or more product quality characteristics of the conditioned UF / DF pool using a machine learning model trained to receive Raman measurements of the conditioned UF / DF pool as input and generate one or more predicted product quality characteristics of the conditioned UF / DF pool as output. For example, the processor may apply real-time Raman measurements to the trained machine learning model to predict the product quality characteristics of the conditioned UF / DF pool. In some examples, the machine learning model may be trained to predict the one or more product quality characteristics using a training dataset including one or more product quality characteristics of conditioned UF / DF pools conditioned with different amounts of buffer added thereto and associated real-time Raman measurements of conditioned UF / DF pools conditioned with different amounts of buffer. In some examples, the trained machine learning model is a trained multivariate linear model, a trained neural network, a trained deep learning model, or a trained ensemble learning model.
[0125] In some cases, the one or more product quality characteristics include the osmolality of the conditioned UF / DF pool. In some cases, the one or more product quality characteristics include the protein concentration of the conditioned UF / DF pool.
[0126] FIG. 4 illustrates a flowchart of a machine learning model-based process 400 for controlling a UF / DF pool conditioning process (e.g., in a collection vessel) according to various embodiments. In various embodiments, process 400 is implemented using a system, such as product quality attribute prediction system 200 of FIG. 2 and / or computing platform 150 of FIG. 1 (e.g., a processor of computing platform 150), to predict one or more product quality attributes (e.g., protein concentration, osmolality concentration, etc.) of the conditioned UF / DF pool in the collection vessel and generate instructions to add or stop adding buffer to the collection vessel based on the predicted product quality attributes. For example, the instructions may include a buffer pump signal to send to a buffer pump, causing the buffer pump to add or stop adding conditioning buffer to the collection vessel to alter the product quality attributes of the conditioned UF / DF pool. The system may determine a weight indicator of the collection vessel or UF / DF pool that corresponds to the new product quality characteristic of the UF / DF pool (e.g., the weight of the collection vessel containing the UF / DF pool, the volume or height of the UF / DF pool in the collection vessel, etc.) and may generate a buffer pump signal that causes the buffer pump to add or stop adding buffer so that the collection vessel or UF / DF pool achieves the weight indicator.
[0127] In step 410, a processor of a computing platform receives a first Raman measurement of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool. The collection vessel is downstream of the UF / DF operation and receives the UF / DF pool for conditioning with a buffer. For example, the processor may receive the first Raman measurement of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel downstream of the UF / DF operation that receives the UF / DF pool for conditioning with a buffer. In some examples, the first Raman measurement is performed in real time during the UF / DF pool conditioning process. In some cases, the computing platform communicates with the Raman spectrometer and the buffer pump remotely and / or wirelessly, with a latency between reception and transmission of approximately 2 seconds or less. For example, the computing platform may be an IoT edge node, an on-premise virtual machine, a cloud server, or a cloud serverless solution. In some examples, the conditioned UF / DF pool comprises a monoclonal antibody (mAb).
[0128] In step 420, the processor predicts one or more product quality characteristics of the conditioned UF / DF pool using a trained machine learning model that receives the first Raman measurements as input. The one or more product quality characteristics include a first protein concentration of the conditioned UF / DF pool. For example, the processor may use the trained machine learning model to predict the first protein concentration of the conditioned UF / DF pool based on analysis of the first Raman measurements. In some cases, the prediction includes predicting the osmolality concentration of the conditioned UF / DF pool based on analysis of the first Raman measurements. In some examples, the trained machine learning model may be a trained multivariate linear model, a trained neural network, a trained deep learning model, a trained ensemble learning model, or a combination thereof.
[0129] In step 430, the processor calculates a collection vessel weight index corresponding to a second protein concentration of the conditioned UF / DF pool that is different from the first protein concentration of the conditioned UF / DF pool.
[0130] In step 440, the processor outputs instructions to add or stop adding buffer to the collection vessel based on the calculated weight index of the collection vessel. For example, the processor may send a buffer pump signal to a buffer pump operably coupled to the collection vessel configured to instruct the buffer pump to add or stop adding buffer to the collection vessel based on the calculated weight index of the collection vessel. In some examples, the buffer is added in an amount sufficient to increase the weight of the collection vessel to within a weight threshold of the calculated weight. In various embodiments, the addition of buffer to the collection vessel occurs without collecting a sample of the conditioned UF / DF pool from the collection vessel to measure the first protein concentration. In some examples, in response to the instruction to add buffer, the buffer may be added in an amount sufficient to increase the weight index of the collection vessel to within a weight threshold of the calculated weight index.
[0131] In various embodiments of method 400, the processor further receives, after transmitting, a second Raman measurement of the conditioned UF / DF pool from the Raman spectrometer. In such a case, the processor is configured to predict a third protein concentration of the conditioned UF / DF pool using a trained machine learning model that incorporates the second Raman measurement as input. Furthermore, the processor compares the second protein concentration to the third protein concentration to determine the effectiveness of adding buffer to the collection vessel.
[0132] In some examples, method 400 does not use any measurements obtained by extracting a sample of the conditioned UF / DF pool from the collection vessel. For example, the addition of buffer to the collection vessel may be performed without the need to collect or extract a sample of the conditioned UF / DF pool from the collection vessel to measure (e.g., offline) any product quality attribute (e.g., any protein concentration or osmolality concentration).
[0133] IV. Machine Learning Modeling of UF / DF Quality Attribute Monitoring and Control 5-11B show diagrams illustrating exemplary implementations of the techniques disclosed herein for constructing (e.g., constructing, training, testing, etc.) machine learning models configured to predict product quality attributes of conditioned UF / DF pools. The constructed machine learning models were linear partial least squares (PLS) regression models (hereinafter referred to as "PLS models"), deep neural network models, or linear least absolute shrinkage and selection operator (lasso) models. However, it should be understood that such regression models were selected for illustrative reasons only, and that any of the machine learning models disclosed herein may be constructed as described herein to predict product quality attributes of conditioned UF / DF pools. In a first aspect of the exemplary embodiment, a PLS model trained to predict product quality attributes of small conditioned UF / DF pools (e.g., approximately 25 mL) was found to accurately predict product quality attributes of conditioned UF / DF pools that were more than an order of magnitude larger, demonstrating the scalability of the constructed PLS model. Furthermore, two types of PLS models were constructed. The "DOE PLS model" was trained on the design of experiments (DOE) training dataset, and the "TA PLS model" was trained on the trend analysis (TA) training dataset. It was shown that the DOE PLS model predicted poorly on the TA test dataset, and the TA PLS model predicted poorly on the DOE test dataset. It was recommended that the TA PLS model should be used for end-to-end monitoring and control of the conditioning phase, and the DOE PLS model should be used for tighter control of target conditions to enable real-time release testing of the drug substance. Alternatively, a comprehensive model using a combination of DOE and TA training datasets could be developed to demonstrate the ability to simultaneously predict quality attributes at both end-to-end and target conditions.
[0134] FIG. 5 illustrates two categories of experiments and analyses performed to develop a model for real-time monitoring and control of the UF / DF conditioning process and target conditions, according to various embodiments. In various embodiments, design of experiments (DOE) analysis refers to an analysis performed to model multiple protein concentrations over a predetermined concentration range using target matrix conditions at the end of the UF / DF pool conditioning step or other predefined moment in the bioprocess. In some aspects, the predetermined range may encompass all concentrations expected at the endpoint of the UF / DF process or other predefined moment in the bioprocess. For example, for a UF / DF conditioning step beginning with diluted pool 550 at time t_0510 and concluding with conditioned pool 560 at time t_final 520, DOE analysis 570 may refer to an analysis performed to model a range of protein concentrations (e.g., including initial protein concentration 530 and final protein concentration 540) using target matrix conditions for the range of protein concentrations at the completion of the UF / DF pool conditioning step at t_final 520. In some instances, such experiments may achieve tighter control over target conditions than experiments relying on trend analysis (TA), as described below, and may allow for real-time release of product.
[0135] In various embodiments, trend analysis (TA) of a UF / DF conditioning or dilution process refers to an analysis performed to model the entire range of protein concentrations from the beginning to the end (“end-to-end”) of the conditioning phase. For example, for a UF / DF conditioning step beginning with diluted pool 550 at time t_0 510 and concluding with conditioned pool 560 at time t_final 520, process TA 580 may refer to an analysis performed to model the entire range of protein concentrations from the beginning 510 to the end 520 of the conditioning phase (e.g., including initial protein concentration 530, final protein concentration 540, and any additional protein concentrations in between for UF / DF pool 590 under dilution or conditioning). In some examples, such experiments can achieve real-time monitoring and control of UF / DF conditioning.
[0136] The scalability of the trained PLS model was demonstrated using a 25 mL conditioned UF / DF pool and a 700 mL conditioned UF / DF pool. The PLS model was trained using the Raman training dataset and osmolality measurements of the 25 mL conditioned UF / DF pool to predict the osmolality of the 700 mL conditioned UF / DF pool. In setting up the scalability experiment, samples of the conditioned UF / DF pool were created using a combination of two different buffers: a modified diafiltration buffer supplemented with polysorbate 20 and a modified conditioning buffer supplemented with low polysorbate 20. Table 1 below lists the trehalose concentrations (mM) of the samples and the corresponding measured osmolality.
[0137] [Table 1]
[0138] Figure 6A shows a plot of Raman scans or measurements collected for the larger 700 mL conditioned UF / DF pool contained in the minifarm. First, a collection vessel was filled with 25 mL of the UF / DF pool, and Raman scans or measurements were collected using an in-line Raman probe. Raman scans were collected using a Raman spectrometer set at 15 seconds of exposure, 10 counts, and 400 mW of power. The collection mode was set to periodic, with a capture time of 3 minutes. The collection vessel was cleaned and sanitized with isopropyl alcohol (IPA) before conducting scalability experiments for each sample listed in Table 1. The minifarm was then filled with the 700 mL UF / DF pool. The minifarm was disassembled, cleaned, and sanitized with water and IPA before conducting experiments for each sample listed in Table 1. The Raman scans for the larger 700 mL conditioned UF / DF pool and the smaller 25 mL UF / DF pool differed slightly at wavenumbers below 500 and above 3000. Reproducibility was acceptable, as demonstrated by the similarity of the Raman scans (at least for wavenumbers between 500 and 3000) between the larger 700 mL conditioned UF / DF pool and the smaller 25 mL UF / DF pool.
[0139] A PLS model was trained using the Raman training dataset and osmolality measurements of a 25 mL conditioned UF / DF pool, then tested using the Raman measurements of a 700 mL conditioned UF / DF pool as input data to predict the osmolality of a 700 mL conditioned UF / DF pool. Osmolality values were measured offline for all conditioned UF / DF pool samples in Table 1. The PLS model used wavenumbers from 500 to 1500, standard normal variance (SNV) and first derivative (1st derivative) preprocessing, and one latent variable using inline Raman scanning on a 25 mL container. The osmolality predictions of the trained PLS model are shown in Figure 6B, compared to the measured osmolality values. The small root mean square error of prediction (RMSEP) and high correlation coefficient of prediction confirm that the PLS model trained using the training dataset collected from 25 mL conditioned UF / DF pools is able to predict the osmolality concentration of 700 mL conditioned UF / DF pools, demonstrating the scalability of the trained PLS model.
[0140] DOE machine learning models (e.g., including the DOE PLS model) and TA machine learning models (e.g., including the TA PLS model) were also developed using 25 mL and 700 mL conditioned UF / DF pools (e.g., the same ones in the experiments described above). Table 2 below shows the experimental design for developing the DOE and TA PLS models trained to predict the protein and osmolality concentrations of conditioned UF / DF pools. The DOE training and testing datasets were obtained from small-scale collection vessel experiments on conditioned UF / DF pool samples 1–13, which focused on capturing the behavior of the full range of protein concentrations using the target matrix conditions (i.e., DOE analysis). In addition to inline Raman scans, offline measurements, including protein and osmolality concentrations, were also collected for all conditioned UF / DF pool samples. The TA training dataset was obtained from small-scale collection vessel experiments on samples 22–29, which focused on capturing the behavior of the full range of protein concentrations using the target matrix conditions. Samples 14-21 are samples in a large collection vessel containing the same material as samples 14-21 in the small collection vessel. Thus, samples 22-29 and 14-21 have similar chemical properties but at different scales (25 mL vs. 700 mL). Furthermore, although the training and test samples were developed independently at different times and dates, both the training and test datasets corresponded to the same molecules and had consistent buffer properties at all stages of the experimental run.
[0141] [Table 2] TIFF2026502096000004.tif18170
[0142] In Table 2, "DOE analysis" refers to experiments corresponding to the final target specification for the recovery or conditioning stage, and "trend analysis" refers to experiments corresponding to the full specification for the recovery or conditioning stage. DOE samples of conditioned UF / DF pools were made using a combination of two different buffers (modified conditioning buffer supplemented with low polysorbate 20, and formulation buffer (1 part 10X standard conditioning buffer diluted with 9 parts standard diafiltration buffer) and monoclonal antibody raw material. Raman and product quality attribute (e.g., protein concentration and osmolality) measurements for the training dataset for the DOE PLS model were obtained from samples 1–7, and in-line Raman measurements for the testing dataset for the DOE PLS model were obtained from samples 8–13. "Trend" conditioned UF / DF pool samples in large collection vessels (samples 14–21) used to obtain the testing dataset for testing the TA PLS model were made using a combination of two different buffers (10X standard conditioning buffer and diafiltration buffer) and monoclonal antibody raw material to reach the conditions shown in Table 2. The training dataset for training the PLS model was obtained from samples 22–29, which were approximately 25–30 mL of conditioned UF / DF pool samples removed from samples 14–21, respectively, after the Raman scans of the latter samples had been collected. Samples were used to perform inline Raman scans in the 25 mL collection vessels, and offline measurements, including protein concentration, offline Raman scans, and osmolality, were performed. To ensure experimental accuracy and reliability, five Raman spectra were collected per sample and averaged via a tabulation method for averaging. Protein concentration was measured using a protein concentration sensor (SoloVPE System® by C Technologies, Inc.™), and osmolality was measured using an osmometer.
[0143] The amount of buffer needed to achieve a stepwise transition from samples 14-21 or samples 22-29 was calculated, and that amount of buffer was added to the collection vessel via pipette. The impeller was turned on for approximately 5 minutes to allow the buffer to mix into the UF / DF pool. The same steps were repeated until the target protein concentration of 43.5 g / L was reached, as shown in Table 2.
[0144] Figure 7A shows inline Raman scans or measurements of conditioned UF / DF pool samples 8–13. The protein concentrations of these samples were measured using a protein concentration sensor (SoloVPE). The scans are highly consistent, with the majority of variation attributed to changes in protein concentration. The DOE training dataset includes the Raman measurements and their corresponding protein concentration measurements. The DOE training dataset was used to construct a DOE PLS model using first derivatives for wavenumbers 500–1850, standard normal variance (SNV), and UV scaling. As used herein, UV scaling refers to any modeling process that projects the surface of a 3D mesh into 2D space. The DOE PLS model used one latent variable. The DOE PLS model was then tested on the DOE test dataset of inline Raman measurements of samples 8–13. Figure 7B shows the performance of the measured protein concentrations compared to the predicted protein concentrations by the DOE PLS model. The small correlation coefficient of R2 = 0.9721 and RMSEP = 1.31352 confirms that the DOE PLS model has acceptable performance for protein concentration.
[0145] Figure 8A shows the in-line Raman scans or measurements of conditioned UF / DF pool samples 14-21. Figure 8B shows the performance of the TA PLS model predicting protein concentrations for the TA test dataset. The TA PLS model trained using Raman and protein concentration measurements of UF / DF conditioned pool samples 22-29 was tested. It was used to predict protein concentrations for UF / DF conditioned pool samples 14-21. The TA PLS model used wavenumbers 500-1850, first derivative, SNV, and UV scaling for data preprocessing. One latent variable was selected. The R2 = 0.9989 and RMSEP = 0.6938 shown in Figure 8B confirm acceptable performance and the feasibility of scaling up the model (e.g., from 25 mL to 750 mL).
[0146] Figure 9 shows the performance of the TA PLS model in predicting the osmolality concentration of the TA test dataset. The TA PLS model, trained using Raman and osmolality measurements of UF / DF-conditioned pools of samples 22–29, was tested. It was used to predict the osmolality concentration of UF / DF-conditioned pools of samples 14–21. The TA PLS model used wavenumbers 500–1500, first derivative, SNV, and UV scaling for data preprocessing. One latent variable was selected. The R2 = 0.9979 and RMSEP = 4.90182 shown in Figure 9 confirm acceptable performance and feasibility of scale-up (e.g., from 25 mL to 750 mL).
[0147] The performance of the DOE PLS model predicting protein concentrations for the "trend" conditioned UF / DF pool samples (samples 14-29) is typically not as good as that of the TA PLS model (e.g., RMSEP = 6.924). This can be seen from a comparison of the measured protein concentrations with the predictions of the DOE PLS model (e.g., using in-line Raman measurements of samples 14-21 or 22-29 as the TA test data set). The performance of the DOE PLS model predicting protein concentrations for the TA test data set is good when the protein concentration is approximately 50 g / L, an endpoint that shares comparable matrix composition between the DOE and trend analyses. The results show that TA PLS models derived from or trained with a TA training dataset are generally better suited to predicting product quality attributes based on a TA test dataset, i.e., a test dataset including Raman measurements of conditioned UF / DF pools obtained for a range of recovery stage matrix conditions (e.g., start to finish recovery stage matrix conditions), compared to DOE PLS models derived from a DPE training dataset or trained with a DOE training dataset.
[0148] Similar to the performance of the DOE PLS model in predicting the protein concentration of the "trend" conditioned UF / DF pool samples, the TA PLS model's performance in predicting the protein concentration of the DOE-conditioned UF / DF pool samples (samples 8–13) was not as good as the DOE PLS model. Except for a protein concentration of ∼50 g / L, i.e., the endpoint, which shares equivalent matrix composition between DOE and trend analysis, the error of prediction by the TA PLS model can be larger (e.g., RMSEP = 4.713), indicating that TA PLS models derived from or trained with the TA training dataset are similarly generally less suitable than DOE PLS models for predicting product quality attributes based on DOE test datasets, i.e., test datasets containing Raman measurements of conditioned UF / DF pools obtained for a fixed target range of matrix conditions.
[0149] Alternatively, a comprehensive model can be developed that includes training data from both DOE and TA. The comprehensive model can include both TA and target DOE characteristics to accurately predict target and trend conditions. For example, a comprehensive model was developed that included training data from both DOE (samples 1–7) and TA (samples 22–29). The comprehensive model can include both TA and target DOE characteristics to accurately predict target and trend conditions simultaneously. A PLS model based on a combination of wavenumbers 500–1850 for data preprocessing, as well as first-derivative, SNV, and UV scaling, is developed. Three latent variables are selected using 7-fold cross-validation. The performance of the comprehensive model is first evaluated on the DOE test dataset (samples 8–13) in Figure 10A. The correlation coefficient, R2 = 0.987, and a small RMSEP = 0.905926 confirm that the model has acceptable performance for protein concentration prediction. Next, the performance of the comprehensive model is evaluated on the process trend test dataset (samples 14–21) in Figure 10B. The small correlation coefficient of R2 = 0.997 and RMSEP = 1.15969 confirms that the model has acceptable performance for protein concentration prediction.
[0150] The lasso model was constructed as a result of a four-fold cross-validation method using the TA training dataset (conditioned UF / DF pooled samples 22-29) with a regularization parameter α = 0.006. Lasso regression automatically removes model coefficients that do not contribute to error minimization by a penalty term α. The model was tested on the TA testing dataset (samples 14-21) in Figure 11A. R2 = 0.996 and RMSEP = 0.88 confirm that the lasso model can be used as an alternative to the PLS model used in Figure 6B.
[0151] A deep neural network (DNN) model was constructed using the TA training dataset (conditioned UF / DF pool samples 22–29) to evaluate whether deploying a nonlinear model capturing nonlinear correlations could result in better performance. The optimal DNN hyperparameters (as a result of hyperparameter tuning) included an input layer, two hidden layers (four neurons each), a RELU activation function for the input and hidden layers, and a linear activation function for the output with L1 = 0.008 to normalize the input layer to avoid overfitting. The DNN model predictions of protein concentrations for the trend test dataset (conditioned UF / DF pool samples 14–21) compared to the measured protein concentrations are shown in Figure 11B. R2 = 0.997 and RMSEP = 0.52 confirm that the DNN model can be used as an alternative to the PLS model in Figure 6B.
[0152] The above data (e.g., training dataset) and resulting model demonstrate that alternative models can yield similar performance to the PLS model described above.
[0153] In various embodiments, the ability to predict product quality attributes of the conditioned UF / DF pool in the collection vessel allows for real-time monitoring and control of collection operations. An in-line spectrometer may perform Raman scans of the conditioned UF / DF pool and transmit the Raman measurements to a computing platform hosting a machine learning model trained to predict product quality attributes based on the analysis of the Raman measurements. The computing platform may be on-site or at the UF / DF collection vessel site, or it may communicate with the UF / DF collection vessel remotely. For example, the computing platform may be an IoT edge node, an on-premise virtual machine, a cloud server, a cloud serverless solution, or a combination thereof, and may communicate with the UF / DF collection vessel with a latency of less than about 3 seconds, about 2 seconds, or about 1 second, including values and subranges therebetween.
[0154] In various embodiments, the machine learning model may predict product quality attributes of the conditioned UF / DF pool (e.g., but not limited to, protein concentration, osmolality concentration, etc.), which may then be signaled to the UF / DF collection vessel or its control unit with very low latency. For example, the duration between the in-line Raman scan and the arrival of the signal at the collection vessel or its control unit may be less than about 3 seconds, about 2 seconds, or about 1 second, including values and subranges therebetween. In this manner, the prediction of the product quality attribute may be used for real-time or near-real-time monitoring of the progress of the conditioned UF / DF collection operation. Furthermore, the prediction may also be used to control the collection operation. For example, the computing platform may determine that a product quality attribute value different from that predicted may indicate improved quality of, or recovery of, the conditioned UF / DF pool. In such cases, the computing platform may send a signal, such as a signal configured to cause the addition of buffer to the collection vessel or its termination, to change or adjust the protein concentration or osmolality of the UF / DF pool to a value indicative of a target or improved quality of the conditioned UF / DF pool.
[0155] V. Exemplary Experimental Demonstration of Real-Time Monitoring and Control of a Conditioned UF / DF Pool Monitoring of product quality attributes of the conditioned UF / DF pool is illustrated with reference to Figures 12 and 13. Figure 12 shows a time series plot of experimental data comparing real-time protein concentration predictions using the process trending model described above with the protein concentrations of samples taken throughout the experiment and measured using a protein concentration sensor (e.g., SoloVPE). Figure 13 shows a time series plot of experimental data comparing osmolality concentration predictions using Raman spectra collected from the real-time experiment of Figure 12 as input to the process trending model with the osmolality of samples taken throughout the experiment and measured offline using an osmometer. Experiments demonstrating the disclosed monitoring of UF / DF pools were carried out using the conditions shown in Table 3 below, where the first step corresponds to filling a collection vessel with the UF / DF diluted pool (mAb feedstock), and each addition step represents a subsequent dilution with conditioning buffer (20 mM sodium acetate, 1057 mM (40%) trehalose, 0.20% w / v polysorbate 20, pH 5.3). The collection vessel used in the experiment was typically used to house and grow cell cultures and other media and was suitable for storing the UF / DF pool and receiving the amount of conditioning buffer specified in the table below. Additionally, the vessel was configured to accept an in-line or in-situ probe for Raman spectroscopy measurements.
[0156] [Table 3]
[0157] Raman scans of the collection vessel were obtained from the start of filling the collection vessel with conditioning buffer using an in-line Raman spectrometer operably coupled to the collection vessel. Raman scans were performed for 15 minutes at a rate of one scan every 3 minutes, for a total of five Raman scans per condition. Mixing was performed throughout the experiment. The Raman measurements were compared with measurements of product quality attributes and submitted to a trained machine learning model that calculated the predicted protein concentration and osmolality of the conditioned UF / DF pool, shown in Figures 12 and 13, respectively. The trained machine learning model was a supervised learning model trained using a training dataset containing multiple Raman measurements of UF / DF pools with known protein concentrations and known osmolality of the UF / DF pools. The supervised machine learning model was trained to receive Raman measurements as input and produce the protein concentration and osmolality of the UF / DF pool as output. The protein concentration in Figure 12 goes through several steps, starting with an initial concentration of 55.6 g / L. At each step of buffer addition, the predicted protein concentration 1210 at least substantially matches the protein concentrations 1220-1270 measured by the protein concentration sensor, demonstrating that the disclosed methods and systems enable real-time monitoring of product quality attributes of the conditioned UF / DF pool over the course of a recovery operation using in-line Raman spectroscopy.
[0158] For example, during the first 30 minutes of the UF / DF pool collection run, the measured protein concentration 1270 substantially matched the predicted protein concentration 1210 for the same time period. Furthermore, when conditioning buffer was added to the collection vessel, the protein concentration of the UF / DF pool changed, which was measured with the protein concentration sensor. The protein concentration was also predicted, as described herein. The new measured protein concentration 1260 was found to substantially match the predicted protein concentration 1210 for that same time period. The same was true for further additions of conditioning buffer to the collection vessel, where measured protein concentrations 1250, 1240, 1230, and 1220 during the time periods following the addition of conditioning buffer substantially matched the predicted protein concentration 1210 for the corresponding time periods. FIG. 14 shows a scatter plot of actual (true measurements) versus predicted (averaged over a window of the last five measurements) protein concentrations corresponding to the time series plot of FIG. 12 , according to various embodiments. R2=0.98627 and RMSEP=0.46116 in Figure 14 confirm the acceptability of the results in Figure 12.
[0159] Referring to FIG. 13, Raman scans collected during the real-time protein concentration experiment described above were used to perform an analysis to predict the osmolality of the conditioned UF / DF pool. The osmolality of the conditioned UF / DF pool measured by the osmometer for Samples 1-6 in Table 3, which at least substantially matches the osmolality predicted by the machine learning model, demonstrates that the disclosed method and system enables real-time monitoring of the product quality attributes of the conditioned UF / DF pool over the course of a recovery operation. FIG. 15 shows a scatter plot of actual osmolality versus predicted osmolality (averaged over a window of the last five measurements) corresponding to the time series plot in FIG. 13, according to various embodiments. The R2=0.9977 and RMSEP=5.89485 in FIG. 15 confirm the acceptability of the results in FIG. 13.
[0160] An experiment demonstrating the disclosed control of the conditioned UF / DF pool using machine learning model product quality attribute predictions was conducted using the conditions shown in Table 4 below. The first step corresponds to filling the collection vessel with the UF / DF diluted pool (mAb feedstock), and each addition step represents a subsequent dilution with conditioning buffer (20 mM sodium acetate, 1057 mM (40%) trehalose, 0.20% w / v polysorbate 20, pH 5.3). The addition of protein and buffer to the collection vessel was fully automated using a cloud server control system that remotely communicates with an on-site control system (using Kepware™ connectivity solution) for the conditioned UF / DF pool collection operation.
[0161] [Table 4]
[0162] In step 1 of the experiment, a collection vessel was filled with a UF / DF pool having a protein concentration of approximately 55.6 g / L. The protein concentration of the initial UF / DF pool was determined as discussed above. Initially, the UF / DF pool was scanned with an inline Raman spectrometer, and the Raman measurements were provided to a remote computing platform hosting a trained machine learning model and communicating with the collection vessel. The Raman scans of the UF / DF pool were provided to the computing platform, and the resulting trained machine learning model analyzed the Raman measurements and predicted the protein concentration of the UF / DF pool (approximately 55.6 g / L). The trained machine learning model was a supervised learning model previously trained using a training dataset containing multiple Raman measurements of UF / DF pools with known protein concentrations and known osmolality concentrations of the UF / DF pools. The supervised machine learning model was trained to receive the Raman measurements as input and generate the protein concentration and osmolality concentration of the UF / DF pool as output.
[0163] Step 2 corresponds to adjusting the protein concentration of the UF / DF pool in the collection vessel by adding conditioning buffer. Mixing was performed throughout the experiment. For example, if a technician operating the collection operation wanted to achieve a target protein concentration of the UF / DF pool at approximately 50 g / L, the remote computing platform determined the amount of conditioning buffer that needed to be added to the collection vessel to adjust the protein concentration of the UF / DF pool to the desired value (50 g / L). The computing platform accessed a mapping (e.g., a formula, table, model, etc.) relating weight indicators of the collection vessel or UF / DF pool to protein concentration values.
[0164] As indicated by the weight indicator, weight generally correlated positively with protein concentration values. In this example, the weight indicator used was a typical real-time weight sensor used in various bioprocesses. This sensor provided a direct, continuous signal measuring the weight of the collection vessel and UF / DF pool. However, in some embodiments, it is contemplated that weight can also be indirectly indicated (inferred) from other sensors. For example, a real-time measurement of vessel height when multiplied by a constant area and density of the material can be used to provide a real-time estimate of weight. It is also contemplated that in some embodiments, the weight indicator can be used to stop the process when a target weight value is achieved. The target weight value can be directly related to the target protein concentration given the initial protein concentration and a simple mass balance, taking into account the starting and target conditions.
[0165] Thus, the computing platform used the mapping to determine or calculate a weight indicator of the collection vessel or UF / DF pool corresponding to a protein concentration of 50 g / L. Examples of weight indicators include the weight of the collection vessel (including the UF / DF pool contained therein), the height or volume of the UF / DF pool, etc. Using the weight of the collection vessel as a non-limiting illustrative example, the computing platform determined or calculated the weight of the collection vessel corresponding to a protein concentration of 50 g / L. The computing platform then determined the amount of buffer that needed to be added to the collection vessel so that its weight at least substantially matched the determined or calculated weight. It should be understood that the computing platform can similarly calculate the weight of the collection vessel for other protein concentrations and determine the amount of buffer that needs to be added to the collection vessel so that its weight increases to the calculated weight. For example, the computing platform can calculate the weight of the collection vessel when the protein concentration is approximately 55.6 g / L and determine the amount of buffer that needs to be added to the collection vessel so that its weight increases to the calculated weight (e.g., corresponding to a protein concentration of 50 g / L). After the desired weight was reached, buffer addition was automatically stopped by stopping the feed pump, and a signal was automatically sent from the computing platform to the feed pump to set the flow rate set point to zero.
[0166] Step 3 corresponds to adjusting the protein concentration of the UF / DF pool in the collection vessel from 50 g / L to a second target concentration of 45 g / L using the same process described above with reference to steps 1 and 2.
[0167] In some examples, Raman measurements of the conditioned UF / DF pool may be performed to assess the effectiveness of buffer addition. For example, these Raman measurements may be processed or analyzed by a machine learning model to predict the protein concentration of the conditioned UF / DF pool, which can then be compared to a target protein concentration to assess the effectiveness of buffer addition. As will be appreciated, predicted quality attributes at target conditions are used to determine release of the drug substance, allowing for immediate release of the product. For example, the predicted product quality attribute may be compared to a predetermined threshold value of the predicted product quality attribute to determine whether the conditioned UF / DF pool is ready for immediate release.
[0168] 16-17 show plots illustrating real-time control of the conditioned UF / DF pool in the collection vessel based on real-time adjustment of the collection vessel's weight. Furthermore, as an illustrative example, FIGS. 16 and 17 show results of experimental demonstrations in which predicted protein concentrations using the disclosed techniques substantially matched measured protein concentrations at the end of a collection run, within 0.5% (an offset of approximately 0.23373 g / L) in the former case (i.e., the experimental run of FIG. 16) and within 0.9% (an offset of approximately 0.40063 g / L) in the latter case (i.e., the experimental run of FIG. 17). As previously described, upon determining the amount of buffer to be added to the collection vessel, the computing platform generates and sends a signal to the buffer pump, causing the buffer pump to add the determined amount of buffer and automatically stop when the determined amount is reached. Therefore, with respect to steps 1 and 2 in Table 4, the computing platform generated and sent a buffer pump signal to the buffer pump, causing it to add approximately 146.3 g of conditioning buffer and automatically stop after reaching the target weight. Raman measurements of the UF / DF pool after buffer addition were acquired in real time, and the trained machine learning model was used to analyze the Raman measurements and calculate the protein concentration. Figure 16 shows that upon buffer addition, the conditioned UF / DF pool achieved a target protein concentration of approximately 50 g / L (with an offset of approximately 0.23373 g / L). Furthermore, Figure 16 shows the signals to add buffer and automatically stop after reaching the target weight, as indicated by the high and low modes of the lines corresponding to Feed_Flow_Setpoint (lane 1) and Feed_Pump_Analog_Output (lane 2), respectively. Additionally, Figure 16 shows the addition of buffer and the corresponding increase in vessel weight until the vessel weight reaches the target weight, which is reflected by the lines corresponding to Totalized_Flow (lane 3) and Vessel_Weight reaching the target weight (lane 4).
[0169] Similarly, Figure 17 is another run showing that adding buffer from step 2 to step 3 of Table 4 resulted in a conditioned UF / DF pool that achieved a target protein concentration of approximately 45 g / L (with an offset of approximately 0.40063 g / L). Similar to Figure 16, Figure 17 also shows the signal to add buffer, followed by automatic halting of addition after the target weight is reached, as indicated by the high and low modes of the lines corresponding to Feed_Flow_Setpoint (lane 1) and Feed_Pump_Analog_Output (lane 2), respectively. Furthermore, Figure 17 shows the addition of buffer and the corresponding increase in vessel weight until the vessel weight reaches the target weight, reflected by the lines corresponding to Totalized_Flow (lane 3) and Vessel_Weight (lane 4). Figures 16 and 17 demonstrate the accuracy of the disclosed technique for predicting target product quality attributes. In the former case, the final protein concentration was predicted within approximately 99.5% accuracy, and in the latter case, the final protein concentration was predicted within approximately 99.1% accuracy. Furthermore, note that the entire process, from Raman scanning of the UF / DF pool in step 1 to the machine learning model's determination of the conditioned UF / DF pool's protein concentration (approximately 45 g / L) after buffer addition in step 3, occurred within approximately 15 minutes. Furthermore, it was observed that the latency between the Raman measurement provided to the machine learning model and the buffer pump signal received at the buffer pump was approximately 1 second or less. Thus, Figures 16-17 demonstrate real-time control of the recovery behavior of the UF / DF pool, as its protein concentration was conditioned with buffer from approximately 55.6 g / L to approximately 50 g / L and then to approximately 45 g / L (the spikes in Figures 16-17 were due to experimental disturbances). While the above discussion relates to controlling recovery operations by predicting the protein concentration of the UF / DF pool, it should be understood that the disclosed techniques apply equally when product quality attributes, such as the osmolality of the UF / DF pool, are used to characterize the UF / DF pool.
[0170] VI. Computer-Implemented Systems FIG. 18 is a block diagram of a computer system according to various embodiments. Computer system 1800 may be an example of one implementation of product quality attribute prediction system 200 described above in FIG. 2. In one or more examples, computer system 1800 may include a bus 1802 or other communication mechanism for communicating information and a processor 1804 coupled with bus 1802 for processing information. In various embodiments, computer system 1800 may also include memory, which may be random access memory (RAM) 1806 or other dynamic storage device, coupled to bus 1802 for determining instructions to be executed by processor 1804. Memory may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 1804. In various embodiments, computer system 1800 may further include read-only memory (ROM) 1808 or other static storage device coupled to bus 1802 for storing static information and instructions for processor 1804. A storage device 1810, such as a magnetic disk or optical disk, may be provided and coupled to bus 1802 for storing information and instructions.
[0171] In various embodiments, computer system 1800 may be coupled via bus 1802 to a display 1812, such as a cathode ray tube (CRT) or liquid crystal display (LCD), for displaying information to a computer user. An input device 1814, including alphanumeric and other keys, may be coupled to bus 1802 for communicating information and command selections to processor 1804. Another type of user input device is a cursor control device 1816, such as a mouse, joystick, trackball, gesture input device, eye-gaze-based input device, or cursor direction keys, for communicating directional information and command selections to processor 1804 and controlling cursor movement on display 1812. This input device 1814 typically has two degrees of freedom, a first axis (e.g., x) and a second axis (e.g., y), that allow the device to specify a position in a plane. However, it should be understood that input devices 1814 that allow three-dimensional (e.g., x, y, and z) cursor movement are also contemplated herein.
[0172] Consistent with a particular implementation of the present teachings, results may be provided by computer system 1800 in response to processor 1804 executing one or more sequences of one or more instructions contained in RAM 1806. Such instructions may be read into RAM 1806 from another computer-readable medium or computer-readable storage medium, such as storage device 1810. Execution of the sequences of instructions contained in RAM 1806 may cause processor 1804 to perform the processes described herein. Alternatively, hard-wired circuitry may be used in place of or in combination with software instructions to implement the present teachings. Thus, implementation of the present teachings is not limited to any specific combination of hardware circuitry and software.
[0173] As used herein, the terms “computer-readable medium” (e.g., data store, data storage, data storage, etc.) or “computer-readable storage medium” refer to any medium that participates in providing instructions to processor 1804 for execution. Such a medium may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical, solid-state, and magnetic disks, such as storage device(s) 1810. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM 1806. Examples of transmission media include, but are not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 1802.
[0174] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or any other magnetic medium, CD-ROMs, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROMs, and EPROMs, flash EPROMs, any other memory chip or cartridge, or any other tangible medium from which a computer can read.
[0175] In addition to computer-readable media, instructions or data may be provided as signals on a transmission medium included in a communication device or system to provide one or more sequences of instructions to the processor 1804 of the computer system 1800 for execution. For example, a communication device may include a transceiver having signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in the disclosure herein. Representative examples of data communication transmission connections may include, but are not limited to, a telephone modem connection, a wide area network (WAN), a local area network (LAN), an infrared data connection, an NFC connection, an optical communication connection, etc.
[0176] It should be understood that the methodologies, flowcharts, diagrams, and accompanying disclosure described herein can be implemented using computer system 1800 as a standalone device or over a distributed network of shared computer processing resources, such as a cloud computing network.
[0177] The methodologies described herein may be implemented by various means depending on the application. For example, the methodologies may be implemented in hardware, firmware, software, or any combination thereof. In the case of a hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, or combinations thereof.
[0178] In various embodiments, the methods of the present teachings may be implemented as firmware and / or software programs and applications written in conventional programming languages such as C, C++, Python, etc. When implemented as firmware and / or software, the embodiments described herein may be implemented on a non-transitory computer-readable medium having stored thereon a program for causing a computer to perform the above-described methods. It should be understood that the various engines described herein may be provided on a computer system, such as computer system 1800, whereby processor 1804 performs the analyses and decisions provided by these engines according to instructions provided by memory components RAM 1806, ROM 1808, or storage device 1810, and / or user input provided via input device 1814.
[0179] VII. Example of Machine Learning Model Implementation FIG. 19 illustrates an exemplary neural network that can be used to implement a machine learning model according to various embodiments of the present disclosure. For example, in various embodiments, the machine learning model 210 may be implemented using a neural network 1900. As shown, the artificial neural network 1900 includes three layers: an input layer 1902, a hidden layer 1904, and an output layer 1906. Each layer 1902, 1904, 1906 may include one or more nodes. For example, the input layer 1902 includes nodes 1908-1914, the hidden layer 1904 includes nodes 1916-1918, and the output layer 1906 includes node 1922. In this example, each node in a layer is connected to all nodes in adjacent layers. For example, node 1908 in the input layer 1902 is connected to both nodes 1916 and 1918 in the hidden layer 1904. Similarly, node 1916 in the hidden layer is connected to all of nodes 1908-1914 in the input layer 1902 and node 1922 in the output layer 1906. Although only one hidden layer is shown in the artificial neural network 1900, it is contemplated that the artificial neural network 1900 used to implement the machine learning model 210 may include any number of hidden layers as needed or desired.
[0180] In this example, the artificial neural network 1900 receives a set of input values 1922-1928 and generates an output value 1930. Each node in the input layer 1902 may correspond to a distinct input value. For example, nodes 1908-1914 in the input layer 1902 may correspond to input values 1922-1928, respectively. In some examples, the input values 1922-1928 may correspond to parameters or values provided as inputs to the machine learning model 210. In some embodiments, each of the nodes 1916-1918 in the hidden layer 1904 generates an expression, which may include a mathematical calculation (or algorithm) that results in a value based on the input values received from the nodes 1908-1914. The mathematical calculation may include assigning a different weight to each of the data values received from the nodes 1908-1914. Nodes 1916 and 1918 may include different algorithms and / or different weights assigned to data variables from nodes 1908-1914, such that each of nodes 1916-1918 may produce different values based on the same input values received from nodes 1908-1914. In some embodiments, the weights initially assigned to features (or input values) for each of nodes 1916-1918 may be randomly generated (e.g., using a computer randomizer). The values generated by nodes 1916 and 1918 may be used by node 1922 in output layer 1906 to generate output values for artificial neural network 1900. When artificial neural network 1900 is used to implement machine learning model 210, output values 1930 generated by artificial neural network 1900 may correspond to product quality attribute 212 predictions.
[0181] The artificial neural network 1900 may be trained using training data. For example, the training data herein may be the model training data set 208. By providing the artificial neural network 1900 with training data, the nodes 1916-1918 in the hidden layer 1904 may be trained (tuned) to generate optimal outputs in the output layer 1906 based on the training data. The artificial neural network 1900 (specifically, the representations of the nodes in the hidden layer 1904) may be trained (tuned) to improve its performance by successively providing different training data sets and penalizing the artificial neural network 1900 when its output is incorrect (e.g., if the difference between the predicted product quality characteristic 212 and the measured product quality characteristic exceeds a certain threshold). In some cases, tuning the artificial neural network 1900 may include adjusting the weights associated with each node in the hidden layer 1904.
[0182] Although the above description relates to artificial neural networks as an example of machine learning, it is understood that other types of machine learning methods may also be suitable for implementing various aspects of the present disclosure. For example, machine learning can be implemented using decision trees, random forests, support vector machines, Bayesian networks, regression models, multivariate linear models, ensemble models, etc., or combinations thereof. The regression model may be a linear regression model (e.g., including a multivariate linear regression model), a logistic regression model, a polynomial regression model, a ridge regression model, a least absolute shrinkage and selection operator (lasso) regression model, a partial least squares (PLS) regression model, a principal component regression model, etc. Other types of machine learning algorithms will not be described in detail herein for simplicity, and it is understood that the present disclosure is not limited to a particular type of machine learning.
[0183] While the present teachings will be described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those skilled in the art.
[0184] In describing various embodiments, the specification may present a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular order of steps set forth, and as one skilled in the art will readily appreciate, the order may be changed and still remain within the spirit and scope of the various embodiments.
Claims
1. 1. A computer-implemented method for predicting one or more product quality attributes of a conditioned ultrafiltration / diafiltration (UF / DF) pool in a bioprocess, comprising: receiving Raman measurements of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of a UF / DF operation and receiving a UF / DF pool for conditioning with a buffer to form the conditioned UF / DF pool; predicting one or more product quality attributes of the conditioned UF / DF pool using a machine learning model trained to receive Raman measurements of the conditioned UF / DF pool as input and generate one or more predicted product quality attributes of the conditioned UF / DF pool as output; 11. A computer-implemented method comprising:
2. 10. The method of claim 1, wherein the one or more product quality characteristics comprises the osmolality of the conditioned UF / DF pool.
3. 3. The method of claim 1 or 2, wherein the conditioned UF / DF pool comprises one or more proteins in solution, and the one or more product quality attributes comprise protein concentration of the conditioned UF / DF pool.
4. 4. The method of any one of claims 1 to 3, wherein the UF / DF pool is a purified UF / DF pool received in the collection vessel from an upstream ultrafiltration and diafiltration process.
5. 5. The method of claim 1, wherein the conditioned UF / DF pool comprises an antibody.
6. The method of claim 5 , wherein the antibody is a monoclonal antibody (mAb).
7. the machine learning model is trained to predict the one or more product quality attributes using a training dataset that includes the one or more product quality attributes of conditioned UF / DF pools conditioned with different amounts of the buffer added into the conditioned UF / DF pools and associated Raman measurements of the conditioned UF / DF pools conditioned with the different amounts of the buffer; 7. The method according to any one of claims 1 to 6.
8. 8. The method of claim 1, wherein the trained machine learning model is a trained multivariate linear model, a trained neural network, a trained deep learning model, or a trained ensemble model.
9. 9. The method of any one of claims 1 to 8, wherein the trained machine learning model is one of a design of experiments (DOE) partial least squares (PLS) regression model, a trend analysis (TA) partial least squares (PLS) regression model, or a combination thereof.
10. generating an indication of whether the conditioned UF / DF pool is ready for immediate release based on the one or more predicted product quality characteristics.
10. The method according to any one of claims 1 to 9.
11. 11. The method of claim 1, wherein the steps of receiving the Raman measurements and predicting the one or more product quality attributes are repeated one or more times at a predetermined frequency.
12. 1. A computer-implemented method for controlling an ultrafiltration / diafiltration (UF / DF) pool conditioning process, comprising: receiving, by a processor, a first Raman measurement of the conditioned UF / DF pool from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of a UF / DF operation and receiving the UF / DF pool for conditioning with a buffer to form the conditioned UF / DF pool; predicting, by the processor, one or more product quality attributes of the conditioned UF / DF pool using a trained machine learning model that receives the first Raman measurement as input, wherein the one or more product quality attributes comprise a first protein concentration of the conditioned UF / DF pool; calculating, by the processor, a weight index of the collection vessel corresponding to a second protein concentration of the conditioned UF / DF pool that is different from the first protein concentration of the conditioned UF / DF pool; outputting, by the processor, instructions to add or stop adding the buffer to the collection container based on the calculated weight indicator of the collection container; 11. A computer-implemented method comprising:
13. 13. The method of claim 12, wherein the first Raman measurement corresponds to a current point in time of the conditioned UF / DF pool.
14. the processor is in wired or wireless communication with the Raman spectrometer; a latency between said receiving and said outputting is about 2 seconds or less; 14. The method of claim 12 or 13.
15. 15. The method of claim 12, wherein the processor is an edge node, a processor running a virtual machine, a processor of a cloud server, or a processor of a cloud serverless solution.
16. 16. The method of any one of claims 12 to 15, wherein the one or more product quality attributes of the conditioned UF / UD pool further comprise an osmolality of the conditioned UF / DF pool, and wherein predicting comprises predicting an osmolality of the conditioned UF / DF pool based on the analysis of the first Raman measurements.
17. receiving, at the processor, from the Raman spectrometer, a second Raman measurement of the conditioned UF / DF pool after the transmitting; using the processor to predict a third protein concentration of the conditioned UF / DF pool using the trained machine learning model that takes as input the second Raman measurement; and comparing, by the processor, the second protein concentration to the third protein concentration to determine the effectiveness of the addition of the buffer to the collection vessel; 17. The method of any one of claims 12 to 16, further comprising:
18. 18. The method of any one of claims 12 to 17, wherein the method does not use any measurements obtained by extracting a sample of the conditioned UF / DF pool from the collection vessel.
19. 19. The method of any one of claims 12 to 18, wherein the conditioned UF / DF pool comprises antibodies.
20. 20. The method of claim 19, wherein the antibody comprises a monoclonal antibody (mAb).
21. in response to an instruction to add the buffer, adding the buffer in an amount sufficient to increase the weight index of the collection vessel to within the calculated weight threshold; 21. The method of any one of claims 12 to 20, further comprising:
22. 22. The method of any one of claims 12 to 21, wherein the trained machine learning model is a trained multivariate linear model, a trained neural network, a trained deep learning model, or a trained ensemble model.
23. 23. The method of any one of claims 12 to 22, wherein the trained machine learning model is one of a design of experiments (DOE) partial least squares (PLS) regression model, a trend analysis (TA) partial least squares (PLS) regression model, or a combination thereof.
24. receiving, at the processor, from the Raman spectrometer, a second Raman measurement of the conditioned UF / DF pool after the output; predicting, by the processor, a third protein concentration of the conditioned UF / DF pool using the trained machine learning model that takes as input the second Raman measurement; and generating, by the processor, an indication of whether the conditioned UF / DF pool is ready for immediate release based on the predicted third protein concentration; 24. The method of any one of claims 12 to 23, further comprising:
25. calculating, by the processor, a weight index of the collection vessel corresponding to the first protein concentration of the conditioned UF / DF; determining, by the processor, an amount of buffer to add to the collection vessel based on the weight index of the collection vessel corresponding to the first protein concentration and the weight index of the collection vessel corresponding to the second protein concentration; 25. The method of any one of claims 12 to 24, further comprising:
26. 1. A system comprising:
1. An ultrafiltration / diafiltration (UF / DF) pool recovery system comprising: a collection vessel downstream of the UF / DF operation, the collection vessel configured to receive and store the UF / DF pool for conditioning with the buffer to form a conditioned UF / DF pool; a Raman spectrometer operably coupled to the collection vessel and configured to perform Raman measurements on the conditioned UF / DF pool; an ultrafiltration / diafiltration (UF / DF) pool collection system comprising: a communications module operably coupled to a remote computing platform and the Raman spectrometer, configured to: (i) receive the Raman measurements from the Raman spectrometer and upload the Raman measurements to the remote computing platform for predicting one or more product quality characteristics by the remote computing platform based on the Raman measurements; and (ii) transmit signals received from the remote computing platform and related to the one or more product quality characteristics to the UF / DF pool recovery system; a latency between the receipt of the Raman measurement at the communications module and the transmission of the signal to the UF / DF pool collection system satisfies a predetermined latency threshold. a communication module; Including, the system.
27. the remote computing platform; the remote computing platform includes a processor configured to receive the Raman measurements from the communication module, predict one or more product quality characteristics, and output a signal related to the one or more product quality characteristics.
27. The system of claim 26.
28. the one or more product quality attributes comprise a first protein concentration of the conditioned UF / DF pool; the processor is further configured to calculate a weight index of the collection vessel corresponding to a second protein concentration of the conditioned UF / DF pool that is different from the first protein concentration of the conditioned UF / DF pool; the UF / DF pool recovery system further comprising a buffer pump operably coupled to the recovery vessel; the signal is a buffer pump signal configured to instruct the buffer pump to add or stop adding the buffer to the collection container based on a weight indicator of the collection container; 28. The system of claim 27.
29. 27. The system of claim 26, wherein the one or more product quality characteristics comprises osmolality of the conditioned UF / DF pool.
30. 30. The system of any one of claims 26 to 29, wherein the latency threshold is about 2 seconds.
31. 31. The system of any one of claims 26 to 30, wherein the conditioned UF / DF pool comprises an antibody.
32. 32. The system of claim 31, wherein the antibody comprises a monoclonal antibody (mAb).
33. 1. A computer-implemented method for providing a tool for predicting product quality attributes and / or controlling an ultrafiltration / diafiltration (UF / DF) pool conditioning process, comprising: A training data set, a plurality of Raman measurements of the conditioned UF / DF pool obtained from a Raman spectrometer operably coupled to a collection vessel containing the conditioned UF / DF pool, the collection vessel being downstream of an ultrafiltration / diafiltration (UF / DF) operation and receiving the UF / DF pool for conditioning with a buffer; and corresponding measurements of one or more product quality attributes of the conditioned UF / DF pool; and obtaining a training dataset including: using the training dataset to train a machine learning model to predict the one or more product quality attributes of the conditioned UF / DF pool using Raman measurements of the conditioned UF / DF pool as input; 11. A computer-implemented method comprising:
34. 1. A system comprising: one or more processors; one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 25 or 33; and Including, the system.
35. 34. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 25 or 33.