Customized raman modeling
Patent Information
- Application Number
- CN202580011945.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2025-04-22
- Publication Date
- 2026-08-28
AI Technical Summary
这就是为什么此类模型通常具有有限的适用性,并且无法容易地用于新过程、新规模等
Smart Images

Figure CN122663533A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a computer-implemented method for generating a spectral model for controlling and / or monitoring processes producing chemicals, biopharmaceuticals, or biotechnology products. It also relates to a computer program product comprising computer-readable instructions that, when loaded onto and executed, cause the computer system to perform operations according to the method, and to a computer system operable to control and / or monitor processes producing chemicals, biopharmaceuticals, or biotechnology products. Background Technology
[0002] Examples of processes according to this disclosure are industrial processes, specifically biopharmaceutical processes, such as those for the production of monoclonal antibodies (mAbs). The processes disclosed may involve the chemical transformation of materials accompanied by mass, heat, and / or energy transfer. The process is particularly likely to be scale-dependent; in other words, the process may behave differently on a small scale (e.g., in a laboratory) compared to a large-scale process (e.g., in production). The process may include multiphase chemical reactions.
[0003] The process disclosed herein may be a batch process, an example of an industrial biopharmaceutical process. This batch process may involve producing small quantities of product and gradually or incrementally scaling up to larger quantities. The batch process may involve chemical and / or biological reactions that require time to complete. In batch processing, multiple batches can be produced, and these can be carried out at different scales (e.g., about five liters, about ten liters, about one hundred liters). There may be pauses between each batch, for example, to set up a new batch.
[0004] The process disclosed herein may be a fed-batch process. In a fed-batch process, the growth rate can be regulated compared to a simple batch process. This process can be carried out in a bioreactor designed to accommodate increased volumes. The production system of a fed-batch process can, in particular, always be in a quasi-steady state. During cultivation, one or more nutrients can be fed into the bioreactor, and the products can be retained in the culture medium until the end of the operation. A fed-batch process may involve cultivation in which a basal medium supports the initial cell culture, and a supplemental medium is added to prevent nutrient depletion. The basal and supplemental media can be considered as parts of the culture medium within the vessel.
[0005] The process disclosed herein may also be a perfusion process or a continuous process.
[0006] In a continuous process, a portion of the mixed culture medium (i.e., culture medium and cells) is harvested, which essentially produces a constant cell density. In a perfusion process, only the culture medium is harvested, and the cells are retained in the culture via a cell retention device. The process can be designed such that growth is limited by the availability of one or two (or more) limiting components of the culture medium. Growth depends not only on external components in the culture medium but also on substances excreted by the cells into the medium (which is considered a major cause of cell death in fed-batch processes). For example, during perfusion, desired components can be added to the culture medium, but simultaneously, the old culture medium can be at least partially replaced via a harvest line (in a controlled manner) along with the product harvest.
[0007] Developing new production processes is both time-consuming and expensive, primarily due to the sheer number and variability of operating conditions affecting their end results. This limitation is particularly challenging for small companies that have identified the products they intend to produce and some or all of their quality attributes, but lack the methods available for successfully monitoring and controlling key process parameters and product attributes.
[0008] Spectroscopic techniques have long been implemented in process monitoring within the biopharmaceutical industry due to their ability to simultaneously and non-destructively measure multiple analytes and their potential for online or in-line integration with processes. In particular, the use of Raman, near-infrared (NIR), and ultraviolet (UV) spectroscopy, along with chemometric tools, for monitoring mammalian cell cultures has been well-demonstrated and documented. Major metabolites and nutrients in cell cultures, as well as several cellular characteristics and critical quality attributes (CQAs), can be predicted using spectral data and partial least squares (PLS) models, with prediction root mean square error (RMSEP) very close to the accuracy of offline reference measurements. Real-time monitoring of critical process parameters (CPPs) has also been used to implement glucose feed control loops, thereby improving productivity. These results demonstrate that spectroscopic methods can be effectively used for real-time and in-situ monitoring of cell cultures, automating processes, and even opening the door to batch real-time release (RTR).
[0009] This disclosure addresses the challenges of developing models for process monitoring and / or control (for CPP, CQA, and / or other attributes) using spectral data. Developing such models is a time-consuming and substantial undertaking, as users need to begin by collecting data in small-scale bioreactors, such as from 15 ml to 20 L, specifically 15 ml, 250 ml, 1 L, 2 L, 5 L, 10 L, 20 L, and then gradually progress to larger-scale bioreactors, such as 50 L, 100 L, up to 1,000 L or 2,000 L. Developing spectral models with acceptable performance in terms of accuracy and precision is also very complex, as users need to test a variety of conditions and design options. Applying spectral data and other relevant data generated from previous processes can support model development, but a technical solution for how to achieve this remains lacking. Another issue is that any implementation of the solution will require access to a broad process database (with the spectral data recorded for it), which is typically not feasible for most companies.
[0010] To date, most of the completed work has focused on models built on the same Raman analyzer hardware and within the same limited laboratory environment. Therefore, key issues regarding the use of Raman spectroscopy in a biopharmaceutical setting have not yet been addressed by these existing models: instrument variability, variability between laboratories, process scale, type (feedback, continuous feeding, perfusion, etc.), cell lines and culture media, products, etc. This is why such models often have limited applicability and cannot be readily applied to new processes, new scales, etc.
[0011] According to one aspect, the object of the present invention is to allow the generation of spectral models for controlling and / or monitoring processes for producing chemicals, biopharmaceuticals, or biotechnology products, wherein reliable spectral models with enhanced applicability specific to the chemical, biopharmaceutical, or biotechnology process can be generated.
[0012] This objective is achieved through the subject matter of the independent claims. Specific embodiments can be derived from the dependent claims. Summary of the Invention
[0013] According to one aspect of the present invention, a computer-implemented method for generating a spectral model for controlling and / or monitoring processes for producing chemicals, biopharmaceuticals, or biotechnology products is provided, the method comprising: A database is provided that stores multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology. Each observation dataset includes stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution of the corresponding past process. Obtain user-based input data from the user, which includes at least one of the following: input spectral trajectory data and one or more input process descriptors; A subset of stored spectral trajectory data is determined by querying a database using user-input data based on at least one selection criterion; and A spectral model is generated using a subset of the stored spectral trajectory data, which provides process parameter values for the overall duration of the execution process.
[0014] A database is provided that stores multiple observation datasets associated with corresponding observations of past processes in chemicals, biopharmaceuticals, or biotechnology. Each observation dataset includes stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution or operation of the corresponding past process. Therefore, the database specifically stores distinct data from past processes that have already been run. These past processes differ from those to be controlled and / or monitored using a spectral model generated by the method. In other words, past processes are processes that are different from those to be controlled and / or monitored. The processes to be controlled and / or monitored are future processes that differ from past processes on a different timescale. Specifically, the observation datasets can be stored in the database at time points prior to the start of operation of the process to be controlled and / or monitored. Therefore, the database specifically stores historical data used to generate spectral models for controlling and / or monitoring future processes.
[0015] The database can be a relational database, where data is stored in related tables. Communication can be conducted, for example, using a database query language (such as Structured Query Language SQL). The database can be hosted on a server system and / or accessible to users via a network (such as the cloud). The database can reside on a server cluster under the control of a vendor.
[0016] The database stores observational datasets containing stored spectral trajectory data. Spectral trajectory data specifically refers to or includes spectral data or spectra with associated temporal information, such as Raman spectroscopy, infrared (IR) spectroscopy, fluorescence spectroscopy, and / or ultraviolet-visible spectroscopy. For example, the associated temporal information could be information about process maturity, i.e., process maturity information. Such spectral trajectory data is beneficial for creating models, i.e., spectral models, where the spectrum can be used as input data (x-values), and maturity and / or time are output or target values (y-values).
[0017] Therefore, a trajectory can be a time-based curve of measurements recorded during the execution or operation of a corresponding process. In other words, a trajectory can be understood as a summary and an overview of the associated process. A trajectory can be implemented as a curve or graph describing how the associated process changes over time. A trajectory can also be referred to as a control chart (or a batch control chart in the context of an associated batch process). In the context of a batch process, each batch can have its own corresponding trajectory.
[0018] Specifically, the database is a broad database capable of storing spectral trajectory data from different types of processes, as well as one or more process descriptors (such as one or more of the following: process type, culture medium platform, product, Chinese hamster ovary (CHO) cell line, scale, etc.). More generally, process descriptors may include information describing one or more characteristics of the process, specifically... a priori (That is, characteristics known prior to the execution of the process, such as characteristics related to the process setup.) Therefore, a process descriptor can include information about the characteristics of the process that are not measured during the process or otherwise derived from a running or completed process. For example, a process descriptor can be all metadata that helps describe how the process operates and anything else related to that process.
[0019] A process descriptor can be relevant metadata describing the background of the spectral data source, such as cell lines, culture media, media composition, and / or strains. Relevant metadata may also provide information about scale, culture media platform, and / or feeding strategy. Specifically, a process descriptor can provide information about one or more of the following: the equipment used to run the biological process, the scale of the process, the type of process, the type and / or name of the product, and / or the biological system used to produce the product, the culture media system used, etc. Process descriptors can be user-defined or entered.
[0020] The database can be accessed by multiple users. It can store spectral data from various biological processes. In addition to time evolution data (i.e., spectral trajectory data) from most or all of the process, the database can also store spectral data without time evolution, i.e., single-point spectral data. Within the context of the process, each process batch (or perfusion experiment) can have its own corresponding trajectory. Therefore, the database content can cover CHO-based protein production from different culture medium platforms, cell lines, product types, process types (batch, perfusion), and / or process scales.
[0021] The observation dataset may also include at least one corresponding actual analytical measurement of the process parameter recorded during the execution of the corresponding past process. Specifically, one or more actual analytical measurements of the process parameter refer to data measured during the corresponding past process. Therefore, the “actual analytical measurement of the process parameter” is actually measured, while the process parameter values provided by the spectral model described herein are not measured but derived from spectral data.
[0022] The analytical measurements of process parameters can be, for example, data typically obtained by an offline analyzer, which are not spectral data, such as the levels of glucose or other compounds (e.g., ammonia, salts and / or dissolved oxygen) in cell culture media, cell parameters (e.g., viable cell density, total cell density, cell viability and / or cell diameter, etc.), process performance parameters (e.g., titer), and product quality attributes (e.g., glycosylation, charge variants and / or aggregates, etc.).
[0023] Therefore, the database may contain spectral data of relevant analytical measurements of process parameters (e.g., analytical measurements of process parameters and / or product quality attributes) and one or more process descriptors, wherein one or more process descriptors may be used for the relevant spectra of a pre-selected model, as described below.
[0024] Analytical measurements may be or may include additional process measurements (CPP, CQA, process KPI) from measuring instruments (such as one or more sensors, one or more analyzers, etc.) that were achieved during past processes by online, inline, or nearline measurements, such as pH, titer, viable cell density (VCD), intact mass, temperature, dissolved oxygen (DO), culture medium composition, osmolarity, etc.
[0025] Nearline measurements can be measurements in which samples are taken from the process flow, isolated from the process flow, and analyzed near the process flow. In-line measurements can be measurements in which samples are not taken from the process flow and can be invasive or non-invasive. In-line measurements can be measurements in which samples are removed from the manufacturing process and can be returned to the process flow.
[0026] For continuous upstream processes such as perfusion, continuous online sampling and / or testing can be performed to ensure that process inputs are specifically maintained within target ranges. The analytical requirements for continuous upstream treatments using online spectroscopic methods to determine bioreactor conditions and / or product quantities are relatively well-defined, while a significant gap remains in rapid online product quality measurements. Online product quality testing at defined intervals can complement online analytical outputs, thereby ensuring that product quality (such as charge or glycosylation profiles) is maintained.
[0027] During the process of clearly defining the parameters, processing conditions can be adjusted as needed to compensate for online / near-line analysis output. For continuous downstream processing (DSP), the challenge is much greater because there are currently no online instruments available to monitor purity and / or charge or size heterogeneity to demonstrate the success of the unit operation.
[0028] Therefore, the database specifically stores historical data of spectral trajectories, one or more process descriptors, and / or corresponding actual analytical measurements of process parameters. Optionally, the database may also store process parameters. This allows for the development of spectral models for process control and / or monitoring based on access to extensive historical data. This can be beneficial for most users who do not otherwise have access to extensive historical data. In particular, a large number of different processes can be used as the basis for the database, which will produce more robust models that can be easily transferred between processes of various sizes and / or types; for example, a model developed in enhanced fed-batch processes can be applied to infusion processes.
[0029] User-based input data is obtained (directly or indirectly) from the user, wherein the user-based input data includes at least one of input spectral trajectory data and / or one or more input process descriptors. Specifically, the user-based input data may include input spectral trajectory data and one or more input process descriptors. However, it is also possible for the user-based input data to include input spectral trajectory data but not input process descriptors. Obtaining both input spectral trajectory data and one or more input descriptors is advantageous because a subset of the stored spectral trajectory data that better reflects the conditions of the process believed to be controlled and / or monitored can be identified. Furthermore, in the case of both input spectral trajectory data and one or more input descriptors, a more robust spectral model will be ensured.
[0030] Similar to the stored spectral trajectory data, the input spectral trajectory data specifically refers to or includes spectral data or spectra with associated time information, such as Raman spectroscopy, IR spectroscopy, fluorescence spectroscopy, or UV-Vis spectroscopy. For example, the associated time information could be information about process maturity, i.e., process maturity information. This spectral trajectory data is beneficial for creating models, i.e., spectral models, where the spectrum is used as input data (x-values), and maturity and / or time are output or target values (y-values). Therefore, a trajectory can be a time-based curve of measurements recorded during the execution or operation of the corresponding process. In other words, a trajectory can be understood as summarizing and providing an overview of the associated process. A trajectory can be implemented as a curve or graph describing how the associated process changes over time. A trajectory can also be referred to as a control chart (or a batch control chart in the context of an associated batch process). In the context of a batch process, each batch can have its own corresponding trajectory.
[0031] The user-based input database can include spectral trajectory data from different types of processes and one or more input process descriptors, such as process type, culture medium platform, product, CHO cell line, and / or scale. As mentioned above, the input process descriptor can be all metadata that helps describe how the process operates and any other content related to the process, such as the equipment used to run the biological process, the scale of the process, the type of process, the type and name of the product, and a description of the biological system used to produce the product, the culture medium system used, etc.
[0032] User-based input data can be obtained from multiple users. User-based input data can include spectral trajectory data from different biological processes. In addition to time evolution data from most or all processes (i.e., spectral trajectory data), user-based input data can also include single-point spectral data. Within the context of a process, each process batch (or perfusion experiment) can have its own corresponding trajectory. Therefore, the content of user-based input data can cover CHO-based protein production from different culture medium platforms, cell lines, product types, process types (batch, perfusion), and / or process scales.
[0033] As previously mentioned, input data is obtained from at least one user (directly or indirectly). Specifically, a user can define one or more input process descriptors. Input process descriptors can include similar or identical process descriptors as outlined above regarding stored process descriptors. Defining one or more input process descriptors helps the user define the process of interest.
[0034] In addition, users can provide input spectral trajectory data. To acquire this input spectral trajectory data, at least one run of the desired process can be performed. However, it is not necessary to run the desired process until its completion. The process can be interrupted after a certain period of time, specifically after sufficient information about the desired process has been obtained. Interruption can mean stopping the process after a certain period of time and not continuing to run the same process. In this respect, the process that is run and interrupted is different from future processes that will be controlled and / or monitored by the generated spectral model. Instead, the interrupted process is only used to provide and / or acquire input spectral trajectory data for subsequent steps to determine a subset of the stored spectral trajectory data. By doing so, the spectral model to be generated in subsequent steps can be used for the overall duration of executing one or more future processes.
[0035] Obtaining user-based input data from at least one user allows for the selection of highly relevant (potentially the most relevant) stored spectral trajectory data to generate the spectral model described below. Furthermore, it ensures appropriate spectral preprocessing and the generation of empirically validated models suitable for process control and / or monitoring. This improved selection of relevant stored spectral trajectory data for generating the spectral model can be achieved through, for example, the combined application of phase comparison, spectral preprocessing, and / or multiple search algorithms.
[0036] After providing a database that stores multiple observation datasets and after obtaining user-based input data, a subset of the stored spectral trajectory data is determined by querying the database using the user-based input data, based on at least one selection criterion.
[0037] This method is specifically based on a mathematical algorithm for selecting the most similar stored spectral trajectory data from a database, whereby the algorithm can select the most relevant spectral trajectory data for the spectral model to be generated.
[0038] Therefore, the step of determining a subset of stored spectral trajectory data based on at least one selection criterion by querying the database using user input data can be accomplished by using the input spectral trajectory data. For example, this would result in a so-called batch evolution model (BEM), which can be compared with existing models for similarity.
[0039] The stored spectral trajectory data can be selected to generate a spectral model, such as a multivariate spectral model, that can be used to control and / or monitor processes. The selection algorithm applied will depend on the type and / or quantity of user-based input data. To select the most suitable stored spectral trajectory data for the spectral model, the user can provide information as user-based input data for the selection algorithm and use this user-based input data to query the database.
[0040] According to one example, querying a database using user-input data can be based on comparing spectral trajectory data stored in the database from completed biopharmaceutical processes (i.e., from past or historical processes) with user-input spectral trajectory data that can be derived from at least a portion of the processes the user has run. Thus, time-related data are compared to each other.
[0041] When querying the database using time-related data, the user can provide a batch of user-input spectral data for several cultures (e.g., N cultures) and / or additional relevant process descriptors. The input spectral batch data can be data with a time component, i.e., input spectral trajectory data. Therefore, by comparing the input spectral process trajectories obtained from the user with the spectral process trajectories stored in the database, relevant spectral trajectory data from the database can be obtained for generating the spectral model.
[0042] Spectral preprocessing can be used to improve database searches, such as asymmetric least squares correction (AsLS), standard normal variable (SNV), multiplicative signal correction (MSC), derivatives, peak area, peak height and / or water band normalization and / or multivariate curve resolution (MCR), to deconvolve the signals of known culture medium components (such as glucose, glutamine) from the rest of the spectrum, and then compare the deconvoluted spectral “background”. Multiple preprocessing and / or selection algorithms can be used in parallel to improve the robustness of the selection.
[0043] To obtain input spectral trajectory data from the user (and subsequently for querying a database), the user may have already run a bioreactor (e.g., in a multi-parallel bioreactor system) and may have recorded Raman spectra and corresponding glucose values for at least a portion of that bioreactor operation. In other words, the user does not need to run the required bioreactor process completely. Instead, the user can interrupt or stop the bioreactor process at some point after having sufficient information to create the input spectral trajectory data. Instead of running the entire bioreactor process, the user can use a database to supplement available data with one or more data sources from the database, providing readily available spectral models, such as Raman calibration models, or models used for glucose concentration prediction in the user's bioprocess.
[0044] Users can input Raman spectral trajectory data and / or one or more process descriptors into the database. One or more process descriptors can be the process size, the equipment used to run the biological process, the type of Raman analyzer, the type of process (e.g., fed-batch, perfusion), the name of the product, the type of culture medium, and / or the type of cell line.
[0045] Database algorithms can use user-provided descriptors to pre-select a set of spectral trajectories from a database. Specifically, the algorithm can select only spectral trajectories from those using the same culture medium and / or product type as the user.
[0046] Therefore, one or more database descriptors from the user can be used for pre-selection of spectral trajectory data in the database. For example, additional metadata can be used to narrow the selection to relevant culture medium platforms, cell lines, process types, and / or scales. In particular, relevant spectral trajectory data for spectral models can be selected by comparing process trajectories for specific stages of a process (e.g., cell development stages) rather than the entire process. This is specifically to improve model performance by building models based on the most relevant spectral trajectory data, and also to overcome the difficulties associated with comparing process trajectories with different dynamics (e.g., different total process times, different switching points between process stages, different feeding strategies, etc.).
[0047] The spectral trajectory data can then be processed using one or more methods and process time-trajectory comparison algorithms to fine-tune the spectral selection to multiple spectral trajectories that best match the user's spectral trajectory data. This spectral trajectory data (e.g., with associated glucose values), along with the spectral trajectory data from the user, can then be included in, for example, a spectral model (e.g., a Raman calibration model) used to predict glucose concentration. This model can then be applied or output, for example, provided to the user.
[0048] Depending on the concept, a method is disclosed where the data being compared is not time-related. Specifically, the user-input data to be compared with data stored in a database can be based on the culture medium prior to the start of the biopharmaceutical process, or a search spectrum used as input to the database at a given time point in the biopharmaceutical process. Depending on the concept, the user can provide spectral data of the culture medium at a given time point before or after inoculation, along with relevant information about the user's desired process or any other relevant metadata. Metadata can include, for example, cell line, product name, and / or batch type. Therefore, the data can be data without a time component. It is possible for the user to provide spectral data of the culture medium before or after feeding.
[0049] Depending on the specific concept, users do not need to have batch spectral data. Instead, data without a time component is possible. Furthermore, a similarity search can be used between the input spectrum or between the input spectrum and spectra in the database to select a relevant set of spectra for generating the spectral model. Possible methods for applying similarity searches include spectral similarity measurement algorithms, principal component analysis (PCA), dendrograms / Mahathani distance, Euclidean distance, support vector machines (SVM), random forests, etc. Additional metadata provided by the user can be used for pre-selection of spectral data, narrowing the selection to relevant culture medium platforms, cell lines, process types, scales, etc. Spectral preprocessing can be used to improve database searches, such as AsLS, SNV, MSC, derivatives, peak area, peak height, and / or band normalization or multivariate curve resolution (MCR) to deconvolve signals of known culture medium components (e.g., glucose, glutamine) from the remainder of the spectrum, then comparing the deconvoluted spectral "background".
[0050] Multiple preprocessing and selection algorithms can be used in parallel to improve the robustness of selection. Depending on the concept, users can use one or more Raman spectra from their culture media and / or processes at specific time points to obtain spectral models, such as Raman-based models for glucose prediction in a user's process. Users can submit this spectrum along with one or more descriptors of their process (e.g., bioprocess scale, equipment used in the process, Raman analyzer type, process type (e.g., fed-batch), product name / category (e.g., IgG1), culture medium type (e.g., Cellca platform), and / or cell line type (e.g., CHO DG44)) to the database.
[0051] According to aspects of the invention, the database algorithm can pre-classify spectra in a database based on a given process descriptor. For example, the database algorithm can pre-select only spectra collected on the Cellca platform for the CHO process of producing IgG1 products. The algorithm can then compare the input spectra with spectra in the database, specifically selecting the most similar spectra based on principal component analysis (PCA) similarity search, and can provide the user with a set of spectra with relevant glucose values and / or a multivariate model built based on them for glucose prediction. Multiple spectral preprocessing algorithms can be applied or tested before completing the similarity search. Furthermore, multiple similarity searches can also be tested.
[0052] The query according to aspects of the invention allows selection of the most similar spectral trajectory data based on at least one selection criterion. Therefore, a subset of the stored trajectory data can be determined based on user-input data (i.e., input spectral trajectory data and / or at least one of one or more input process descriptors), as further described below. This ensures the selection of the stored spectral trajectory data for fine-tuning in order to generate a spectral model.
[0053] In the next step, a spectral model can be generated using a selected subset of the stored spectral trajectory data. This spectral model provides process parameter values for the overall duration of future execution processes. This spectral model can be provided to the user for monitoring and / or controlling the desired process. Specifically, the desired process parameters and preferred quality properties can be monitored and / or controlled using the generated spectral model.
[0054] Once similar data is selected from the database based on user-input data, this data can be processed to generate a spectral calibration model. Thus, a spectral calibration model, such as a Raman-based monitoring model, can be obtained. This calibration model can be a partial least squares (PLS) model, although other methods, such as multiple linear regression models, can also be used. In particular, partial least squares (PLS) and / or multiple linear regression modeling methods can be used to correlate Raman spectra with analytical measurements of process parameters.
[0055] To construct a calibration model, the spectrum can be processed using a combination of one or more spectral preprocessing methods (e.g., AsLS, SNV, MSC, derivative, peak area, peak height, and / or band normalization). Multivariate curve resolution (MCR) can be used to deconvolve signals of known culture medium components (e.g., glucose, glutamine) from the rest of the spectrum, and this approach helps compare the deconvoluted spectral "background" to improve similarity searches. Therefore, spectral preprocessing can both improve database searches and facilitate obtaining a good spectral model. To improve the model, only portions of the spectrum can be selected, for example, only bands corresponding to the Raman signals of the molecules at the predicted concentrations. The most relevant portions of the spectrum used for model construction can also be obtained by applying different variable selection methods. Spectral processing can involve taking a portion of the spectrum, and this portion can be determined based on process knowledge or by applying some variable selection method.
[0056] Furthermore, different scaling can be applied to the data during model building; for example, centering can be used on spectral models.
[0057] According to this method, a model based on spectral data is provided, such as input spectral trajectory data obtained from a user and stored spectral trajectory data selected from a database. This model can be used to provide process parameter values for the overall duration of future execution processes. Furthermore, process parameters and / or product attributes can be predicted or obtained from this model. The generated spectral model can then be provided to the user.
[0058] Specifically, a globally universal spectral model is advantageously generated, i.e., a spectral calibration model that is not a local spectral calibration model used to obtain responses related to the spectrum used as input for database searches. In particular, using the entire dataset to predict new measurement points is referred to as a global model, and like any model, it needs to strike a balance between robustness and accuracy. "The entire dataset" can mean using the entire time process of the dataset, not just a single point in time or a single time interval of the dataset to predict new measurement points. According to this disclosure, the entire observation dataset of the database and the entire user-based input data of past processes are specifically used to generate a spectral model that predicts from the perspective of providing process parameter values for the overall duration of performing a new (future) process. "The entire observation dataset" of the database and "the entire user-based data" of past processes can respectively mean using the entire time process of the observation dataset or the user-based data, and not just a single point in time or a single time interval of that data to generate a spectral model.
[0059] In contrast to this disclosure and described only for better understanding of the invention, a different approach is to use only local data, i.e., not the entire dataset, based on similarity or distance methods to predict new measurement points. Using only local data can mean that new measurement points can be predicted using a single point in time or a single time interval of that data. This approach is called a local model. Local models provide an alternative to global models because they can dynamically respond to process conditions, account for process drift, and simplify model maintenance. Such local models may provide only process parameter values for the same process and may not provide process parameter values for the overall duration of executing new (future) processes. For example, Just-In-Time (JITL) learning can be integrated to develop local models for glucose, lactate, glutamine, glutamate, calcium, sodium, viability, and live cell density (VCD).
[0060] In other words, a globally applicable spectral calibration model is best suited for making future predictions within a single experiment or process, without querying the input spectrum to predict the output response associated with that input spectrum. A globally applicable spectral calibration model can be used to monitor and / or control specific process parameters, such as glucose concentration or titer. This contrasts with local spectral models that are used for predictions within the same experiment and process but not for future predictions in individual experiments.
[0061] In consideration of aspects of the present invention, it is noted that all statistical descriptors, performance data, etc., as well as data from multiple observation datasets in the database and user-input-based data, are determined and made accessible to the user before the process is run. In other words, the data is... a priori This general approach differs from local approaches, which can provide data in real time during the process (e.g., by acquiring current values from spectra or other sources and one or more time intervals prior to the current time to predict subsequent values), where the data is not deterministic or indeterminate before the process runs. In contrast, the general approach considers the evolution of the entire process, for example, by using a database with stored time-correlated trajectory data and optional input spectral trajectory data.
[0062] The method for generating spectral models for controlling and / or monitoring processes (specifically, processes using spectral systems such as Raman spectroscopy) to produce chemicals, biopharmaceuticals, or biotechnology products allows for the application of a large number of different processes as the basis for a database. This results in more robust models that can be easily transferred between processes of various scales and / or types; for example, a model developed in an enhanced fed-batch process can then be applied to a filling process. New production processes can be developed, and the method can be used to successfully monitor and / or control key process parameters and / or product properties in a time- and cost-effective manner. The method also allows for improved selection of relevant spectral data used to generate the spectral model, for example, through the combined application of phase comparison, spectral preprocessing, and / or multiple search algorithms. This leads to an improved output spectral model that provides process parameter values for the overall duration of the executed process.
[0063] In summary, the method according to one aspect of the invention allows for improved selection of relevant spectral data to generate more robust spectral models for process monitoring and / or control, specifically multivariate spectral models generated without any (or only limited) experimental datasets on the user side. Alternatively or additionally, the method according to one aspect of the invention allows for easier transfer of spectral models between various process steps, scales, and / or process types. This model allows for faster and less costly process development, resulting in faster release of chemicals, biopharmaceuticals, or biotechnology products, and improved user monitoring and / or control of biological processes. Therefore, a technical solution is provided that allows the use of historical process data from multiple processes to generate one or more reliable models based on limited input data from a new process.
[0064] The method may also include using spectral models to control and / or monitor the process.
[0065] After generating a spectral model that provides process parameter values for the overall duration of the process, the spectral model can be used by the user to control and / or monitor the process.
[0066] The steps of using a subset of stored spectral trajectory data to generate a spectral model that provides process parameter values for the substantially overall duration of the process, and optionally, the subsequent use of this spectral model to control and / or monitor the process, indicate that past or historical datasets of past or historical processes are used to generate the spectral model, specifically a globally generalized spectral model. This spectral model can be used to control and / or monitor one or more future or new processes that differ from past or historical processes.
[0067] Therefore, to obtain process parameter values for the overall duration of these processes, it is not necessary to run one or more future or new processes. Instead, using the spectral model allows for faster and less costly process development, resulting in more rapid releases of chemicals, biopharmaceuticals, or biotechnology products, and improved user monitoring and / or control of biological processes.
[0068] Each observation dataset may also include stored process parameter data, and may also include input process parameter data based on user input data.
[0069] Users input Raman spectral trajectory data and / or one or more input process descriptors into the database. Additionally, users can input process parameter data. The database algorithm can then use the user-provided process parameter data to pre-select a set of spectral trajectories from the database.
[0070] Process parameter data can be measured online, inline, or nearline during the process via sensors and / or analyzers, such as pH, glucose, lactate, titer, viable cell density (VCD), intact mass, aggregation, etc. Process parameter data can be user-defined and identified based on the user's process knowledge. Therefore, process parameter data can include information about one or more parameters in the process that the user can control or record, such as pH, dissolved oxygen, gas flow rate, temperature, pressure, or glucose level.
[0071] Process parameter data may include critical process parameters (CPPs) and / or significant process parameters (KPPs). CPPs are process parameters that can have variability that affects critical quality attributes (CQAs) and therefore should be monitored and / or controlled to ensure the process can achieve the desired quality. Here, a critical quality attribute is or includes physical, chemical, biological, and / or microbiological characteristics or features that should be within appropriate limits, ranges, or distributions to ensure the desired product quality. Significant process parameters (KPPs) are adjustable parameters of the process, preferably adjustable from a variable perspective, which may ensure optimal process performance when maintained within a narrow range. The range of significant process parameters can be established during process development, and variations in the operating range can be managed within the quality system. Furthermore, process parameter data may include well-controlled parameters. Well-controlled parameters are process parameters that can be controlled through process design and / or standardized procedures or automated control systems to ensure they remain within the process's design space. Variations beyond the design space are only possible in the event of a failure in the control system, and such failure modes can be mitigated.
[0072] Preferably in a perfusion bioreactor paired with a macroporous membrane filter, but in different processes, key process parameter data may include tangential flow filtration (TFF) tangential flow rate, glucose feed rate, nutrient feed rate, antifoamer concentration, dissolved oxygen, maximum pCO2, pH, temperature, agitation rate, culture duration, and / or culture medium osmolarity. Important process parameter data may include culture medium feed rate, harvest / TFF permeation rate, biomass removal and / or effluent rate, glucose concentration, gas flow rate, and / or bioreactor weight. Key quality attributes that may be affected may include antibody-dependent cell-mediated cytotoxicity (ADCC) activity, deamidation isomers, charge variants, size variants, oligosaccharides (such as defucosylated and / or galactosylated glycans), glycosylation-related (such as sialic acid content, mannose content, and non-glycosylated heavy chains), host cell protein (HCP) levels, and / or DNA levels.
[0073] Preferably during the protein A CMCC process, but also in different processes, key process parameter data may include additional considerations such as loading UV penetration (post-primary loading column), loading UV flow-through (post-secondary loading column), start / stop (initial / final cycle before / after steady state), switching time, protein loading ratio (g / L resin), loading pH, start elution collection, bioload reduction solution contact time, and / or impurity wash volume. Important process parameter data may include loading concentration, end elution collection, column contamination / product cycle number, column bed height, column packing quality (HETP and asymmetry), and / or flow rate, such as loading and / or non-loading. Key quality attributes that may be affected may include deamidation isomers, glycosylation-related factors (such as sialic acid content, mannose content, and / or non-glycosylated heavy chains), methotrexate, antifoaming agent C, protein A ligands, HCP, DNA, and / or exogenous viral factors (AVA) / viral validation statements.
[0074] Preferably during a low-pH virus inactivation process, but also in different processes, key process parameter data may include mixing model (plug flow reactor) or time / solution homogeneity (chamber / semi-batch), homogeneity (plug flow reactor), incubation time, low pH / acid titrant addition, neutral pH / base titrant addition, and / or temperature. Important process parameter data may include the titration dose added over time. Key quality attributes that may be affected may include size exclusion chromatography (SEC) (aggregates / high molecular weight (HMW)), HCP, DNA, exogenous viral factors (AVA) / virus validation statements.
[0075] Preferably during anion exchange chromatography, but also in different processes, key process parameters may include sample loading concentration, protein loading ratio (g / L resin), sample loading conductivity, sample loading pH, product cycle number, column height / membrane area, and / or sample loading flow rate. Important process parameters may include initiation and / or termination of flow-through collection. Potentially affected key quality attributes may include SEC (aggregates / HMW), protein A ligand, HCP, DNA, and / or exogenous viral agents (AVA) / viral validation statements.
[0076] Preferably in a process involving pure binding-elution chromatography, but also in other processes, key process parameters may include sample loading UV penetration (post-primary loading column), sample loading UV flow-through (post-secondary loading column), protein loading ratio (g / L resin), start / stop (initial / final cycle before / after steady state), switching time, start elution collection, end elution collection, sample loading conductivity, sample loading pH, bioload reduction solution contact time, and / or impurity wash volume. Important process parameters may include column contamination / product cycle count, sample concentration, column bed height, column packing quality (HETP and asymmetry), and / or flow rate, such as loading and / or non-loading. Potentially affected key quality attributes may include SEC (aggregates / HMW), charge variants, protein A ligands, HCP, DNA, and / or exogenous viral agents (AVA) / viral validation statements.
[0077] Preferably in processes involving virus filtration, but also in other processes, key process parameter data may include feed flow rate, sample concentration, filter volume / volume loading, sample aggregate / particle level, sample conductivity, sample pH, process pauses (duration), and / or process pauses (number). Important process parameter data may include filter usage duration. Potentially affected key quality attributes may include SEC (aggregates / HMW) and / or exogenous viral agents (AVA) / virus validation statements.
[0078] Preferably in processes with single-pass tangential flow filtration, but also in other processes, key process parameter data may include feed flux, sample concentration, membrane mass loading, membrane usage duration, retentate flux, and / or transmembrane pressure (TMP). In particular, there may be no important process parameter data, and key quality attributes that may be affected may include active pharmaceutical ingredient protein concentration and / or SEC (aggregates / HMW).
[0079] Preferably in processes involving online dialysis, but also in other processes, key process parameter data may include buffer feed flow rate, membrane mass loading, membrane usage duration, dialysis volume factor, permeate flux, product concentration, product feed flow rate, and / or TMP. In particular, there may be no important process parameter data, and key quality attributes that may be affected may include drug substance protein concentration, SEC (aggregates / HMW), drug substance osmolarity, and / or drug substance pH.
[0080] Preferably during routine flow / bioload reduction filtration processes with high concentrations of ultrafiltration products, but also in other processes, key process parameter data may include extended filter operating times (days / weeks rather than hours) and / or filter usage duration. Important process parameter data may include volumetric loading. Potentially affected key quality attributes may include drug substance bioload.
[0081] As described above, the step of querying the database using user input data to determine a subset of the stored spectral data based on at least one selection criterion can be accomplished by using spectral trajectory data and / or one or more process descriptors.
[0082] Furthermore, the step of querying the database using user input data to determine a subset of the stored spectral data based on at least one selection criterion can be accomplished using process parameter data.
[0083] Process parameter data can be used as a pre-selection criterion. For example, specific observation datasets can be pre-selected based on process parameter data so that spectral trajectory data can be used from these pre-selected datasets. In addition to spectral trajectory data and one or more process descriptors, process parameter data can be used. This ensures an efficient method for selecting data in a time-saving manner, resulting in a more robust spectral model. This model allows for faster and less costly process development, leading to faster releases of chemicals, biopharmaceuticals, or biotechnology products, and improving user monitoring and / or control of biological processes.
[0084] The method may also include defining at least one selection criterion regarding the similarity between the stored observation dataset and user-based input data, wherein optionally, defining the selection criterion includes a combination of at least one of the steps (i) to (iii): (i) Match user-input data with the stored observation dataset; (ii) Spectral preprocessing (applied to the spectral data in the database (202) and the provided user-based spectral data) by applying at least one of the following methods: -AsLS correction (asymmetric least squares correction, which uses a symmetric least squares algorithm to calculate a nonlinear baseline for each spectrum and then subtracts the baseline from the spectrum). - Moving window (removing noise by applying a moving average or median window to the spectrum), -Derivative, -SNV filtering (Standard Normal Transform, which "normalizes" each observation (spectrum) by subtracting the mean and dividing by the standard deviation). -MSC (multiplicative signal correction, which normalizes each spectrum by regressing the average spectrum on a selected set of spectra). -Average the (processed) signal within the selected signal range. - Peak area, peak height, or water zone normalization - Multivariate curve resolution (MCR); (iii) Compare all or part / stages of the process trajectory by applying at least one of the following methods: - Compare the spectral similarity at specific time points or time intervals during the comparison process. -Dynamic time warping -Model array, - Two-tailed one-sided test (TOST) for time series analysis, -Calculation of the (spectral) similarity index for the entire process. - Mechanistic modeling of metabolic behavior for comparison processes. - Batch Modeling — Comparing batch time trajectories using PLS models; At least one of the selection criteria is defined by at least one of the following algorithms: - For each time point in the trajectory of the batch evolution model, maximize the correlation coefficient and minimize the average distance. - Minimize the total RMSEP between database batches and user trajectories in the batch evolution model. - Minimize the RMSEP at each time point between the database batch and the user trajectory in the batch evolution model to identify the similarity between various time points related to the process stage, and - Minimize the distance to the center of the batch-level model.
[0085] The chosen standard can be either a 0-1 standard or a linear standard.
[0086] The choice of criterion can depend on the method applied. Specifically, if defined as a 0-1 criterion, stored spectral trajectory data from the database that produces a 1:1 similarity result to the input spectral trajectory data will be included in the subset of stored spectral trajectory data used to generate the final spectral model. Alternatively, if the criterion is continuous, i.e., a linear criterion, some range will be applied, for example, based on a 95% confidence interval. Therefore, some tolerance can be allowed between the stored spectral trajectory data and the input spectral trajectory data when determining the subset of stored spectral trajectory data used to generate the final spectral model.
[0087] By including the most similar spectral trajectory data in the 0-1 standard case, and also by including spectral trajectory data that are further apart in similarity in the model in the linear standard case, the robustness of the spectral model is further improved.
[0088] The analytical measurements of process parameters may include product quality attributes or process performance parameters, such as glucose, lactic acid, and / or titer levels.
[0089] Product quality attributes can be physical, chemical, biological, or microbiological characteristics or features that should be within appropriate limits, ranges, or distributions to ensure the desired product quality. Exemplary product quality attributes can be glycosylation, charge variants, or aggregation patterns.
[0090] For example, online analytics tools can be used to monitor product quality attributes.
[0091] Therefore, further consideration of product quality attributes helps ensure that manufacturing operations are carried out within specified specifications. This helps improve the quality of stored data in the database, thereby improving the spectral models that provide process parameter values for the overall duration of the execution process.
[0092] According to another aspect of the present invention, a computer program product is provided, comprising computer-readable instructions that, when loaded and executed on a computer system, cause the computer system to perform operations according to the aforementioned method. Therefore, any features and technical effects of the aforementioned method can also be applied to the computer program product of other aspects of the present invention.
[0093] According to another aspect of the invention, a computer system is provided operable to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products, the computer system comprising a database and one or more processors associated with the database, wherein the database is configured to: Storing multiple observation datasets associated with corresponding observations of past processes involving chemicals, biopharmaceuticals, or biotechnology, each observation dataset including stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution of the corresponding past process; wherein one or more processors are configured to: Obtain user-based input data from the user, which includes input spectral trajectory data and / or at least one of one or more input process descriptors; A subset of stored spectral trajectory data is determined by querying a database using user-input data based on at least one selection criterion; and A spectral model is generated using a subset of the stored spectral trajectory data, which provides process parameter values for the overall duration of the execution process.
[0094] In particular, the methods and all the features of the computer program products described above, as well as their corresponding effects, can also be applied to computer systems.
[0095] One or more processors can be configured to use a spectral model to control and / or monitor the process. Each observation dataset may also include stored process parameter data. User-based input data may also include input process parameter data.
[0096] One or more processors may also be configured to define at least one selection criterion regarding the similarity between the stored observation dataset and user-based input data, wherein optionally, one or more processors are configured to define the selection criterion including a combination of at least one of the steps (i) to (iii): (i) Match user-input data with the stored observation dataset; (ii) Spectral preprocessing (applied to the spectral data in the database (202) and the provided user-based spectral data) by applying at least one of the following methods: -AsLS correction (asymmetric least squares correction, which uses a symmetric least squares algorithm to calculate a nonlinear baseline for each spectrum and then subtracts the baseline from the spectrum). - Moving window (removing noise by applying a moving average or median window to the spectrum), -Derivative, -SNV filtering (Standard Normal Transform, which "normalizes" each observation (spectrum) by subtracting the mean and dividing by the standard deviation). -MSC (multiplicative signal correction, which normalizes each spectrum by regressing the average spectrum on a selected set of spectra). -Average the (processed) signal within the selected signal range. - Peak area, peak height, or water zone normalization - Multivariate curve resolution (MCR); (iii) Compare all or part / stages of the process trajectory by applying at least one of the following methods: - Compare the spectral similarity at specific time points or time intervals during the comparison process. -Dynamic time warping -Model array, - TOST test for time series analysis -Calculation of the (spectral) similarity index for the entire process. - Mechanistic modeling of metabolic behavior for comparison processes. - Batch Modeling — Comparing batch time trajectories using PLS models; At least one of the selection criteria is defined by at least one of the following algorithms: - For each time point in the trajectory of the batch evolution model, maximize the correlation coefficient and minimize the average distance. - Minimize the total RMSEP between database batches and user trajectories in the batch evolution model. - Minimize the RMSEP at each time point between the database batch and the user trajectory in the batch evolution model to identify the similarity between various time points related to the process stage, and - Minimize the distance to the center of the batch-level model.
[0097] The selection criteria can be a 0-1 standard or a linear standard. Analytical measurements can include product quality attributes or process performance parameters, such as glucose, lactic acid, and / or titers. Attached Figure Description
[0098] Further aspects, features, and embodiments of this disclosure will be described by way of example in conjunction with the following drawings.
[0099] Figure 1 A computer implementation method for generating a spectral model, according to one aspect, is illustrated. This spectral model is used to control and / or monitor processes for producing chemical, biopharmaceutical, or biotechnology products. Figure 2 A more detailed general workflow is shown according to one aspect of a computer implementation method for generating a spectral model used to control and / or monitor processes for producing chemical, biopharmaceutical, or biotechnology products. Figure 3 The illustration shows exemplary spectral trajectory data stored in a database or provided by the user as input spectral trajectory data. Figure 4 A computer system is shown that is operable, according to one aspect, for controlling, monitoring, and / or predicting processes for producing chemical, biopharmaceutical, or biotechnology products. Figure 5 A computer-implemented method according to a first embodiment of one aspect is shown for generating spectral models to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products. Detailed Implementation
[0100] Figure 1 An exemplary workflow is shown for a computer implementation method 100 for generating a spectral model according to one aspect, the spectral model being used to control and / or monitor a process (e.g., a process performed by a spectral system such as a Raman spectroscopy system) to produce chemicals, biopharmaceuticals, or biotechnology products.
[0101] Method 100 includes step S120: providing a database 202 that stores multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology, each observation dataset including stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution of the corresponding past process.
[0102] Method 100 further includes step S140: obtaining user-based input data from at least one user, the input data including input spectral trajectory data and / or at least one of one or more input process descriptors.
[0103] As a next step, method 100 includes step S160: determining a subset of the stored spectral trajectory data by querying database 202 using user-based input data according to at least one selection criterion.
[0104] Method 100 further includes step S180: generating a spectral model using a subset of the stored spectral trajectory data, which provides process parameter values for the overall duration of the execution process.
[0105] Method 100 may include step S200: using a spectral model to control and / or monitor the process.
[0106] Each observation dataset may also include stored process parameter data, and based on user input data, it may also include input process parameter data.
[0107] Method 100 may include defining at least one selection criterion regarding the similarity between a stored observation dataset and user-based input data. Optionally, defining the selection criterion includes a combination of at least one of the following steps: matching the user-based input data with the stored observation dataset, applying spectral preprocessing to the spectral data in database 202 and the provided user-based spectral data, spectral similarity assessment, and comparison of all or part / stages of the process trajectory. The selection criterion may be a 0-1 criterion or a linear criterion. Analytical measurements of process parameters may include product quality attributes or process performance parameters, such as glucose, lactic acid, and / or titers.
[0108] Figure 2 The general workflow of a computer implementation method 100 for generating a spectral model according to one aspect is shown, the spectral model being used to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products.
[0109] Method 100 includes the previously described steps: S120: providing a database, S140: obtaining user-based input data, S160: determining a subset of stored spectral trajectory data, and S180: using the subset of stored spectral trajectory data to generate a spectral model.
[0110] Specifically, method 100 can begin from step S120.
[0111] In step S120 of providing the database, database 202 is generated by storing multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology. The observation datasets include stored spectral trajectory data, one or more process descriptors, and at least one corresponding analytical measurement of the process parameter.
[0112] After step S120 of providing the database, there may be a user's process document recording, which in turn allows step S140, namely, obtaining user-based input data including at least one of input spectral trajectory data and / or one or more input process descriptors, which may include step S145, namely, recording spectral trajectory data (i.e., data with time components), and / or step S150, namely, the user uploading spectral trajectory data to database 202, and optionally updating (S125) database 202 with the spectral trajectory data recorded by the user, and / or step S155, namely, pre-selecting spectral trajectory data in database 202 based on the user's process parameter data.
[0113] According to different concepts, a method is disclosed that includes the step of providing a database, which includes storing multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology. These observation datasets include stored spectral data without a time component, one or more process descriptors, and at least one corresponding analytical measurement of a process parameter. In these different concepts, the step of providing the database may be followed by a user's process documentation record, which in turn allows for the acquisition of user-based input data, including input spectral data without a time component and / or at least one of one or more input process descriptors. The user's process documentation record may include the steps of recording spectral data without a time component and / or uploading spectral data without a time component to the database by the user, and optionally, updating the database with the user-recorded spectral data without a time component, and / or pre-selecting spectral data in the database based on the user's process descriptor data. Therefore, in these different concepts, the data being compared to each other is not time-related. Specifically, the user-based input data to be compared with the data stored in the database may be based on a culture medium prior to the start of a biopharmaceutical process, or a search spectrum used as input to the database at a given time point in the biopharmaceutical process.
[0114] According to aspects of the invention, and as Figure 2 As shown, method 100 proceeds to step S160: determining a subset of stored spectral trajectory data by querying a database using user-based input data according to at least one selection criterion, followed by step S180: generating a spectral model using the subset of stored spectral trajectory data, which provides process parameter values for the overall duration of the execution process. Thus, a spectral calibration model, such as a Raman-based monitoring model, can be obtained. This calibration model can be a partial least squares (PLS) model, although other methods, such as multiple linear regression models, can also be used. In particular, Raman spectra can be correlated with analytical measurements using partial least squares (PLS) and / or multiple linear regression modeling methods. To construct the calibration model, the spectrum can be processed using a combination of one or more spectral preprocessing methods (e.g., AsLS, SNV, MSC, derivative, peak area, peak height, and / or band normalization). Multivariate curve resolution (MCR) can be used to deconvolve signals of known culture medium components (e.g., glucose, glutamine) from the remainder of the spectrum, and this approach helps to compare the deconvoluted spectral “background” to improve similarity searches. Therefore, spectral preprocessing can be used to improve database searches and also facilitate the acquisition of good spectral models. To improve the model, only a portion of the spectrum can be selected, for example, only the bands corresponding to the Raman signals of the molecules whose concentrations are to be predicted can be included. The most relevant part of the spectrum used for model building can also be obtained by applying different variable selection methods. Spectral processing can involve taking a portion of the spectrum, and this portion can be determined based on process knowledge or by applying some variable selection method. Furthermore, in model building, different scaling can be applied to the data; for example, centering can be used on the spectral model.
[0115] After generating the spectral model, the user can download the spectral model from the database (S165). Optionally, the user can upload the spectral model to the location in the system that will be used to generate process parameter values. The user can record the process and spectral measurements, including spectral trajectory data, and process them via the spectral model (S170).
[0116] Furthermore, the method may include an optional step of assessing whether model confidence and / or trajectory and / or prediction accuracy are met. If the accuracy is met, method 100 proceeds to step S200: using a spectral model to control and / or monitor the process. If the accuracy is not met, method 100 returns to step S150: uploading the spectral trajectory data to a database. However, if the method may not include the optional step of assessing whether model confidence and / or trajectory and / or prediction accuracy are met, step S170 (i.e., recording the process and spectral measurements including the spectral trajectory data and processing them via a spectral model) will immediately follow step S200 (i.e., using a spectral model to control and / or monitor the process).
[0117] Regarding the steps for assessing whether model confidence and / or trajectory and / or prediction accuracy are met, RMSEP values can be used to check whether the prediction errors of process parameter values are within predefined acceptance criteria. Alternatively or additionally, predictions of specific process parameter values obtained from a new process using the applied model can be compared with offline reference values obtained from the same sample from that process. The evaluation can be based on the following criteria: if the predicted values are within predefined acceptance criteria (e.g., within approximately + / - 10%), the model is considered good.
[0118] After step S200, the method may end in step S220.
[0119] Any one of the aforementioned steps S110 to S220 can be performed at least partially as a cloud application and / or as a desktop application. Regarding Figure 2 The different steps of the described method 100 can be followed as follows Figure 2 The steps are performed in the order indicated by the arrows. However, alternative orders of these process steps are also possible.
[0120] Figure 3 Exemplary spectral trajectory data is shown, either stored in a database or provided by the user as input spectral trajectory data.
[0121] The exemplary process trajectory summarizes the evolution of a biological process monitored over time using spectral data in one component. Figure 3 Twelve different trajectories are shown (see Figure 3 The solid lines in the diagram (1 to 12) represent the trajectories, each of which can be derived from different experiments. Spectral data (Y-axis) are plotted against time (X-axis). Figure 3 The spectral data marked "t[1]" can refer to the first score of the BEM model, for example, as a "summary" of the temporal spectral trajectories for each different experiment. The average trajectory of these twelve different trajectories and the three standard deviations around the average are shown in dashed lines.
[0122] When querying the database using time-related input data, users can provide batches of spectral data from several cultivation cycles (e.g., N cultivation cycles) based on user-inputted data and / or additional relevant process parameter data. Each of the N cultivation cycles can be generated by, for example... Figure 3 The trajectory shown is represented by [the image / data].
[0123] The input spectral batch data can be data with a time component, i.e., input spectral trajectory data. Therefore, by comparing the input spectral process trajectory obtained from the user with the spectral process trajectory stored in the database, relevant spectral trajectory data from the database can be obtained for generating the spectral model.
[0124] Figure 4 A computer system 200 is shown that is operable, according to one aspect, for controlling, monitoring, and / or predicting processes for producing chemicals, biopharmaceuticals, or biotechnology products. In particular, the computer system 200 is adapted to perform the computer-implemented method 100 described above.
[0125] Computer system 200 includes database 202 and one or more processors 204 associated with database 202. The one or more processors 204 may be part of control system 210 (more specifically, biological process control system 210), which may also include at least one control device 206 that can be connected to and / or communicate with the one or more processors 204.
[0126] Database 202 is configured to store multiple observation datasets associated with corresponding observations of past processes involving chemicals, biopharmaceuticals, or biotechnology. Each observation dataset includes stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution of the corresponding past process. Specifically, regarding... Figure 1 and Figure 2 The database 202 provided in step S120 described may be about Figure 4 The database shown and described.
[0127] One or more processors 204 execute about Figure 1 and Figure 2Steps S140 to S200 are described. Specifically, one or more processors 204 are configured to obtain (S140) user-based input data from a user, which includes at least one of input spectral trajectory data and / or one or more input process descriptors. The one or more processors 204 are also configured to determine (S160) a subset of stored spectral data by querying a database using the user-based input data according to at least one selection criterion, and to generate a spectral model using (S180) the subset of stored spectral trajectory data, which provides process parameter values for the overall duration of the execution process. Furthermore, the one or more processors 204 may be configured to use the spectral model to control and / or monitor (S200) the process.
[0128] Computer system 200 may also include a recommendation system 214 and / or be able to connect to and / or communicate with the recommendation system 214. The generated spectral model can be transferred from the recommendation system 214, which may be provided by a supplier of the biological process control system 210, to the biological process control system 210. The recommendation system 214 may be connected to and / or communicate with database 202.
[0129] The biological process control system 210 can communicate with a recommendation system 214 via one or more processors 204, which in turn can communicate with a database 202, both of which can be provided by the supplier of the process control system 210. The recommendation system 214 can be software configured to connect to the database 202 and allow users to query the database, for example via a control device 206 (e.g., a personal computer), which is connected to one or more processors 204. Furthermore, the recommendation system 214 can be configured to extract relevant stored spectral trajectory data from the database 202, construct a spectral model, and send the spectral model back to the process control system 210 and / or the Raman analyzer 212.
[0130] As an apparatus for generating a computer-implemented method for controlling and / or monitoring processes that produce chemicals, biopharmaceuticals, or biotechnology products, the user may have several components of a technology platform, such as the recommender system 214. These components may all be supplied by a single vendor, or they may have necessary communication interfaces with each other. This description is a specific example but not a general-purpose apparatus. Other apparatuses may be possible.
[0131] Computer system 200 can be connected to and / or communicate with bioreactor 208. For example, bioreactor 208 can be implemented as a 50 L reusable bioreactor or a single-use bioreactor, such as one equipped with a BioPAT for Raman measurements. ® Biostat with spectral port and FlexSafe 50 L bag® STR50 (Sartorius Stetti Biotechnology Co., Ltd.). Alternatively, culture vessels with working volumes of 5 L, 10 L, 15 L, 20 L, and 30 L can be used.
[0132] Bioreactor 208 may include control and / or online and / or inline measurement capabilities. Specifically, bioreactor 208 may include or be connected to a bioprocess control system. The bioreactor may be connected to a Raman analyzer 212 and / or a Raman spectrometer, wherein at least one Raman immersion probe is placed in or on bioreactor 208 to continuously measure Raman spectra.
[0133] Bioreactor 208 may be capable of controlling and / or measuring one or more of the following: temperature, pH, dissolved oxygen (DO) concentration, or cell density. Bioreactor 208 may be capable of recording various measurements using the above-mentioned measuring devices and other measuring devices.
[0134] In conjunction with at least one control device 206, the bioreactor 208 is particularly capable of performing various forms of online measurements and / or analyses, wherein the culture medium is measured while being maintained in the bioreactor 208.
[0135] At least one control device 206 may be connected to and / or communicate with the bioprocess control system 210 and / or constitute the bioprocess control system 210, which in turn may be connected to and / or communicate with a spectrometer control system including and / or operating a Raman spectrometer and / or Raman analyzer 212, and preferably connected to and / or communicate with the bioreactor 208 via the Raman analyzer 212. The Raman analyzer 212 may be connected to and / or communicate with the bioreactor 208.
[0136] Accordingly, it can be done at the local level (e.g., based on BioBrain) ® The control unit of Sartorius Steyr Biotechnology Co., Ltd. provides a system for controlling biological processes, namely a biological process control system 210. A spectrometer control system for operating the Raman spectrometer and / or Raman analyzer 212 and receiving spectra may also be provided. Furthermore, a monitoring and control system (e.g., BioPAT) may be present. ® The MFCS Scada system connects to database 202 and advisory subsystem (e.g., recommendation system 214), and also to bioprocess control system 210 (e.g., bioprocess controller (Biobrain)). ® Sartorius Steyr Biotechnology Co., Ltd.) and the spectrometer control system receive data and initiate device phases and control loops at the local level.
[0137] The Raman analyzer 212 can be configured to record time-correlated spectra of the biological processes processed in the bioreactor 208, and / or can be used by a spectral model as the aforementioned spectral trajectory data to provide, for example, real-time predictions of the concentrations of the substrate glucose and the metabolite lactate in the bioreactor 208. The real-time predictions of the concentrations of the substrate glucose and the metabolite lactate in the bioreactor 208 can be used, for example, to control the parameters via a proportional-integral-derivative (PID) feedback loop.
[0138] Database 202 may be a relational database, where data is stored in related tables. Communication may be conducted, for example, using a database query language (e.g., Structured Query Language SQL). Database 202 may be hosted on a server system and / or accessible to users via a network (e.g., the cloud). Database 202 may reside on a server cluster under the control of a vendor.
[0139] Figure 5 A computer-implemented method 300 for generating spectral models to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products is shown according to a first embodiment of one aspect.
[0140] exist Figure 5 The embodiments shown exemplarily illustrate the monitoring and / or control of biopharmaceutical processes.
[0141] The purpose of Method 300 is to obtain a spectral (calibration) model to monitor and / or control process parameters and / or monitor product concentration and quality based on cell lines, culture media, products, and process parameters. The user knows what they want to produce and may have the platform technology to achieve that goal. However, the user may not have, or only have a limited amount of time-correlated spectral data, i.e., spectral trajectory data, to generate the spectral model. Method 300, however, can be used to generate the spectral model.
[0142] This disclosure provides an exemplary description of monoclonal antibodies (mAbs) in CHO cell lines within bioreactor 208. This disclosure is also applicable to all other biopharmaceutical processes that manufacture biopharmaceutical products (i.e., recombinant and non-recombinant proteins; vaccines; gene vectors; DNA; RNA; antibiotics, secondary metabolites, growth factors, cells for cell therapy or regenerative medicine; semi-synthetic products, such as artificial organs) in manufacturing systems such as: animal cells (e.g., CHO, HEK, PerC6, VERO, MDCK, etc.); insect cells (e.g., SF9, SF21); microorganisms (e.g., Escherichia coli, Saccharomyces cerevisiae, Pichia pastoris, etc.); algae; plant cells; cell-free expression systems (cell extracts, recombinant ribosome systems, etc.); primary cells; stem cells; natural and transgenic cells; matrix-based cell systems, etc.
[0143] Not only about Figure 5 The bioreactor process mentioned in the first embodiment shown provides significant user benefits for all production stages in which at least one process critical process parameter (CPP) is measured and / or controlled. A process critical process parameter is a parameter that affects the process flow (process key performance indicator - process KPI) and / or product titer, yield, or quality (critical quality attribute - CQA).
[0144] Method 300 can begin from step S302.
[0145] In step S305, the user can record Raman spectral trajectory data during one or more culture cycles. Additionally, the user can record input process parameter data. Accordingly, time-dependent trajectory data is sampled during process execution. This distinguishes method 300 from a different concept that uses non-time-dependent data, where the user records Raman spectra at specific time points (e.g., specific Raman spectra of the cell culture medium before inoculation).
[0146] In step S310, the user can access or switch to the recommendation system 214.
[0147] In step S315, if necessary, the user can log in to the recommendation system 214 using their username.
[0148] In step S320, the user can input the necessary data characterizing the process into the recommender system 214. Examples of the necessary data characterizing the process are one or more of the following: a. Biological type of the process: based on CHO b. Process scaling: 10 L c. Types of processes i. Product type: mAb ii. Process type: Feed-in batch d. Biological systems i.Cellca 2 cell line e. Culture medium platform i. Smart CHO f. Key Quality Attributes (CQA) to be monitored i. Product concentration ii. Polysaccharide profile g. Key process parameters (CPP) to be monitored i. Glucose concentration ii. Lactic acid concentration Specifically, users can input at least one of these necessary data into the recommender system 214. The data under points a through e are parameters that can be used to pre-select relevant spectral trajectory data from the database, and are not all mandatory. The data under points f and g can be input so that the system knows what model should be generated, and are mandatory for extracting spectra with associated relevant CQA or CPP values.
[0149] In step S325, the user can upload at least one Raman spectral trajectory data recorded in step 305 to the recommendation system 214.
[0150] Method 300 may also include the previously described steps S120, S140, S160 and / or S180, namely: - Provide (S120) a database 202 that stores multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology. Each observation dataset includes stored spectral trajectory data, one or more stored process descriptors, and at least one corresponding actual analytical measurement of a process parameter recorded during the execution of the corresponding past process. - (S140) Obtain user-based input data, which includes input spectral trajectory data and / or at least one of one or more input process descriptors. - Based on at least one selection criterion, a subset of the stored spectral trajectory data is determined (S160) by querying the database 202 using user input data. - A spectral model is generated using a subset of the spectral trajectory data stored in (S180), which provides process parameter values for the overall duration of the execution process.
[0151] Specifically, performing steps S120, S140, S160, and S180 may include performing steps S330 to S340 as follows: In step S330, and after the data input is complete, the recommendation system 214 can connect to a database 202 that stores Raman spectral trajectory data, one or more process descriptors, and at least one corresponding actual analytical measurement of process parameters recorded during the execution of the respective past processes. In this regard, database 202 stores one or more reference data, such as product concentration, glucose concentration, lactic acid concentration, etc.
[0152] In step S335, the recorded spectral trajectory data can be compared with multiple existing spectral trajectory data in database 202 (specifically, with the totality of all existing spectra). This comparison may include: a. Raman spectra can be selected from all available spectra based on one or more input parameters, for example, Raman spectra from CHO processes, 10 L fed-batch scales, mAb products, Cellca 2 cell lines, and smart CHO media can be selected.
[0153] b. The selected spectrum, along with its associated data—product, glucose, lactate concentrations, and / or glycan spectra—is transmitted to the recommendation system 214.
[0154] c. The transmitted spectrum can be compared with the input spectrum by the recommendation system 214, as follows: i. At least one spectral preprocessing method can be tested on the input and transmitted spectra, and the general similarity between the unprocessed and processed spectra can be measured or determined by at least one spectral similarity measurement algorithm. The optimal preprocessing / similarity algorithm combination can be selected based on the closest similarity between the input spectrum and the transmitted spectrum.
[0155] ii. For the selected combination of preprocessing / similarity algorithms, the algorithm can select the most similar spectra that can simultaneously cover the maximum concentration range of a given critical process parameter (CPP) or critical quality attribute (CQA).
[0156] In step S340, the recommendation system 214 can use the selected spectral trajectory data and its related information to generate a spectral model that provides process parameter values for the overall duration of the execution process. The related information may be products, glucose, lactate concentrations, or glycan spectra to automatically construct partial least squares (PLS) and / or orthogonal partial least squares (OPLS). ® A calibration model—abbreviated as (O)PLS—is used to predict product, glucose, lactate concentration, or glycan spectra. The automated model building process may include applying the spectral preprocessing steps selected in step S335.
[0157] Method 300 may also include the previously described step S200, namely, using a spectral model to control and / or monitor the process. Specifically, performing step S200 may include performing steps S345 to S375 as follows: In step S345, the spectral models from the recommendation system 214 can be sent to the biological process control system 210. For example, (O)PLS Raman calibration models for product concentration, (O)PLS Raman calibration models for glucose concentration, (O)PLS Raman calibration models for lactate concentration, and / or (O)PLS Raman calibration models for the content of specific polysaccharide species can be sent from the recommendation system 214 to the biological process control system 210.
[0158] The biological process control system 210 now has the necessary information to initiate batch processing and to control and / or monitor the process.
[0159] In step S350, the user can prepare the bioreactor 208 for the process by aseptically filling the reactor with culture medium and establishing the necessary connections.
[0160] In step S355, the user can start batch processing and / or Raman spectrometer.
[0161] Once all setpoints in bioreactor 208 are reached, the user can receive the following information from the bioprocess control system 210: Bioreactor 208 can be incubated with the Cellca 2 cell line. This can be performed by the user in step S360. Alternatively, incubation with the Cellca 2 cell line can be started automatically.
[0162] In step S365, Raman spectra can be continuously measured using a Raman probe connected to a Raman spectrometer that transmits the Raman spectra to a Raman calibration model implemented in the bioprocess control system 210.
[0163] In step S370, the bioprocess control system 210 can control process parameters (e.g., lactate and / or glucose concentrations) and / or monitor CQA (product concentration and / or glycosylation profile) via output from the Raman calibration model. Additional near-line or offline measurements of glucose, lactate, product concentration, and / or glycosylation can be performed daily. Once the formulation is fully finalized, the batch processing is considered complete.
[0164] In step S375, the user can harvest the contents of bioreactor 208 and extract monoclonal antibodies (mAbs) from them. The process control system can then transmit available near-line or offline data, including Raman spectroscopy along with glucose, lactate, product concentration, and / or glycosylation, to database 202 via recommendation system 214. In database 202, the amount of available data has been increased, making the data available to other users.
[0165] about Figure 5 The different steps of the described method 300 can be performed as previously described and Figure 5 The arrows indicate the order in which the steps are performed. However, alternative orders of these process steps are also possible.
[0166] List of reference numerals
[0167] 100 Computer Implementation Methods
[0168] S110 Start Computer Implementation Method
[0169] S120 provides database steps
[0170] S125 Steps to update the database
[0171] S140 Step to obtain user-based input data
[0172] S145 Steps for recording spectral trajectory data
[0173] S150 Steps for uploading spectral trajectory data to the database
[0174] Steps for obtaining spectral trajectory data from the S155 pre-selected database
[0175] S160 The step of determining a subset of the stored spectral trajectory data
[0176] S165 Steps for downloading the calibration model from the database
[0177] S170 The steps of recording the process and spectral measurements, including spectral trajectory data, and processing them via a spectral model.
[0178] S180 is the step of generating a spectral model using a subset of the stored trajectory data.
[0179] S200 uses a spectral model to control and / or monitor the process.
[0180] S220 End Computer Implementation Method
[0181] 200 computer systems
[0182] 202 Database
[0183] 204 One or more processors
[0184] 206 At least one control device
[0185] 208 Bioreactor
[0186] 210 Biological Process Control System
[0187] 212 Raman Analyzer
[0188] 214 Recommendation Systems
[0189] 300 Computer implementation method according to the first embodiment
[0190] S302 Starting with Computer Implementation Methods
[0191] The steps for recording Raman spectral trajectory data during N cultivation cycles of S305
[0192] S310 Steps to enter the recommendation system 214
[0193] Steps to log in to the recommendation system 214 (S315)
[0194] S320 is the step of inputting the necessary data characterizing the process into the recommender system 214.
[0195] S325 Upload the Raman spectral trajectory data recorded in step 305 to the recommendation system 214.
[0196] S330 The step of connecting the recommendation system 214 to the database 202
[0197] S335 The step of comparing the recorded spectral trajectory data with the total existing spectral trajectory data in database 202.
[0198] S340 is the step of using the selected spectral trajectory data and its related information to generate a spectral model.
[0199] S345 The step of sending the spectral model from the recommendation system 214 to the biological process control system 210
[0200] S350 involves preparing bioreactor 208 for this process by aseptically filling the reactor with culture medium and establishing all necessary connections.
[0201] Steps for starting batch processing and Raman spectrometer on S355
[0202] S360 receives information from the bioprocess control system 210 regarding the incubation of bioreactor 208 with Cellca 2 cell line.
[0203] S365 Steps for continuously measuring Raman spectra using a Raman probe connected to a Raman spectrometer
[0204] S370 Steps for controlling process parameters by biological process control system 210
[0205] S375 The step of harvesting the contents of bioreactor 208 and extracting monoclonal antibodies (mAbs) from it.
Claims
1. A computer-implemented method (100) for generating a spectral model, said spectral model being used to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products, said computer-implemented method comprising: - Provide (S120) a database (202) that stores multiple observation datasets associated with corresponding observations of past processes of chemicals, biopharmaceuticals, or biotechnology, each of the observation datasets including: --Stored spectral trajectory data, --One or more stored procedure descriptors, and --At least one corresponding actual analytical measurement of the process parameter recorded during the execution of the corresponding past process; - Obtain user-based input data (S140), the input data including at least one of the following: -- Input spectral trajectory data, and --One or more input procedure descriptors; - Based on at least one selection criterion, a subset of the stored spectral trajectory data is determined (S160) by querying the database (202) using the user-based input data; and - The spectral model is generated using the subset of spectral trajectory data stored in (S180), the spectral model providing process parameter values for the overall duration of the process.
2. The method of claim 1, further comprising using the spectral model to control and / or monitor (S200) the process.
3. The method of claim 1 or 2, further comprising: The selection criterion is defined based on the similarity between the stored observation dataset and the user-based input data, wherein optionally, defining the selection criterion includes a combination of at least one of the steps (i) to (iii): (i) Match the user-based input data with the stored observation dataset; (ii) Spectral preprocessing (applied to the spectral data in the database (202) and the provided user-based spectral data), by applying at least one of the following methods: -AsLS correction (asymmetric least squares correction, which uses a symmetric least squares algorithm to calculate a nonlinear baseline for each spectrum and then subtracts the baseline from the spectrum). - Moving window (removing noise by applying a moving average or median window to the spectrum), -Derivative, -SNV filtering (Standard Normal Transform, which "normalizes" each observation (spectrum) by subtracting the mean and dividing by the standard deviation). -MSC (multiplicative signal correction, which normalizes each spectrum by regressing the average spectrum on a selected set of spectra). -Average the (processed) signal within the selected signal range. - Peak area, peak height, or water zone normalization - Multivariate curve resolution (MCR); (iii) Compare all or part / stages of each process trajectory by applying at least one of the following methods: - Compare the spectral similarity at specific time points or time intervals in the process. -Dynamic time warping -Model array, - TOST test for time series analysis -Calculation of the (spectral) similarity index for the entire process. - Mechanistic modeling for comparing the metabolic behavior of various processes. - Batch modeling — Using PLS models to compare batch time trajectories; The at least one selection criterion is defined by at least one of the following algorithms: - For each time point in the trajectory of the batch evolution model, maximize the correlation coefficient and minimize the average distance. - Minimize the total root mean square error (RMSEP) of prediction between the database batch and the user trajectory in the batch evolution model. - Minimize the RMSEP of the database batch and user trajectory at each time point in the batch evolution model to identify the similarity between time points related to process stages, and - Minimize the distance to the center of the batch-level model.
4. The method as described in any of the preceding claims, wherein, The selection criteria are either the 0-1 standard or the linear standard.
5. The method as described in any one of the preceding claims, wherein, The analytical measurements of the process parameters include product quality attributes or process performance parameters, such as glucose, lactic acid, and / or titers.
6. A computer program product comprising computer-readable instructions that, when loaded onto and executed on a computer system, cause the computer system to perform the operation of the method (100) as described in any of the preceding claims.
7. A computer system (200) operable to control and / or monitor processes for producing chemicals, biopharmaceuticals, or biotechnology products, said computer system (200) including a database (202) and one or more processors (204) associated with said database (202). in, The database (202) is configured as follows: -Storing (S120) multiple observation datasets associated with corresponding observations of past processes involving chemicals, biopharmaceuticals, or biotechnology, each of said observation datasets comprising: --Stored spectral trajectory data, --Stored procedure descriptor, and --At least one corresponding actual analytical measurement of the process parameter recorded during the execution of the corresponding past process; The one or more processors (204) are configured to: - Obtain user-based input data (S140), the input data including at least one of the following: -- Input spectral trajectory data, and -- Input procedure descriptor; - Based on at least one selection criterion, a subset of the stored spectral trajectory data is determined (S160) by querying the database (202) using the user-based input data; and - A spectral model is generated using the subset of spectral trajectory data stored in (S180), the spectral model providing process parameter values for the overall duration of the process.
8. The method of claim 7, wherein, The one or more processors (204) are configured to use the spectral model to control and / or monitor (S200) the process.
9. The computer system as claimed in claim 7 or 8, wherein, The one or more processors (204) are also configured to: The selection criterion is defined based on the similarity between the stored observation dataset and the user-based input data, wherein optionally, the one or more processors (204) are configured to define the selection criterion, the selection criterion comprising a combination of at least one of the steps (i) to (iii): (i) Match the user-based input data with the stored observation dataset; (ii) Spectral preprocessing (applied to the spectral data in the database (202) and the provided user-based spectral data), by applying at least one of the following methods: -AsLS correction (asymmetric least squares correction, which uses a symmetric least squares algorithm to calculate a nonlinear baseline for each spectrum and then subtracts the baseline from the spectrum). - Moving window (removing noise by applying a moving average or median window to the spectrum), -Derivative, -SNV filtering (Standard Normal Transform, which "normalizes" each observation (spectrum) by subtracting the mean and dividing by the standard deviation). -MSC (multiplicative signal correction, which normalizes each spectrum by regressing the average spectrum on a selected set of spectra). -Average the (processed) signal within the selected signal range. - Peak area, peak height, or water zone normalization - Multivariate curve resolution (MCR); (iii) Compare all or part / stages of each process trajectory by applying at least one of the following methods: - Compare the spectral similarity at specific time points or time intervals in the process. -Dynamic time warping -Model array, - TOST test for time series analysis -Calculation of the (spectral) similarity index for the entire process. - Mechanistic modeling for comparing the metabolic behavior of various processes. - Batch modeling — Using PLS models to compare batch time trajectories; The at least one selection criterion is defined by at least one of the following algorithms: - For each time point in the trajectory of the batch evolution model, maximize the correlation coefficient and minimize the average distance. - Minimize the total root mean square error (RMSEP) of prediction between the database batch and the user trajectory in the batch evolution model. - Minimize the RMSEP of the database batch and user trajectory at each time point in the batch evolution model to identify the similarity between time points related to process stages, and - Minimize the distance to the center of the batch-level model.
10. The computer system as claimed in any one of claims 7 to 9, wherein, The selection criteria are either the 0-1 standard or the linear standard.
11. The computer system as claimed in any one of claims 7 to 10, wherein, The analytical measurements of the process parameters include product quality attributes or process performance parameters, such as glucose, lactic acid, and / or titers.