Automatic material supplementing method based on real-time dynamic measurement result

Through real-time dynamic detection results, the consumption or generation rate of cell culture medium is calculated, and the markers are selected for automatic feeding is solved, which solves the problem of insufficient or excessive feeding in traditional methods, and achieves the stability and efficient production of the cell culture process.

WO2025140186A1PCT designated stage expired Publication Date: 2025-07-03WUXI BIOLOGICS CO LTD +1

Patent Information

Application Number
PCT/CN2024/141814
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the existing cell culture process, the determination of feed quantity and time points mainly depends on experience, resulting in insufficient or excessive feed, affecting the stability and yield of cell culture. In addition, traditional methods require a lot of manual operation and offline monitoring, which has lag and error.

Method used

Through real-time dynamic detection results, calculate the consumption or generation rate of components, select the fastest-changing components as markers, and calculate the feeding strategy based on historical data to achieve automated and real-time feeding control to reduce human intervention and errors.

Benefits of technology

It improves the monitoring and control accuracy of the cell culture process, ensures that the nutrient level meets cell needs, stabilizes the process and improves yield, reduces artificial errors and pollution risks, and is suitable for production processes of different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141814_03072025_PF_FP_ABST
    Figure CN2024141814_03072025_PF_FP_ABST
Patent Text Reader

Abstract

An automatic material supplementing method based on a real-time dynamic measurement result, the method comprising: collecting a real-time dynamic measurement result, and calculating a change rate of each component within a unit time, wherein the change rate is a consumption rate or a generation rate; selecting a component having the highest change rate as a marker within a current time period; collecting historical measurement data, and calculating a cell growth rate in a logarithmic growth phase and a target product expression rate in a cultivation and harvesting stage; taking, as a set component concentration value within the current time period of the logarithmic growth phase, a concentration level of the marker when the cell growth rate in the logarithmic growth phase is the highest; taking, as a set component concentration value within the current time period of the cultivation and harvesting stage, a concentration level of the marker when the target product expression rate in the cultivation and harvesting stage is the highest; and calculating an addition dosage of the marker within the current time period on the basis of the change rate of the marker and / or the difference between a real-time measured concentration of the marker within the current time period and the set component concentration value within the current time period, and calculating an addition dosage of a culture medium on the basis of the addition dosage of the marker, so as to obtain a material supplementing strategy.
Need to check novelty before this filing date? Find Prior Art

Description

An automated feeding method based on real-time dynamic detection results

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on the application with CN application number 202311837262.9 and application date December 28, 2023, and claims its priority. The disclosed content of the CN application is hereby introduced as a whole into this application. Technical Field

[0003] The present disclosure belongs to the field of biotechnology, and particularly relates to an automated feeding method based on real-time dynamic detection results. Background Art

[0004] Cell culture has been widely used in the research and development and production of biopharmaceuticals, and an extremely critical factor in ensuring the success of cell culture is the cell culture process. The quality of the cell culture process depends largely on the control of the culture conditions in the bioreactor. In addition to maintaining important physical and chemical conditions such as pH, temperature, and dissolved oxygen (DO), meeting the nutritional needs of cells plays an important role in establishing a suitable growth environment. The metabolism of mammalian cells involves complex metabolic pathways, and the cells consume a large amount of nutrients, such as carbon sources, nitrogen sources, and trace elements. At present, the industry generally uses comprehensive feed medium, which combines nutrients such as carbon sources, nitrogen sources, and trace elements according to the metabolic needs of cells to form a comprehensive feed medium. Summary of the Invention

[0005] A brief overview of the present disclosure is provided below to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts of the present disclosure in a simplified form as a prelude to the more detailed description that will be given later.

[0006] In a first aspect, the present disclosure provides an automated feeding method based on real-time dynamic detection results, the automated feeding method comprising:

[0007] Collect real-time dynamic detection results and calculate the change rate of each component per unit time, which is the consumption rate or generation rate; select the component with the fastest change rate as the marker for the current time period;

[0008] Collect historical test data, calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; use the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration setting value for the current time period in the logarithmic growth phase; use the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum as the component concentration setting value for the current time period in the culture harvest phase;

[0009] The amount of marker added in the current time period is calculated according to the rate of change of the marker, and the amount of culture medium added is calculated according to the amount of marker added, so as to obtain the feeding strategy.

[0010] In some embodiments, the step of selecting the marker comprises:

[0011] Collect the current real-time dynamic detection results and calculate the change rate of each component in the current unit time;

[0012] Collect historical real-time dynamic detection results and calculate the change rate of each component in historical unit time;

[0013] The change rate of each component in the current unit time is combined with the change rate of each component in the historical unit time, and the component with the fastest change rate is selected as the marker of the current time period.

[0014] In some embodiments, the historical detection data is historical detection data of a fermentation process that is the same as the current cell culture process mode.

[0015] In some embodiments, the step of calculating the amount of culture medium additives comprises: 补 =(G 设 -G 实 )×V 培 / G 补

[0016] Among them, m 补 Indicates the amount of feed medium to be added, G 设 Indicates the set marker concentration, G 实 represents the current marker concentration detected by the Raman spectrum in real time, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium.

[0017] In some embodiments, the feed medium includes a marker and other components in a specific ratio, and the amount of other components added is obtained according to the amount of the marker added.

[0018] In some embodiments, after the feeding strategy is derived, the method further includes writing the feeding strategy into a control system.

[0019] In some embodiments, the control system is configured to receive numerical instructions for feeding amount, set value, and feeding time point, and then associate with the peristaltic pump hardware to instantly feed in a dosage manner.

[0020] In a second aspect, the present disclosure provides an automatic feeding device based on real-time dynamic detection results, the automatic feeding device comprising:

[0021] The marker screening unit is used to collect real-time dynamic detection results, calculate the change rate of each component per unit time, which is the consumption rate or generation rate; and select the component with the fastest change rate as the marker for the current time period;

[0022] The marker concentration setting unit is used to: collect historical detection data, calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; use the marker concentration level when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration setting value for the current time period in the logarithmic growth phase; use the marker concentration level when the target product expression rate in the culture harvest phase is the maximum as the component concentration setting value for the current time period in the culture harvest phase;

[0023] The feed amount calculation unit is used to calculate the amount of marker added in the current time period according to the change rate of the marker, calculate the amount of culture medium added according to the amount of marker added, and obtain the feeding strategy.

[0024] In some embodiments, in the marker screening unit, the marker selection step includes:

[0025] Collect the current real-time dynamic detection results and calculate the change rate of each component in the current unit time;

[0026] Collect historical real-time dynamic detection results and calculate the change rate of each component in historical unit time;

[0027] The change rate of each component in the current unit time is combined with the change rate of each component in the historical unit time, and the component with the fastest change rate is selected as the marker of the current time period.

[0028] In some embodiments, in the marker concentration setting unit, the historical detection data is historical detection data of a fermentation process that is the same as the current cell culture process mode.

[0029] In some embodiments, in the feed amount calculation unit, the step of calculating the amount of culture medium added comprises: 补 =(G 设 -G 实 )×V 培 / G 补

[0030] Among them, m 补 Indicates the amount of feed medium to be added, G 设 Indicates the set marker concentration, G 实 represents the current marker concentration detected by the Raman spectrum in real time, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium.

[0031] In some embodiments, in the feed amount calculation unit, the feed medium includes a marker and other components in a specific ratio, and the addition amount of the other components is obtained according to the addition amount of the marker.

[0032] In some embodiments, after the feeding amount calculation unit obtains the feeding strategy, it also includes writing the feeding strategy into the control system.

[0033] In some embodiments, the control system is configured to receive numerical instructions for feeding amount, set value, and feeding time point, and then associate with the peristaltic pump hardware to instantly feed in a dosage manner.

[0034] In a third aspect, the present disclosure provides an automated feeding method based on real-time dynamic detection results, comprising:

[0035] Obtaining historical real-time dynamic detection results, calculating the historical change rate of each component based on the historical real-time dynamic detection results, and selecting a component or a combination of multiple components as a marker based on the historical change rate of each component;

[0036] Collect the current real-time dynamic detection results, and calculate the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection results;

[0037] The feeding strategy is derived from at least one of the following:

[0038] Calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase based on historical real-time dynamic detection results, and use the marker concentration level when the cell growth rate in the logarithmic growth phase is maximum as the marker component concentration set value for the current time period if the current time period enters the logarithmic growth phase; use the marker concentration level when the target product expression rate in the culture harvest phase is maximum as the marker component concentration set value for the current time period if the current time period enters the culture harvest phase, and calculate the additive dosage of the feed medium in the current time period based on the difference between the real-time detection concentration of the marker in the current time period and the component concentration set value in the current time period, to obtain a first feeding strategy; or

[0039] The current change rate of the marker is calculated according to the real-time detection concentration of the marker in the current time period, and the added dosage of the feed medium in the current time period is calculated according to the current change rate of the marker to obtain the second feeding strategy.

[0040] In some embodiments, the automated feeding method further comprises:

[0041] Based on the historical real-time dynamic detection results, calculate multiple historical change rates of each component in multiple historical time periods;

[0042] The multiple historical change rates of each component are combined to obtain a combined historical change rate of each component, and one component or a combination of multiple components is selected as a marker based on the combined historical change rate of each component.

[0043] In some embodiments, the historical real-time dynamic detection results and the current real-time dynamic detection results come from different fermentation processes with the same cell culture process mode.

[0044] In some embodiments, calculating the amount of feed medium added in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the set value of the component concentration in the current time period comprises one or more of the following:

[0045] When the rate of change of the marker is expressed as a consumption rate, m 补 =(G 设 -G 实 )×V 培 / G 补 ,

[0046] Among them, m 补 Indicates the amount of feed medium added in the current time period, G 设 Indicates the component concentration setting value of the marker in the current time period, G 实 Indicates the real-time detection concentration of the marker in the current time period, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium; or

[0047] In the case where the rate of change of the marker is expressed as the rate of generation, m 补 =k×(A 实 -A 设 )×V 培 / B 补 ,

[0048] Among them, m 补 Indicates the amount of feed medium added in the current time period, A 设 Indicates the component concentration setting value of the marker in the current time period, A实 Indicates the real-time detection concentration of the marker in the current time period, V 培 represents the volume of cell culture medium in the bioreactor, B 补 represents the concentration of the medium component from which the marker in the feed medium is converted, and k represents the conversion factor from the concentration of the marker to the concentration of the medium component.

[0049] In some embodiments, calculating the amount of feed medium to be added during the current time period based on the current rate of change of the marker comprises one or more of the following:

[0050] When the rate of change of the marker is expressed as a consumption rate, m 补 =R×V 培 / G 补 , R=G t0 -G 实 ,

[0051] Among them, m 补 represents the amount of feed medium added in the current time period, R is the current rate of change of the marker, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium, G t0 Indicates the real-time detection concentration of the marker in the previous time period, G 实 Indicates the real-time detection concentration of the marker in the current time period; or

[0052] In the case where the rate of change of the marker is expressed as the rate of generation, m 补 =k×R×V 培 / B 补 , R=A 实 -A t0 ,

[0053] Among them, m 补 represents the amount of feed medium added in the current time period, R is the current rate of change of the marker, V 培 represents the volume of cell culture medium in the bioreactor, A t0 Indicates the real-time detection concentration of the marker in the previous time period, A 实 Indicates the real-time detection concentration of the marker in the current time period, B 补 represents the concentration of the medium component from which the marker in the feed medium is converted, and k represents the conversion factor from the concentration of the marker to the concentration of the medium component.

[0054] In some embodiments, the automated feeding method includes: determining a feeding strategy to be applied based on the first feeding strategy and the second feeding strategy.

[0055] In some embodiments, the feeding strategy to be applied is determined as the average of the first feeding strategy and the second feeding strategy.

[0056] In some embodiments, the feeding strategy to be applied is selected from a first feeding strategy and a second feeding strategy based on the phase into which the current time period falls.

[0057] In some embodiments, when the current time period enters the rapid production phase of the target product, the first feeding strategy is selected as the feeding strategy to be applied; when the current time period enters the logarithmic growth phase, the second feeding strategy is selected as the feeding strategy to be applied.

[0058] In some embodiments, calculating the amount of feed medium to be added during the current time period comprises:

[0059] Calculating the amount of the marker or the medium component from which the marker is converted in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the set value of the component concentration in the current time period, or based on the current rate of change of the marker; and

[0060] The amount of feed medium added in the current time period is calculated based on the amount of the marker or the medium component from which the marker is converted added in the current time period.

[0061] In some embodiments, the feed medium comprises a marker and other components in a specific ratio, and wherein the amount of the other components added is obtained according to the amount of the marker added.

[0062] In some embodiments, the automated feeding method includes: writing the feeding strategy into a control system; wherein the control system is configured to, after receiving a numerical instruction regarding the feeding strategy, associate with the peristaltic pump hardware to instantaneously feed in a dosage manner.

[0063] In some embodiments, the real-time detected concentration of the reference component in the current time period is compared with a set value; in response to the real-time detected concentration of the reference component in the current time period being not lower than the set value, the automated feeding method is not executed; or in response to the real-time detected concentration of the reference component in the current time period being lower than the set value, the automated feeding method is executed.

[0064] In some embodiments, the current real-time dynamic detection result is online spectral data, wherein calculating the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a trained text convolution model to obtain the real-time detection concentration of the marker in the current time period, wherein the text convolution model is trained using online spectral data of an indicator of the cell culture fluid in a bioreactor at multiple different time points as samples and offline target values ​​of the indicator measured by sampling the cell culture fluid as labels, and the text convolution model is configured to receive the real-time detection spectral data of the indicator of the cell culture fluid to be predicted and output a predicted value of the indicator.

[0065] In some embodiments, the text convolution model is configured to: set a word vector mapping layer, perform feature dimensionality reduction; and after taking the average in the spatial dimension, use a fully connected structure to obtain the predicted value of the indicator.

[0066] In some embodiments, the text convolution model is configured to: after setting the word vector mapping layer and performing feature dimensionality reduction, further perform neighbor feature sampling.

[0067] In some embodiments, the online spectral data is Raman spectral signal data.

[0068] In some embodiments, the indicators include one or more of the following: pH value, carbon dioxide partial pressure, sodium ion concentration, potassium ion concentration, glucose concentration, lactate concentration, ammonium ion concentration, glutamine concentration, glutamate concentration, target protein concentration, lactate dehydrogenase concentration, terbium ion concentration, phosphate ion concentration, osmotic pressure, viable cell density, cell viability, average viable cell diameter and amino acid concentration.

[0069] In some embodiments, calculating the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a machine learning combination model for predicting the concentration of cell culture fluid components to obtain the real-time detection concentration of the marker in the current time period.

[0070] The machine learning combination model is established by the following method:

[0071] obtaining a dataset of component concentrations of a bioreactor cell culture fluid, the dataset comprising a training dataset, a validation dataset, and a test dataset;

[0072] Using multiple machine learning algorithms to establish multiple single prediction models, wherein each of the multiple single prediction models predicts the test data set after being trained with the training data set and verified with the validation data set;

[0073] Comparing the prediction results of the multiple single prediction models with the test data set to obtain the sum of squares of the prediction errors of the multiple single prediction models, and determining the weights of the multiple single prediction models in the machine learning combination model according to the size of the sum of squares of the prediction errors;

[0074] The multiple single prediction models are combined through a weight assignment method to obtain the machine learning combined model.

[0075] In some embodiments, calculating the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a machine learning combination model for predicting the concentration of cell culture fluid components to obtain the real-time detection concentration of the marker in the current time period.

[0076] The machine learning combined model is a migrated model obtained by migrating the original model.

[0077] The original model is established by the following method:

[0078] obtaining a dataset of component concentrations of a bioreactor cell culture fluid, the dataset comprising a training dataset, a validation dataset, and a test dataset,

[0079] Multiple machine learning algorithms are used to establish multiple single prediction models, wherein each of the multiple single prediction models predicts the test data set after being trained with the training data set and verified with the validation data set.

[0080] Compare the prediction results of the multiple single prediction models with the test data set to obtain the sum of squares of the prediction errors of the multiple single prediction models, and determine the weights of the multiple single prediction models in the original model according to the size of the sum of squares of the prediction errors.

[0081] Combining the multiple single prediction models to obtain the original model through a weight assignment method;

[0082] The model migration includes:

[0083] Obtaining an original data set used to establish the original model and a new batch training data set of cell culture fluid component concentrations of a new batch of biological reactions, wherein the original data set includes an original training data set, an original validation data set, and an original test data set;

[0084] Perform scale correction or scale matching on the new batch training data set and the original training data set to obtain a new training data set.

[0085] Using the new training dataset, the original validation dataset, and the original test dataset, multiple single prediction models are established using multiple machine learning algorithms. After each of the multiple single prediction models is trained with the new training dataset and verified with the original validation dataset, predictions are made on the original test dataset.

[0086] Comparing the prediction results of the multiple single prediction models with the original test data set to obtain the sum of squares of the prediction errors of the multiple single prediction models, and determining the weights of the multiple single prediction models in the migrated model according to the size of the sum of squares of the prediction errors,

[0087] Combining the multiple single prediction models to obtain the migrated model through a weight assignment method,

[0088] The scale correction includes incorporating a specified proportion of data from the new batch training data set into the original training data set used to establish the original model;

[0089] The scale matching includes merging new batches of training data whose numerical differences with the original training data in the original training data set collected at the same time as the new batches of training data into the original training data set.

[0090] In some embodiments, the prescribed ratio is a value selected from 1% to 10%, or a value selected from 1.5% to 7.5%, or a value selected from 2% to 5%. In some embodiments, the prescribed threshold is a value selected from 1% to 10%, or a value selected from 3% to 8%, or a value selected from 4% to 6%, or 5%.

[0091] In some embodiments, the plurality of machine learning algorithms include at least two selected from the following: partial least squares, cubic tree, random forest, support vector machine, time series.

[0092] In some embodiments, the component concentration dataset includes online Raman spectroscopy data and its corresponding offline detection data, and the sampling time of the offline detection data matches the corresponding online Raman spectroscopy data.

[0093] In some embodiments, the component concentration is selected from the group consisting of viable cell density, glucose concentration, lactate concentration, target product concentration, and amino acid concentration.

[0094] In some embodiments, the method further includes performing data preprocessing on the Raman spectroscopy data, wherein the data preprocessing includes at least one of the following: screening out abnormal data points, removing spike peaks, correcting Raman shifts and light intensity, baseline calibration, smoothing, and derivation.

[0095] In some embodiments, selecting a component or a combination of components as a marker based on the historical rate of change of each component includes:

[0096] One component or a combination of multiple components among the N components with the fastest historical change rates is used as a marker, where N is a positive integer.

[0097] In a fourth aspect, the present disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the automated feeding method based on real-time dynamic detection results described in the first aspect or the third aspect are implemented.

[0098] In a fifth aspect, the present disclosure provides a computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, and when the computer program is executed by the processor, the steps of the automated feeding method based on real-time dynamic detection results described in the first aspect or the third aspect are implemented.

[0099] In a sixth aspect, the present disclosure provides a computer program product, comprising instructions, which, when executed by a processor, implement the automated feeding method based on real-time dynamic detection results described in the first aspect or the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] The foregoing and other features and advantages of the present disclosure will become apparent from the following description of the embodiments of the present disclosure taken in conjunction with the accompanying drawings, which are incorporated herein and form a part of the specification and serve to further explain the principles of the present disclosure and enable those skilled in the art to make and use the present disclosure.

[0101] FIG1 is a flow chart of the automated feeding method.

[0102] Figure 2 is a flow chart of the feedback control application.

[0103] Figure 3 is a diagram of manual timed feeding.

[0104] Figure 4 is a real-time dynamic feeding diagram.

[0105] FIG5 is a comparison diagram of live cell density.

[0106] Figure 6 is a comparison chart of cell viability.

[0107] FIG7 is a graph showing the comparison of glucose concentration effects.

[0108] Figure 8 is a graph comparing the effects of lactic acid concentration.

[0109] FIG9 is a graph showing the comparison of lactate dehydrogenase concentrations.

[0110] FIG10 is a graph showing the comparison of target protein expression levels.

[0111] Figure 11 is a graph showing the comparison of protein expression levels in individual cells.

[0112] FIG12 is a diagram showing the comparison of ammonium ion concentrations.

[0113] FIG13 is a graph (I) comparing the amino acid concentrations of manual feeding and dynamic feeding when the marker is glucose.

[0114] FIG14 is a graph (II) comparing the amino acid concentrations of manual feeding and dynamic feeding when the marker is glucose.

[0115] FIG15 is a comparison diagram of live cell density.

[0116] FIG16 is a diagram showing a comparison of cell viability.

[0117] FIG17 is a graph showing the comparison of glucose concentration effects.

[0118] FIG18 is a graph showing the comparison of lactic acid concentration effects.

[0119] FIG19 is a graph showing the comparison of target protein concentrations.

[0120] FIG20 is a schematic diagram of a cell culture feeding control device.

[0121] FIG21 is a flow chart illustrating an automated feeding method based on real-time dynamic detection results according to some embodiments of the present disclosure.

[0122] FIG22 is a flow chart illustrating an automated feeding method based on real-time dynamic detection results according to some embodiments of the present disclosure.

[0123] FIG23 is a schematic block diagram illustrating an automated feeding device based on real-time dynamic detection results according to some embodiments of the present disclosure.

[0124] FIG24 is a schematic block diagram illustrating an automated feeding device based on real-time dynamic detection results according to some embodiments of the present disclosure.

[0125] Figure 25 is a schematic diagram of the training and application of a text convolutional model according to some embodiments of the present disclosure.

[0126] Figure 26 is a schematic diagram of the training and application of a machine learning combination model according to some embodiments of the present disclosure.

[0127] FIG27 is a flowchart of data preprocessing according to some embodiments of the present disclosure.

[0128] FIG28 is a data learning workflow diagram according to some embodiments of the present disclosure.

[0129] FIG29 illustrates two methods of model migration according to some embodiments of the present disclosure.

[0130] FIG30 shows the effect of viable cell density prediction.

[0131] FIG31 shows the glucose concentration prediction effect.

[0132] FIG. 32 shows the lactate concentration prediction effect.

[0133] FIG33 shows the effect of osmotic pressure prediction.

[0134] FIG34 shows the target protein concentration prediction effect.

[0135] FIG35 shows the predicted effect of histidine concentration.

[0136] FIG36 shows the effect of the model on predicting viable cell density when migrating between different scales.

[0137] FIG37 shows the glucose concentration prediction effect of the model when migrating between different scales.

[0138] FIG38 shows a schematic block diagram of a computer device according to some embodiments of the present disclosure;

[0139] FIG39 illustrates a schematic block diagram of a computer system on which embodiments of the present disclosure may be implemented.

[0140] Note that in the embodiments described below, the same reference numerals are sometimes used in common across different drawings to denote the same parts or parts having the same functions, and their repeated descriptions are omitted. In some cases, similar reference numerals and letters are used to denote similar items, so once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0141] For ease of understanding, the positions, sizes, and ranges of various structures shown in the drawings and the like may not represent actual positions, sizes, and ranges, etc. Therefore, the present disclosure is not limited to the positions, sizes, and ranges disclosed in the drawings and the like. DETAILED DESCRIPTION

[0142] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0143] The following description of at least one exemplary embodiment is merely illustrative and is not intended to limit the present disclosure, its application, or use. In other words, the structures and methods herein are presented in an exemplary manner to illustrate various embodiments of the structures and methods of the present disclosure. However, those skilled in the art will appreciate that these are merely exemplary of the disclosure that may be implemented, and are not exhaustive. Furthermore, the drawings are not necessarily drawn to scale, and some features may be exaggerated to illustrate details of specific components.

[0144] In addition, technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0145] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0146] The following describes in detail, in conjunction with the accompanying drawings, various embodiments of the present disclosure, including automated feeding methods based on real-time dynamic detection results. It should be understood that the actual method may also include other additional steps, which are not discussed herein and are not shown in the accompanying drawings to avoid obscuring the key points of the present disclosure.

[0147] At present, in order to maintain the nutrients required for the cell growth environment during the cell culture process, the cell culture medium needs to be fed. However, the determination of the amount of feed medium to be added and the feeding time point is often based on the experience of related processes. This may lead to insufficient or excessive feeding amount per unit time and confusion in feeding time points, thereby affecting process development and ultimately causing variations in yield between batches.

[0148] Cell culture fluid is a crucial factor influencing the success of cell culture and has a significant impact on how to regulate cell culture process parameters. Biopharmaceutical R&D requires extensive cell culture experiments, with the most reliable cell culture process flow screened from these experimental results to ensure long-term stable production processes and yields. Ideally, cell culture experiments would explore the effects of the concentrations of a wide range of cell culture fluid components (such as sugars, metal ions, or amino acids) on cell lines at all times. However, this presents a challenge for the entire biopharmaceutical industry, requiring not only significant time, manpower, and material resources, but also very high technical capabilities. The application of Raman spectroscopy analysis models enables real-time and rapid monitoring of various parameters in the cell culture process, shortening process development time and reducing experimental costs during small-scale laboratory process development, and reducing differences in product quality and yield between batches during large-scale production.

[0149] Feeding is a critical process in the cell culture process, with irreversible consequences. Therefore, developing a real-time, dynamic, and automated feeding strategy is crucial.

[0150] To address one or more of the aforementioned issues, the present disclosure provides an automated feeding method based on real-time dynamic monitoring results. This method utilizes real-time feedback from process analytical technology on cell culture fluid component concentrations to provide timely, on-demand, and automated feedback without human intervention. This addresses the complex and inefficient process development strategies typically associated with large-scale parallel feeding experiments and manual, periodic sampling during cell culture production of recombinant proteins.

[0151] 21 , which shows an automated feeding method 100 based on real-time dynamic detection results (hereinafter referred to as method 100 ) according to some embodiments of the present disclosure. Method 100 includes steps S102 to S106 .

[0152] In step S102, historical real-time dynamic detection results are obtained, historical change rates of each component are calculated based on the historical real-time dynamic detection results, and a component or a combination of multiple components is selected as a marker based on the historical change rates of each component.

[0153] In this article, the rate of change may refer to the change in concentration per unit time. It will be understood that, depending on the specific type of component, its rate of change may be expressed as a consumption rate or a generation rate. For example, for consumable components such as nutrients such as glucose, its rate of change is generally expressed as a consumption rate, while for generating components such as products such as target proteins, its rate of change is generally expressed as a generation rate. In this article, for ease of description, the consumption rate and the generation rate are collectively referred to as the rate of change, and those skilled in the art can clearly determine whether its rate of change is a consumption rate or a generation rate based on whether the specific component is consumable or generating.

[0154] In some embodiments, calculating the historical rate of change of each component based on the historical real-time dynamic detection results may include calculating the real-time detected concentration of each component in the historical time period based on the historical real-time dynamic detection results, and calculating the historical rate of change of each component based on the real-time detected concentration of each component in the historical time period. Various exemplary implementations of calculating the real-time detected concentration of each component (including a marker) based on the real-time dynamic detection results will be described in detail below.

[0155] In some embodiments, step S102 may include: calculating multiple historical change rates of each component in multiple historical time periods based on historical real-time dynamic detection results; merging the multiple historical change rates of each component to obtain a combined historical change rate of each component, and selecting a component or a combination of multiple components as a marker based on the combined historical change rate of each component.

[0156] For example, in the experimental process of therefrom collecting real-time dynamic detection result, every 5 minutes can be used as a time period, so that in the experimental process of for example 14 days by a definite date, multiple such 5 minute time periods can be arranged.The rate of change of component in this time period can be determined based on the real-time detection concentration of component in each time period and the real-time detection concentration in the previous time period.For example, the rate of change of component in multiple time periods can be merged (including but not limited to averaging, seeking weighted average, seeking maximum value, seeking minimum value, seeking median etc.) to obtain the merging rate of change of this component in any suitable manner.Then, it is possible to screen markers based on merging rate of change.

[0157] In some embodiments, selecting a component or a combination of multiple components as a marker based on the historical change rate of each component may include: using one component or a combination of multiple components among the N components with the fastest historical change rate as a marker, where N is a positive integer. In some examples, N=1. In some examples, N>1. For example, N=2 or 3 or 5 or 7 or 10. In some embodiments, the components can be sorted from fast to slow according to the historical change rate, and one component or a combination of multiple components in some of the top-ranked components can be selected as a marker. The components selected as markers may or may not include the components with the fastest historical change rate. In some embodiments, each component and each combination of components in these top-ranked components can be used as candidate markers for culture control experiments, and the candidate markers that perform better in the culture control experiments can be ultimately determined as markers, so that they can be used to determine the feeding strategy in subsequent culture processes.

[0158] In step S104, the current real-time dynamic detection results are collected, and the real-time detection concentration of the marker in the current time period is calculated based on the current real-time dynamic detection results. As mentioned above, various exemplary implementations of calculating the real-time detection concentration of a component (such as a marker) based on the real-time dynamic detection results will be described in detail below.

[0159] In step S106 , a feeding strategy is derived through at least one of sub-steps S1062 and S1064 .

[0160] Specifically, in sub-step S1062, based on the historical real-time dynamic detection results, the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase are calculated, and when the current time segment enters the logarithmic growth phase, the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum is used as the component concentration set value of the marker in the current time segment. When the current time segment enters the culture harvest phase, the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum is used as the component concentration set value of the marker in the current time segment, and the amount of feed medium added in the current time segment is calculated based on the difference between the real-time detection concentration of the marker in the current time segment and the component concentration set value in the current time segment, thereby obtaining a first feeding strategy. For example, the historical real-time dynamic detection results and the current real-time dynamic detection results can be from different fermentation processes with the same cell culture process mode. Specifically, the historical real-time dynamic detection results can be collected during the historical fermentation process, and the current real-time dynamic detection results can be collected during the current fermentation process, and the historical fermentation process and the current fermentation process can have the same cell culture process mode.

[0161] As a non-limiting example, calculating the amount of feed medium to be added in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the set value of the component concentration in the current time period may include one or more of the following:

[0162] When the rate of change of the marker is a consumption rate, in other words, when the marker is a consumable component, m 补 =(G 设 -G 实 )×V 培 / G 补 ,

[0163] Among them, m 补 Indicates the amount of feed medium added in the current time period, G 设 Indicates the component concentration setting value of the marker in the current time period, G 实 Indicates the real-time detection concentration of the marker in the current time period, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium; or

[0164] When the rate of change of the marker is expressed as a production rate, in other words, when the marker is a production-type component, m 补 =k×(A 实 -A 设 )×V 培 / B 补 ,

[0165] Among them, m 补 Indicates the amount of feed medium added in the current time period, A 设 Indicates the component concentration setting value of the marker in the current time period, A 实 Indicates the real-time detection concentration of the marker in the current time period, V 培 represents the volume of cell culture medium in the bioreactor, B 补 represents the concentration of the medium component from which the marker in the feed medium is converted, and k represents the conversion factor from the concentration of the marker to the concentration of the medium component.

[0166] It can be understood that the productive component is generally converted from one or more culture medium components, so the additive amount of the feed medium can be calculated based on the concentration relationship coefficient between the productive component and the one or more culture medium components.

[0167] In sub-step S1064, the dosage of the feed medium to be added in the current time period is calculated according to the current rate of change of the marker, and a second feeding strategy is obtained.

[0168] As non-limiting examples, calculating the amount of feed medium to be added during the current time period based on the current rate of change of the marker may include one or more of the following:

[0169] When the rate of change of the marker is a consumption rate, in other words, when the marker is a consumable component, m 补 =R×V 培 / G 补 , R=G t0 -G 实 ,

[0170] Among them, m 补 represents the amount of feed medium added in the current time period, R is the current rate of change of the marker, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium, G t0 Indicates the real-time detection concentration of the marker in the previous time period, G 实 Indicates the real-time detection concentration of the marker in the current time period; or

[0171] When the rate of change of the marker is expressed as a production rate, in other words, when the marker is a production-type component, m 补 =k×R×V 培 / B 补 , R=A 实 -A t0 ,

[0172] Among them, m 补represents the amount of feed medium added in the current time period, R is the current rate of change of the marker, V 培 represents the volume of cell culture medium in the bioreactor, A t0 Indicates the real-time detection concentration of the marker in the previous time period, A 实 Indicates the real-time detection concentration of the marker in the current time period, B 补 represents the concentration of the medium component from which the marker in the feed medium is converted, and k represents the conversion factor from the concentration of the marker to the concentration of the medium component.

[0173] Substeps S1062 and S1064 can be used independently or in combination. When used in combination, one of substeps S1062 and S1064 can be selected for use according to conditions, such as in stages. Alternatively, the average, weighted, or other combined value of the feed amounts calculated by the two methods can be used.

[0174] Therefore, in some embodiments, method 100 may include determining a feeding strategy to be applied based on the first feeding strategy and the second feeding strategy. In some examples, the feeding strategy to be applied is determined to be the average of the first feeding strategy and the second feeding strategy. In other examples, the feeding strategy to be applied is selected from the first feeding strategy and the second feeding strategy based on the stage in which the current time period falls. For example, when the current time period enters the rapid production period of the target product (e.g., the rapid production period of protein), the first feeding strategy is selected as the feeding strategy to be applied; and when the current time period enters the logarithmic growth period, the second feeding strategy is selected as the feeding strategy to be applied.

[0175] Specifically, calculating the additive amount of the feed medium in the current time period may include: calculating the additive amount of the marker or the medium component converted from the marker in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the set value of the component concentration in the current time period, or based on the current rate of change of the marker; and calculating the additive amount of the feed medium in the current time period based on the additive amount of the marker or the medium component converted from the marker in the current time period. In other words, whether in sub-step S1062 or S1064, the additive amount of the marker or the medium component converted from the marker in the current time period may be calculated first, and then the additive amount of the feed medium in the current time period may be calculated based on the additive amount of the marker or the medium component converted from the marker in the current time period.

[0176] For markers that are consumable components, the amount of feed medium added in the current time period can be calculated directly based on the amount of marker added in the current time period. For example, the feed medium may include a specific ratio of marker and other components, so the amount of the other components added can be calculated based on the amount of marker added. For markers that are productive components, the amount of the medium component from which the marker is converted can be determined based on the concentration change of the standard, and then the amount of feed medium added in the current time period can be calculated based on the amount of the medium component from which the marker is converted.

[0177] In some embodiments, the method 100 may further include writing a feeding strategy into a control system. The control system may be configured to, upon receiving a numerical instruction regarding the feeding strategy, associate the peristaltic pump hardware to instantaneously feed the material in a dosage manner.

[0178] Exemplary, real-time dynamic detection results can be collected every time period (e.g., 5 minutes) and the real-time detection concentration of each component in this time period can be calculated based on the real-time dynamic detection results collected in this time period. One or more or all of these components can be selected as reference components, for example, the marker screened by the above method or the culture medium component converted from it can be used as the reference component. The real-time detection concentration of the reference component in the current time period is compared with the set value, thereby determining whether to execute method 100 based on the result of the comparison. For example, in response to the real-time detection concentration of the reference component in the current time period not being lower than the set value (indicating that no feed supplement is needed at this time), method 100 is not executed (thus saving computing resources), or in response to the real-time detection concentration of the reference component in the current time period being lower than the set value (indicating that feed supplement may be needed at this time), method 100 is executed (thus allowing timely feed supplement). The set value here can be specifically set according to actual needs, and it may have different values ​​for different components, and can be expressed as not only a single value, but also multiple values ​​or ranges.

[0179] 22 shows an automated feeding method 100 ′ (hereinafter referred to as method 100 ′) based on real-time dynamic detection results according to some embodiments of the present disclosure. The method 100 ′ includes steps S102 ′ to S106 ′.

[0180] In step S102 ′, real-time dynamic detection results are collected, and the change rate of each component per unit time is calculated, which is the consumption rate or the generation rate; the component with the fastest change rate is selected as the marker of the current time period.

[0181] In step S104', historical detection data is collected to calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum is used as the component concentration setting value for the current time period in the logarithmic growth phase; the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum is used as the component concentration setting value for the current time period in the culture harvest phase.

[0182] In step S106 ′, the amount of marker added in the current time period is calculated based on the rate of change of the marker, and the amount of culture medium added is calculated based on the amount of marker added to derive a feeding strategy.

[0183] In some embodiments, step S102' may include: collecting real-time dynamic detection results, and calculating the real-time detected concentration of each component in the current time period based on the real-time dynamic detection results; calculating the rate of change of each component per unit time based on the real-time detected concentration of each component in the current time period, and selecting a component or a combination of components based on the rate of change of each component (the selection method can be similar to the above embodiment of method 100) as a marker for the current time period. In some embodiments, step S104' may include: collecting historical detection data, calculating the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase, and if the current time period enters the logarithmic growth phase, using the marker concentration level when the cell growth rate in the logarithmic growth phase is the maximum as the marker component concentration set value for the current time period; if the current time period enters the culture harvest phase, using the marker concentration level when the target product expression rate in the culture harvest phase is the maximum as the marker component concentration set value for the current time period. In some embodiments, step S106' may alternatively include calculating the amount of feed medium to be added in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the component concentration set value for the current time period, and / or based on the rate of change of the marker.

[0184] Method 100' differs from method 100 in that method 100 statically selects markers for feed control during the current culture process based on historical test data, while method 100' dynamically selects markers for feed control during the current culture process based on current test data. Method 100 facilitates stable and reliable feed control. Method 100' also facilitates flexible and adaptive feed control. Other various embodiments of method 100' are similar to those of method 100, and thus, reference can be made to the above description and will not be elaborated upon here.

[0185] As mentioned above, cell culture is now widely used in the development and production of biopharmaceuticals. Maintaining an ideal level of nutrients at each stage of the culture process is crucial, as it significantly impacts the yield, quality, and cost of biopharmaceuticals. To ensure cell growth in an optimal environment, optimizing the feeding control method is crucial. Traditional feeding methods rely on experience with the culture process, setting different ratios for daily or alternate-day feeding. This method not only requires significant manpower for sampling and testing, as well as calculating the required feed volume based on offline monitoring data, but also requires equipment (balances, peristaltic pumps). Furthermore, this feeding method cannot be monitored in a timely manner, which can lead to nutrient deficiencies or excesses at certain points during the cell culture process.

[0186] In contrast, in the present disclosure, the consumption rate or generation rate of the component is calculated based on (historical or current) time increment data, and the marker is selected from the component based on the consumption rate or generation rate of the component; the current detection concentration or the current consumption rate or generation rate of the marker is calculated based on the current time increment data, and the feed amount is dynamically calculated in real time based on the current detection concentration or the current consumption rate or generation rate of the marker. The present disclosure avoids large-scale sampling, reduces human operation, and reduces human error by reversely controlling the feed operation through the marker concentration in the cell culture fluid. At the same time, through real-time monitoring and automatic control, it is ensured that the nutrient level in the cell culture fluid is maintained at a reasonable level that the cells actually need during the entire culture process.

[0187] Figure 1 is a flow chart showing a non-limiting example process in which an automated feeding method according to an embodiment of the present disclosure is applied. As shown in Figure 1, the output parameters of the process analysis technology are first digitally entered into the software, the culture time is matched, and the consumption / generation rate of each component per unit time is calculated. According to the time increment data, the consumption / generation rate of different components in different time periods is calculated. A component or a combination of multiple components with the fastest consumption / generation rate is selected as a marker. According to the cell culture process mode and data, the cell growth rate in the logarithmic growth phase and the target protein expression rate in the culture harvest phase are calculated respectively. When these two parameters are maximum, the concentration level of each component is recorded respectively. The concentration level of the marker when each of these two parameters is maximum is selected as the component concentration set value of the marker when the current time period enters the corresponding stage. When the feed medium contains a marker, the marker addition amount in the current time period is calculated by the consumption rate of the marker, and the amount of other nutrient components added to the bioreactor is calculated according to the concentration ratio of each component in the feed medium. When the feed medium does not contain a marker, the marker generation rate is used to calculate the amount of the medium component to be added during the current time period. Simultaneously, the amount of other nutrients added to the bioreactor is calculated based on the concentration ratios of the components in the feed medium. All numerical values ​​for the feeding strategy are then written to the control system using Object Linking and Embedding (OLE) Process Control (OPC).

[0188] The present disclosure also provides a cell culture feeding control device, which can realize real-time cell culture feeding control. Each unit in the device implements the following operating steps, including:

[0189] (1) Use the concentration of cell culture fluid components fed back by process analysis tools as feedback control markers.

[0190] 1.1 Process analysis tools, including but not limited to Raman spectrometers, infrared spectrometers, online biochemical analyzers, online high performance liquid chromatographs, and online mass spectrometers.

[0191] 1.2 Select feedback control markers using the aforementioned method.

[0192] (2) Map the component concentration value to the process value position of the controller control module through the data transmission protocol (OPC / Modbus).

[0193] 2.1 Controller, including but not limited to Finesse, GE or STR.

[0194] (3) Load the multivariate analysis and processing logic of component concentration into the control module to calculate the feed amount. Identify the process characteristics and determine the feed time.

[0195] 3.1 The multivariate analysis and processing control module is developed based on the following software, including but not limited to SIMCA, R studio, and Matlab.

[0196] (4) Set the correlation formula between feeding amount and feeding time.

[0197] (5) Complete all action operations based on the feeding amount and feeding time.

[0198] (6) Or, each time a new component concentration value is input, the amount of nutrients required by the cells at that moment is calculated, and the calculated data value is mapped to the peristaltic pump of the controller.

[0199] (7) Complete all action operations, complete the fermentation control process, and output the parameters of the fermentation process control.

[0200] Figure 23 shows an automated feeding device 200 (referred to as device 200) based on real-time dynamic detection results according to some embodiments of the present disclosure. The device 200 can be used to implement any embodiment of the above-mentioned method 100. Specifically, the device 200 includes a marker screening unit 202, a marker concentration determination unit 204 and a feeding strategy determination unit 206. The marker screening unit 202 can be configured to obtain historical real-time dynamic detection results, calculate the historical change rate of each component based on the historical real-time dynamic detection results, and select a component or a combination of multiple components as a marker based on the historical change rate of each component. The marker concentration determination unit 204 can be configured to collect the current real-time dynamic detection results, and calculate the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection results. The feeding strategy determination unit 206 can be configured to derive the feeding strategy by at least one of the following: calculating the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase based on historical real-time dynamic detection results, and using the concentration level of the marker when the cell growth rate in the logarithmic growth phase is maximum as the component concentration setting value of the marker in the current time period when the current time period enters the logarithmic growth phase; using the concentration level of the marker when the target product expression rate in the culture harvest phase is maximum as the component concentration setting value of the marker in the current time period when the current time period enters the culture harvest phase, and calculating the additive amount of the feed medium in the current time period based on the difference between the real-time detected concentration of the marker in the current time period and the component concentration setting value in the current time period, thereby deriving the first feeding strategy; or calculating the current change rate of the marker based on the real-time detected concentration of the marker in the current time period, and calculating the additive amount of the feed medium in the current time period based on the current change rate of the marker, thereby deriving the second feeding strategy.

[0201] Various embodiments of the apparatus 200 may be similarly described with reference to various embodiments of the method 100 , and are not described in detail herein.

[0202] Figure 24 shows an automated feeding device 300 (referred to as device 300) based on real-time dynamic detection results according to some embodiments of the present disclosure. The device 300 can be used to implement any embodiment of the above-mentioned method 100'. Specifically, the device 300 includes a marker screening unit 302, a marker concentration setting unit 304 and a feeding amount calculation unit 306. The marker screening unit 302 can be configured to: collect real-time dynamic detection results, calculate the change rate of each component per unit time, which is the consumption rate or the generation rate; select the component with the fastest change rate as the marker for the current time period. The marker concentration setting unit 304 can be configured to collect historical detection data, calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; use the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration setting value for the current time period of the logarithmic growth phase; use the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum as the component concentration setting value for the current time period of the culture harvest phase. Feeding amount calculation unit 306 can be configured to calculate the amount of marker added in the current time period based on the rate of change of the marker, and calculate the amount of culture medium added based on the amount of marker added, thereby determining a feeding strategy. Various embodiments of apparatus 300 can be similarly described with reference to various embodiments of method 100 ′, and will not be further described here.

[0203] Compared with the related art, the present disclosure has the following beneficial effects:

[0204] The methods and devices disclosed herein can improve monitoring and control of the cell culture process, responding to the dynamically changing nutritional needs of cells. They can adjust the addition of mixed feed media based on the concentrations of the components in the cell culture environment, providing an optimal culture environment for cells, maximizing process stability and increasing yield. They can intelligently generate feeding strategies to improve efficiency, and automatically perform feeding operations based on the generated feeding strategies, reducing human error and contamination risks. These methods and devices improve the speed and accuracy of cell culture process development, provide customized feed media for different cell lines, clones, and processes, and can be rapidly applied to process development and production processes of varying scales.

[0205] The technical solutions of the present disclosure are further illustrated below by specific examples. Those skilled in the art should understand that the examples are only for facilitating understanding of the present disclosure and should not be regarded as specific limitations of the present disclosure.

[0206] If no specific techniques or conditions are specified in the examples, the experiments were carried out according to the techniques or conditions described in the literature in the field or according to the product instructions. If no manufacturer is specified for the reagents or instruments used, they are all conventional products that can be purchased through regular channels.

[0207] Comparative Example 1

[0208] This comparative example utilized a traditional manual timed feeding strategy. Traditional feeding is categorized by feed volume or physical and chemical properties. Typically, feeds are divided into three types: "large feeds" (weakly acidic feeds with larger volumes), "small feeds" (alkaline feeds with smaller volumes), and separate glucose mother solution feeds. The amounts and timing of these feeds are determined based on statistical analysis of nutrient consumption rates obtained from cell culture experiments.

[0209] The manual timed feeding strategy in this comparative example is to feed every other day, that is, to manually add feed medium to the reactor on the 2nd, 4th, 6th, 8th, 10th, and 12th days of culture. The calculation formula for the feed amount is:

[0210] Day 2, 4, 6, m 大补 =V 培 ×3%;m 小补 =V 培 ×0.3%

[0211] Day 8, 10, m 大补 =V 培 ×4%;m 小补 =V 培 ×0.4%

[0212] Day 12, m 大补 =V 培 ×1%;m 小补 =V 培 ×0.1%

[0213] The strategy for glucose supplementation is generally based on experience or experimental data to control the glucose concentration within an appropriate range. The residual glucose concentration in the cell culture medium is detected by offline biochemical detection equipment, and the amount of glucose required for daily supplementation is calculated. The formula for manual glucose supplementation is as follows: m 补糖 =(G 糖设 -G 糖残 )×V 培 / G 补

[0214] Among them, m 补糖 Indicates the amount of glucose mother solution to be added, G 糖设 Indicates the set glucose concentration, G 糖残 Represents the current residual glucose concentration measured by offline equipment, V 培represents the volume of cell culture medium in the bioreactor, G 补 Indicates the glucose concentration in glucose mother solution.

[0215] Example 1

[0216] The automated feeding method based on real-time dynamic detection results was used to automatically control feeding in Case 1. Simultaneously, the traditional manual timing feeding strategy in Comparative Example 1 was used for a comparative experiment.

[0217] In this embodiment, the feedback control application flow chart is shown in FIG2 .

[0218] The feedback control flow chart covers a variety of cell culture processes, such as fed-batch culture (batch feeding) and perfusion (continuous feeding). In both processes, process analytical technology can be used to monitor the concentration of components in the cell culture fluid. After analyzing and calculating each concentration, the disclosed method is used to control each process operation control point in the flow chart. These include, but are not limited to, batch feeding operation points, continuous feeding operation points, cell discharge operation points (to stabilize the perfusion process, a certain amount of starved cell fluid must be discharged from the reactor in real time or periodically), pH control operation points (alkali addition), ventilation control operation points, and harvest rate control operation points.

[0219] The experimental materials and related experimental methods in Case 1 are as follows:

[0220] In this example, a CHO-K1 cell line producing a monoclonal antibody was used in the cell culture experiments. The cells were cultured in a Kuhner shaker at 36.5°C, 110 rpm, and 6% CO2 in the seeding phase. The culture medium used was a self-developed medium (MagniCHO). A 3L reactor was used in the production phase, with an initial culture volume of 1.5L, a culture temperature of 36.5°C, a pH setting of 6.90±0.25, a dissolved oxygen saturation of 40%, and an initial inoculation density of 1.0×10 6 Cell number / mL, feed medium is also self-developed medium (MagniCHO Feed1), materials are shown in Table 1.

[0221] Table 1

[0222] High-dimensional data analysis and logical calculations of the process were implemented using the open source programming tool R (version 4.0.5).

[0223] The automated feeding method based on real-time dynamic detection results includes the following steps:

[0224] (1) Obtain historical real-time dynamic detection results collected during historical fermentation processes with the same cell culture process mode as the current fermentation process, and calculate multiple historical consumption / generation rates of each component in multiple historical time periods; merge the multiple historical consumption / generation rates of each component to obtain a combined historical consumption / generation rate of each component, and select the component with the fastest combined historical consumption / generation rate as a marker.

[0225] (2) Collecting the current real-time dynamic detection results during the current fermentation process, and calculating the real-time detection concentration of the marker in the current time period based on the current real-time dynamic detection results.

[0226] (3) Based on the historical real-time dynamic detection results, the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase are calculated, and the concentration level of each component when these two parameters are maximum is counted. The concentration level corresponding to the screened marker is used as the component concentration setting value for the current time period.

[0227] (4) Calculating the amount of the marker or the medium component converted from the marker in the current time period, and calculating the amount of the feed medium based on the amount of the marker or the medium component converted from the marker, to obtain a feeding strategy. The feeding strategy is written into the control system; after receiving the numerical instructions of the feeding amount, set value, and feeding time point, the control system is associated with the peristaltic pump hardware to instantaneously feed the feed in a dosage manner. After completing the current operation, the control operation is stopped until the next feeding instruction is issued.

[0228] In the case where the marker is a consumable component, the steps for calculating the amount of feed medium to be added include: 补 =(G 设 -G 实 )×V 培 / G 补

[0229] Among them, m 补 Indicates the amount of feed medium to be added, G 设 Indicates the set marker concentration, G 实 represents the current marker concentration detected by the Raman spectrum in real time, V 培 represents the volume of cell culture medium in the bioreactor, G 补 represents the marker concentration in the feed medium.

[0230] In the case where the marker is a productive component, the steps for calculating the amount of feed medium to be added include: 补 =k×(A 实 -A 设 )×V 培 / B 补 ,

[0231] Among them, m 补 Indicates the amount of feed medium added in the current time period, A 设 Indicates the component concentration setting value of the marker in the current time period, A 实 Indicates the real-time detection concentration of the marker in the current time period, V 培 represents the volume of cell culture medium in the bioreactor, B 补 represents the concentration of the medium component from which the marker in the feed medium is converted, and k represents the conversion factor from the concentration of the marker to the concentration of the medium component.

[0232] The feeding strategies of each time are summarized to obtain the automated feeding strategy of the current fermentation process.

[0233] Based on the real-time dynamic detection results collected during the historical fermentation process, glucose (a consumable component) was selected as a marker for fermentation control. The feeding time and feeding amount of glucose are shown in Figures 3 and 4. Figure 3 shows the manual timed feeding diagram, and Figure 4 shows the real-time dynamic feeding diagram.

[0234] The diagrams of manual and real-time dynamic feeding show that a total of seven feeding operations occurred during the manual feeding process (indicated by the light-colored line at the bottom), and the corresponding marker (glucose) concentration also changed at the moment of feeding (indicated by the dark-colored line at the top). Real-time dynamic feeding, on the other hand, evenly distributes the feeding amount throughout the cell culture process, and the marker process value is always maintained near the set value.

[0235] Fermentation control was performed according to the automated feeding strategy. In addition, other unselected components were used as markers for fermentation control and comparative experiments were conducted. The unselected components included aspartic acid, asparagine, histidine, methionine or a combination thereof.

[0236] The fermentation results of the above automated feeding strategy are as follows:

[0237] Figure 5 compares viable cell density. The results demonstrate that manual timed feeding and automatic real-time dynamic feeding have little impact on viable cell density during cell culture. The growth trends and peak densities of the viable cell density parameters under both feeding methods are comparable.

[0238] Figure 6 shows a comparison of cell viability. The results show that although manual timed feeding can control the glucose level, the lactate concentration is still high due to the long lag and fluctuating glucose concentration, resulting in low cell viability.

[0239] Figure 7 shows a comparison of glucose concentrations. The results indicate that manual timed feeding, which uses daily offline data to control the amount of glucose added on the second or third day, exhibits significant hysteresis and large fluctuations in glucose concentration in the cell culture medium.

[0240] Figure 8 shows a comparison of lactate concentrations. The results indicate that the traditional method, which relies on daily offline monitoring data to adjust the amount of glucose added to feed the second or third day, controls lactate levels. Manually scheduled feeding has significant lag, and the glucose concentration in the cell culture medium fluctuates widely. Therefore, controlling glucose concentration to control lactate levels is difficult, resulting in elevated lactate concentrations.

[0241] Figure 9 is a comparison of lactate dehydrogenase concentrations. The results show that the lactate dehydrogenase concentration in the traditional manual feeding operation is higher than that in the real-time dynamic feeding operation. Due to the higher lactic acid concentration, the pH of the cell culture medium is higher, and the rapid decline in cell viability causes the culture medium state to deteriorate further, more cells die and rupture, and release more lactate dehydrogenase.

[0242] Figure 10 shows a comparison of target protein expression, and Figure 11 shows protein expression in individual cells. Both figures demonstrate that the yields of real-time dynamic feeding and manual timed feeding are comparable. All dynamic feeding methods are superior to manual feeding, with dynamic feeding using methionine and glucose as markers significantly outperforming manual feeding.

[0243] Figure 12 shows a comparison of ammonium ion concentrations, demonstrating that real-time dynamic feeding also controls ammonium ion levels, a metabolic waste product, to lower levels than manual feeding. This performance also increases viable cell density and cell viability, thereby ensuring high target protein production.

[0244] Figure 13 is a graph comparing the amino acid concentrations of scheduled manual feeding and real-time dynamic feeding when the marker is glucose (I), and Figure 14 is a graph comparing the amino acid concentrations of scheduled manual feeding and real-time dynamic feeding when the marker is glucose (II). The results show that different types of amino acids have different concentration levels, and the concentration variation range of dynamic feeding is more stable and smaller than that of manual feeding, which may provide a more stable nutritional environment that is more suitable for cell growth and protein synthesis. In Figure 13, the 5-6mM level is proline; the 2-4mM level is leucine; and the 0-2mM level is cysteine. In Figure 14, the 3-6mM level is threonine; the 2-4mM level is lysine; and the 0-1mM level is tryptophan. Scheduled manual feeding is indicated by square marks; real-time dynamic feeding is indicated by circular marks.

[0245] In summary, in Case 1, after calculation and screening, glucose was ultimately used as a marker. Furthermore, in Case 1, there was a problem of lactic acid accumulation during the culture process. Selecting glucose as a marker not only enabled subsequent automatic feeding control, but also allowed lactic acid levels to be controlled by controlling the glucose concentration. The traditional method is to control lactic acid by regulating the amount of glucose added on the second or third day based on daily offline test data. This method has a serious hysteresis and the glucose concentration in the cell culture medium fluctuates greatly (as shown in the circular curve in FIG7 ). Although the glucose level is controlled, the lactic acid concentration remains high due to the long lag and fluctuating glucose concentration, resulting in low cell density and viability. The real-time dynamic feeding action in this embodiment effectively controls the glucose concentration within a relatively low range throughout the culture process (as shown in the diamond curve in FIG7 ), while maintaining the lactic acid level within a low concentration range at the end of the cell culture (as shown in the diamond curve in FIG8 ), thereby maintaining a high cell viability (as shown in the diamond curve in FIG6 ). Ultimately, at the end of the cell culture, the cells can more efficiently synthesize more target protein ( FIG10 ). The lactate dehydrogenase concentration curve also supports the low viability at the end of the culture (Figure 9). The trends in the amino acid comparison diagrams in Figures 13 and 14 also demonstrate that when glucose is used as the feed marker, real-time dynamic feeding can reduce fluctuations in amino acid concentrations in the cell culture medium, maintaining amino acid concentrations within a stable range.

[0246] In the comparative experiment, the sum of aspartic acid and asparagine, or histidine, or methionine was selected as a marker. During the cell culture process, asparagine and aspartic acid will achieve dynamic equilibrium through acyl transfer. When selecting a marker, choosing only one of them will cause an error in the feeding operation, thereby leading to cell culture failure. Therefore, the sum of the concentrations of the two substances in the cell culture fluid is used as a marker to control the feeding operation. It can be seen from Figure 12 that the toxic byproduct ammonium ion in another cell culture fluid is also reduced by 100% by the use of amino acid concentration as a counter-control marker for real-time dynamic feeding operation. This phenomenon also helps to improve cell viability and cell status, and helps to increase the yield of target protein. At the same time, it can be found from the curve controlled by amino acids that the performance of the other process parameters is similar.

[0247] Example 2

[0248] This example investigates the effects of different ranges of glucose concentration on fermentation based on the automated feeding strategy in Example 1, further verifying the advantages of the automated feeding strategy disclosed herein.

[0249] Using glucose as a marker, the same real-time dynamic feeding action was used to feed at high and low concentration levels. The fermentation test results are shown below:

[0250] FIG15 is a graph showing a comparison of viable cell density. The results show that higher glucose concentration levels can maintain higher biomass (viable cell density levels).

[0251] FIG16 is a graph showing a comparison of cell viability. The results show that higher glucose levels have a poorer effect on maintaining viability in the later stages of culture. This may be because higher glucose concentrations induce more lactic acid production.

[0252] FIG17 is a graph showing a comparison of glucose concentrations. The results show that the diamond curve has a higher glucose concentration, which is approximately 5-6 g / L.

[0253] Figure 18 is a graph comparing the effects of lactate concentration. The results show that high concentrations of glucose cannot maintain lower lactate levels. This may be because the enrichment of glucose in the glycolysis pathway of the cell causes excessive lactate accumulation, while glucose does not enter the tricarboxylic acid cycle pathway.

[0254] Figure 19 is a comparison chart of the target protein concentrations. The results show that excessively high lactic acid levels cause the pH of the culture medium to be too low. The cells change their state in such low pH, and the protein synthesis pathway is blocked, resulting in lower yields.

[0255] In summary, when glucose was selected as the feeding marker, given different ranges of glucose concentrations, and using the same real-time dynamic feeding action, it was also demonstrated that lower glucose concentrations could effectively control lactate concentration, thereby ensuring cell viability and stable yield. The above comparative experiments show that the automated feeding strategy disclosed herein can automatically generate a feeding strategy based on real-time detection results, and that this feeding strategy achieves better fermentation results than traditional feeding strategies.

[0256] Example 3

[0257] A cell culture feeding control device, comprising:

[0258] (1) Use the concentration of cell culture fluid components fed back by process analysis tools as feedback control markers.

[0259] 1.1 Process analysis tools, including but not limited to Raman spectrometers, infrared spectrometers, online biochemical analyzers, online high performance liquid chromatographs, and online mass spectrometers.

[0260] 1.2 Select feedback control markers using the method of Example 1.

[0261] (2) Map the component concentration value to the process value position of the controller control module through the data transmission protocol (OPC / Modbus).

[0262] 2.1 Controller, including but not limited to Finesse, GE or STR.

[0263] (3) Load the multivariate analysis and processing logic of component concentration into the control module to calculate the feed amount. Identify the process characteristics and determine the feed time.

[0264] 3.1 The multivariate analysis and processing control module is developed based on the following software, including but not limited to SIMCA, R studio, and Matlab.

[0265] (4) Set the correlation formula between feeding amount and feeding time.

[0266] (5) Complete all action operations based on the feeding amount and feeding time.

[0267] (6) Or, each time a new component concentration value is input, the amount of nutrients required by the cells at that moment is calculated, and the calculated data value is mapped to the peristaltic pump of the controller.

[0268] (7) Complete all actions. See Figures 4 to 19 for the specific effects.

[0269] FIG20 is a schematic diagram of a cell culture feeding control device.

[0270] In summary, the present disclosure utilizes reverse control of nutrient concentration in the cell culture fluid to feed the cells, thereby avoiding extensive sampling, reducing labor costs, and preventing human error. Furthermore, through real-time monitoring and automated control, the nutrient levels in the cell culture fluid are maintained at a reasonable level consistent with the actual needs of the cells throughout the entire culture process.

[0271] Various embodiments for calculating the real-time concentration of a component based on real-time dynamic detection results are described in detail below. It will be appreciated that the following embodiments can be applied regardless of whether the real-time concentration of each component in a historical time period is calculated based on historical real-time dynamic detection results, or the real-time concentration of each component or marker in a current time period is calculated based on current real-time dynamic detection results.

[0272] In some embodiments, the real-time dynamic detection result is online spectral data (such as but not limited to Raman spectral signal data, etc.), wherein calculating the real-time detection concentration of the component based on the real-time dynamic detection result includes inputting the real-time dynamic detection result into a trained text convolution model to obtain the real-time detection concentration of the component. Specifically, the text convolution model is trained by using the online spectral data of the indicators of the cell culture fluid in the bioreactor at multiple different time points as samples and the offline target values ​​of the indicators measured by sampling the cell culture fluid as labels. The text convolution model is configured to receive the real-time detection spectral data of the indicators of the cell culture fluid to be predicted and output the predicted value of the indicator. In some examples, the text convolution model is configured to: set a word vector mapping layer, perform feature dimensionality reduction; and after taking the average in the spatial dimension, use a fully connected structure to obtain the predicted value of the indicator. Further, in some examples, the text convolution model is configured to: after setting the word vector mapping layer and performing feature dimensionality reduction, further perform neighbor feature sampling. Specifically, the indicators may include one or more of the following: pH value, carbon dioxide partial pressure, sodium ion concentration, potassium ion concentration, glucose concentration, lactate concentration, ammonium ion concentration, glutamine concentration, glutamate concentration, target protein concentration, lactate dehydrogenase concentration, terbium ion concentration, phosphate ion concentration, osmotic pressure, viable cell density, cell viability, average viable cell diameter, and amino acid concentration. For example, FIG25 schematically illustrates the training process of the text convolutional model based on historical Raman spectra and historical component data, and the application process of predicting the concentrations of various components in cell culture fluid based on input online Raman spectral data.

[0273] Text convolution is the application of convolutional neural networks (CNN) in text classification. Convolutional neural networks are a deep learning method commonly used in image processing. When used for text processing, a specialized text convolution model, TextCNN, is derived. Yoon Kim proposed TextCNN in his paper (2014 EMNLP) Convolutional Neural Networks for Sentence Classification. The core idea of ​​convolutional neural networks is to capture local features. For text, local features are sliding windows consisting of several words, similar to N-grams. The advantage of convolutional neural networks is that they can automatically combine and filter N-gram features to obtain semantic information at different levels of abstraction. When convolutional neural networks (CNN) are applied to text classification tasks, multiple convolution kernels of different sizes are used to extract key information in sentences, thereby better capturing local correlations.

[0274] This property makes it well-suited for use in the context of the present disclosure. This is because the relationships among the various parameters of the culture medium in this disclosure are often interrelated. For example, an increase in carbon dioxide partial pressure may lead to a decrease in pH while also increasing the overall pressure of the system. This parameter property is similar to natural language processing using convolutional neural networks.

[0275] The present disclosure can use spectral data as detection data. In some examples, the present disclosure uses a Raman spectrometer to detect Raman spectral signals in cell culture fluid during real-time cell culture as detection data (training set), and uses a machine learning model to analyze the Raman spectral data in real time and establish a relationship model with various indicators of the culture fluid detected offline at the corresponding time.

[0276] Raman spectroscopy is a vibrational spectroscopy technique that detects and identifies molecular vibrations by measuring the Raman scattering effect of excitation light on a sample. It enables non-destructive analysis of chemical composition and molecular structure. The number, frequency shift, intensity, and shape of the Raman spectral bands generated by Raman scattering are directly related to the vibration and rotation of the molecules. In particular, under certain conditions, their intensity is linearly correlated with the concentration of the substance. This allows for the detection of a substance's structure, composition, and concentration. Compared to other spectral analysis methods such as infrared, near-infrared, and ultraviolet fluorescence, Raman spectroscopy offers significant advantages, including a wide detection range; non-destructive, rapid, and pollution-free operation; remote detection technology; and high sensitivity.

[0277] Consequently, with improvements in laser sampling and detector technology, the application of Raman spectroscopy in polymer, pharmaceutical, biomanufacturing, and biomedical analysis has surged over the past three decades. Thanks to these technological advances, Raman spectroscopy has become a practical analytical technique for use both inside and outside the laboratory. In the pharmaceutical bioreactor sector, Raman spectroscopy is often used for online monitoring. Since the first reported application of in situ Raman measurements in biomanufacturing, it has been used to provide online, real-time predictions of several key process states, such as glucose, lactate, glutamate, glutamine, ammonia, and VCD.

[0278] However, due to the principles and characteristics of Raman spectroscopy, certain values ​​within certain bands are highly sensitive to changes in component concentrations and can significantly impact predictions. Some floating-point changes are subtle, requiring scaling and mapping to a two-dimensional spatial structure to facilitate extraction. Furthermore, the affected bands can be fragmented, and the fragments themselves are not fixed. Therefore, convolutional neural networks with fixed convolution kernels often fail to accurately capture specific features.

[0279] Based on the above reasons, the inventors of the present invention referred to the model structure of text convolution-TextCNN, which is often used in text processing, and used one-dimensional convolution with different kernel sizes to extract features of different step-length neighborhoods in the mapping feature map, and used the spatial correlation constraints of the intermediate feature extraction layer to guide the full connection operation to preserve the feature relationship of the original position neighbors, and accurately predicted the component indicators.

[0280] In the present disclosure, items such as concentration, pressure, density, temperature, and content in a bioreactor are collectively referred to as indicators.

[0281] The following is a further detailed description of the method using the text convolution model through specific examples.

[0282] Example 4

[0283] First, Raman spectral data and offline target values ​​were collected. A Raman spectrometer was used to obtain online spectral data of various indicators over time from the cell culture medium in the bioreactor. Offline target values ​​for these indicators were calculated based on time. Samples were taken periodically throughout the experiment and offline detection equipment was used to obtain offline target values ​​for each batch to calibrate the calculated offline target values.

[0284] Specifically, a monoclonal antibody-producing cell line (Chinese hamster ovary cells) was used in the cell culture. The cells were cultured in a Kuhner shaker at 36.5°C, 110 rpm, and 6% CO2 in the seed phase using Cytiva's Hyclone Actipro medium. During the production phase, 3L and 250L reactors were used, with initial culture volumes of 1.5L and 140L, respectively. The culture temperature was 36.5°C, the pH setting was 6.90 ± 0.25, the dissolved oxygen saturation was 40%, and the initial inoculation density was 1.0 × 10 6 cells / mL, and the feed medium was also Cytiva's Hyclone Cell boost7a / 7b (10% / 1%).

[0285] Online spectral data refers to the full Raman shift spectral data collected by the Raman spectrometer. The present disclosure uses a Raman Rxn2 analyzer (Kaiser Optical Systems) equipped with an immersion optical probe. The probe is installed in a 3L bioreactor (Applikon) and directly immersed in the cell culture suspension. The Raman spectra of different bioreactor objects are recorded throughout the experiment. For a single recorded spectrum, 30 subsequent spectra are captured with a 10s exposure time and averaged, resulting in a collection interval of approximately 5min for each bioreactor. The laser excitation wavelength is 785nm, providing a 100-3425cm -1 spectral coverage (Raman shift).

[0286] Offline target values ​​refer to cell culture fluid metrics obtained through real-time on-site sample testing. During the production culture phase, the Raman instrument probe is placed in the culture fluid. Samples are collected five times daily and viable cell density is measured using a Beckman Vi-Cell XR. Glucose, lactate, and target protein concentrations are measured using a Roche Cedex Bio analyzer. Amino acid concentrations are measured using an Agilent HPLC.

[0287] Secondly, build a machine learning model. The process of building a machine learning model includes: data processing, model building, data parameter adjustment, and model verification.

[0288] Data processing: Spectral data is read. Each sample contains a spectrum file and offline target value. The spectrum file is converted into spectral values. The detection time of the spectrum is mapped to the detection time of the offline target value. The modeling feature data (spectral values ​​corresponding to different Raman shifts) and target values ​​are obtained. Data pre-processing randomly splits the training data into validation data and test data, and the feature values ​​of the training data are standardized and normalized. The test data is standardized according to the training data standardization rules.

[0289] Model building: Spectral data of different sizes and lengths were used as input to construct various machine learning models. The machine learning models established in this embodiment include partial least squares regression (PLS), Xgboost, convolutional neural network (CNN), residual network (ResNet), and text convolution (TextCNN) models.

[0290] Data parameter adjustment: The model is trained using the training set, and the model hyperparameters are adjusted based on the performance of the validation set. The specific parameters of each model are as follows:

[0291] PLS - Apply Savitzky-Golay filter to filter and then find the second-order derivative, input the derivative into the model, and n_components is 5.

[0292] Xgboost--First normalize the spectral data, PCA reduce the dimension to 50, and set n_components to 500.

[0293] Convolutional Neural Network (CNN) model - 1D convolution superimposed on a maximum pooling layer, repeated 4 times and then connected to two fully connected layers to output the result.

[0294] Residual Neural Network (ResNet) - at least 2 residual network layers, each of which contains multiple one-dimensional convolutional layers, multiple maximum pooling layers, and fully connected layers.

[0295] Text Convolution (TextCNN) - When building a text convolution model, the following methods are further adopted:

[0296] 1) Set up the word embedding layer: This maps the input one-dimensional floating-point vector to a two-dimensional vector of a specified length and feature dimension. Because each Raman shift value is a single floating-point number, feature extraction requires considering its specific magnitude, making it impossible to use a word embedding matrix. Therefore, a fully connected architecture is used to achieve feature mapping. This establishes a fixed mapping between input and output, while also preserving the spatial structure to a certain extent.

[0297] 2) Perform feature dimensionality reduction: Directly mapping the input floating-point vector to a 3000-length spectral value will produce a 3000×N (N is the feature mapping dimension) feature vector. Therefore, the resulting dimension needs to be compressed to prevent the introduction of large computational overhead at the beginning of the operation. In practice, the feature dimension after mapping is predefined to be A×N, where A is the length of the mapped feature vector and N is the feature dimension. Using a fully connected structure, the original input vector is directly mapped to a length A×N vector, and then a reshape operation is used to generate the mapping features. This compresses the feature length while simultaneously mapping the features.

[0298] 3) Neighbor feature sampling: as needed, after setting the word vector mapping layer and performing feature dimensionality reduction, further neighbor feature sampling can be performed. Since it is known that different substances correspond to one or several characteristic regions in the spectrum, not all feature extractions of input dimensions are valid in feature extraction and operation, and effective feature extraction requires the ability to perform feature operations between adjacent spectra. Since the full connection operation as a global feature integration disrupts the original spatial structure distribution in setting the word vector mapping layer and performing feature dimensionality reduction, it is necessary to be able to reintroduce spatial association in the subsequent structure. Specifically, referring to the textCNN structure, one-dimensional convolutions with different kernel sizes are used to extract features of different step-length neighborhoods in the mapping feature map, and the spatial association constraints of the intermediate feature extraction layer are used to guide the full connection operation to preserve the feature relationship of the original position neighbors.

[0299] 4) After taking the mean value in the spatial dimension, the fully connected structure is used to predict the corresponding target value.

[0300] First, we validate the model and obtain the validation results of the test set. In addition, to ensure the production applicability of the model, we select a forward-looking wet experiment to evaluate the effectiveness of all models.

[0301] Specifically, the wet experiment process is described as follows: Under the same culture conditions as the historical training data, two Raman sensors periodically scanned two bioreactors to collect Raman spectral data. Four points were collected daily, and samples were collected simultaneously for offline data collection. A total of 140 data points were collected over the entire culture cycle. The root mean square error (RMSE) between the offline and online data was calculated. The results are shown in Table 2. A comparison of multiple models with various parameters revealed that TextCNN significantly outperformed most metrics.

[0302] Table 2 RMSE of model-predicted values ​​and offline measured values ​​in cell culture medium

[0303] In other embodiments, calculating the real-time detected concentration of a component based on the real-time dynamic detection results includes inputting the real-time dynamic detection results into a machine learning combination model for predicting the concentration of a cell culture fluid component to obtain the real-time detected concentration of the component. The machine learning combination model is established by the following method: obtaining a dataset of component concentrations of a bioreactor cell culture fluid, the dataset comprising a training dataset, a validation dataset, and a test dataset; using multiple machine learning algorithms to establish multiple single prediction models, wherein each of the multiple single prediction models predicts the test dataset after being trained with the training dataset and verified with the validation dataset; comparing the prediction results of the multiple single prediction models with the test dataset to obtain the sum of squared prediction errors of the multiple single prediction models, and determining the weights of the multiple single prediction models in the machine learning combination model based on the size of the sum of squared prediction errors; and combining the multiple single prediction models to obtain the machine learning combination model through a weight assignment method.

[0304] In some other embodiments, calculating the real-time detected concentration of a component based on the real-time dynamic detection results includes inputting the real-time dynamic detection results into a machine learning combination model for predicting the concentration of a cell culture fluid component to obtain the real-time detected concentration of the component. Unlike the aforementioned embodiments, the machine learning combination model here is a migrated model obtained by performing model migration relative to the original model. The original model here can, for example, be the machine learning combination model of the aforementioned embodiment, specifically established by the following method: obtaining a dataset of component concentrations of a bioreactor cell culture fluid, the dataset comprising a training dataset, a validation dataset, and a test dataset; using multiple machine learning algorithms to establish multiple single prediction models, wherein each of the multiple single prediction models predicts the test dataset after being trained with the training dataset and verified with the validation dataset; comparing the prediction results of the multiple single prediction models with the test dataset to obtain the sum of squared prediction errors of the multiple single prediction models, and determining the weights of the multiple single prediction models in the original model based on the size of the sum of squared prediction errors; combining the multiple single prediction models to obtain the original model by a weight assignment method. In addition, model migration includes: obtaining the original data set used to establish the original model, and a new batch training data set of the concentrations of cell culture fluid components of the new batch biological reaction, the original data set including the original training data set, the original verification data set, and the original test data set; performing scale correction or scale matching on the new batch training data set and the original training data set to thereby obtain a new training data set; using the new training data set, the original verification data set, and the original test data set, respectively, using a variety of machine learning algorithms to establish multiple single prediction models, wherein each of the multiple single prediction models predicts the original test data set after being trained with the new training data set and verified with the original verification data set; comparing the prediction results of the multiple single prediction models with the original test data set to obtain the sum of squares of the prediction errors of the multiple single prediction models, and determining the weights of the multiple single prediction models in the migrated model according to the size of the sum of squares of the prediction errors; and combining the multiple single prediction models to obtain the migrated model through a weight assignment method.

[0305] Specifically, the scale correction may include incorporating a specified proportion of data from the new batch training data set into the original training data set used to establish the original model, and the scale matching may include incorporating new batch training data from the new batch training data set whose numerical difference with the original training data in the original training data set collected at the same time is less than a specified threshold into the original training data set. In some examples, the specified proportion is a value selected from 1% to 10%, or a value selected from 1.5% to 7.5%, or a value selected from 2% to 5%. In some examples, the specified threshold is a value selected from 1% to 10%, or a value selected from 3% to 8%, or a value selected from 4% to 6%, or 5%.

[0306] For example, the plurality of machine learning algorithms may include at least two selected from the following: partial least squares factorial, cubic tree, random forest, support vector machine, time series. The dataset of the component concentration may include online Raman spectroscopy data and its corresponding offline detection data, and the sampling time of the offline detection data matches the corresponding online Raman spectroscopy data. As a non-limiting example, the component concentration may be selected from the group consisting of viable cell density, glucose concentration, lactate concentration, target product concentration, and amino acid concentration. In some examples, the Raman spectroscopy data may also be subjected to data preprocessing, and the data preprocessing may include at least one of the following: screening out abnormal data points, removing spike peaks, Raman shift correction and light intensity correction, baseline calibration, smoothing, and derivation.

[0307] For example, Figure 26 schematically depicts the training process of the machine learning combination model as the original model based on Raman spectral data and offline detection data, as well as the application process of predicting the concentrations of various components of cell culture fluid based on the input online Raman spectral data. Figure 26 also schematically depicts the training process of the machine learning combination model as the migrated model based on scale correction or scale matching, as well as the application process of predicting the concentrations of various components of cell culture fluid based on the input online Raman spectral data. The process shown in Figure 26 can be divided into three major steps: data preparation, model establishment, and verification and adjustment. At the same time, taking into account model migration, the present disclosure also has the steps of correcting data and matching data. The following are exemplary explanations of each.

[0308] (1) Data preparation

[0309] In the embodiments of the present disclosure, although other methods that meet the requirements of the present disclosure and can detect the concentrations of various components of the cell culture fluid can also be used, the Raman spectral signals in the cell culture fluid measured by a Raman spectrometer are mainly used as detection data (training set) and verification data (validation set).

[0310] After obtaining Raman spectroscopy data, the embodiment of the present disclosure first pre-processes the data to better apply it to the machine learning model. Referring to Figure 27, in the embodiment of the present disclosure, the data is further processed through the process shown in Figure 27.

[0311] After entering the data preprocessing process, perform the following steps to preprocess the data:

[0312] (a) Screening out outliers: Simple statistical methods are used to preliminarily screen out outliers. Specifically, the mean, median, standard deviation, etc. can be calculated to remove outliers that significantly deviate from the main data.

[0313] (b) Spike removal: Cosmic spikes in Raman spectra originate from electrons generated by high-energy cosmic particles on CCD or complementary metal oxide semiconductor detectors. They appear randomly in the Raman spectrum and appear as very narrow but extremely intense spectral features. Due to their high intensity, spike addition makes data analysis difficult. If interfering spikes are present, the results of normalization and feature extraction are meaningless. After the spike is detected, it can be removed by linear interpolation based on the two boundary points of the spike. Alternatively, the spike can be replaced by its continuous measurement at the same wavenumber position of the spike. In this case, the fluorescence difference and intensity change between the two measurements must be taken into account.

[0314] (c) Raman shift and intensity correction: Raman spectroscopy should produce identical results regardless of the measurement environment, equipment, or other conditions. However, this is often not the case. Instead, spectral variations are observed between instruments over time and under varying measurement conditions. Well-designed normalization methods are needed to eliminate these unwanted spectral variations and normalize all measured Raman spectra to the same reference material. One of the most fundamental methods for this standardization in Raman spectroscopy is spectrometer calibration, which consists of wavenumber and intensity corrections. Using a stable, identical spectrometer for shift and intensity corrections, followed by validation using the same optical standards, is the ideal approach to standardizing Raman spectra. Based on this calibration, the wavenumber axis is calibrated by fitting a (polynomial) function between the measured and theoretical positions of the well-defined Raman bands of the wavenumber standards. The intensity axis is calibrated by dividing the measured Raman intensity by the instrument's intensity response function, which is derived as the ratio between the measured and theoretical emission of the intensity standard within the wavenumber range of interest.

[0315] (d) Baseline Calibration: Baseline calibration can be applied in two ways: removing the substrate spectrum or removing the fluorescence baseline. The former is used to remove the substrate Raman signal from the measured Raman spectrum; the latter aims to remove the sample fluorescence, which appears as a slowly varying baseline in the Raman spectrum. If the substrate has a large number of Raman bands, especially if these bands overlap with those of the sample, the substrate contribution needs to be removed. To this end, a reference spectrum of the substrate is often required to estimate the substrate contribution in the recorded Raman spectrum. For heterogeneous substrates, statistical methods can be useful; for example, multivariate curve resolution can be used to address this heterogeneous substrate contribution. Fluorescence baseline removal is generally more complex than substrate calibration because the fluorescence baseline depends on the sample and setup. This fluorescence baseline is often removed using mathematical methods such as calculated derivative spectra, sensitive nonlinear iterative peak clipping algorithms, asymmetric least squares (ALS) smoothing, modified polynomial fitting, standard normal variates, multiplicative scatter correction, and extended multiplicative signal correction (EMSC). These methods are flexible, easy to use, require no instrument modification, and perform adequately in most cases. However, if the fluorescence intensity is too strong to be mathematically calibrated, methods based on instrument modifications may be required. This class of techniques includes time-series Raman spectroscopy, modulation Raman spectroscopy, and shifted excitation difference Raman spectroscopy.

[0316] (e) Smoothing: In Raman spectral analysis, smoothing or filtering can be performed using spectral and / or spatial filtering. Spectral filtering removes noise by applying a low-pass filter along the wavenumber axis. The filter can be a mean, median, Gaussian, or polynomial function. Spatial filtering has a similar concept to spectral filtering, but it applies the low-pass filter to the spatial domain. Both methods have their advantages and disadvantages. Spectral filtering reduces spectral resolution but preserves spatial resolution, and vice versa.

[0317] (f) Derivative: Derivative is optional and its main purpose is to further improve the signal-to-noise ratio.

[0318] (g) Normalization: Normalization aims to eliminate the effects of excitation intensity fluctuations or focus changes, which can be performed by conventional normalization methods.

[0319] As shown in Figure 27, after the steps (b) spike removal, (c) Raman shift correction and intensity correction, (d) baseline calibration, (e) smoothing, and optionally (f) derivation, checkpoints are set to verify the data preprocessing results of the previous steps to determine whether the previous steps were over-processed. If over-processing is found, the previous step is returned to, adjusted, and then repeated.

[0320] Specifically, at checkpoint 1, it is determined whether the spike peaks have been excessively removed. At checkpoint 2, it is determined whether Raman shifts and intensity drifts are still present. At checkpoint 3, it is determined whether the characteristic spectra of the fluorescence and substrate are still present. At checkpoint 4, it is determined whether there is any loss of shift and intensity after smoothing. At checkpoint 5, it is determined whether the original spectrum is distorted after the derivative is taken. Based on the judgment results of each checkpoint, not every preprocessing operation requires a derivative operation. Therefore, if a derivative operation is not performed, no judgment is made at checkpoint 5.

[0321] (2) Model establishment

[0322] Typically, after obtaining data from Raman spectroscopy, technicians attempt to convert the Raman signals into digital information and identifiable corresponding data, and further identify the substances based on the similarity between the measured spectra and spectra of known substances in a spectral database. For example, in the present disclosure, this method identifies the concentrations / contents of various components in a cell culture fluid. However, this is not feasible in the case of Raman spectroscopy detection of cell culture processes in a bioreactor, as in the present disclosure. Typically, technicians can identify the detected substances and their concentrations / contents by comparing the detected Raman spectral data with data from a database of known substances. However, the cell culture process is a complex one, potentially producing a vast variety of substances, and the detectable sample components present in the cell culture fluid are extremely complex. In particular, when the measured spectrum contains signals from the substrate, the model prediction will be biased towards the substrate. That is, the detection results largely reflect the components of the substrate and / or culture medium, which account for the majority of the mass of the cell culture fluid, rather than the concentrations / contents of certain characteristic components that influence the cell culture process.

[0323] Therefore, in an embodiment of the present disclosure, a method based on machine learning is used to extract characteristic spectra according to an algorithm and assign offline detection values ​​to the characteristic spectra to obtain analytical measurement spectra.

[0324] In the embodiments of the present disclosure, for preprocessing data, for example, five algorithms, including partial least squares (PLS), cubic tree (Cubist), random forest (RF), support vector machine (SVM), and time series, can be used to establish single prediction models respectively, and then the prediction results are used to use the inverse variance (RV) to determine the weight of each algorithm in the combined model, calculate to obtain new prediction results, and establish a machine learning combined model.

[0325] The following introduces the exemplary algorithms used in the present disclosure.

[0326] Partial Least Squares (PLS):

[0327] Partial least squares regression ≈ multiple linear regression analysis + canonical correlation analysis + principal component analysis

[0328] Step 1: Center the original data X and Y to obtain X0 and Y0, and select a column from Y0 as u1, generally the column with the largest variance. The sample covariance formula of the standardized data is:

[0329] Step 2: Iterate and solve the transformation weights (w1, c1) and factors (u1, t1) of X and Y until convergence. Using the information u1 of Y, calculate the transformation weight w1 of X (w1 transforms from X0 to the factor t1, t1 = X0 * w1) and the factor t1, thereby approximating the information of X0 using t1. ||w1||→1 t1=X0w1

[0330] Using the information t1 of X, calculate the transformation weight c1 of Y (c1 realizes the transformation from Y0 to factor u1, u1=Y0*c1) and update the factor u1, so that the information of Y0 is approximately expressed by t1. |||c1||→1

[0331] Determine whether a reasonable solution has been found. If △u < threshold (such as 10 -6 ) then continue with the following steps; otherwise take u1=u1* and return to step 2.

[0332] Step 3: Find the residual matrix of X and Y, and find the load p1 of X (p1 reflects the direct relationship between X0 and factor t1).

[0333] Step 4: Using X1 and Y1, repeat the above steps to solve the next batch of PLS ​​parameters.

[0334] Cubist:

[0335] The Cubic Tree model selects an ensemble learning algorithm based on model trees. When the first tree model is built following the M5 model tree rules, the next model tree is an adaptation of the training set results. If the model overestimates the target value, the next model response is to adjust it downward, and so on. The final estimate is the average of the values ​​calculated by the models of each tree.

[0336] The nodes of the model tree are not constants, but linear function models. The criterion for segmenting the space is not to reduce the square error, but to reduce the sample standard deviation.

[0337] M5 model tree:

[0338] The standard deviation of the Y value (i.e., the target attribute value) of the samples covered by a node is regarded as a measure of error.

[0339] T is the set of instances that reach the node, |T| represents the size of the set, sd represents the standard deviation, and Ti is the set of instances on the i-th subtree.

[0340] The node pruning of the tree model is a bottom-up recursive process. The linear regression method is used to fit the regression equation of each node and calculate the root mean square error of the regression function prediction.

[0341] Calculate the reduction E of the MSE from each node to its child nodes R =|N|R MSE -∑i|N i |R MSEi

[0342] Random Forest (RF):

[0343] The random forest model is a comprehensive learning model that uses the Bagging algorithm to build multiple decision tree models and uses the average value to calculate the results of all decision tree models.

[0344] Enhancements in the Bagging Algorithm

[0345] Support Vector Machine (SVM):

[0346] Support vector machines are a method for classifying linear and nonlinear data.

[0347] Linearly separable support vector machine:

[0348] W is the weight vector, b is the bias scalar, and the training target formula is when the parameters of the training target data reach the minimum value:

[0349] If the constraints of the above formula are expressed as

[0350] The maximum interval separating hyperplane problem is expressed as

[0351] Non-linear Separable Support Vector Machine:

[0352] Introduce the kernel function method.

[0353] The polynomial kernel function is

[0354] The Gaussian kernel function is

[0355] Time Series Index:

[0356] Exponential smoothing is used to calculate the exponential smoothing value and combine it with the time series forecasting model for prediction and judgment. The principle is that the exponential smoothing value of any value is the weighted average of the actual value and the previous value.

[0357] When the time series has obvious trend changes, the exponential smoothing method is used for prediction.

[0358] Secondary exponential smoothing is a smoothing of primary exponential smoothing and is suitable for time series with linear trends.

[0359] The prediction formula for the future T cycle is y t+T =A t +B t T

[0360] Reciprocal variance method:

[0361] Combine the results of the above model algorithms and use the inverse variance method to determine the weight of each algorithm in the combined model based on the size of the sum of squares of the prediction errors.

[0362] Where Qi is the sum of squares of the differences between the true value and the predicted value

[0363] Five prediction models were established using the above modeling process, and then the five models were combined into a new machine learning combination model using the weight assignment method.

[0364] The present disclosure also establishes a data learning workflow, as shown in FIG28 , which shows the process of establishing a model and comparing and screening the prediction effects of each model.

[0365] In the data learning workflow, there are at least the following steps:

[0366] (I) Sample classification: In the first stage of data learning, statistical sampling is performed to prepare data for statistical modeling. An effective statistical model with a limited sample size is achieved. The accessible dataset is divided into three subsets: training, validation, and test datasets. The three subsets are used for model training, model optimization, and model evaluation, respectively. In many cases, data splitting is repeated multiple times, such as in cross-validation (CV) or bootstrap methods. Each repetition generates a different separation of the three subsets, and statistical modeling is performed multiple times. This provides an opportunity to verify model stability and reproducibility, for example, based on the mean and standard deviation of accuracy or the root mean square error (RMSE) calculated from multiple predictions on different test datasets.

[0367] (II) Single Prediction Model Construction: Statistical modeling and machine learning begin with dimensionality reduction. This is particularly important for Raman spectroscopy, where datasets consist of a large number of correlated features and the sample size is limited. The benefits of dimensionality reduction are twofold: first, it makes visualization simpler and clearer, thus helping to better outline the characteristics of the dataset; second, it can improve and accelerate subsequent modeling by removing redundant information and extracting useful features from the data. In the second stage of data learning, the dimensionality reduction output is input into a subsequent model, which may be a clustering, classification, or regression model. Models can be linear or nonlinear, parametric or nonparametric, and supervised or unsupervised. While the choice of model to use is data-dependent, it should be remembered that model generalization is likely to decrease with increasing model complexity. Models should be as parsimonious as possible without sacrificing performance. This means that linear and parametric models are preferred over nonlinear and nonparametric models in terms of generalization. Another important part of model construction is variable importance or significance. These coefficients are calculated based on the trained model and indicate the significance of each variable to the model and task. Variables with high coefficients for variable importance are considered more important to the model, and this interpretation should be considered in conjunction with the spectral model of the data at hand. These values ​​can be further used for feature selection, resulting in a more parsimonious model. However, it should be recognized that variables with excessively high coefficients or excessive noise in the model should be removed from the model, as they are likely unreliable.

[0368] (III) Single prediction model evaluation: The model usually predicts unknown samples worse than the training / validation data. This is called model shrinkage. In extreme cases, the statistical model can perfectly predict the training / validation data, but due to overfitting, it cannot predict unknown samples at all, that is, the model fits the training data too perfectly and loses its generalizability. Therefore, it is very important to check the prediction of unknown samples and control the error rate to ensure that the statistical model can be used in practice, that is, model evaluation. Here, the model built in the previous step is used to predict the test data generated by statistical sampling. If the prediction error given a predefined threshold is too large, the statistical model should be re-modeled with modifications. Regression models calculate the deviation between predicted values ​​and actual values ​​for model evaluation; classification and clustering models use the confusion matrix of predicted values ​​and actual values ​​as the judgment basis. The confusion matrix can calculate a variety of features, including accuracy, sensitivity, specificity, etc.

[0369] (IV) Combining Single Forecast Models: The predicted values ​​of the evaluated single forecast models are used as new input data in the inverse variance formula. The weight coefficients of each single forecast model within the new combined model are then assigned through calculations to establish a combined model. Based on the principles of single forecast modeling, to prevent a single forecast model from having an excessively large weight that affects the final forecast, the weight allocation algorithm is still primarily based on a simple one.

[0370] (V) Combined model evaluation: The same method as (III) was used to evaluate the combined model.

[0371] (VI) Model storage: After establishing a qualified combination model, the model is stored and the data preprocessing is stored together.

[0372] As shown in Figure 28, after (II) establishing a single prediction model, (III) evaluating a single prediction model, and (V) evaluating a combined model, checkpoints are set up to check whether the previous steps are appropriate. Specifically, at checkpoint 1, the model's suitability is determined based on the verification results. If it is not appropriate, a model using another modeling method needs to be added or replaced. At checkpoint 2, the model's applicability is checked. If it is not appropriate, a model using another modeling method needs to be added or replaced and re-modeled. At checkpoint 3, the combined model's prediction results are checked for suitability. If not, the weights assigned to each single model need to be recalculated.

[0373] (3) Model migration

[0374] As previously mentioned, the combined model ultimately obtained by this disclosure is intended to be highly versatile, meaning that the same combined model can accurately predict different batches, substrates, processes, scales, and even when there are differences in spectral variations. This characteristic is sometimes referred to as model portability in this disclosure.

[0375] If a model has good transferability, then, provided all the data learning procedures are executed correctly, a well-tuned model will also be able to predict new data well in the future. However, for various reasons, this is not always the case in reality. A model may fail to predict new data, or further adjustments to the model parameters may be needed.

[0376] This phenomenon is particularly severe when using Raman spectroscopy data. This is because Raman spectroscopy is extremely sensitive, and even slight variations in the instrument, measurement conditions, or sample preparation can manifest as substantial shifts in Raman shifts or changes in Raman intensity. These spectral variations are unavoidable in practice, resulting in poor performance of existing models when predicting new data in novel biological reactions. Consequently, model transferability may be poor between different bioreactors, or even between different cell culture batches.

[0377] Of course, if possible, it would be ideal to build a new model from scratch for each batch of cell culture. However, this obviously requires re-obtaining a large amount of training data and re-performing the combined model building process, which is uneconomical in all aspects.

[0378] In this context, the present disclosure establishes two model migration methods. FIG29 illustrates the two methods of model migration, wherein the upper half of FIG29 illustrates the scale matching method, and the lower half of FIG29 illustrates the scale correction method, and wherein a checkpoint is used to check whether erroneous spectra are introduced during the data matching and normalization process. Those skilled in the art will appreciate that although the two methods of model migration are referred to herein as scale correction and scale matching, the need for model migration is not limited to changes in the scale of the biological reaction. Model migration when substrates, processes, etc., change can also be performed using scale correction or scale correction methods.

[0379] In the present disclosure, "model migration" refers to adjusting the combination model established according to the method for establishing a machine learning combination model disclosed in the present disclosure to make it applicable to a new batch of biological reactions. The "new batch of biological reactions" refers to different batches of biological reactions, including different operating batches of the same biological reaction, and also includes different biological reactions (such as biological reactions with different substrates, processes and / or scales). The data of the new batch of biological reactions are referred to as "new batch data sets". The new batch data set is different from the original training data set used to establish the machine learning combination model. The training data set used when establishing the machine learning combination model is combined with the new batch data set, and the machine learning combination model trained thereby is referred to as the "migrated model".

[0380] Specifically, there are two methods for model migration.

[0381] Model transfer method 1: Scale correction. A specified proportion of a new batch training dataset (including spectral data and corresponding offline detection data) is added to the existing model training set, and the model is retrained to overcome differences in batches, substrates, processes, scales, and spectral variations, achieving accurate predictions. The specified proportion can be selected from a value between 1% and 10%, a value between 1.5% and 7.5%, or a value between 2% and 5%.

[0382] Model migration method 2: scale matching. According to the described experimental method, a new batch of a new round of experimental data (new batch, substrate, process, scale) is collected, and the new batch of training data whose numerical difference with the original training data in the original training data set with the same collection time is less than the specified threshold value is incorporated into the original training data set. Specifically, the spectral data in the new batch training data set is compared with the spectral data in the original training set with the same collection time, thereby obtaining the spectral data in the new batch training data set whose spectral numerical difference is less than the specified threshold value. The spectral data in the new batch training data set and the corresponding offline detection data are added to the original training set for retraining the model, and the new model is applied to a brand new experimental environment. The specified threshold value is a value selected from 1% to 10%, or a value selected from 3% to 8%, or a value selected from 4% to 6%, or 5%.

[0383] In model migration, the retraining model refers to adding the new batch training set to the original training set to form a new training set, so that the data set of the component concentration of the bioreactor cell culture fluid includes the new training set, the original verification data set, and the original test set; using multiple machine learning algorithms to establish multiple single prediction models, wherein after the prediction model is established with the new training data set, the original verification data set is predicted; the prediction results are compared with the original test data set to obtain the sum of squares of prediction errors of multiple single prediction models, and the weights of multiple single prediction models in the combined model are determined by the size of the sum of squares of prediction errors; through the weight assignment method, multiple single prediction models are combined to obtain a machine learning combined model.

[0384] The retraining model can be automatically performed using a pre-established process or procedure, such as automatic calibration, one-time calibration, or timed calibration.

[0385] The following is a further detailed description of the method of using machine learning combination models through specific examples.

[0386] Example 5

[0387] First, offline detection data and Raman spectral data are obtained through the cell culture process of the wet experiment.

[0388] The cell culture experiments used a monoclonal antibody-producing CHO-K1 cell line. Cytiva's Hyclone Actipro medium was used in a Kuhner shaker at 36.5°C, 110 rpm, and 6% CO2 during the seeding phase. The production phase involved 3L and 200L reactors with initial culture volumes of 1.5L and 140L, respectively. The culture temperature was 36.5°C, the pH was set at 6.90 ± 0.25, the dissolved oxygen saturation was 40%, and the initial inoculum density was 1.0 × 10 6 cells / mL, and the feed medium was also Cytiva's Hyclone Cell boost 7a / 7b (10% / 1%).

[0389] Obtaining offline data: During the production culture phase, the Raman instrument probe was placed in the culture medium. Samples were taken five times daily and viable cell density was measured using a Beckman Vi-Cell XR. Glucose, lactate, and target protein concentrations were measured using a Roche Cedex Bio analyzer. Amino acid concentrations were measured using an Agilent HPLC.

[0390] Raman spectroscopy data were obtained using a Raman Rxn2 analyzer (Kaiser Optical Systems) equipped with an immersion optical probe. The probe was installed in a 3 L bioreactor (Applikon) and immersed directly in the cell culture suspension. Raman spectra were recorded for different bioreactors throughout the experiment. For a single recorded spectrum, 30 subsequent spectra were captured with a 10 s exposure time and averaged, resulting in an acquisition interval of approximately 5 min for each bioreactor. The laser excitation wavelength was 785 nm, providing a range of 100 to 3425 cm -1 After reading the spectral data, the spectral file is converted into spectral values. Each sample contains a spectral file and the offline target value is mapped one by one according to the time change. The modeled characteristic data (spectral values ​​corresponding to different Raman shifts) and target values ​​are obtained, and then data preprocessing begins.

[0391] After obtaining the data, the following method is used for data preprocessing.

[0392] In this embodiment, a simple statistical method is used to preliminarily screen out abnormal data points.

[0393] Perform SG filtering and smoothing, use polynomials to smooth the data, and remove "spectral burrs" and eliminate random noise based on the least squares method.

[0394] SG smoothing is used to remove spike peaks, where the width of the smoothing window is n = 2m + 1

[0395] Fit n = 2m + 1 isocentric data in a window to a scaled K-order polynomial for SG smoothing.

[0396] Then, Fourier transform is used to perform Raman shift correction and light intensity correction.

[0397] Baseline calibration can be performed by using second-order derivative, polynomial difference, or first-order derivative. In this embodiment, second-order derivative is used to perform baseline calibration.

[0398] Afterwards, the SG smoothing first-order convolution formula is used for smoothing.

[0399] In this embodiment, spectral derivation is used as the derivation to eliminate baseline drift, smooth background noise, and improve resolution.

[0400] Finally, normalization is performed using the standard normal distribution.

[0401] As mentioned above, in the present disclosure, for the preprocessed data, five algorithms, namely partial least squares (PLS), cubic tree (Cubist), random forest (RF), support vector machine (SVM), and time series, are used to establish a single prediction model respectively. The prediction results are then used to determine the gap between the prediction results and the validation set using the inverse variance (RV), thereby determining the weight of each algorithm in the combined model, calculating to obtain a new prediction result, and establishing a machine learning combined model.

[0402] The data used to build a single prediction model is identical. The training dataset is derived from three fed-batch bioreactors cultured under identical conditions, and the prediction dataset is derived from a fourth bioreactor cultured under identical conditions. The training dataset contains 160 spectral data points and their corresponding offline detection data, the validation dataset contains 50 spectral data points and their corresponding offline detection data, and the test dataset contains 70 spectral data points and their corresponding offline detection data.

[0403] After performing predictions using a combined model of the five algorithms described in this disclosure—partial least squares (PLS), cubic tree (Cubist), random forest (RF), support vector machine (SVM), and time series—the predicted results were compared with the validation set. Specific prediction results and comparisons are shown in Figures 30-35 and Table 3. The RMSEP in Tables 3 and 4 refers to the root mean square error of the predictions.

[0404] Table 3 Amino acid RMSEP data

[0405] In order to obtain quantitative comparison results, Table 4 shows the results of comparing the PLS model and the machine learning combination model of the present disclosure using the leave-one-out interaction validation method.

[0406] Table 4

[0407] The following further verifies the model transferability of the disclosed method. Figures 36 and 37 show the prediction results of two indicators (viable cell density and glucose concentration) obtained after the model is transferred between different scales.

[0408] Table 5 compares the two migration methods and finds that the scale-corrected model migration method is more suitable for practical applications.

[0409] Table 5

[0410] The present disclosure also provides a kind of computer equipment, including memory and processor, on the memory is stored a computer program that can run on the processor, and the computer program realizes the step of the automated feeding method based on real-time dynamic detection result according to any embodiment of the present disclosure when being executed by the processor.Specifically, with reference to Figure 38, it shows a schematic block diagram of a computer equipment 500 according to some embodiments of the present disclosure.As shown in Figure 38, computer equipment 500 includes a processor 502 and a memory 504.Processor 502 can be, for example, a central processing unit (CPU) of computer equipment 500.Processor 502 can be any type of general-purpose processor, or can be a processor specially designed for automated feeding, such as an application-specific integrated circuit ("ASIC").Memory 504 can be coupled to processor 502 and can include various computer-readable media that can be accessed by processor 502.In various embodiments, memory 504 described herein can include volatile and non-volatile media, removable and non-removable media. For example, the memory 504 may include any combination of the following: random access memory ("RAM"), dynamic RAM ("DRAM"), static RAM ("SRAM"), read-only memory ("ROM"), flash memory, cache memory, and / or any other type of non-transitory computer-readable medium. The memory 504 may store a computer program that is executable on the processor 502. When executed by the processor 502, the computer program implements the automated feeding method based on real-time dynamic detection results (e.g., methods 100 and 100') according to any embodiment of the present disclosure.

[0411] The present disclosure also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the automated feeding method based on real-time dynamic detection results according to any embodiment of the present disclosure are implemented.

[0412] The present disclosure also provides a computer program product, which may include instructions that, when executed by a processor, may implement the automated feeding method based on real-time dynamic detection results according to any embodiment of the present disclosure. The instructions may be any instruction set to be executed directly by the processor, such as machine code, or any instruction set to be executed indirectly, such as a script. The instructions may be stored in an object code format for direct processing by the processor, or in any other computer language, including a script or collection of independent source code modules that are interpreted on demand or compiled in advance.

[0413] Figure 39 shows a schematic block diagram of a computer system 700 on which an embodiment of the present disclosure can be implemented. The computer system 700 includes a bus 702 or other communication mechanism for transmitting information, and a processing device 704 coupled to the bus 702 for processing information. The computer system 700 also includes a memory 706 coupled to the bus 702 for storing instructions to be executed by the processing device 704. The memory 706 can be a random access memory (RAM) or other dynamic storage device. The memory 706 can also be used to store temporary variables or other intermediate information during the execution of the instructions to be executed by the processing device 704. The computer system 700 also includes a read-only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions for the processing device 704. A storage device 710 such as a magnetic disk or optical disk is provided and coupled to the bus 702 for storing information and instructions. The computer system 700 can be coupled via bus 702 to an output device 712 for providing output to a user, such as, but not limited to, a display (such as a cathode ray tube (CRT) or liquid crystal display (LCD)), speakers, etc. Input devices 714, such as a keyboard, mouse, microphone, etc., are coupled to bus 702 for communicating information and command selections to the processing device 704. The computer system 700 can perform embodiments of the present disclosure. Consistent with certain implementations of the present disclosure, the computer system 700 provides results in response to the processing device 704 executing one or more sequences of one or more instructions contained in the memory 706. Such instructions can be read into the memory 706 from another computer-readable medium, such as storage device 710. Execution of the sequence of instructions contained in the memory 706 causes the processing device 704 to perform the methods described herein. Alternatively, hard-wired circuitry can be used in place of or in combination with software instructions to implement the present teachings. Therefore, implementations of the present disclosure are not limited to any specific combination of hardware circuitry and software. In various embodiments, computer system 700 can be connected to one or more other computer systems like computer system 700 across a network via network interface 716 to form a networked system. The network can include a private network or a public network such as the Internet. In a networked system, one or more computer systems can store data and supply data to other computer systems. As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to processing device 704 for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks such as storage device 710. Volatile media include dynamic memory such as memory 706. Transmission media include coaxial cables, copper wire, and optical fibers, including the wiring comprising bus 702.Common forms of computer-readable media or computer program products include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, digital video disks (DVDs), Blu-ray discs, any other optical media, thumb drives, memory cards, RAM, PROMs and EPROMs, flash EPROMs, any other memory chips or cartridges, or any other tangible media from which a computer can read. Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processing device 704 for execution. For example, the instructions may initially be carried on a diskette of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 700 may receive the data on the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to bus 702 may receive the data carried in the infrared signal and place the data on bus 702. Bus 702 carries the data to memory 706, from which processing device 704 retrieves the instructions and executes them. Optionally, the instructions received by memory 706 may be stored on storage device 710 either before or after execution by processing device 704 .

[0414] According to various embodiments, instructions configured to be executed by a processing device to perform a method are stored on a computer-readable medium. A computer-readable medium can be a device that stores digital information. For example, a computer-readable medium includes a compact disk read-only memory (CD-ROM) as known in the art for storing software. The computer-readable medium is accessed by a processor adapted to execute the instructions configured to be executed.

[0415] The foregoing description describes one or more exemplary embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the exemplary embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0416] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a server system. Of course, the present disclosure does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0417] Although one or more embodiments of the present disclosure provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When an actual device or terminal product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment).

[0418] The terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes the elements is not precluded. For example, if words such as "first," "second," etc. are used to indicate names, they do not imply any particular order.

[0419] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0420] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.

[0421] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture comprising instruction means that implement the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.

[0422] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0423] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0424] The same or similar parts between the various embodiments of the present disclosure can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of the present disclosure, the reference terms "one embodiment", "some embodiments", "embodiments", "examples", "specific examples", or "some examples" are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present disclosure, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the present disclosure and the features of different embodiments or examples without contradiction.

[0425] In addition, when used in this disclosure, the words "herein," "above," "below," "hereunder," "supra," and words of similar meaning shall refer to the disclosure as a whole and not to any particular portions of the disclosure. Furthermore, unless expressly stated otherwise or otherwise understood in the context of use, conditional language used herein, such as "may," "might," "for example," "such as," and the like, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or states. Thus, such conditional language is generally not intended to imply that one or more embodiments in any way require features, elements, and / or states, or whether such features, elements, and / or states are included or performed in any particular embodiment.

[0426] The foregoing is merely an example of one or more embodiments of the present disclosure and is not intended to limit the one or more embodiments of the present disclosure. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims.

Claims

1. An automated feeding method based on real-time dynamic detection results, comprising: Collecting real-time dynamic detection results, calculating the change rate of each component per unit time, where the change rate is the consumption rate or the generation rate; Selecting the component with the fastest change rate as the marker for the current time period; Collecting historical detection data, and calculating the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; Taking the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration setting value for the current time period in the logarithmic growth phase; taking the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum as the component concentration setting value for the current time period in the culture harvest phase; Calculating the addition dose of the marker in the current time period according to the change rate of the marker, and calculating the addition dose of the culture medium according to the addition dose of the marker, so as to obtain a feeding strategy.

2. The automated feeding method based on real-time dynamic detection results according to claim 1, wherein, The step of selecting the marker includes: Collecting the current real-time dynamic detection results, and calculating the change rate of each component per unit time currently; Collecting the historical real-time dynamic detection results, and calculating the change rate of each component per unit time historically; Combining the change rate of each component per unit time currently with the change rate of each component per unit time historically, and screening out the component with the fastest change rate as the marker for the current time period.

3. The automated feeding method based on real-time dynamic detection results according to claim 1 or 2, wherein, The historical detection data is the historical detection data of the fermentation process with the same cell culture process mode as this time.

4. The automated feeding method based on real-time dynamic detection results according to any one of claims 1-3, wherein, The step of calculating the addition dose of the culture medium includes: m 补 = (G 设 - G 实 ) × V 培 / G 补 Among them, m 补 represents the amount of the fed-batch medium to be added, G 设 represents the set marker concentration, G 实 represents the current marker concentration detected by the model for real-time Raman spectroscopy analysis, V 培 represents the volume of the cell culture medium in the bioreactor, G 补 represents the marker concentration in the fed-batch medium.

5. The automated feeding method based on real-time dynamic detection results according to any one of claims 1-4, wherein, The feeding culture medium includes a specific proportion of the marker and other components, and the addition amount of other components is obtained according to the addition dose of the marker.

6. The automated feeding method based on real-time dynamic detection results according to any one of claims 1-5, wherein, After obtaining the feeding strategy, it further includes writing the feeding strategy into the control system; Optionally, the control system is configured to, after receiving numerical instructions of the feeding amount, the setting value, and the feeding time point, associate with the peristaltic pump hardware and perform instantaneous feeding in a dosage manner.

7. An automated feeding device based on real-time dynamic detection results, comprising: A marker screening unit, configured to: collect real-time dynamic detection results, calculate the change rate of each component per unit time, where the change rate is the consumption rate or the generation rate; Selecting the component with the fastest change rate as the marker for the current time period; A marker concentration setting unit, configured to: collect historical detection data, and calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase; Taking the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration setting value for the current time period in the logarithmic growth phase; taking the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum as the component concentration setting value for the current time period in the culture harvest phase; A feeding amount calculation unit, configured to: calculate the addition dose of the marker in the current time period according to the change rate of the marker, and calculate the addition dose of the culture medium according to the addition dose of the marker, so as to obtain a feeding strategy.

8. The automated feeding device based on real-time dynamic detection results according to claim 7, wherein, In the marker screening unit, the step of selecting the marker includes: Collecting the current real-time dynamic detection results, and calculating the change rate of each component per unit time currently; Collect the real-time dynamic detection results of history, and calculate the change rates of each component within the historical unit time; Merge the change rates of each component within the current unit time with the change rates of each component within the historical unit time, and screen out the component with the fastest change rate as the marker for the current time period; Optionally, in the marker concentration setting unit, the historical detection data is the historical detection data of the fermentation process with the same cell culture process mode as this time; Optionally, in the feeding amount calculation unit, the step of calculating the addition dose of the culture medium includes: m 补 = (G 设 - G 实 ) × V 培 / G 补 Among them, m 补 represents the amount of the feeding medium to be added, G 设 represents the set marker concentration, G 实 represents the current marker concentration detected by the model for real-time analysis of Raman spectroscopy, V 培 represents the volume of the cell culture medium in the bioreactor, G 补 represents the marker concentration in the feeding medium; Optionally, in the feeding amount calculation unit, the feeding culture medium includes a specific proportion of the marker and other components, and the addition amounts of other components are obtained based on the addition dose of the marker; Optionally, in the feeding amount calculation unit, after obtaining the feeding strategy, it further includes writing the feeding strategy into the control system; Optionally, the control system is configured to receive the numerical instructions of the feeding amount, set value, and feeding time point, and then associate with the peristaltic pump hardware to perform instantaneous feeding in a dosage manner.

9. An automated feeding method based on real-time dynamic detection results, including: Obtain the real-time dynamic detection results of history, calculate the historical change rates of each component based on the real-time dynamic detection results of history, and select a component or a combination of multiple components as the marker based on the historical change rates of each component; Collect the real-time dynamic detection results of the current time, and calculate the real-time detection concentration of the marker in the current time period based on the real-time dynamic detection results of the current time; Obtain the feeding strategy through at least one of the following: Based on the real-time dynamic detection results of history, calculate the cell growth rate in the logarithmic growth phase and the target product expression rate in the culture harvest phase, and use the concentration level of the marker when the cell growth rate in the logarithmic growth phase is the maximum as the component concentration set value of the marker in the current time period when the current time period falls into the logarithmic growth phase, and use the concentration level of the marker when the target product expression rate in the culture harvest phase is the maximum as the component concentration set value of the marker in the current time period when the current time period falls into the culture harvest phase, and calculate the addition dose of the feeding culture medium in the current time period based on the difference between the real-time detection concentration of the marker in the current time period and the component concentration set value in the current time period, so as to obtain the first feeding strategy; Or Calculate the current change rate of the marker based on the real-time detection concentration of the marker in the current time period, and calculate the addition dose of the feeding culture medium in the current time period based on the current change rate of the marker, so as to obtain the second feeding strategy.

10. The automated feeding method based on real-time dynamic detection results according to claim 9, further including: Calculate the multiple historical change rates of each component in multiple historical time periods based on the real-time dynamic detection results of history; Merge the multiple historical change rates of each component to obtain the combined historical change rate of each component, and select a component or a combination of multiple components as the marker based on the combined historical change rate of each component.

11. The automated feeding method based on real-time dynamic detection results according to claim 9 or 10, wherein, The real-time dynamic detection results of the history and the current real-time dynamic detection results come from different fermentation processes with the same cell culture process mode.

12. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-11, wherein, Calculating the addition dosage of the feeding medium in the current time period based on the difference between the real-time detection concentration of the marker in the current time period and the set value of the component concentration in the current time period includes one or more of the following: When the change rate of the marker shows a consumption rate, m 补 = (G 设 - G 实 ) × V 培 / G 补 , Among them, m 补 represents the addition dose of the feeding medium in the current time period, G 设 represents the set value of the component concentration of the marker in the current time period, G 实 represents the real-time detected concentration of the marker in the current time period, V 培 represents the volume of the cell culture medium in the bioreactor, G 补 represents the concentration of the marker in the feeding medium; or When the change rate of the marker shows a production rate, m 补 = k × (A 实 - A 设 ) × V 培 / B 补 , Among them, m 补 represents the addition dose of the feeding medium in the current time period, A 设 represents the set value of the component concentration of the marker in the current time period, A 实 represents the real-time detected concentration of the marker in the current time period, V 培 represents the volume of the cell culture medium in the bioreactor, B 补 represents the concentration of the medium component from which the marker in the feeding medium is transformed, and k represents the conversion coefficient from the concentration of the marker to the concentration of the medium component.

13. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-12, wherein, Calculating the addition dosage of the feeding medium in the current time period based on the current change rate of the marker includes one or more of the following: When the change rate of the marker shows a consumption rate, m 补 = R × V 培 / G 补 where R = G t0 - G 实 , where m 补 represents the addition dosage of the feeding medium in the current time period, R is the current change rate of the marker, V 培 represents the volume of the cell culture solution in the bioreactor, G 补 represents the marker concentration in the feeding medium, G t0 represents the real-time detection concentration of the marker in the previous time period, G 实 represents the real-time detection concentration of the marker in the current time period; or When the change rate of the marker shows a production rate, m 补 = k × R × V 培 / B 补 where R = A 实 - A t0 , Among them, m 补 represents the addition dose of the feeding medium in the current time period, R is the current change rate of the marker, V 培 represents the volume of the cell culture medium in the bioreactor, A t0 represents the real-time detection concentration of the marker in the previous time period, A 实 represents the real-time detection concentration of the marker in the current time period, B 补 represents the concentration of the medium component from which the marker in the feeding medium is transformed, and k represents the conversion coefficient from the concentration of the marker to the concentration of the medium component.

14. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-13, comprising: Determining the feeding strategy to be applied based on the first feeding strategy and the second feeding strategy, Optionally, the feeding strategy to be applied is determined as the average of the first feeding strategy and the second feeding strategy, Optionally, the feeding strategy to be applied is selected from the first feeding strategy and the second feeding strategy based on the stage in which the current time period falls, Optionally, when the current time period falls into the rapid production period of the target product, the first feeding strategy is selected as the feeding strategy to be applied; When the current time period falls into the logarithmic growth phase, the second feeding strategy is selected as the feeding strategy to be applied.

15. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-14, wherein, Calculating the addition dosage of the feeding medium in the current time period includes: Calculating the addition dosage of the marker or the culture medium component converted from the marker in the current time period according to the difference between the real-time detection concentration of the marker in the current time period and the set value of the component concentration in the current time period, or according to the current change rate of the marker; and Calculating the addition dosage of the feeding medium in the current time period according to the addition dosage of the marker or the culture medium component converted from the marker in the current time period.

16. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-15, wherein, The feeding medium includes a specific proportion of the marker and other components, and wherein the addition amounts of other components are obtained according to the addition dosage of the marker, Optionally, the automated feeding method includes writing the feeding strategy into the control system, wherein the control system is configured to, after receiving a numerical instruction regarding the feeding strategy, associate with the peristaltic pump hardware and perform instantaneous feeding in a dosage manner, Optionally, comparing the real-time detection concentration of the reference component in the current time period with the set value, and in response to the real-time detection concentration of the reference component in the current time period not being lower than the set value, not performing the automated feeding method, or in response to the real-time detection concentration of the reference component in the current time period being lower than the set value, performing the automated feeding method.

17. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-16, wherein, The current real-time dynamic detection result is online spectral data, wherein calculating the real-time detection concentration of the marker in the current time period according to the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a trained text convolutional model to obtain the real-time detection concentration of the marker in the current time period, The text convolution model is trained with the online spectral data of the indicators of the cell culture medium in the bioreactor at multiple different time points as samples and the offline target values of the indicators measured by sampling the cell culture medium as labels. The text convolution model is configured to receive the real-time detection spectral data of the indicators of the cell culture medium to be predicted and output the predicted values of the indicators.

18. The automated feeding method based on real-time dynamic detection results according to claim 17, wherein, The text convolution model is configured to: Set up a word vector mapping layer and perform feature dimensionality reduction; and After taking the mean in the spatial dimension, use a fully connected structure to obtain the predicted value of the indicator.

19. The automated feeding method based on real-time dynamic detection results according to claim 18, wherein, The text convolution model is configured to: After setting up the word vector mapping layer and performing feature dimensionality reduction, further perform neighbor feature sampling.

20. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-16, wherein, Calculating the real-time detection concentration of the marker in the current time period according to the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a machine learning combined model for predicting the component concentration of the cell culture medium to obtain the real-time detection concentration of the marker in the current time period. Among them, the machine learning combined model is established by the following method: Obtain a data set of the component concentrations of the cell culture medium in the bioreactor, and the data set includes a training data set, a validation data set, and a test data set; Respectively adopt a variety of machine learning algorithms to establish multiple single prediction models. Among them, after each of the multiple single prediction models is trained with the training data set and validated with the validation data set, it predicts the test data set. Compare the prediction results of the multiple single prediction models with the test data set to obtain the sum of the squared prediction errors of the multiple single prediction models, and determine the weights of the multiple single prediction models in the machine learning combined model according to the magnitude of the sum of the squared prediction errors. Through the weight assignment method, combine the multiple single prediction models to obtain the machine learning combined model.

21. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-16, wherein, Calculating the real-time detection concentration of the marker in the current time period according to the current real-time dynamic detection result includes inputting the current real-time dynamic detection result into a machine learning combined model for predicting the component concentration of the cell culture medium to obtain the real-time detection concentration of the marker in the current time period. Among them, the machine learning combined model is a migrated model obtained by migrating the original model. Among them, the original model is established by the following method: Obtain a data set of the component concentrations of the cell culture medium in the bioreactor, and the data set includes a training data set, a validation data set, and a test data set. Respectively adopt a variety of machine learning algorithms to establish multiple single prediction models. Among them, after each of the multiple single prediction models is trained with the training data set and validated with the validation data set, it predicts the test data set. Compare the prediction results of the multiple single prediction models with the test data set to obtain the sum of the squared prediction errors of the multiple single prediction models, and determine the weights of the multiple single prediction models in the original model according to the magnitude of the sum of the squared prediction errors. Through the weight assignment method, combine the multiple single prediction models to obtain the original model; Among them, the model migration includes: Obtain the original dataset used to establish the original model and the new batch training dataset of the component concentrations in the cell culture medium for the new batch of biological reactions. The original dataset includes the original training dataset, the original validation dataset, and the original test dataset. Perform scale correction or scale matching on the new batch training dataset and the original training dataset to obtain a new training dataset. Use the new training dataset, the original validation dataset, and the original test dataset, and establish multiple single prediction models respectively using a variety of machine learning algorithms. Among them, after each of the multiple single prediction models is trained with the new training dataset and verified with the original validation dataset, it makes predictions on the original test dataset. Compare the prediction results of the multiple single prediction models with the original test dataset to obtain the sum of the squared prediction errors of the multiple single prediction models, and determine the weights of the multiple single prediction models in the migrated model based on the magnitudes of the sum of the squared prediction errors. Combine the multiple single prediction models through a weight assignment method to obtain the migrated model. Among them, the scale correction includes incorporating a specified proportion of the data in the new batch training dataset into the original training dataset used to establish the original model. The scale matching includes incorporating the new batch training data in the new batch training dataset whose numerical difference is less than a specified threshold compared with the original training data with the same collection time in the original training dataset into the original training dataset.

22. The automated feeding method based on real-time dynamic detection results according to any one of claims 20-21, wherein, The dataset of the component concentrations includes online Raman spectroscopy data and its corresponding offline detection data, and the sampling time of the offline detection data matches the corresponding online Raman spectroscopy data.

23. The automated feeding method based on real-time dynamic detection results according to any one of claims 9-22, wherein, Selecting one component or a combination of multiple components as a biomarker based on the historical change rates of each component includes: Taking one component or a combination of multiple components from the N components with the fastest historical change rates as the biomarker, where N is a positive integer.

24. A non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the automated feeding method based on real-time dynamic detection results according to any one of claims 1-6, 9-23.

25. A computer device, including a memory and a processor. A computer program capable of running on the processor is stored on the memory. When the computer program is executed by the processor, it implements the steps of the automated feeding method based on real-time dynamic detection results according to any one of claims 1-6, 9-23.

26. A computer program product, the computer program product includes instructions. When the instructions are executed by a processor, it implements the automated feeding method based on real-time dynamic detection results according to any one of claims 1-6, 9-23.

Citation Information

Patent Citations

  • Method and apparatus for monitoring and automated control of bioreactor

    CN115985404A

  • Method for monitoring component concentration of cell culture fluid in bioreactor in real time

    CN116052778A

  • Predictive modeling and control of cell culture

    CN116457453A

  • Method, system and controller for process control in a bioreactor

    US20170253848A1

  • Dynamic nutrient control processes

    WO2022221269A1

Cited By

  • Adaptive cell culture method and device based on adjustable culture environment

    CN120700213A

  • CAR-T cell culture method and system based on sensing calculation cooperative controller

    CN120738402A

  • Method and system for preparing and controlling low-temperature efficient straw decay-promoting microbial inoculum

    CN120866584A

  • Haematococcus culture dissolved oxygen coupling regulation and control method for corrosion-resistant photoreactor

    CN120945134A

  • Multi-parameter integrated non-contact real-time monitoring and control cell culture system and method

    CN120988837A