A cell culture process abnormality early warning method and system
Patent Information
- Application Number
- CN202610992235.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-06
AI Technical Summary
[0003]本申请提供了一种细胞培养过程异常预警方法及系统,用于针对解决现有技术中细胞培养过程异常预警适应性差的技术问题
本申请提出了一种细胞培养过程异常预警方法及系统,通过预构建的数据时序预测器生成未来时间窗口内的预测数据序列,计算实际监测数据与预测数据的偏差并采用连续累积超限计数生成初步预警信号,再通过第一异常验证插件和第二异常验证插件对偏差序列进行双重验证,最终依据决策矩阵综合判定输出预警信号,显著提高了细胞培养过程异常预警的准确性和鲁棒性。相比传统方法,本申请提供的技术方案显著降低了对固定阈值的依赖,通过时序预测适应了培养过程的时变特性,利用连续累积计数机制有效滤除了随机噪声引起的瞬时波动,并结合历史成功模式匹配与生成式模型重构误差的双重视角,实现了对细胞培养过程异常状态的多角度交叉验证,大幅减少了虚警与漏警,达到了在复杂动态培养环境下识别真实异常趋势并输出细胞培养过程预警的技术效果。
Smart Images

Figure CN122490007B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a method and system for early warning of abnormalities in cell culture processes. Background Technology
[0002] In cell culture, real-time monitoring of key parameters and early warning of anomalies are crucial for ensuring successful culture. Current technologies typically employ fixed threshold methods, triggering an alarm when a parameter exceeds preset upper or lower limits. However, cell culture is a highly dynamic, nonlinear, and time-varying process with complex coupling relationships between different parameters, and the normal fluctuation range varies with the culture stage. Fixed threshold methods struggle to adapt to these time-varying characteristics, easily leading to missed alarms in the early stages of culture due to overly wide thresholds, or frequent false alarms in the later stages due to overly narrow thresholds. Summary of the Invention
[0003] This application provides a method and system for early warning of abnormalities in cell culture processes, which addresses the technical problem of poor adaptability of early warning systems for abnormalities in cell culture processes in existing technologies.
[0004] In view of the above problems, this application provides a method and system for early warning of abnormalities in the cell culture process.
[0005] In a first aspect, this application provides a method for early warning of abnormalities in the cell culture process, the method comprising: Input the historical monitoring data sequence during cell culture into a pre-constructed data time series predictor to generate a predicted data sequence within a future preset time window; Within the preset time window, the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point is calculated. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified. Based on the preliminary warning signal, the first and second anomaly verification plugins are activated. The deviation sequence to be verified is input into the first anomaly verification plugin. By calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database, the first verification result is output. The deviation sequence to be verified is input into the second anomaly verification plugin, the anomaly score of the deviation sequence to be verified is calculated by the variational autoencoder, and the second verification result is output based on the anomaly score; Based on a preset decision matrix, the first and second verification results are comprehensively judged to generate a final early warning signal, which provides an early warning of abnormalities in the cell culture process.
[0006] Secondly, this application provides an early warning system for abnormal cell culture processes, comprising: The prediction sequence generation module is used to input historical monitoring data sequences during cell culture into a pre-built data time series predictor to generate prediction data sequences within a preset future time window; The preliminary warning module is used to calculate the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point within the preset time window. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified. The first verification result acquisition module is used to activate the first and second anomaly verification plugins based on the preliminary warning signal, input the deviation sequence to be verified into the first anomaly verification plugin, and output the first verification result by calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database. The second verification result acquisition module is used to input the deviation sequence to be verified into the second anomaly verification plug-in, calculate the anomaly score of the deviation sequence to be verified through a variational autoencoder, and output the second verification result based on the anomaly score. The anomaly warning module is used to comprehensively judge the first verification result and the second verification result based on a preset decision matrix, generate a final warning signal, and provide anomaly warning for the cell culture process.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application proposes a method and system for early warning of anomalies in cell culture processes. It generates predicted data sequences within future time windows using a pre-constructed data time-series predictor, calculates the deviation between actual monitored data and predicted data, and generates an initial early warning signal using continuous cumulative over-limit counting. The deviation sequence is then double-verified using a first and second anomaly verification plugin. Finally, an early warning signal is output based on a comprehensive decision matrix, significantly improving the accuracy and robustness of early warning for anomalies in cell culture processes. Compared to traditional methods, the technical solution provided in this application significantly reduces reliance on fixed thresholds, adapts to the time-varying characteristics of the culture process through time-series prediction, effectively filters out instantaneous fluctuations caused by random noise using a continuous cumulative counting mechanism, and combines the dual perspectives of historical success pattern matching and generative model reconstruction errors to achieve multi-angle cross-validation of abnormal states in the cell culture process. This greatly reduces false alarms and missed alarms, achieving the technical effect of identifying true abnormal trends and outputting early warnings for cell culture processes in complex and dynamic culture environments. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating an abnormal early warning method for cell culture process provided in an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of the structure of a cell culture process abnormality early warning system provided in an embodiment of this application.
[0011] The components represented by each number in the attached diagram are explained below: The system includes a prediction sequence generation module 100, a preliminary early warning module 200, a first verification result acquisition module 300, a second verification result acquisition module 400, and an anomaly early warning module 500. Detailed Implementation
[0012] This application provides a method and system for early warning of abnormalities in cell culture processes, which addresses the technical problem of poor adaptability of early warning systems for abnormalities in existing cell culture processes.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0015] Example 1, as Figure 1 As shown, this application provides a method for early warning of abnormalities in the cell culture process, wherein the method includes: S10: Input the historical monitoring data sequence during cell culture into the pre-constructed data time series predictor to generate the predicted data sequence within the preset time window in the future.
[0016] During cell culture, various monitoring parameters such as pH, dissolved oxygen, and glucose concentration exhibit nonlinear time-varying characteristics, and there are complex couplings between different parameters. Traditional fixed threshold methods cannot capture the dynamic change trends of these parameters, leading to normal biological fluctuations in culture being misjudged as abnormal.
[0017] Step S10 in the method provided in this application embodiment includes: The construction steps of the data time series predictor include: Collect complete monitoring data for each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success criteria, and organize the monitoring data of each successful batch into a time series sample set in chronological order; Constrained by the time span of a preset time window, several sample monitoring data sequences are collected, and the historical actual monitoring data sequences of different sample monitoring data sequences within the historical time window are obtained as sample prediction data sequences, thus obtaining several sample prediction data sequences. An initial time series prediction model is constructed using a long short-term memory network. The input layer of the initial time series prediction model is configured to receive monitoring data sequences from multiple consecutive time points in the past, and the output layer of the initial time series prediction model is configured to output predicted data sequences from multiple consecutive time points in the future. Using the aforementioned sample monitoring data sequences and sample prediction data sequences as training data, the initial time series prediction model is iteratively trained until the model converges, generating a data time series predictor; The monitoring data includes at least five of the following parameters: pH value of the cell culture environment, dissolved oxygen concentration, culture temperature, glucose concentration, lactate concentration, ammonium ion concentration, live cell density, cell viability, osmotic pressure, and cell-specific yield.
[0018] In this embodiment, historical monitoring data sequences during cell culture are input into a pre-constructed data time series predictor to generate predicted data sequences within a future preset time window.
[0019] Specifically, first, the timing predictor is constructed.
[0020] Collect complete monitoring data for each successful batch that has completed the culture process and whose final quality indicators meet the preset success criteria. Organize the monitoring data of each successful batch into a time-series sample set in chronological order. For example, collect complete monitoring data sequences for batches that have completed the culture process and met the final quality indicators, and organize them into a sample prediction data sequence in chronological order. For example, for a successful batch, samples are taken every 2 hours from the start to the end of culture, obtaining monitoring data at 120 time points, each containing multiple monitoring data points. Organize the monitoring data of all successful batches into a time-series sample set in chronological order. For example, the final quality indicator may be a cell viability greater than 90% or a target product concentration meeting requirements. The monitoring data includes at least five of the following parameters: pH value, dissolved oxygen concentration, culture temperature, glucose concentration, lactate concentration, ammonium ion concentration, viable cell density, cell viability, osmotic pressure, and cell-specific yield.
[0021] Furthermore, constrained by the time span of a preset time window, several sample monitoring data sequences are collected, and the historical actual monitoring data sequences of different sample monitoring data sequences within the historical time window are obtained as sample prediction data sequences, resulting in several sample prediction data sequences. For example, the time window is set to 12 hours, constrained by the time span of the preset time window. Specifically, for each successful batch, starting from the first time point, a monitoring data sequence of length 6 is sequentially extracted from the sliding window as the input sequence, and the actual monitoring data sequences of the next 6 time points are used as the corresponding sample prediction data sequences.
[0022] Furthermore, an initial time-series prediction model is constructed using a Long Short-Term Memory (LSTM) network. The input layer of this model is configured to receive monitoring data sequences from multiple consecutive past time points, and the output layer is configured to output predicted data sequences from multiple consecutive future time points. For example, the input layer receives monitoring data sequences from six consecutive past time points, where each time point contains, for instance, 10 monitoring parameters, resulting in an input dimension of 6 × 10 = 60. The hidden layer contains two LSTM layers, each with 128 memory units. The output layer is a fully connected layer with an output dimension of 6 × 10 = 60, corresponding to the predicted values of the 10 parameters for the six future time points. The activation function is a linear function. Mean squared error is used as the loss function, the optimizer is Adam, the learning rate is 0.001, the batch size is 32, and the training epochs are 200.
[0023] Furthermore, using the aforementioned sample monitoring data sequences and sample prediction data sequences as training data, the initial time-series prediction model is iteratively trained until the model converges, generating a data time-series predictor. For example, all collected sample monitoring data sequences are used as input, and the corresponding sample prediction data sequences are used as output. Supervised training is performed on the model until the loss function converges, generating a data time-series predictor. In practical applications, the historical monitoring data sequences of the six most recent consecutive time points in the current cell culture process, such as the actual data from hour 1 to hour 6, are input into the predictor. The model outputs predicted data sequences for the next six time points, i.e., hours 7 to 12, with each time point containing predicted values for 10 monitoring parameters. The length of the time window can be set based on actual needs, and the sample data is organized based on the time window length.
[0024] By using a data time series predictor trained with historical monitoring data sequences, the predicted values of each parameter within a future time window can be output based on the actual change patterns over a period of time, providing a dynamic benchmark for subsequent deviation calculations.
[0025] S20: Within the preset time window, calculate the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, generate a preliminary warning signal and obtain the current deviation sequence as the deviation sequence to be verified.
[0026] Single or occasional deviations may originate from random factors such as measurement noise and instantaneous disturbances. Traditional methods that directly trigger alarms will generate a large number of false alarms, while relying solely on single-point over-limits may miss the true abnormal trend of continuous deviations and lack consideration for the continuity of the deviation time series.
[0027] Step S20 in the method provided in this application embodiment includes: At each sampling time point, the difference between the actual monitored value of each monitoring parameter and the predicted value at the corresponding time point in the predicted data sequence is calculated. The difference is used as the single parameter deviation value at the current time point, and the single parameter deviation values of all monitoring parameters are combined to form the multidimensional deviation vector at the current time point. Calculate the norm value of the multidimensional deviation vector at each time point, and compare the norm value with a preset deviation threshold. If the norm value is greater than the preset deviation threshold, the deviation state at the current time point is determined to be an out-of-limit state; otherwise, it is determined to be a normal state. Set up a fixed-length deviation state cache queue, and store the deviation states at each time point in chronological order. Whenever a new deviation state is stored in the deviation state cache queue, traverse backward from the end of the deviation state cache queue and count the number of deviation states that are consecutively in the out-of-limit state as the cumulative count. When the cumulative number of times reaches the preset value, a preliminary warning signal is generated, and a sequence composed of multidimensional deviation vectors from multiple consecutive time points before and after the trigger time is extracted as a deviation sequence to be verified.
[0028] In this embodiment of the application, within the preset time window, the deviation value between the actual monitoring data at the current time point and the predicted data sequence at the same time point is calculated. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified.
[0029] Specifically, firstly, at each sampling time point, the difference between the actual monitored value of each monitoring parameter and the predicted value at the corresponding time point in the predicted data sequence is calculated. This difference is used as the single-parameter deviation value at the current time point, and the single-parameter deviation values of all monitoring parameters are combined to form a multi-dimensional deviation vector for the current time point. For example, at each sampling time point, the actual monitoring data vector for the current time point is obtained, which includes the actual values of 10 parameters such as pH, dissolved oxygen, and glucose. Simultaneously, the predicted value vector for the same time point is extracted from the predicted data sequence generated in step S10. For each monitoring parameter, the difference between the actual value and the predicted value is calculated to obtain the single-parameter deviation value for that parameter. All 10 single-parameter deviation values are arranged in a fixed order to form a multi-dimensional deviation vector for the current time point, with a dimension of 10. For example, at a certain time point, the actual pH is 6.8, the predicted pH is 7.0, and the deviation is -0.2; the actual dissolved oxygen is 60%, the predicted dissolved oxygen is 65%, and the deviation is -5%. All single-parameter deviation values are combined into a vector form according to the parameter category order in the actual monitoring data vector to obtain the multi-dimensional deviation vector.
[0030] Further, the norm value of the multidimensional deviation vector at each time point is calculated, and the norm value is compared with a preset deviation threshold. If the norm value is greater than the preset deviation threshold, the deviation state at the current time point is determined to be in an out-of-limit state; otherwise, it is determined to be in a normal state. For example, the norm value of the multidimensional deviation vector can be calculated using the 2-norm, i.e., the Euclidean norm: each single-parameter deviation value is dimensionless, squared, summed, and then the square root is taken. The dimensionless processing is obtained by dividing the single-parameter deviation value of each parameter by its corresponding rated range value, so that the deviation values of each parameter are converted to the same dimension level. The preset deviation threshold is set to 5.0, and for example, the preset deviation threshold is determined based on twice the maximum deviation in historical normal batches. If the calculated norm value is greater than 5.0, the deviation state at the current time point is determined to be in an out-of-limit state; otherwise, it is in a normal state. For example, the deviations are as follows: pH -0.2, dissolved oxygen concentration -5.0%, culture temperature 0.3℃, glucose concentration -1.2 mmol / L, lactate concentration 0.8 mmol / L, ammonium ion concentration 0.1 mmol / L, and viable cell density deviation. Cells / mL, cell viability deviation -2.0%, osmotic pressure deviation 3.0 mOsm / kg, cell-specific yield deviation -0.01 pg / cell / day. The multidimensional deviation vector is then [-0.2, -5.0, 0.3, -1.2, 0.8, 0.1, -0.5, -2.0, 3.0, -0.01]. First, the square of each deviation value is calculated, resulting in a sequence of 0.04, 25.0, 0.09, 1.44, 0.64, 0.01, 0.25, 4.0, 9.0, and 0.0001. These squared values are then summed, yielding a total of 40.4701. Finally, the square root is taken, resulting in a norm of approximately 6.36. With a preset deviation threshold of 5.0, and since 6.36 is greater than 5.0, the deviation at the current time point is determined to be out of limit.
[0031] Furthermore, a fixed-length deviation state cache queue is set up to store the deviation states at each time point in chronological order. Whenever a new deviation state is stored in the deviation state cache queue, the queue is traversed backward from the end to count the number of consecutive deviation states in the out-of-limit state as the cumulative count. For example, a fixed-length deviation state cache queue, such as a queue length of 10, is set up to record the deviation states at the most recent 10 sampling time points. In chronological order, at each new sampling point, its deviation state, such as out-of-limit or normal, is stored at the tail of the queue; if the queue is full, the element at the head of the queue is removed. Whenever a new state is stored, the queue is traversed backward from the end, i.e., the latest state, to count the number of consecutive out-of-limit states as the cumulative count. For example, if the most recent 5 states in the queue are [normal, out-of-limit, out-of-limit, out-of-limit, out-of-limit], counting backward from the tail, the number of consecutive out-of-limit states is 4, i.e., the cumulative count is 4.
[0032] Furthermore, when the accumulated number of times reaches the preset value, a preliminary warning signal is generated, and a sequence composed of multidimensional deviation vectors from multiple consecutive time points before and after the trigger time is extracted as the deviation sequence to be verified. For example, the preset value is set to 3; when the accumulated number of times reaches 3, a preliminary warning signal is generated. When generating the preliminary warning signal, a sequence composed of multidimensional deviation vectors from multiple consecutive time points before and after the trigger time is extracted as the deviation sequence to be verified. For example, the multidimensional deviation vectors from the 5 time points before the trigger time and the trigger time itself, a total of 6 time points, are taken, arranged in chronological order to form a 6×10 matrix, which serves as the deviation sequence to be verified. If the trigger time is located in the early stage of cultivation, resulting in insufficient preceding time points, the missing preceding time points are filled with zero vectors.
[0033] By calculating the deviation between the actual and predicted values and counting the cumulative number of consecutive exceedances, a preliminary warning is only triggered when the deviation continues for a certain length, effectively filtering out instantaneous fluctuations caused by random noise. At the same time, the deviation sequence before and after the triggering time is extracted as the deviation sequence to be verified, preserving the complete temporal pattern of the deviation evolution and providing structured input for subsequent multi-angle verification.
[0034] S30: Activate the first anomaly verification plugin and the second anomaly verification plugin based on the preliminary warning signal, input the deviation sequence to be verified into the first anomaly verification plugin, calculate the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database, and output the first verification result.
[0035] Even if an initial warning can identify persistent deviations, it is difficult to determine whether the deviation pattern has appeared in successful batches based solely on the deviation amplitude of the current batch. For example, some deviation patterns may be part of normal fluctuations.
[0036] Step S30 in the method provided in this application embodiment includes: Feature extraction is performed on the deviation sequence to be verified to construct a feature vector that can characterize the core attributes of the deviation sequence. The feature vector includes the single-parameter deviation value of each monitoring parameter at each sampling time point, the first-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, the second-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, and the cross ratio between the deviation values of different monitoring parameters. The cosine distance algorithm is used to calculate the directional distance between the feature vector and each historical feature vector in the historical successful batch database. Then, the Euclidean distance is used to calculate the spatial distance between the feature vector and each historical feature vector. The directional distance and the spatial distance are weighted and fused to obtain the final similarity score. The historical feature vector with the highest similarity score to the feature vector is selected from the historical successful batch database, and the highest similarity score is taken as the best matching score of the current deviation sequence to be verified. The best matching score is compared with a preset similarity threshold. If the best matching score is less than or equal to the similarity threshold, the deviation sequence to be verified is determined to be highly similar to a deviation pattern in a historical successful batch, and a similarity determination result is output. If the best matching score is greater than the similarity threshold, it is determined that the deviation sequence to be verified is not similar to any deviation pattern in the historical successful batch, and the dissimilarity determination result is output. The process of constructing and updating the historical successful batch database and the process of setting the similarity threshold include: Collect complete monitoring data of each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success standard, extract the multidimensional deviation vector of each successful batch at each time point, and organize all the multidimensional deviation vectors of each successful batch into a historical deviation pattern in chronological order. For each successful batch, a feature extraction operation is performed on the historical deviation pattern to obtain the historical feature vector corresponding to the batch, and the historical feature vector is associated with the batch identification information and stored in the historical successful batch database. Calculate the pairwise similarity scores between all historical feature vectors in the historical successful batch database to obtain a similarity score set, statistically analyze the distribution characteristics of the similarity score set, and determine the 75th percentile of the similarity score set as the initial similarity threshold; Whenever a new successful batch completes the cultivation process, the incremental historical feature vector corresponding to the historical deviation pattern of the new successful batch is added to the historical successful batch database, and the pairwise similarity scores between all historical feature vectors are recalculated. The similarity threshold is then updated to the 75th percentile of the new similarity score set.
[0037] In this embodiment of the application, the first anomaly verification plugin and the second anomaly verification plugin are activated based on the preliminary warning signal. The deviation sequence to be verified is input into the first anomaly verification plugin. The first verification result is output by calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database.
[0038] Specifically, firstly, feature extraction is performed on the deviation sequence to be verified to construct a feature vector that can characterize the core attributes of the deviation sequence. The feature vector includes the single-parameter deviation value of each monitoring parameter at each sampling time point, the first-order rate of change of the deviation value of each monitoring parameter over the entire time range of the deviation sequence to be verified, the second-order rate of change of the deviation value of each monitoring parameter over the entire time range of the deviation sequence to be verified, and the cross ratio between the deviation values of different monitoring parameters. For example, the deviation sequence to be verified is obtained, and feature extraction is performed on the sequence to construct a feature vector characterizing the core attributes of the deviation sequence. The feature vector comprises four parts: The first part flattens the original single-parameter deviation values of all 10 parameters at 6 time points, resulting in 60 values; the second part calculates the first-order rate of change of each monitoring parameter's deviation value over the entire 6-time-point sequence, i.e., the difference between adjacent time points, resulting in 5 values (50 values for 10 parameters); the third part calculates the second-order rate of change of each monitoring parameter's deviation value, i.e., the difference between the first-order rates of change, resulting in 4 values (40 values for 10 parameters); and the fourth part calculates the cross ratio between the deviation values of different monitoring parameters. (Example) Specifically, since lactic acid accumulation usually leads to a decrease in pH, the first group calculates the ratio of pH deviation to lactic acid concentration deviation; since glucose consumption and lactic acid production are closely related during glycolysis, the second group calculates the ratio of glucose concentration deviation to lactic acid concentration deviation; since high-density cells consume more oxygen, the third group calculates the ratio of dissolved oxygen concentration deviation to viable cell density deviation; since ammonium ion accumulation inhibits cell viability, the fourth group calculates the ratio of ammonium ion concentration deviation to cell viability deviation; and since high glucose may lead to increased osmotic pressure, the fifth group calculates the ratio of osmotic pressure deviation to glucose concentration deviation. At each time point, these five ratios are calculated separately, yielding five values. Since the deviation sequence to be verified contains six consecutive time points, the fourth part generates a total of 6 × 5 = 30 values. Concatenating all parts yields a feature vector with a dimension of 60 + 50 + 40 + 30 = 180.
[0039] Further, the cosine distance algorithm is used to calculate the directional distance between the feature vector and each historical feature vector in the historical successful batch database. Then, Euclidean distance is used to calculate the spatial distance between the feature vector and each historical feature vector. The directional distance and spatial distance are weighted and fused to obtain the final similarity score. For example, the cosine distance algorithm is used to calculate the directional distance between the feature vector and each historical feature vector in the historical successful batch database, where cosine distance = 1 - cosine similarity, to reflect the difference in direction between the two vectors. Further, Euclidean distance is used to calculate the spatial distance between the feature vector and each historical feature vector. The cosine distance and Euclidean distance are weighted and fused to obtain the final similarity score, where a smaller score indicates a more similarity between the two deviation patterns. For example, the weights can be set to 0.5 and 0.5 respectively.
[0040] Furthermore, the historical feature vector with the highest similarity score to the current feature vector is selected from the historical successful batch database, and this highest similarity score is used as the best matching score for the current deviation sequence to be verified. For example, the historical feature vector with the highest similarity score to the current feature vector is selected from the historical successful batch database, and this highest similarity score is used as the best matching score.
[0041] Furthermore, the best matching score is compared with a preset similarity threshold. If the best matching score is less than or equal to the similarity threshold, it is determined that the deviation sequence to be verified is highly similar to a certain deviation pattern in the historical successful batch, and a similarity determination result is output.
[0042] Furthermore, if the best matching score is greater than the similarity threshold, it is determined that the deviation sequence to be verified is not similar to any deviation pattern in the historical successful batch, and the dissimilarity determination result is output.
[0043] The process of constructing and updating the historical successful batch database and the process of setting the similarity threshold include: First, complete monitoring data are collected for each successful batch that has completed the culture process and whose final quality indicators meet the preset success criteria. Multidimensional deviation vectors are extracted for each successful batch at each time point, and all multidimensional deviation vectors for each successful batch are organized into a historical deviation pattern in chronological order. For example, complete monitoring data are collected for each successful batch that has completed the culture process and whose final quality indicators meet the preset success criteria, such as cell viability greater than 90% or product concentration meeting the standard. For each successful batch, multidimensional deviation vectors are extracted at each time point as described in step S20, and all multidimensional deviation vectors for that batch at all time points are organized into a historical deviation pattern in chronological order. For example, if a batch has 120 time points, the historical deviation pattern is a 120×10 matrix.
[0044] Furthermore, feature extraction is performed on the historical deviation pattern of each successful batch to obtain the historical feature vector corresponding to the batch. This historical feature vector is then associated with batch identification information and stored in the historical successful batch database. For example, the same feature extraction as in step S30 is performed on the historical deviation pattern of each successful batch to obtain the historical feature vector corresponding to the batch. This historical feature vector is then associated with batch identification information, such as the batch number, and stored in the historical successful batch database to construct the historical successful batch database.
[0045] Further, the pairwise similarity scores between all historical feature vectors in the historical successful batch database are calculated to obtain a similarity score set. The distribution characteristics of the similarity score set are statistically analyzed, and the 75th percentile of the similarity score set is determined as the initial similarity threshold. For example, the same method as in step S30 is used to calculate the pairwise similarity scores between all historical feature vectors in the database to obtain a similarity score set. The distribution characteristics of this set are statistically analyzed, the 75th percentile is calculated, and this value is determined as the initial similarity threshold. For example, all pairwise similarity scores are sorted from smallest to largest, and the value at the 75th percentile is taken as the initial similarity threshold. The 75th percentile is chosen as the initial similarity threshold because if a new deviation pattern to be verified has a similarity score less than or equal to this threshold with a certain historical successful batch, then the degree of difference of the new pattern is at least not less than the degree of difference between 75% of historical normal batch pairs, and therefore it can be considered to fall within the normal variation range.
[0046] Furthermore, whenever a new successful batch completes its cultivation process, the incremental historical feature vector corresponding to the historical deviation pattern of the new successful batch is added to the historical successful batch database, and the pairwise similarity scores between all historical feature vectors are recalculated. The similarity threshold is then updated to the 75th percentile of the new similarity score set. For example, whenever a new successful batch completes its cultivation process, the incremental historical feature vector corresponding to the historical deviation pattern of the new batch is added to the database, and the pairwise similarity scores between all historical feature vectors are recalculated. The similarity threshold is then updated to the 75th percentile of the new similarity score set. By determining and dynamically updating the similarity threshold, it is possible to determine whether the current deviation pattern belongs to normal fluctuations based on historical success experience.
[0047] The first anomaly verification plugin matches the current deviation sequence with typical deviation patterns in the historical successful batch database. If it is highly similar to a certain successful pattern, it is initially judged as normal fluctuation; otherwise, it is judged as an anomaly. This introduces historical experience knowledge as a verification basis, improving the ability to identify false alarm anomalies.
[0048] S40: Input the deviation sequence to be verified into the second anomaly verification plugin, calculate the anomaly score of the deviation sequence to be verified through the variational autoencoder, and output the second verification result based on the anomaly score.
[0049] Historical successful batch databases may not cover all normal deviation patterns, relying solely on similarity matching carries the risk of missed detections, and anomalies under high-dimensional parameter coupling are difficult to fully characterize with simple distance metrics.
[0050] Step S40 in the method provided in this application embodiment includes: A pre-trained variational autoencoder model is obtained. The variational autoencoder model consists of an encoder network and a decoder network. The encoder network compresses and maps the input high-dimensional feature vector to a low-dimensional latent space and outputs the mean vector and log-variance vector of the latent encoding. The decoder network maps the sampling points in the latent space back to the original high-dimensional space and outputs the reconstructed feature vector. The variational autoencoder model is trained using the historical feature vectors of each successful batch in the historical successful batch database as training samples. The feature vector of the bias sequence to be verified is input into the variational autoencoder model. The encoder network outputs the latent coding mean vector and the latent coding log variance vector corresponding to the feature vector, and outputs the reconstructed feature vector based on the latent coding mean vector and the latent coding log variance vector. Calculate the reconstruction error between the feature vector and the reconstructed feature vector, and calculate the KL divergence between the coding distribution determined by the latent coding mean vector and the latent coding log variance vector and the standard normal distribution; The reconstruction error value and the KL divergence value are weighted and summed to obtain the comprehensive anomaly score of the deviation sequence to be verified; The process of setting the threshold for the abnormal score and the process of outputting the second verification result based on the abnormal score include: Obtain the feature vectors of all training samples in the historical successful batch database, input the feature vector of each training sample into the variational autoencoder model in sequence, calculate the anomaly score corresponding to each training sample, and obtain the historical normal and anomaly score set. Calculate the statistical parameters of the historical normal and abnormal score set, including the arithmetic mean of the historical normal and abnormal score set and the standard deviation of the historical normal and abnormal score set. Take the sum of the arithmetic mean and the standard deviation as a first abnormal score threshold, and take the sum of the arithmetic mean and the standard deviation as a second abnormal score threshold, wherein the second multiple is greater than the first multiple. The comprehensive anomaly score of the deviation sequence to be verified is compared sequentially with the first anomaly score threshold and the second anomaly score threshold. If the comprehensive anomaly score is less than the first anomaly score threshold, a normal judgment result is output. If the comprehensive anomaly score is greater than or equal to the first anomaly score threshold and less than the second anomaly score threshold, a boundary judgment result is output. If the comprehensive anomaly score is greater than or equal to the second anomaly score threshold, an anomaly judgment result is output.
[0051] In this embodiment of the application, the deviation sequence to be verified is input into the second anomaly verification plugin, the anomaly score of the deviation sequence to be verified is calculated by the variational autoencoder, and the second verification result is output based on the anomaly score.
[0052] Specifically, firstly, a pre-trained variational autoencoder (VAE) model is obtained. This model consists of an encoder network and a decoder network. The encoder network compresses and maps the input high-dimensional feature vector to a low-dimensional latent space and outputs the mean vector and log-variance vector of the latent encoding. The decoder network maps the sampling points in the latent space back to the original high-dimensional space and outputs the reconstructed feature vector. The VAE model is trained using historical feature vectors from each successful batch in a historical successful batch database as training samples. For example, a pre-trained VAE model is obtained. This model consists of an encoder network and a decoder network. The encoder network contains three fully connected layers: the first layer has an input dimension of 180 (corresponding to the feature vector dimension) and outputs 128 nodes, with a linear rectified function as the activation function; the second layer has an input of 128 and outputs 64 nodes, with a linear rectified function as the activation function; the third layer is divided into two branches, outputting the mean vector and log-variance vector of the latent encoding respectively. Each branch has an input of 64 and an output of 16 (latent space dimension), with no activation function. The decoder network also contains three fully connected layers: the first layer has 16 inputs and 64 output nodes, with a linear rectified function (RCF) activation function; the second layer has 64 inputs and 128 output nodes, with a RCF activation function; and the third layer has 128 inputs and 180 output nodes, with a linear function activation function, used to reconstruct input features. Furthermore, the variational autoencoder model is trained using historical feature vectors from each successful batch in the historical successful batch database as training samples. During training, a reparameterization technique is employed, and the loss function is obtained by weighted summation of mean squared error and KL divergence. The Adam optimizer is used with a learning rate of 0.001, a batch size of 32, and 200 training epochs. After training, the variational autoencoder model can learn the distribution of normal deviation sequences in the latent space.
[0053] Further, the feature vector of the deviation sequence to be verified is input into the variational autoencoder model. The encoder network outputs the latent encoding mean vector and the latent encoding log-variance vector corresponding to the feature vector, and outputs a reconstructed feature vector based on the latent encoding mean vector and the latent encoding log-variance vector. For example, the feature vector of the deviation sequence to be verified extracted in step S30 is input into the variational autoencoder model. The encoder network outputs the latent encoding mean vector (16-dimensional) and the latent encoding log-variance vector (16-dimensional) corresponding to the feature vector. When the decoder network samples from the latent space, it first samples a random noise vector of the same dimension as the mean vector from a standard normal distribution, where each element independently follows a normal distribution with a mean of 0 and a variance of 1. Then, the standard deviation vector is calculated; for each dimension, the standard deviation = e^(-1 / 2). 0.5 ×对数方差 Finally, the sampling point is calculated as: Sampling Point = Mean Vector + Standard Deviation Vector × Random Noise Vector. For example, assuming the mean = 0.5 and the logarithmic variance = -0.8 on a certain dimension, then the standard deviation = e 0.5×(-0.8) =0.67, and the random noise sampling value is 0.3, then the sampling points in this dimension = 0.5 + 0.67 × 0.3 = 0.701. The sampling points are input into the decoder network, which outputs the reconstructed feature vector.
[0054] Further, the reconstruction error value between the original feature vector and the reconstructed feature vector is calculated, as well as the KL divergence value between the coding distribution determined by the latent coding mean vector and the latent coding log-variance vector and the standard normal distribution is calculated. For example, the squares of the differences between the components of the original feature vector and the corresponding reconstructed feature vector are calculated, and these squares are summed; finally, the result is divided by 180, i.e., divided by the dimension of the feature vector, to obtain the reconstruction error value. For example, if the first three components of the original feature vector are [0.5, -0.2, 0.3], and the first three components of the reconstructed feature vector are [0.6, -0.1, 0.2], then the difference is [-0.1, -0.1, 0.1], and the sum of squares is 0.01 + 0.01 + 0.01 = 0.03. The remaining 177 components are calculated similarly. Assuming the total sum of squares is 5.4, then the reconstruction error = 5.4 / 180 = 0.03. Further, the KL divergence value between the coding distribution determined by the latent coding mean vector and the latent coding log-variance vector and the standard normal distribution is calculated. For example, the encoder network outputs the latent coding mean vector μ, with a dimension of 16, and each component is denoted as... ; and the logarithmic variance vector logσ 2 Each component is denoted as logσ1 2 ,logσ2 2 ,...,logσ 16 2 For each potential dimension i, calculate Then, the results for all dimensions are summed to obtain the KL divergence value. For example, assuming the latent space has only one dimension, μ=0.2, logσ 2 =-0.5, then Therefore, the KL divergence = 0.5 × [0.6065 + 0.04 - 1 - (-0.5)] = 0.07325. For 16 dimensions, the results of all dimensions are summed. The KL divergence value reflects the degree of deviation between the distribution position of the input feature vector in the latent space and the center of the normal sample distribution.
[0055] Further, the reconstruction error value and the KL divergence value are weighted and summed to obtain the comprehensive anomaly score of the deviation sequence to be verified. For example, the reconstruction error weight is preset to 0.6, and the KL divergence weight is preset to 0.4. That is, the comprehensive anomaly score = 0.6 × reconstruction error + 0.4 × KL divergence. Using the above numerical example, if the reconstruction error = 0.03 and the KL divergence = 0.073, then the comprehensive anomaly score = 0.6 × 0.03 + 0.4 × 0.073 = 0.0472.
[0056] The process of setting the threshold for the abnormal score and the process of outputting the second verification result based on the abnormal score include: First, the feature vectors of all training samples in the historical successful batch database are obtained. The feature vector of each training sample is then sequentially input into the variational autoencoder model to calculate the anomaly score corresponding to each training sample, resulting in a historical set of normal and anomaly scores. For example, the feature vectors of all training samples in the historical successful batch database are obtained. For instance, 200 successful batches are obtained, and a 180-dimensional feature vector is extracted from each batch. The feature vector of each training sample is then sequentially input into the pre-trained variational autoencoder model, and the anomaly score corresponding to each sample is calculated using the method described above, resulting in a historical set of normal and anomaly scores.
[0057] Further, the statistical parameters of the historical normal and abnormal score set are calculated, including the arithmetic mean and standard deviation of the historical normal and abnormal score set. The sum of the arithmetic mean and a first multiple of the standard deviation is used as a first abnormal score threshold, and the sum of the arithmetic mean and a second multiple of the standard deviation is used as a second abnormal score threshold, wherein the second multiple is greater than the first multiple. For example, the arithmetic mean and standard deviation of the historical normal and abnormal score set are calculated. The sum of the arithmetic mean and the first multiple of the standard deviation is used as the first abnormal score threshold. For example, the first multiple can be 1.5, the arithmetic mean is 0.35, and the standard deviation is 0.12, then the first threshold = 0.35 + 1.5 × 0.12 = 0.53; the sum of the arithmetic mean and the second multiple of the standard deviation is used as the second abnormal score threshold. The second multiple is 2.5, then the second threshold = 0.35 + 2.5 × 0.12 = 0.65. The first multiplier of 1.5 and the second multiplier of 2.5 can be adjusted according to the actual tolerance for false alarms and missed alarms.
[0058] Further, the comprehensive anomaly score of the deviation sequence to be verified is compared sequentially with the first anomaly score threshold and the second anomaly score threshold. If the comprehensive anomaly score is less than the first anomaly score threshold, a normal judgment result is output; if the comprehensive anomaly score is greater than or equal to the first anomaly score threshold and less than the second anomaly score threshold, a boundary judgment result is output; if the comprehensive anomaly score is greater than or equal to the second anomaly score threshold, an anomaly judgment result is output. For example, the comprehensive anomaly score of the deviation sequence to be verified is compared sequentially with the first anomaly score threshold and the second anomaly score threshold: if the comprehensive anomaly score is less than 0.53, a normal judgment result is output, indicating that the deviation pattern is close to a normal distribution; if the comprehensive anomaly score is greater than or equal to 0.53 and less than 0.65, a boundary judgment result is output, indicating that the deviation pattern is in a critical region and needs further confirmation; if the comprehensive anomaly score is greater than or equal to 0.65, an anomaly judgment result is output, indicating that the deviation pattern significantly deviates from the normal distribution and is highly likely to be a real anomaly. Through the second anomaly verification plugin, anomaly quantification scoring and hierarchical judgment of deviation sequences based on a generative model are realized.
[0059] By using the features of historical successful batches as training samples, the variational autoencoder learns the latent distribution of normal deviation sequences, calculates the reconstruction error and KL divergence for the current sequence to be verified, and obtains anomaly scores. It can detect deviation patterns that are not exactly the same as historical successful patterns but are also abnormal, which makes up for the shortcomings of the first verification plugin in misjudging the unseen normal patterns and provides a complementary verification perspective.
[0060] S50: Based on the preset decision matrix, the first verification result and the second verification result are comprehensively judged to generate a final warning signal and provide an abnormal warning for the cell culture process.
[0061] The results of a single verification plugin may be uncertain. For example, the first verification result may be dissimilar, but the second verification result may be a boundary state. If an alarm is triggered directly or ignored, it may lead to a misjudgment.
[0062] Step S50 in the method provided in this application embodiment includes: A decision matrix is established with the similarity and dissimilarity judgment results of the first verification result, the normal judgment result, the boundary judgment result, and the abnormal judgment result of the second verification result as inputs. The output of the decision matrix includes a suppression warning signal, a first-level warning signal, and a second-level warning signal. The decision matrix has the following rules: when the first verification result is a similarity determination result and the second verification result is a normal determination result, a suppression warning signal is output to cancel the preliminary warning signal; when the first verification result is a dissimilar determination result and the second verification result is a boundary determination result or an abnormal determination result, a second-level warning signal is output to indicate the true abnormal state; in addition, the other four determination combinations all output a first-level warning signal to indicate potential risks.
[0063] In this embodiment, the first verification result and the second verification result are comprehensively judged based on a preset decision matrix to generate a final early warning signal and provide an early warning of abnormalities in the cell culture process.
[0064] Specifically, firstly, a decision matrix is established with the similarity and dissimilarity judgment results of the first verification result, the normal judgment result, the boundary judgment result, and the anomaly judgment result of the second verification result as input. The output of the decision matrix includes a suppression warning signal, a first-level warning signal, and a second-level warning signal. For example, a decision matrix is established where the rows correspond to the first verification result output by the first anomaly verification plugin, including similarity judgment results (denoted as "similar") and dissimilarity judgment results (denoted as "dissimilarity"); the columns correspond to the second verification result output by the second anomaly verification plugin, including normal judgment results (denoted as "normal"), boundary judgment results (denoted as "boundary"), and anomaly judgment results (denoted as "anomaly"). Each cell of the decision matrix outputs a warning signal, including three types: a suppression warning signal (used to cancel the preliminary warning signal generated in step S20), a first-level warning signal (indicating potential risk), and a second-level warning signal (indicating a true anomaly state).
[0065] Furthermore, the decision matrix's judgment rules are as follows: A suppression warning signal is output only when the first verification result is a similarity judgment result and the second verification result is a normal judgment result; a second-level warning signal is output to indicate a true abnormal state when the first verification result is a dissimilarity judgment result and the second verification result is a boundary judgment result or an abnormal judgment result; otherwise, the remaining four judgment combinations all output a first-level warning signal to indicate potential risks. Specifically, a suppression warning signal is output only when the first verification result is "similar" and the second verification result is "normal"; a second-level warning signal is output when the first verification result is "dissimilar" and the second verification result is "boundary" or "abnormal"; the remaining four judgment combinations all output a first-level warning signal. Specifically, assuming that in a certain verification, the first verification result output by the first abnormality verification plugin is "similar" and the second verification result output by the second abnormality verification plugin is "boundary," then according to the decision matrix, a first-level warning signal is output, indicating that there is a potential risk, and operators are advised to pay attention but not to intervene immediately. If the first verification result is "similar" and the second verification result is "normal", a suppression warning signal is output. Upon receiving this signal, the central server cancels the preliminary warning signal generated in step S20 and does not issue any alarm. If the first verification result is "dissimilar" and the second verification result is "abnormal", a second-level warning signal is output, triggering an audible and visual alarm and forcibly stopping the cultivation process or notifying the operator for emergency handling. If the first verification result is "dissimilar" and the second verification result is "boundary", a second-level warning signal is also output. Because dissimilarity indicates that the deviation pattern does not match the historical successful batches, and boundary further indicates deviation from the normal distribution, it needs to be handled as a real anomaly. If the first verification result is "dissimilar" and the second verification result is "normal", a first-level warning signal is output, indicating that although the historical pattern does not match, the variational autoencoder considers it to still be within the normal distribution range, which may be a new normal pattern, and it should be handled as a potential risk. If the first verification result is "similar" and the second verification result is "abnormal", a first-level warning signal is output, that is, the historical pattern matches similarly, but the generated model determines it to be abnormal, which may be an abnormal pattern that has not yet been covered by the historical successful batches, and it should be handled as a potential risk. The verification results are fused using a decision matrix to generate a final early warning signal, which is then output to the cell culture process monitoring system to achieve tiered early warning.
[0066] By combining the similarity / dissimilarity judgment of the first verification plugin with the normal / boundary / abnormal judgment of the second verification plugin through a preset decision matrix, the output is a suppression warning, a first-level warning, or a second-level warning signal. This achieves the complementarity and trade-off between the two verification results. The suppression warning is only given when both indicate normality, the second-level warning is given when both indicate severe abnormality, and the first-level warning is given in other cases. This significantly improves the accuracy and rationality of the final warning signal.
[0067] Example 2, as Figure 2 As shown, based on the same inventive concept as the cell culture process abnormality early warning method provided in Embodiment 1, this embodiment of the invention also provides a cell culture process abnormality early warning system, including: The prediction sequence generation module 100 is used to input the historical monitoring data sequence during cell culture into a pre-constructed data time series predictor to generate a prediction data sequence within a future preset time window. The preliminary warning module 200 is used to calculate the deviation value between the actual monitoring data at the current time point and the predicted data sequence at the same time point within the preset time window. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified. The first verification result acquisition module 300 is used to activate the first and second anomaly verification plug-ins according to the preliminary warning signal, input the deviation sequence to be verified into the first anomaly verification plug-in, and output the first verification result by calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database. The second verification result acquisition module 400 is used to input the deviation sequence to be verified into the second anomaly verification plug-in, calculate the anomaly score of the deviation sequence to be verified through a variational autoencoder, and output the second verification result based on the anomaly score. The anomaly warning module 500 is used to comprehensively judge the first verification result and the second verification result based on a preset decision matrix, generate a final warning signal, and provide anomaly warning for the cell culture process.
[0068] In one embodiment, the prediction sequence generation module 100 is further configured to: The construction steps of the data time series predictor include: Collect complete monitoring data for each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success criteria, and organize the monitoring data of each successful batch into a time series sample set in chronological order; Constrained by the time span of a preset time window, several sample monitoring data sequences are collected, and the historical actual monitoring data sequences of different sample monitoring data sequences within the historical time window are obtained as sample prediction data sequences, thus obtaining several sample prediction data sequences. An initial time series prediction model is constructed using a long short-term memory network. The input layer of the initial time series prediction model is configured to receive monitoring data sequences from multiple consecutive time points in the past, and the output layer of the initial time series prediction model is configured to output predicted data sequences from multiple consecutive time points in the future. Using the aforementioned sample monitoring data sequences and sample prediction data sequences as training data, the initial time series prediction model is iteratively trained until the model converges, generating a data time series predictor; The monitoring data includes at least five of the following parameters: pH value of the cell culture environment, dissolved oxygen concentration, culture temperature, glucose concentration, lactate concentration, ammonium ion concentration, live cell density, cell viability, osmotic pressure, and cell-specific yield.
[0069] In one embodiment, the preliminary warning module 200 is further configured to: At each sampling time point, the difference between the actual monitored value of each monitoring parameter and the predicted value at the corresponding time point in the predicted data sequence is calculated. The difference is used as the single parameter deviation value at the current time point, and the single parameter deviation values of all monitoring parameters are combined to form the multidimensional deviation vector at the current time point. Calculate the norm value of the multidimensional deviation vector at each time point, and compare the norm value with a preset deviation threshold. If the norm value is greater than the preset deviation threshold, the deviation state at the current time point is determined to be an out-of-limit state; otherwise, it is determined to be a normal state. Set up a fixed-length deviation state cache queue, and store the deviation states at each time point in chronological order. Whenever a new deviation state is stored in the deviation state cache queue, traverse backward from the end of the deviation state cache queue and count the number of deviation states that are consecutively in the out-of-limit state as the cumulative count. When the cumulative number of times reaches the preset value, a preliminary warning signal is generated, and a sequence composed of multidimensional deviation vectors from multiple consecutive time points before and after the trigger time is extracted as a deviation sequence to be verified.
[0070] In this embodiment of the application, within the preset time window, the deviation value between the actual monitoring data at the current time point and the predicted data sequence at the same time point is calculated. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified.
[0071] In one embodiment, the first verification result acquisition module 300 is further configured to: Feature extraction is performed on the deviation sequence to be verified to construct a feature vector that can characterize the core attributes of the deviation sequence. The feature vector includes the single-parameter deviation value of each monitoring parameter at each sampling time point, the first-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, the second-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, and the cross ratio between the deviation values of different monitoring parameters. The cosine distance algorithm is used to calculate the directional distance between the feature vector and each historical feature vector in the historical successful batch database. Then, the Euclidean distance is used to calculate the spatial distance between the feature vector and each historical feature vector. The directional distance and the spatial distance are weighted and fused to obtain the final similarity score. The historical feature vector with the highest similarity score to the feature vector is selected from the historical successful batch database, and the highest similarity score is taken as the best matching score of the current deviation sequence to be verified. The best matching score is compared with a preset similarity threshold. If the best matching score is less than or equal to the similarity threshold, the deviation sequence to be verified is determined to be highly similar to a deviation pattern in a historical successful batch, and a similarity determination result is output. If the best matching score is greater than the similarity threshold, it is determined that the deviation sequence to be verified is not similar to any deviation pattern in the historical successful batch, and the dissimilarity determination result is output. The process of constructing and updating the historical successful batch database and the process of setting the similarity threshold include: Collect complete monitoring data of each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success standard, extract the multidimensional deviation vector of each successful batch at each time point, and organize all the multidimensional deviation vectors of each successful batch into a historical deviation pattern in chronological order. For each successful batch, a feature extraction operation is performed on the historical deviation pattern to obtain the historical feature vector corresponding to the batch, and the historical feature vector is associated with the batch identification information and stored in the historical successful batch database. Calculate the pairwise similarity scores between all historical feature vectors in the historical successful batch database to obtain a similarity score set, statistically analyze the distribution characteristics of the similarity score set, and determine the 75th percentile of the similarity score set as the initial similarity threshold; Whenever a new successful batch completes the cultivation process, the incremental historical feature vector corresponding to the historical deviation pattern of the new successful batch is added to the historical successful batch database, and the pairwise similarity scores between all historical feature vectors are recalculated. The similarity threshold is then updated to the 75th percentile of the new similarity score set.
[0072] In one embodiment, the second verification result acquisition module 400 is further configured to: A pre-trained variational autoencoder model is obtained. The variational autoencoder model consists of an encoder network and a decoder network. The encoder network compresses and maps the input high-dimensional feature vector to a low-dimensional latent space and outputs the mean vector and log-variance vector of the latent encoding. The decoder network maps the sampling points in the latent space back to the original high-dimensional space and outputs the reconstructed feature vector. The variational autoencoder model is trained using the historical feature vectors of each successful batch in the historical successful batch database as training samples. The feature vector of the bias sequence to be verified is input into the variational autoencoder model. The encoder network outputs the latent coding mean vector and the latent coding log variance vector corresponding to the feature vector, and outputs the reconstructed feature vector based on the latent coding mean vector and the latent coding log variance vector. Calculate the reconstruction error between the feature vector and the reconstructed feature vector, and calculate the KL divergence between the coding distribution determined by the latent coding mean vector and the latent coding log variance vector and the standard normal distribution; The reconstruction error value and the KL divergence value are weighted and summed to obtain the comprehensive anomaly score of the deviation sequence to be verified; The process of setting the threshold for the abnormal score and the process of outputting the second verification result based on the abnormal score include: Obtain the feature vectors of all training samples in the historical successful batch database, input the feature vector of each training sample into the variational autoencoder model in sequence, calculate the anomaly score corresponding to each training sample, and obtain the historical normal and anomaly score set. Calculate the statistical parameters of the historical normal and abnormal score set, including the arithmetic mean of the historical normal and abnormal score set and the standard deviation of the historical normal and abnormal score set. Take the sum of the arithmetic mean and the standard deviation as a first abnormal score threshold, and take the sum of the arithmetic mean and the standard deviation as a second abnormal score threshold, wherein the second multiple is greater than the first multiple. The comprehensive anomaly score of the deviation sequence to be verified is compared sequentially with the first anomaly score threshold and the second anomaly score threshold. If the comprehensive anomaly score is less than the first anomaly score threshold, a normal judgment result is output. If the comprehensive anomaly score is greater than or equal to the first anomaly score threshold and less than the second anomaly score threshold, a boundary judgment result is output. If the comprehensive anomaly score is greater than or equal to the second anomaly score threshold, an anomaly judgment result is output.
[0073] In one embodiment, the anomaly warning module 500 is further configured to: A decision matrix is established with the similarity and dissimilarity judgment results of the first verification result, the normal judgment result, the boundary judgment result, and the abnormal judgment result of the second verification result as inputs. The output of the decision matrix includes a suppression warning signal, a first-level warning signal, and a second-level warning signal. The decision matrix has the following rules: when the first verification result is a similarity determination result and the second verification result is a normal determination result, a suppression warning signal is output to cancel the preliminary warning signal; when the first verification result is a dissimilar determination result and the second verification result is a boundary determination result or an abnormal determination result, a second-level warning signal is output to indicate the true abnormal state; in addition, the other four determination combinations all output a first-level warning signal to indicate potential risks.
[0074] In summary, the embodiments of this application have at least the following technical effects: This application proposes a method and system for early warning of anomalies in cell culture processes. It generates predicted data sequences within future time windows using a pre-constructed data time-series predictor, calculates the deviation between actual monitored data and predicted data, and generates an initial early warning signal using continuous cumulative over-limit counting. The deviation sequence is then double-verified using a first and second anomaly verification plugin. Finally, an early warning signal is output based on a comprehensive decision matrix, significantly improving the accuracy and robustness of early warning for anomalies in cell culture processes. Compared to traditional methods, the technical solution provided in this application significantly reduces reliance on fixed thresholds, adapts to the time-varying characteristics of the culture process through time-series prediction, effectively filters out instantaneous fluctuations caused by random noise using a continuous cumulative counting mechanism, and combines the dual perspectives of historical success pattern matching and generative model reconstruction errors to achieve multi-angle cross-validation of abnormal states in the cell culture process. This greatly reduces false alarms and missed alarms, achieving the technical effect of identifying true abnormal trends and outputting early warnings for cell culture processes in complex and dynamic culture environments.
[0075] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0076] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0077] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A method for early warning of abnormalities in cell culture process, characterized in that, The methods include: Input the historical monitoring data sequence during cell culture into a pre-constructed data time series predictor to generate a predicted data sequence within a future preset time window; Within the preset time window, the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point is calculated. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified. Based on the preliminary warning signal, the first and second anomaly verification plugins are activated. The deviation sequence to be verified is input into the first anomaly verification plugin. By calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database, the first verification result is output. The deviation sequence to be verified is input into the second anomaly verification plugin, the anomaly score of the deviation sequence to be verified is calculated by the variational autoencoder, and the second verification result is output based on the anomaly score; Based on a preset decision matrix, the first and second verification results are comprehensively judged to generate a final early warning signal, providing early warning of abnormalities in the cell culture process, including: A decision matrix is established with the similarity and dissimilarity judgment results of the first verification result, the normal judgment result, the boundary judgment result, and the abnormal judgment result of the second verification result as inputs. The output of the decision matrix includes a suppression warning signal, a first-level warning signal, and a second-level warning signal. The decision matrix has the following rules: when the first verification result is a similarity determination result and the second verification result is a normal determination result, a suppression warning signal is output to cancel the preliminary warning signal; when the first verification result is a dissimilar determination result and the second verification result is a boundary determination result or an abnormal determination result, a second-level warning signal is output to indicate the true abnormal state. In addition, the other four judgment combinations all output a first-level warning signal to indicate potential risks.
2. The method for early warning of abnormalities in cell culture process according to claim 1, characterized in that, The construction steps of the data time series predictor include: Collect complete monitoring data for each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success criteria, and organize the monitoring data of each successful batch into a time series sample set in chronological order; Constrained by the time span of a preset time window, several sample monitoring data sequences are collected, and the historical actual monitoring data sequences of different sample monitoring data sequences within the historical time window are obtained as sample prediction data sequences, thus obtaining several sample prediction data sequences. An initial time series prediction model is constructed using a long short-term memory network. The input layer of the initial time series prediction model is configured to receive monitoring data sequences from multiple consecutive time points in the past, and the output layer of the initial time series prediction model is configured to output predicted data sequences from multiple consecutive time points in the future. Using the aforementioned sample monitoring data sequences and sample prediction data sequences as training data, the initial time series prediction model is iteratively trained until the model converges, thereby generating a data time series predictor.
3. The method for early warning of abnormalities in cell culture process according to claim 1, characterized in that, The monitoring data includes at least five of the following parameters: pH value of the cell culture environment, dissolved oxygen concentration, culture temperature, glucose concentration, lactate concentration, ammonium ion concentration, live cell density, cell viability, osmotic pressure, and cell-specific yield.
4. The method for early warning of abnormalities in cell culture process according to claim 1, characterized in that, Calculate the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point. If the cumulative number of times the deviation value exceeds a preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified, including: At each sampling time point, the difference between the actual monitored value of each monitoring parameter and the predicted value at the corresponding time point in the predicted data sequence is calculated. The difference is used as the single parameter deviation value at the current time point, and the single parameter deviation values of all monitoring parameters are combined to form the multidimensional deviation vector at the current time point. Calculate the norm value of the multidimensional deviation vector at each time point, and compare the norm value with a preset deviation threshold. If the norm value is greater than the preset deviation threshold, the deviation state at the current time point is determined to be an out-of-limit state; otherwise, it is determined to be a normal state. Set up a fixed-length deviation state cache queue, and store the deviation states at each time point in chronological order. Whenever a new deviation state is stored in the deviation state cache queue, traverse backward from the end of the deviation state cache queue and count the number of deviation states that are consecutively in the out-of-limit state as the cumulative count. When the cumulative number of times reaches the preset value, a preliminary warning signal is generated, and a sequence composed of multidimensional deviation vectors from multiple consecutive time points before and after the trigger time is extracted as a deviation sequence to be verified.
5. The method for early warning of abnormalities in cell culture process according to claim 1, characterized in that, The deviation sequence to be verified is input into the first anomaly verification plugin. By calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database, a first verification result is output, including: Feature extraction is performed on the deviation sequence to be verified to construct a feature vector that can characterize the core attributes of the deviation sequence. The feature vector includes the single-parameter deviation value of each monitoring parameter at each sampling time point, the first-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, the second-order rate of change of the deviation value of each monitoring parameter in the entire time range of the deviation sequence to be verified, and the cross ratio between the deviation values of different monitoring parameters. The cosine distance algorithm is used to calculate the directional distance between the feature vector and each historical feature vector in the historical successful batch database. Then, the Euclidean distance is used to calculate the spatial distance between the feature vector and each historical feature vector. The directional distance and the spatial distance are weighted and fused to obtain the final similarity score. The historical feature vector with the highest similarity score to the feature vector is selected from the historical successful batch database, and the highest similarity score is taken as the best matching score of the current deviation sequence to be verified. The best matching score is compared with a preset similarity threshold. If the best matching score is less than or equal to the similarity threshold, the deviation sequence to be verified is determined to be highly similar to a deviation pattern in a historical successful batch, and a similarity determination result is output. If the best matching score is greater than the similarity threshold, it is determined that the deviation sequence to be verified is not similar to any deviation pattern in the historical successful batch, and the dissimilarity determination result is output.
6. The method for early warning of abnormalities in cell culture process according to claim 5, characterized in that, The process of constructing and updating the historical successful batch database and the process of setting the similarity threshold include: Collect complete monitoring data of each successful batch that has completed the cultivation process and whose final quality indicators have reached the preset success standard, extract the multidimensional deviation vector of each successful batch at each time point, and organize all the multidimensional deviation vectors of each successful batch into a historical deviation pattern in chronological order. For each successful batch, a feature extraction operation is performed on the historical deviation pattern to obtain the historical feature vector corresponding to the batch, and the historical feature vector is associated with the batch identification information and stored in the historical successful batch database. Calculate the pairwise similarity scores between all historical feature vectors in the historical successful batch database to obtain a similarity score set, statistically analyze the distribution characteristics of the similarity score set, and determine the 75th percentile of the similarity score set as the initial similarity threshold; Whenever a new successful batch completes the cultivation process, the incremental historical feature vector corresponding to the historical deviation pattern of the new successful batch is added to the historical successful batch database, and the pairwise similarity scores between all historical feature vectors are recalculated. The similarity threshold is then updated to the 75th percentile of the new similarity score set.
7. The method for early warning of abnormalities in cell culture process according to claim 1, characterized in that, The deviation sequence to be verified is input into the second anomaly verification plugin, the anomaly score of the deviation sequence to be verified is calculated by the variational autoencoder, and the second verification result is output based on the anomaly score, including: A pre-trained variational autoencoder model is obtained. The variational autoencoder model consists of an encoder network and a decoder network. The encoder network compresses and maps the input high-dimensional feature vector to a low-dimensional latent space and outputs the mean vector and log-variance vector of the latent encoding. The decoder network maps the sampling points in the latent space back to the original high-dimensional space and outputs the reconstructed feature vector. The variational autoencoder model is trained using the historical feature vectors of each successful batch in the historical successful batch database as training samples. The feature vector of the bias sequence to be verified is input into the variational autoencoder model. The encoder network outputs the latent coding mean vector and the latent coding log variance vector corresponding to the feature vector, and outputs the reconstructed feature vector based on the latent coding mean vector and the latent coding log variance vector. Calculate the reconstruction error between the feature vector and the reconstructed feature vector, and calculate the KL divergence between the coding distribution determined by the latent coding mean vector and the latent coding log variance vector and the standard normal distribution; The reconstruction error value and the KL divergence value are weighted and summed to obtain the comprehensive anomaly score of the deviation sequence to be verified.
8. The method for early warning of abnormalities in cell culture process according to claim 7, characterized in that, The process of setting the threshold for the abnormal score and the process of outputting the second verification result based on the abnormal score include: Obtain the feature vectors of all training samples in the historical successful batch database, input the feature vector of each training sample into the variational autoencoder model in sequence, calculate the anomaly score corresponding to each training sample, and obtain the historical normal and anomaly score set. Calculate the statistical parameters of the historical normal and abnormal score set, including the arithmetic mean of the historical normal and abnormal score set and the standard deviation of the historical normal and abnormal score set. Take the sum of the arithmetic mean and the standard deviation as a first abnormal score threshold, and take the sum of the arithmetic mean and the standard deviation as a second abnormal score threshold, wherein the second multiple is greater than the first multiple. The comprehensive anomaly score of the deviation sequence to be verified is compared sequentially with the first anomaly score threshold and the second anomaly score threshold. If the comprehensive anomaly score is less than the first anomaly score threshold, a normal judgment result is output. If the comprehensive anomaly score is greater than or equal to the first anomaly score threshold and less than the second anomaly score threshold, a boundary judgment result is output. If the comprehensive anomaly score is greater than or equal to the second anomaly score threshold, an anomaly judgment result is output.
9. A cell culture process abnormality early warning system, characterized in that, For implementing the cell culture process abnormality early warning method according to any one of claims 1-8, the system comprises: The prediction sequence generation module is used to input historical monitoring data sequences during cell culture into a pre-built data time series predictor to generate prediction data sequences within a preset future time window; The preliminary warning module is used to calculate the deviation between the actual monitoring data at the current time point and the predicted data sequence at the same time point within the preset time window. If the cumulative number of times the deviation value exceeds the preset deviation threshold reaches a preset value, a preliminary warning signal is generated, and the current deviation sequence is obtained as the deviation sequence to be verified. The first verification result acquisition module is used to activate the first and second anomaly verification plugins based on the preliminary warning signal, input the deviation sequence to be verified into the first anomaly verification plugin, and output the first verification result by calculating the similarity between the deviation sequence to be verified and each historical deviation pattern in the historical successful batch database. The second verification result acquisition module is used to input the deviation sequence to be verified into the second anomaly verification plug-in, calculate the anomaly score of the deviation sequence to be verified through a variational autoencoder, and output the second verification result based on the anomaly score. An anomaly warning module is used to comprehensively determine the first and second verification results based on a preset decision matrix, generate a final warning signal, and provide anomaly warnings for the cell culture process, including: A decision matrix is established with the similarity and dissimilarity judgment results of the first verification result, the normal judgment result, the boundary judgment result, and the abnormal judgment result of the second verification result as inputs. The output of the decision matrix includes a suppression warning signal, a first-level warning signal, and a second-level warning signal. The decision matrix has the following rules: when the first verification result is a similarity determination result and the second verification result is a normal determination result, a suppression warning signal is output to cancel the preliminary warning signal; when the first verification result is a dissimilar determination result and the second verification result is a boundary determination result or an abnormal determination result, a second-level warning signal is output to indicate the true abnormal state; in addition, the other four determination combinations all output a first-level warning signal to indicate potential risks.
Citation Information
Patent Citations
Abnormal data detection method and device and server
CN110377447A
Time series prediction execution based on deviation risk evaluation
US20240202579A1