Cell culture parameter prediction method and system based on time sequence fusion attention

By employing a temporal fusion attention-based cell culture parameter prediction method, which combines self-attention mechanism and temporal convolution method, the problem of predicting high-dimensional and complex cell culture data is solved, improving prediction accuracy and training efficiency, and adapting to rapid prediction in new scenarios.

CN120977394APending Publication Date: 2025-11-18UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410604315.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively process high-dimensional, complex cell culture data with significant dimensional differences. Traditional deep learning methods exhibit poor prediction accuracy in multivariate cell culture scenarios, making them unsuitable for optimizing vaccine production processes.

Method used

A cell culture parameter prediction method based on temporal fusion attention is adopted. By combining a temporal prediction backbone network, an adaptive feature matcher and a self-attention layer, the temporal and long-term correlations in cell data are captured. A simplified LSTM unit and an adaptive feature matcher are designed to adapt to new scenario data.

Benefits of technology

It improves the training efficiency and accuracy of cell culture parameter prediction, can quickly adapt to data changes in new scenarios, and achieves efficient prediction in small sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977394A_ABST
    Figure CN120977394A_ABST
Patent Text Reader

Abstract

The invention discloses a cell culture parameter prediction method and system based on time sequence fusion attention. The method comprises the steps that influence factor data and target data in the cell culture process are collected and preprocessed; inputting the preprocessed influence factor data and target data into a time sequence prediction backbone network composed of a variable selection structure and a time sequence unit, and performing time sequence prediction on the influence factor data and the target data; inputting the hidden state output by the time sequence prediction backbone network into an adaptive feature matcher, and storing or matching the input hidden state; inputting the time sequence hidden states output by the adaptive feature matcher into a self-attention layer suitable for time sequence prediction, and carrying out attention score calculation on all the time sequence hidden states; and after the attention score is obtained, outputting a prediction result through layer normalization and a full connection layer. By adopting the method, the cell culture parameters can be predicted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cell culture parameter prediction, in particular to a cell culture parameter prediction method and system based on time sequence fusion attention. BACKGROUND

[0002] In modern vaccine production processes, pathogens (such as viruses or bacteria) often need to be cultured and propagated in cell cultures. This technique of culturing pathogens with cells can better control the growth conditions of the pathogens, ensure that their quantity is sufficient for the preparation of vaccines, and also evaluate the effectiveness, safety and stability of candidate vaccines. By conducting various experiments in cultured cells, the progress of vaccine research can be accelerated. Currently, the parameters set during the cell culture process (such as initial culture density, type of culture medium, etc.) and the recorded data (such as cell viability, alkali consumption, etc.) can be analyzed simultaneously to determine whether the vaccine production is successful, but how to determine it still depends on the experience, seniority and knowledge of the operator. How to achieve data-driven cell culture parameter prediction, realize cell culture parameter analysis and timely output of optimization schemes is a hot topic in the current vaccine manufacturing industry.

[0003] In order to complete the prediction of various cell culture parameters, it is necessary to quantify the experience and knowledge of cell culture operators, to mine the mathematical laws between cell culture parameters and the results provided by operators, to form feature data that machines can recognize and calculate, and to complete the prediction of future cell parameters through mathematical laws and known parameters. Common methods include establishing a gray model to extract the differentiated change law in the data, using an autoregressive model to extract the time sequence change law in the data, and using deep learning methods to learn deep features in the data. However, with the increasing number of strains and drugs, the internal mechanisms of biochemical reactions and the metabolic mechanisms of microbial biochemical processes in different production processes become increasingly complex, the influencing factors gradually increase, and the difficulty of analyzing data laws gradually increases. At the same time, some new cell culture scenarios can provide a small amount of data, and traditional deep learning methods will suffer from underfitting when facing this problem, resulting in poor prediction accuracy. These problems indicate that traditional prediction methods cannot be applied to the prediction and control of current cell culture parameters, and cannot provide reference value for optimization decisions of vaccine production processes. SUMMARY

[0004] The present application provides a cell culture parameter prediction method and system based on time sequence fusion attention to solve the problems existing in the prior art. The technical solution is as follows:

[0005] On the one hand, a cell culture parameter prediction method based on time sequence fusion attention is provided, comprising:

[0006] S1. Collect and preprocess data on influencing factors and target data in the cell culture process;

[0007] S2. Input the preprocessed influencing factor data and target data into the temporal prediction backbone network composed of variable selection structure and temporal unit to perform temporal prediction of influencing factor data and target data. The temporal prediction backbone network is the first part of the temporal fusion attention prediction model.

[0008] S3. Input the hidden state output by the time-series prediction backbone network into the adaptive feature matcher to store or match the input hidden state;

[0009] S4. Input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model.

[0010] S5. After obtaining the attention score, the prediction result is output through layer normalization and fully connected layers. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

[0011] Optionally, the preprocessing in S1 specifically includes:

[0012] S11. Divide the influencing factor data according to the stage of cell culture to obtain the processed cell data table;

[0013] S12. Use linear interpolation to fill the data and obtain the filled cell data table.

[0014] S13. Manually add time step indexes and convert them into CSV tables to obtain preprocessed cell CSV tables.

[0015] Optionally, the variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization, one of which is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows:

[0016]

[0017] That This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

[0018] Optionally, the adaptive feature matcher summarizes the input vector and also modifies the input vector. The algorithm formula for modifying the input vector is as follows:

[0019]

[0020] in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

[0021] Optionally, the temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole unit. The number of whole units is concatenated according to the number of input time steps and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows:

[0022]

[0023] in The matrix consists of all hidden state vectors. The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.

[0024] On the other hand, a cell culture parameter prediction system based on temporal fusion attention is provided, the system comprising:

[0025] The data acquisition and preprocessing module is used to collect and preprocess data on influencing factors and target data in the cell culture process.

[0026] The temporal prediction module is used to input the preprocessed influencing factor data and target data into the temporal prediction backbone network composed of variable selection structure and temporal unit to perform temporal prediction of influencing factor data and target data. The temporal prediction backbone network is the first part of the temporal fusion attention prediction model.

[0027] An adaptive feature matching module is used to input the hidden state output by the temporal prediction backbone network into the adaptive feature matcher, and to store or match the input hidden state.

[0028] The attention score calculation module is used to input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and to calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model.

[0029] The prediction module is used to output the prediction result through layer normalization and fully connected layers after obtaining the attention score. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

[0030] Optionally, the acquisition preprocessing module is specifically used for:

[0031] The data on influencing factors were divided according to the stage of cell culture to obtain the processed cell data table;

[0032] Linear interpolation was used to fill the data, resulting in a filled cell data table.

[0033] The time step index was manually added and converted into a CSV table to obtain a preprocessed cell CSV table.

[0034] Optionally, the variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization, one of which is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows:

[0035]

[0036] in This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

[0037] Optionally, the adaptive feature matcher summarizes the input vector and also modifies the input vector. The algorithm formula for modifying the input vector is as follows:

[0038]

[0039] in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

[0040] Optionally, the temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole unit. The number of whole units is concatenated according to the number of input time steps and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows:

[0041]

[0042] in The matrix consisting of all hidden state vectors The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.

[0043] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method for predicting cell culture parameters based on temporal fusion attention.

[0044] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-described method for predicting cell culture parameters based on temporal fusion attention.

[0045] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0046] 1) Compared with traditional machine learning methods, the temporal fusion attention prediction model of this invention can handle high-dimensional, complex cell culture data with large dimensional differences, so as to maximize the use of every parameter in the cell data.

[0047] 2) The model combines self-attention mechanism and temporal convolution method, which can process high-dimensional cell culture data in parallel, and can also better capture the long-term correlation between variables, thereby improving training efficiency and prediction accuracy.

[0048] 3) An adaptive feature adapter was designed, which, combined with a large amount of historical cell culture data, allows the prediction model to quickly adapt to data in new scenarios and more efficiently predict cell parameters in small sample scenarios. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart of a cell culture parameter prediction method based on temporal fusion attention provided in an embodiment of the present invention;

[0051] Figure 2 This is a schematic diagram of the cell culture data preprocessing process provided in an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of the variable selection structure and timing unit provided in the embodiments of the present invention;

[0053] Figure 4 This is a schematic diagram of the adaptive feature matcher provided in an embodiment of the present invention;

[0054] Figure 5This is a schematic diagram of the temporal fusion attention prediction model provided in an embodiment of the present invention;

[0055] Figure 6 This is a schematic diagram of the cell culture parameter prediction method model training provided in an embodiment of the present invention;

[0056] Figure 7 This is a block diagram of a cell culture parameter prediction system based on temporal fusion attention provided in an embodiment of the present invention;

[0057] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0058] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0059] This invention provides a method for predicting cell culture parameters based on temporal fusion attention. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart shown is a cell culture parameter prediction method based on temporal fusion attention. The processing flow of this method may include the following steps:

[0060] S1. Collect and preprocess data on influencing factors and target data in the cell culture process;

[0061] In this embodiment of the invention, key parameter data and quality control data indicators of the cell culture process are collected, such as data on influencing factors like cell density, pH value, and bacterial strain, as well as target data such as cell viability and protein content. Typical cell culture process data are established to form an original dataset in txt, json, or excel format documents.

[0062] Because the original cell culture process data contains too many types of time-series data, therefore, such as Figure 2As shown, this embodiment of the invention performs necessary preprocessing on the original data: First, this embodiment of the invention divides the influencing factors according to the stage of cell culture to obtain a processed cell data table. For example, the initial culture density, centrifugation time and temperature, concentration factor, etc. are divided into the production stage of concentrated antigen solution. This can effectively help operators to better analyze the degree and cause of the influence of each factor on the target when training and testing the model. In order to avoid the training program encountering null values ​​and generating errors during runtime, it is necessary to fill the missing values ​​of the above data. In this embodiment of the invention, linear interpolation is used to fill the missing values ​​to obtain a filled cell data table. At the same time, although most cell culture data have corresponding recording times, the recording times are mostly in string form and require separate model training, and the recording intervals are not uniform. In order to emphasize the temporal relationship between data, this embodiment of the invention manually adds a time step index from 1 to the data length to the data, thereby replacing the original recording time as a better time index, and converts it into a CSV table to obtain a preprocessed cell CSV table.

[0063] S2. Input the preprocessed influencing factor data and target data into the temporal prediction backbone network composed of variable selection structure and temporal unit to perform temporal prediction of influencing factor data and target data. The temporal prediction backbone network is the first part of the temporal fusion attention prediction model.

[0064] Optionally, such as Figure 3 As shown, the variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization function, one of which is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows:

[0065]

[0066] in This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

[0067] S3. Input the hidden state output by the time-series prediction backbone network into the adaptive feature matcher to store or match the input hidden state;

[0068] Optionally, such as Figure 4 As shown, the adaptive feature matcher summarizes and modifies the input vector. The algorithm formula for modifying the input vector is as follows:

[0069]

[0070] in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

[0071] S4. Input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model.

[0072] S5. After obtaining the attention score, the prediction result is output through layer normalization and fully connected layers. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

[0073] Optionally, such as Figure 5As shown, the temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole. The number of whole units is equal to the number of input and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows:

[0074]

[0075] in The matrix consists of all hidden state vectors. The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.

[0076] like Figure 6 As shown, the training of the model in this embodiment of the invention is divided into two stages: pre-training and actual training. In the pre-training stage, a large amount of historical cell culture data is used as input. During training, the weight parameters of all components (adaptive feature matcher and temporal fusion attention prediction model) are updated. The adaptive feature matcher does not change the hidden state of the input; it only updates its own weights as a storage of historical data features. In the actual training stage, a small sample of new scene cell data is used as input. During training, half of the weight parameters of each component need to be frozen to prevent them from participating in parameter updates. The adaptive feature matcher does not update its own weights; it only changes the hidden state of the input, allowing the self-attention layer to better fine-tune its weights. In the testing stage, the target variable in the input data is treated as unknown, and other variables are treated as influencing factors. These are then input into the model to obtain the predicted value of the future target variable.

[0077] like Figure 7 As shown, this embodiment of the invention also provides a cell culture parameter prediction system based on temporal fusion attention, the system comprising:

[0078] The data acquisition and preprocessing module 710 is used to acquire and preprocess data on influencing factors and target data in the cell culture process.

[0079] The time series prediction module 720 is used to input the preprocessed influencing factor data and target data into the time series prediction backbone network composed of variable selection structure and time series unit to perform time series prediction of influencing factor data and target data. The time series prediction backbone network is the first part of the time series fusion attention prediction model.

[0080] The adaptive feature matching module 730 is used to input the hidden state output by the temporal prediction backbone network into the adaptive feature matcher, and to store or match the input hidden state.

[0081] The attention score calculation module 740 is used to input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and to calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model.

[0082] The prediction module 750 is used to output the prediction result through layer normalization and fully connected layers after obtaining the attention score. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

[0083] Optionally, the acquisition preprocessing module is specifically used for:

[0084] The data on influencing factors were divided according to the stage of cell culture to obtain the processed cell data table;

[0085] Linear interpolation was used to fill the data, resulting in a filled cell data table.

[0086] The time step index was manually added and converted into a CSV table to obtain a preprocessed cell CSV table.

[0087] Optionally, the variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization, one of which is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows:

[0088]

[0089] in This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

[0090] Optionally, the adaptive feature matcher summarizes the input vector and also modifies the input vector. The algorithm formula for modifying the input vector is as follows:

[0091]

[0092] in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

[0093] Optionally, the temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole unit. The number of whole units is concatenated according to the number of input time steps and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows:

[0094]

[0095] in The matrix consists of all hidden state vectors. The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.

[0096] The cell culture parameter prediction system based on temporal fusion attention provided in this embodiment of the invention has a functional structure that corresponds to the cell culture parameter prediction method based on temporal fusion attention provided in this embodiment of the invention, and will not be described again here.

[0097] Figure 8 This is a schematic diagram of the structure of an electronic device 800 provided in an embodiment of the present invention. The electronic device 800 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 801 and one or more memories 802. The memory 802 stores at least one instruction, which is loaded and executed by the processor 801 to implement the steps of the cell culture parameter prediction method based on temporal fusion attention described above.

[0098] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to perform the aforementioned cell culture parameter prediction method based on temporal fusion attention. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0099] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0100] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting cell culture parameters based on temporal fusion attention, characterized in that, The method includes: S1. Collect and preprocess data on influencing factors and target data in the cell culture process; S2. Input the preprocessed influencing factor data and target data into the temporal prediction backbone network composed of variable selection structure and temporal unit to perform temporal prediction of influencing factor data and target data. The temporal prediction backbone network is the first part of the temporal fusion attention prediction model. S3. Input the hidden state output by the time-series prediction backbone network into the adaptive feature matcher to store or match the input hidden state; S4. Input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model. S5. After obtaining the attention score, the prediction result is output through layer normalization and fully connected layers. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

2. The method according to claim 1, characterized in that, The preprocessing in S1 specifically includes: S11. Divide the influencing factor data according to the stage of cell culture to obtain the processed cell data table; S12. Use linear interpolation to fill the data and obtain the filled cell data table. S13. Manually add time step indexes and convert them into CSV tables to obtain preprocessed cell CSV tables.

3. The method according to claim 1, characterized in that, The variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization function. One of these branches is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows: in This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

4. The method according to claim 3, characterized in that, The adaptive feature matcher summarizes and modifies the input vector. The algorithm for modifying the input vector is as follows: in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

5. The method according to claim 4, characterized in that, The temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole. The number of whole units is equal to the number of input and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows: in The matrix consists of all hidden state vectors. The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.

6. A cell culture parameter prediction system based on temporal fusion attention, characterized in that, The system includes: The data acquisition and preprocessing module is used to collect and preprocess data on influencing factors and target data in the cell culture process. The temporal prediction module is used to input the preprocessed influencing factor data and target data into the temporal prediction backbone network composed of variable selection structure and temporal unit to perform temporal prediction of influencing factor data and target data. The temporal prediction backbone network is the first part of the temporal fusion attention prediction model. An adaptive feature matching module is used to input the hidden state output by the temporal prediction backbone network into the adaptive feature matcher, and to store or match the input hidden state. The attention score calculation module is used to input the temporal hidden state output by the adaptive feature matcher into the self-attention layer suitable for temporal prediction, and to calculate the attention score for all temporal hidden states. The self-attention layer suitable for temporal prediction is the second part of the temporal fusion attention prediction model. The prediction module is used to output the prediction result through layer normalization and fully connected layers after obtaining the attention score. The layer normalization and fully connected layers are the third part of the temporal fusion attention prediction model.

7. The system according to claim 6, characterized in that, The acquisition preprocessing module is specifically used for: The data on influencing factors were divided according to the stage of cell culture to obtain the processed cell data table; Linear interpolation was used to fill the data, resulting in a filled cell data table. The time step index was manually added and converted into a CSV table to obtain a preprocessed cell CSV table.

8. The system according to claim 6, characterized in that, The variable selection structure consists of two branches, each consisting of a fully connected layer, a ReLU activation function, and a layer normalization function. One of these branches is responsible for processing the target data. One process transforms the data into vectors that the model can recognize, while the other process handles all the influencing factor data. This approach maps text and numerical quantities with inconsistent units to a unified vector space. The vectors output from the two branches are then concatenated as input to the subsequent model. During training, the parameters of the fully connected layers are continuously updated, allowing for the selection of more important temporal variables as subsequent inputs, thus improving training efficiency. Due to the variable selection structure, only the temporal unit needs to capture temporal correlations. Therefore, a simplified LSTM unit is designed, merging the original forget gate and input gate into a single update gate as the primary means of changing the hidden state of the vector. The amount of previous memory stored at the current time step is defined to capture long-term correlations of variables in the sequence; while the reset gate determines how new input information is combined with previous memory to capture short-term correlations of variables. The formula for calculating the change in the hidden state is as follows: in This represents the hidden state at the current moment after inputting the variables. This is the hidden state from the previous moment. To update the gate weights and input The result after multiplication To reset gate weights and input The result after multiplication For the XOR operation between matrices, the subsequent model will use the hidden states. This is used to capture temporal correlations in the data to output predicted values.

9. The system according to claim 8, characterized in that, The adaptive feature matcher summarizes and modifies the input vector. The algorithm for modifying the input vector is as follows: in The hidden state vector after matching, The historical data feature weights stored for the matcher are randomly initialized before pre-training and need to be updated by the adaptive feature matcher through training on historical data. The weight parameters, This represents the maximum distance between two vectors in the vector space. A small value indicates that the two vectors are very close, while a large value indicates that they are quite different. For historical data vectors, Given a new scene data vector, the above formula can be used to match and modify the input vector with historical data features, making it conform to historical data features, which is beneficial for subsequent predictions.

10. The system according to claim 9, characterized in that, The temporal fusion attention prediction model comprises three parts. The first part is the temporal prediction backbone network, which is composed of the variable selection structure and temporal units as a whole. The number of whole units is equal to the number of input and output time steps, forming the temporal prediction backbone network. The temporal prediction backbone network is connected to the adaptive feature matcher. The hidden states output by the temporal prediction backbone network are input into the adaptive feature matcher. The adaptive feature matcher stores or matches the input hidden states and then inputs them into the second part of the temporal fusion attention prediction model: a self-attention layer suitable for temporal prediction. The self-attention layer suitable for temporal prediction calculates attention scores for all temporal hidden states, and the calculation formula is as follows: in The matrix consists of all hidden state vectors. The attention score for the hidden state matrix. , , There are 3 learnable weights. The feature dimension is the dimension of the time series data. After obtaining the attention score through the above formula, the prediction result is output through the third part of the time series fusion attention prediction model: layer normalization and fully connected layer.