Cell culture dynamic regulation system based on multi-modal time series data
The cell culture dynamic control system based on multimodal time-series data solves the problems of single data utilization, insufficient mining of time-series evolution patterns, and unstable control in existing technologies. It achieves high-precision state identification and stable control of the cell culture process, improving the intelligence level and batch consistency of cell culture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINAN UNIVERSITY
- Filing Date
- 2026-04-24
- Publication Date
- 2026-05-29
AI Technical Summary
In existing cell culture technologies, data utilization methods are limited, the exploration of temporal evolution patterns is insufficient, the utilization depth of microscopic images is limited, and most regulation is weak closed-loop or semi-closed-loop, lacking the ability to predict future trends and control stability.
A cell culture dynamic regulation system employing multimodal time-series data is developed. By acquiring microscopic image sequences, culture environment parameter sequences, and historical control sequences, cell morphological evolution features are extracted using convolutional neural networks and ConvLSTM, and long-term and short-term parameter dependencies are extracted using a dual-scale Transformer. Response curves for proliferation activity, metabolic load, and environmental disturbances are constructed, and migration and abnormality risk indices for the culture stages are generated. Cell states and future trends are predicted using a stage-gated cross-modal attention network, and regulatory commands are output using a model predictive control algorithm with safety barrier constraints.
It achieves high-precision state identification, reliable trend prediction and stable control of the cell culture process, improves the intelligence level and batch consistency of the cell culture process, and avoids possible misjudgment and over-adjustment problems in traditional methods.
Smart Images

Figure CN122105016A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent regulation technology for cell culture, and more specifically to a dynamic regulation system for cell culture based on multimodal time-series data. Background Technology
[0002] Cell culture is a crucial foundational step in biopharmaceutical, regenerative medicine, cell therapy, vaccine development, organoid construction, and bioreactor production. During culture, the cell proliferation status, metabolic state, and responsiveness to external environmental disturbances directly affect cell viability, product expression levels, batch consistency, and final harvest quality. Current cell culture methods typically rely on incubators, bioreactors, and associated monitoring equipment to set or adjust parameters such as temperature, pH, dissolved oxygen, carbon dioxide concentration, feed rate, and medium exchange volume, supplemented by microscopic observation, offline sampling and testing, and empirical interpretation to assess the culture status.
[0003] However, existing technologies generally suffer from the following problems: First, the data utilization methods are relatively simplistic. Some schemes rely solely on a limited number of environmental parameters such as temperature, pH, and dissolved oxygen for threshold control. While some schemes incorporate microscopic images or metabolic indicators, these are typically only used as auxiliary monitoring information and are not closely linked to control decisions. This makes it difficult to comprehensively reflect complex states such as cell proliferation, aggregation, adhesion, metabolic load, and response to culture disturbances. Especially during cell culture, different modalities of data exhibit varying sampling frequencies, time delays, and noise levels. Without a unified temporal reconstruction and quality assessment mechanism, state identification biases are easily generated.
[0004] Secondly, existing methods fail to adequately explore the temporal evolution patterns of the culture process. Cell culture is not a static process but rather encompasses multiple stages, including inoculation adaptation, rapid expansion, homeostasis maintenance, decay, and harvest preparation. The sensitivity of cells to control actions such as feeding, medium replacement, gas supply, and agitation varies significantly across these stages. Existing control methods often employ fixed process curves, empirical rules, or simple feedback regulation, responding based on single-point detection results at the current moment. This makes it difficult to extract long-term dependent features from continuous image sequences, metabolic change trajectories, and control history, thus hindering the accurate assessment of stage migration trends and early signs of abnormal evolution.
[0005] Third, the application of microscopic images in cell state assessment is limited. Most existing methods rely on single-frame images for cell counting, confluence calculation, or morphological observation, which struggles to effectively address issues such as blurred cell boundaries, clustering and occlusion, uneven brightness, and subtle changes between adjacent time points. Furthermore, image features are often simply stitched together with environmental and metabolic parameters, lacking an adaptive fusion mechanism for changes during the culture phase. This results in insufficient mining of multimodal correlation information, affecting the accuracy of state prediction.
[0006] Fourth, most existing culture regulation still falls under weak closed-loop or semi-closed-loop control. In practical applications, although parameters such as feeding, medium replacement, gas supply, and stirring can be adjusted based on detection results, these adjustments usually rely on human experience and lack the ability to predict future culture trends, as well as a unified constraint on the safety boundaries of control actions. For example, when cell metabolic load increases, lactate accumulation accelerates, or external disturbances are too strong, if the state changes in subsequent periods are not comprehensively considered, it may lead to excessive feeding, excessive dissolved oxygen fluctuations, or increased shear damage, thereby affecting cell quality and culture stability.
[0007] Therefore, there is an urgent need for a unified temporal reconstruction and joint modeling system that can extract cell morphological evolution features and multi-parameter long- and short-term dependence features from microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences and historical control sequences, identify migration and abnormal risks in the culture stage based on multiple state feature curves, and further realize a prediction-driven, constraint-controllable closed-loop dynamic regulation system to improve the intelligence level, state recognition accuracy and regulation stability of the cell culture process. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention discloses a cell culture dynamic control system based on multimodal time-series data. This system acquires cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences. After time-series reconstruction, it uses convolutional neural networks and ConvLSTM to extract cell morphological evolution features, employs a dual-scale Transformer to extract long-term and short-term parameter dependencies, and constructs proliferation activity curves, metabolic load curves, and environmental disturbance response curves. Based on these, it generates culture stage migration indices and abnormal risk indices. Furthermore, it uses a stage-gated cross-modal attention network to predict cell state and future culture trends, and utilizes a model predictive control algorithm with safety barrier constraints to output feeding, medium replacement, gas supply, and stirring adjustment commands, achieving closed-loop dynamic control of the cell culture process. This system has the advantages of accurate state recognition, reliable trend prediction, and high control stability.
[0009] A cell culture dynamic regulation system based on multimodal time-series data, comprising: The data acquisition unit is used to acquire cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences. The time series reconstruction unit is used to align each sequence according to a unified time window and perform missing data compensation based on modal confidence to generate standardized time series data. The feature extraction unit is used to extract cell morphological evolution features using convolutional neural networks and ConvLSTM, and to extract long-term and short-term dependence features of environmental parameters, metabolic parameters and historical control quantities using dual-scale Transformer. The state characterization unit is used to construct a first characteristic curve characterizing cell proliferation activity, a second characteristic curve characterizing metabolic load, and a third characteristic curve characterizing environmental disturbance response. Based on the slope difference, inflection point spacing, and phase offset of the three characteristic curves, the culture stage migration index and abnormal risk index are generated. The fusion prediction unit is used to input the culture stage migration index and abnormal risk index into the stage-gated cross-modal attention network to predict cell state and future culture trends. The decision control unit is used to generate feeding, liquid replacement, gas regulation and stirring regulation commands based on the prediction results through a model predictive control algorithm with safety barrier constraints, and output them to the execution equipment to achieve closed-loop dynamic control.
[0010] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) This invention overcomes the problems of incomplete information, sensitivity to fluctuations, and high misjudgment rate caused by relying solely on single sensor data or single frame images to judge the culture state in existing technologies by simultaneously acquiring cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences, and first performing unified time window alignment and modal confidence compensation, and then extracting image-side morphological evolution features and parameter-side long-term and short-term dependence features respectively. In particular, by constructing proliferation activity curves, metabolic load curves, and environmental disturbance response curves, it is possible to comprehensively characterize cell proliferation capacity, metabolic pressure, and external control response features from a dynamic evolution perspective, thereby more accurately identifying culture stages and abnormal risks, and improving the ability to perceive the true state of cells.
[0011] (2) Based on multimodal feature extraction, this invention inputs the culture stage migration index and the abnormal risk index into a stage-gated cross-modal attention network, enabling the network to dynamically adjust the fusion weights of image modality and parameter modality according to different culture stages, thus avoiding the problem that traditional fixed fusion methods are difficult to adapt to changes in culture stages. Compared with the method of statically judging only the current state, this invention can jointly predict cell viability, cell density, lactate accumulation trend, stage transfer probability and abnormal risk level in at least two future prediction time domains. Therefore, it can detect potential problems such as metabolic imbalance, excessive perturbation, activity decay and abnormal contamination earlier, providing a more timely and reliable basis for subsequent regulation and improving the predictability and stability of the culture process.
[0012] (3) After completing the state prediction, this invention further employs a model predictive control algorithm with safety barrier constraints to generate feeding, liquid replacement, gas supply, and stirring adjustment commands based on future culture trends. It also sets preset safety domains for pH, dissolved oxygen, temperature, shear strength, feeding amount per unit time, and liquid replacement ratio per unit time, thereby avoiding problems such as excessive feeding, excessive gas supply fluctuations, or excessive stirring that exist in traditional empirical control. Simultaneously, this invention can also online correct modal confidence, stage gating vectors, and control optimization weights based on the deviation between the predicted and actual curves, and automatically switch to conservative control mode when the abnormal risk index exceeds a threshold. Therefore, it has better adaptive capabilities, control accuracy, and operational safety, which is beneficial for improving cell culture success rate and batch consistency. Attached Figure Description
[0013] Figure 1 This is a flowchart of a method for dynamic regulation of cell culture based on multimodal time-series data according to the present invention; Figure 2 This is a structural diagram of a cell culture dynamic regulation system module based on multimodal time-series data according to the present invention. Detailed Implementation
[0014] Those skilled in the art will understand that, in order to make the above-mentioned objects, features, and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This application illustrates a cell culture dynamic regulation system based on multimodal time-series data, comprising: The data acquisition unit is used to acquire cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences. In one embodiment, the data acquisition unit is deployed outside the cell culture chamber or microbioreactor and is communicatively connected to the microscopic imaging component, environmental sensing component, metabolic detection component, execution recording component and central controller. It is used to synchronously acquire cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences and historical control sequences throughout the cell culture process, and output them to the time sequence reconstruction unit according to a unified timestamp.
[0015] Specifically, the microscopic imaging assembly includes an inverted microscope, an industrial camera, an autofocus mechanism, and a ring-shaped supplementary light source. The inverted microscope is positioned at the bottom observation side of the culture vessel, and the industrial camera is mounted at the microscope eyepiece or imaging interface for periodic imaging of cells within culture dishes, cell culture flasks, microplates, or transparent reaction chambers. The autofocus mechanism is used for fine-tuning near a preset focal plane to reduce defocusing caused by liquid level fluctuations or thermal drift of the device. The ring-shaped supplementary light source provides constant illumination. A central controller controls the industrial camera to acquire images according to a preset sampling period, such as acquiring one frame every 30 seconds, 1 minute, or 5 minutes. During rapid proliferation, this can be increased to one frame every 20 seconds, thus forming a sequence of cell microscopic images. Each frame includes the acquisition time, culture batch number, field of view number, focal length, and exposure parameters. To improve representativeness, 3 to 12 fixed fields of view can be preset within a culture vessel. The central controller sequentially acquires images from each field of view using a polling method to form a multi-field microscopic image sequence.
[0016] The environmental sensing components include a temperature sensor, a pH sensor, a dissolved oxygen sensor, a carbon dioxide concentration sensor, a liquid level sensor, a pressure sensor, and a stirring speed acquisition device. These sensors are respectively installed in the incubator, reactor chamber, gas supply line, or stirring drive end, and are connected to the central controller via RS485 bus, CAN bus, or Ethernet. The central controller collects various environmental parameters according to a preset cycle; for example, temperature, pH, and dissolved oxygen are collected every 5 seconds; carbon dioxide concentration, liquid level, and pressure are collected every 10 seconds; and stirring speed is read in real time or recorded every 1 second, thus forming a sequence of culture environment parameters. To avoid false sampling caused by short-term sensor jitter, the central controller can attach sensor status flags to the raw sampled values, including normal, calibration in progress, drift alarm, and offline status.
[0017] The metabolic detection component includes an online biochemical analysis module or an automated sampling and detection module. The online biochemical analysis module can employ a microfluidic detection chip, ion-selective electrode module, near-infrared detection module, or enzyme reaction detection module to detect glucose concentration, lactate concentration, glutamine concentration, ammonium ion concentration, and osmotic pressure in the culture medium. The automated sampling and detection module includes a micro-sampling pump, sampling needle, detection cell, and waste liquid recovery bottle. A central controller controls it to automatically sample according to preset time intervals, such as collecting metabolic samples every 10, 20, or 30 minutes, completing online detection, and outputting the detection results and corresponding timestamps. For metabolic indicators with longer detection times, the sampling time and result return time can also be recorded for subsequent time-series reconstruction unit to perform time alignment and delay correction.
[0018] The execution recording component is used to collect historical control sequences, specifically including feed pump operation records, liquid exchange valve opening and closing records, air supply flow rate setting records, stirring speed adjustment records, chemical dosing records, and alarm intervention records. The central controller automatically records the action trigger time, duration, target setpoint, actual feedback value, and action source for each control action executed. The action source includes automatic control, manual intervention, and fault protection. Taking feed as an example, the central controller records the start and end times of feed, feed volume, feed flow rate, and pH and dissolved oxygen changes before and after feed; taking liquid exchange as an example, it records the discharge volume, inflow volume, liquid exchange ratio, and valve response time; taking air supply regulation as an example, it records the oxygen flow rate, carbon dioxide flow rate, and air-to-water ratio; taking stirring regulation as an example, it records the speed before adjustment, the speed after adjustment, and the duration of the adjustment. These records, arranged chronologically, form a historical control sequence.
[0019] In a specific culture implementation, taking adherent mammalian cell culture as an example, the central controller is set to acquire microscopic images every 1 minute, temperature, pH, and dissolved oxygen every 5 seconds, liquid level and pressure every 10 seconds, and glucose, lactate, and glutamine every 20 minutes. Feeding, medium replacement, and gas supply actions are recorded immediately upon occurrence. After the culture begins, the data acquisition unit runs continuously for 48 to 120 hours, obtaining a raw multimodal dataset containing multi-field microscopic images, environmental parameters, metabolic parameters, and control action logs. This dataset is then uniformly packaged into timestamped data packets and sent to the time-series reconstruction unit.
[0020] Furthermore, the data acquisition unit also includes a local caching module and a communication verification module. The local caching module temporarily stores the raw data when the network is interrupted or the detection module is briefly offline; the communication verification module adds a device number, checksum, and timestamp consistency mark to each data packet to ensure the integrity and traceability of subsequent multimodal data fusion.
[0021] The time series reconstruction unit is used to align each sequence according to a unified time window and perform missing data compensation based on modal confidence to generate standardized time series data. In one embodiment, the time-series reconstruction unit is located in a central controller or an independent data processing server, and is communicatively connected to the data acquisition unit. It receives cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences. It then performs unified timeline construction, timestamp correction, window alignment, modality confidence calculation, missing data compensation, and standardized encapsulation processing on multimodal sequences with different sampling frequencies, time delays, and data qualities to output standardized time-series data for subsequent feature extraction units. Specifically, the time-series reconstruction unit includes a reference timeline generation subunit, a timestamp correction subunit, a window resampling subunit, a modality confidence assessment subunit, a missing data compensation subunit, and a standardized encapsulation subunit.
[0022] The reference timeline generation subunit generates a global reference timeline based on the start time of the culture task, the total culture duration, and a preset uniform time window length. In one specific implementation, the uniform time window length is set to 1 minute, meaning that starting from the culture start time, the entire culture process is divided into multiple continuous time windows with a step size of 1 minute, and each time window is recorded as a standard time slice. For example, for a 72-hour culture process, 4320 standard time slices can be generated, each with a unique time index. For fast-moving scenarios, the uniform time window can also be set to 30 seconds; for scenarios with slower metabolic changes, it can be set to 5 minutes. Preferably, the uniform time window length is determined jointly based on the image sampling period and the control response time to balance temporal resolution and computational burden.
[0023] The timestamp correction subunit is used to eliminate clock skew and detection delays between different acquisition devices. Specifically, for microscopic image sequences, the image acquisition trigger time is used as the original timestamp; for environmental parameter sequences, the sensor reporting time is used as the original timestamp; for metabolic parameter sequences, both the sampling time and the detection result return time are recorded, with the sampling time as the main timestamp and the detection result return time as the delay marker; for historical control sequences, the control command issuance time, execution start time, and execution completion time are recorded separately. The timing reconstruction unit pre-stores communication delay compensation tables and calibration offset tables for each device. For example, the microscope camera has a fixed delay of 0.2 seconds, the online biochemical analysis module has a fixed delay of 90 seconds, and the valve pump execution feedback delay is 1 to 3 seconds. Based on these, the original timestamps are corrected to unify the multimodal data under the same system time reference.
[0024] The window resampling subunit maps the corrected modal data to the unified time window. For high-frequency environmental parameter data, such as temperature, pH, dissolved oxygen, and stirring speed, one or more of the following can be calculated as window characterization values within each standard time window: average, maximum, minimum, fluctuation amplitude, and slope of change. For microscopic image sequences, if multiple frames exist within a certain time window, the clearest image is selected as the master image, or the mean and variation features of all images within the window are extracted; if no image exists within the time window, it is marked as missing. For metabolic parameter sequences, due to the low sampling frequency, such as once every 20 minutes, each detected value is first mapped to an adjacent time window, and then interpolation estimation is performed by combining the detected values before and after. For historical control sequences, if a control action occurs within a certain time window, the action type, action intensity, duration, and deviation before and after the action are recorded within that time window; if an action spans multiple time windows, it is allocated to the corresponding time windows according to the time proportion. After the above processing, each standard time slice forms a set of aligned multimodal window data.
[0025] The modal confidence assessment subunit is used to calculate the confidence level of each modal data within each time window. In one embodiment, the modal confidence level is jointly determined by sampling integrity, noise level, drift degree, and time delay reliability. Sampling integrity characterizes whether valid data is obtained within the time window; noise level characterizes the degree of sensor value jitter, image blurring, or detection anomalies; drift degree characterizes the deviation between the current data and the recent stable average; and time delay reliability characterizes the availability of the detection result or feedback result relative to the actual sampling time. For example, when the pH sensor has continuous sampling within a certain time window and the fluctuation is within the normal range, its modal confidence level can be set to high; when the edge sharpness of the microscopic image decreases due to defocusing, the image modal confidence level decreases; when the metabolic detection result has a timeout return or the difference between two consecutive detections increases abnormally, the metabolic modal confidence level decreases. Preferably, the modal confidence level ranges from 0 to 1 and is used as a weighting basis for subsequent missing value compensation and feature extraction.
[0026] The missing data compensation subunit employs differentiated compensation strategies for different modalities. For environmental parameter sequences, if the missing duration does not exceed three standard time windows, compensation is performed using linear interpolation, spline interpolation, or Kalman filter prediction of adjacent effective values. If the missing duration exceeds three standard time windows, conditional estimation is performed by combining historical control sequences and associated environmental variables, such as estimating dissolved oxygen values based on changes in gas supply and historical dissolved oxygen response. For metabolic parameter sequences, trend interpolation based on preceding and following detection points is preferred, with compensation values corrected by environmental parameter changes and feeding records superimposed. For example, the compensation value for glucose concentration after feeding is not directly calculated using linear interpolation, but rather estimated by constraint combining feeding volume, feeding concentration, and the previous detection value. For microscopic image sequences, if an image is missing in a single time window, it is filled using the image feature vectors of adjacent effective time windows. If multiple time windows are missing consecutively, images are not directly forged; instead, the image modality is marked as a low-confidence state, and its weight is reduced by subsequent feature extraction units. For historical control sequences, if feedback records are missing, retrospective recording is performed based on controller output logs and actuator status logs. Preferably, the missing compensation subunit does not simply replace the missing value, but simultaneously outputs "compensation value + compensation source label + compensation confidence level", so that the subsequent model can distinguish between real observation data and estimated data.
[0027] The standardized encapsulation subunit performs scale and format unification processing on the aligned and compensated multimodal data within each time window. Specifically, it performs min-max normalization, Z-score normalization, or relative change normalization based on batch baseline for continuous numerical data such as temperature, pH, dissolved oxygen, glucose concentration, and lactate concentration; it performs discrete encoding or continuous quantity normalization encoding for historical control actions; it retains the original image index for image data and outputs the image feature occupancy structure and image confidence markers for the corresponding time window. Finally, each standard time slice outputs a standardized time-series data package containing the following: time window index, image data or image feature index, environmental parameter vector, metabolic parameter vector, control action vector, confidence vector for each modality, missing compensation marker vector, and quality status marker. All time slices are sequentially concatenated to form a standardized time-series data sequence and sent to the feature extraction unit.
[0028] In one specific embodiment, the total cell culture time is 96 hours. Microscopic images are acquired one frame every minute; temperature, pH, and dissolved oxygen are sampled every 5 seconds; liquid level and pressure are sampled every 10 seconds; glucose and lactate are measured every 20 minutes; and feeding and liquid replacement actions are recorded based on event triggers. The temporal reconstruction unit sets a unified time window of 1 minute and performs the following operations for each 1-minute time window: averages and fluctuations of temperature, pH, and dissolved oxygen within that minute are calculated; the clearest microscopic image frame within that minute is selected; glucose and lactate values measured every 20 minutes are interpolated and distributed to 20 adjacent time windows based on preceding and following detection points; if feeding occurs in that minute, the feeding amount, feeding rate, and duration of the action are recorded; if image acquisition fails in that minute, the image is marked as missing and features from the previous valid time window are used as temporary placeholders. Subsequently, confidence levels for image modality, environmental modality, metabolic modality, and control modality are generated based on image clarity, sensor fluctuations, detection latency, and data completeness, respectively, ultimately forming a standardized temporal data sequence suitable for subsequent modeling.
[0029] By using the aforementioned temporal reconstruction method, multimodal raw data with different sampling frequencies, delays, and qualities can be uniformly mapped to the same time base. Modal confidence scores are used to distinguish between high-confidence observations and low-confidence compensation data, thereby providing temporally consistent, structurally standardized, and quality-controllable input data for subsequent convolutional neural networks, ConvLSTM, dual-scale Transformers, and stage-gated cross-modal attention networks. This improves the accuracy and stability of cell state recognition, trend prediction, and closed-loop regulation.
[0030] The feature extraction unit is used to extract cell morphological evolution features using convolutional neural networks and ConvLSTM, and to extract long-term and short-term dependence features of environmental parameters, metabolic parameters and historical control quantities using dual-scale Transformer. In one embodiment, the feature extraction unit is located in a data processing server and is communicatively connected to the temporal reconstruction unit, the state representation unit, and the fusion prediction unit. It is used to extract cell morphological evolution features of image modalities, as well as long- and short-term dependence features of environmental parameters, metabolic parameters, and historical control variables from standardized time-series data. The feature extraction unit includes an image temporal feature extraction subunit and a parametric temporal feature extraction subunit. The image temporal feature extraction subunit employs a convolutional neural network, a boundary-enhanced temporal attention layer, and a ConvLSTM cascaded structure, while the parametric temporal feature extraction subunit employs a dual-scale Transformer structure.
[0031] The overall workflow is as follows. In one specific implementation, the feature extraction unit processes multimodal temporal data using a sliding window approach. An example is taken using eight consecutive standard time windows as one image temporal segment and thirty consecutive standard time windows as one parameter temporal segment. First, the image temporal feature extraction subunit reads the corresponding microscopic image frames from multiple consecutive time windows, performing size unification, brightness normalization, and noise suppression on each frame. Then, a convolutional neural network is used to extract spatial features such as cell boundaries, cluster contours, adherence areas, texture density, and local morphological differences in each frame. Next, the spatial features output by the convolutional neural network are fed into a boundary enhancement temporal attention layer to enhance areas with blurred cell boundaries, cluster occlusion, and subtle changes in adjacent time intervals. Finally, the enhanced frame-by-frame features are input into a ConvLSTM in chronological order to extract the evolutionary features of cell morphology over continuous time, resulting in a cell morphology evolution feature sequence and its aggregated feature vector. Simultaneously, the parameter temporal feature extraction subunit receives environmental parameter sequences, metabolic parameter sequences, and historical control quantity sequences, maps them into parameter vector sequences of a unified dimension, and feeds them into the short-scale Transformer branch and the long-scale Transformer branch, respectively. The short-scale Transformer is used to extract local fluctuations, transient perturbations, and short-term response features after control actions, while the long-scale Transformer is used to extract changes during the cultivation stage, metabolic trend evolution, and the cumulative effect of historical control. The outputs of the two branches are concatenated, linearly mapped, and gated fusion to form the parameter-side long- and short-term dependency features. The image-side output and the parameter-side output are then sent to the state representation unit and the fusion prediction unit.
[0032] In one embodiment, the input sequence of microscopic images to the image temporal feature extraction subunit consists of 8 or 12 consecutive frames, with each frame having a uniform resolution of 256×256 pixels. If the original image is a grayscale image, the input dimension is 256×256×1; if the original image is a color image, the input dimension is 256×256×3. Each image sequence corresponds to a time window segment and includes a corresponding image modality confidence score.
[0033] The specific structure of the convolutional neural network is as follows. In one specific embodiment, the convolutional neural network adopts a four-level convolutional feature extraction structure, including an input convolutional layer, a first convolutional block, a second convolutional block, a third convolutional block, and a fourth convolutional block. The input convolutional layer is used to perform preliminary feature mapping on the original image. It uses a convolution operation with a kernel size of 3×3, a stride of 1, and 32 channels, followed by a batch normalization layer and a ReLU activation layer to output a primary feature map with a size of 256×256×32. The first convolutional block includes two cascaded 3×3 convolutional layers, each followed by a batch normalization layer and a ReLU activation layer. The output channel number of this convolutional block is 32, and the feature map is downsampled to 128×128×32 through a 2×2 max pooling layer for extracting local cell edges, textures, and initial contour information. The second convolutional block consists of two cascaded 3×3 convolutional layers with 64 output channels, followed by a 2×2 max-pooling layer, downsampling the feature map to 64×64×64 for extracting cell cluster contours, adherence boundaries, and local aggregation morphology. The third convolutional block consists of two cascaded 3×3 convolutional layers with 128 output channels, followed by a 2×2 max-pooling layer, downsampling the feature map to 32×32×128 for extracting clonal expansion morphology, intercellular space distribution, and regional texture density variations. The fourth convolutional block consists of two cascaded 3×3 convolutional layers with 256 output channels. Pooling is not performed; instead, the feature size remains 32×32×256 for outputting a high semantic spatial feature map. Preferably, a residual connection is added between the second and third convolutional blocks to reduce gradient decay caused by deep convolutions; a 1×1 convolutional compression layer is added at the end of the fourth convolutional block to compress the number of channels from 256 to 128, reducing the computational cost of subsequent temporal modeling.
[0034] The specific structure of the boundary-enhanced temporal attention layer is as follows: To improve the expression of regions with blurred cell boundaries, clustering occlusion, and subtle morphological changes, a boundary-enhanced temporal attention layer is inserted between the convolutional neural network and the ConvLSTM. This boundary-enhanced temporal attention layer includes a multi-scale spatial feature extraction branch, an adjacent frame difference branch, and a channel-space joint attention branch. The multi-scale spatial feature extraction branch performs two parallel operations on the current frame feature map output by the convolutional neural network: the first path uses a 3×3 standard convolution to extract conventional edge features; the second path uses a 3×3 dilated convolution with a dilation rate of 2 to expand the receptive field and extract the outer edges of cell clusters and local expansion contours. The two outputs are concatenated along the channel dimension and then compressed into a uniform number of channels by a 1×1 convolution to obtain multi-scale boundary features. The adjacent frame difference branch receives the convolutional feature map of the current time t and the convolutional feature map of the previous time t-1, performs element-wise difference operation to obtain the time difference feature map, and then maps it to the time difference response feature map through 1×1 convolution and ReLU activation, which is used to highlight subtle morphological changes such as cell boundary expansion, aggregation and reconstruction, contraction and regression, and local activity changes. The channel-spatial joint attention branch fuses the multi-scale boundary features and the time difference response features. First, it generates channel weights through global average pooling and global max pooling, and then generates spatial weights through 7×7 convolution, finally forming a joint attention map. This joint attention map is multiplied element-wise with the original convolutional feature map and added element-wise with the time difference response feature to obtain the enhanced image feature map. The size of the enhanced image feature map is preferably 32×32×128. The purpose of this structure is to enhance features that are more sensitive to proliferation boundaries, aggregation and separation regions, and short-term change regions before feeding them into ConvLSTM, thereby improving the effectiveness of temporal modeling.
[0035] In one embodiment, the ConvLSTM employs a two-layer stacked structure. The first ConvLSTM layer receives the enhanced image feature sequence, with an input sequence length of 8 frames, an input tensor size of 8×32×32×128, a convolution kernel size of 3×3, 128 hidden state channels, a stride of 1, and a same padding method. This layer primarily captures short-term continuous changes in cell morphology, such as cell expansion, local edge advancement, and clustering / reconstruction trends within adjacent time windows. The second ConvLSTM layer receives the hidden state sequence from the first ConvLSTM layer, also with a 3×3 convolution kernel and 64 hidden state channels. It is used to further aggregate higher-level temporal dependencies, such as the continuation of amplification trends, the recovery process after perturbation, and continuous micro-change signals before stage switching. Preferably, a normalization layer is inserted between the two ConvLSTM layers to reduce the fluctuation of image feature scale between different batches; after the second ConvLSTM layer, temporal global average pooling and spatial global average pooling are added to obtain a one-dimensional image temporal feature vector. This vector can be further mapped through a fully connected layer to a fixed-dimensional cell morphological evolution feature, such as 128-dimensional or 256-dimensional.
[0036] The image temporal feature extraction subunit outputs two types of results: first, image temporal feature vectors that can be directly used by subsequent state characterization units; second, cell morphological statistical features used to construct the first feature curve, including cell confluence increment, cell count growth rate, clonal area expansion rate, cell motility, and cluster compactness. Specifically, the cell confluence increment can be calculated from the difference in cell coverage area between adjacent time-series images; the cell count growth rate can be obtained from the rate of change of instance segmentation or density estimation results within adjacent time windows; the clonal area expansion rate can be obtained from the rate of change of connected region area; cell motility can be characterized by a combination of contour displacement and local texture changes between adjacent frames; and cluster compactness can be determined by indicators such as the ratio of cluster area to its convex hull area and boundary curvature changes.
[0037] In one embodiment, the parameter temporal feature extraction subunit receives environmental parameters, metabolic parameters, and historical control values within 30 consecutive standard time windows. Each time window corresponds to a set of parameter vectors, including, for example, temperature, pH, dissolved oxygen, carbon dioxide concentration, liquid level, glucose concentration, lactate concentration, glutamine concentration, feed rate, liquid exchange ratio, gas supply rate, and stirring speed. The parameters are standardized and concatenated to form a d-dimensional input vector, for example, d = 16 or 24.
[0038] The specific structure of the dual-scale Transformer is as follows: it includes a short-scale Transformer branch and a long-scale Transformer branch. The short-scale Transformer branch receives parameter vector sequences from the most recent 6 to 10 time windows to extract short-term fluctuation features. This branch includes an input embedding layer, a position encoding layer, and two Transformer encoder layers. Each encoder layer includes a multi-head self-attention module and a feedforward network module. The number of multi-heads is preferably 4, and the hidden dimension is preferably 128. This branch focuses on learning the local responses of pH, dissolved oxygen, temperature, and liquid level after a control action, as well as the short-term changes in glucose and lactate. The long-scale Transformer branch receives parameter vector sequences from the most recent 20 to 30 time windows to extract phased trend features. This branch also includes an input embedding layer, a position encoding layer, and three Transformer encoder layers. The number of multi-heads is preferably 8, and the hidden dimension is preferably 128 or 256. This branch focuses on learning metabolic load accumulation, phased expansion trends, the impact of long-term gas supply strategies, and the hysteresis effect of control quantities. The short-scale Transformer branch outputs short-term dependency features, while the long-scale Transformer branch outputs long-term dependency features. The outputs of the two branches are concatenated, mapped to a unified dimension through a linear mapping layer, and then passed through a gated fusion layer that assigns different weights based on the current modality confidence and the activity of the control action, outputting a parameter-side long and short-term dependency feature vector.
[0039] The parameter-side feature outputs are as follows: the parameter temporal feature extraction subunit outputs a fixed-dimensional parameter temporal feature vector, and further outputs intermediate statistics used to construct the second and third feature curves. The intermediate statistics corresponding to the second feature curve include glucose consumption rate, lactate production rate, glutamine consumption rate, pH correction frequency, and dissolved oxygen compensation amount; the intermediate statistics corresponding to the third feature curve include pH recovery time after control actions, dissolved oxygen overshoot, temperature fluctuation amplitude, and recovery delay after stirring disturbances.
[0040] In one specific embodiment, the workflow of the feature extraction unit is as follows: Step S31: Receive standardized temporal data output by the temporal reconstruction unit, and generate image segments and parameter segments according to the image sliding window and parameter sliding window; Step S32: Perform size unification, brightness normalization, and noise suppression on each frame of the microscopic image in the image segment; Step S33: Input each preprocessed frame of the image into a convolutional neural network, and extract multi-level spatial features through a four-level convolutional structure; Step S34: Input the feature map output by the convolutional neural network into the boundary enhancement time difference attention layer, extract boundary enhancement features and time difference response features through multi-scale boundary branches and adjacent frame difference branches, and form an enhanced feature map through joint attention weighting; Step S35: Calculate the enhanced feature map over continuous time... The sequence is input into a two-layer stacked ConvLSTM to extract cell morphological evolution features and output an image temporal feature vector; Step S36: Environmental parameters, metabolic parameters, and historical control variables are concatenated to form a parameter sequence, which is then input into the short-scale Transformer branch and the long-scale Transformer branch respectively to extract short-term and long-term parameter dependency features; Step S37: The short-term and long-term dependency features are concatenated and gated to obtain a parameter-side long-short-term dependency feature vector; Step S38: The image temporal feature vector and the parameter-side long-short-term dependency feature vector are output to the state representation unit and the fusion prediction unit for subsequent construction of three feature curves, generation of migration index and abnormal risk index during culture stage, and prediction of cell state.
[0041] The core principle of this feature extraction structure lies in two aspects: First, microscopic image sequences are data with high spatial correlation and strong local texture dependence. Ordinary temporal networks alone cannot accurately extract cell boundaries, aggregation morphology, and continuous expansion trajectories. Therefore, a convolutional neural network is first used to extract spatial features, followed by a boundary-enhancing temporal attention layer to strengthen the expression of cell boundaries and regions of subtle changes. Finally, a ConvLSTM is used to extract temporal evolution patterns. Second, environmental parameters, metabolic parameters, and historical control variables are low-dimensional, multivariate temporal data, exhibiting both short-term perturbation responses and long-term trends. Therefore, a dual-scale Transformer is used to learn local fluctuations and long-term dependencies separately, and then fused to more comprehensively characterize the temporal changes of these parameters. This structure improves the ability to identify cell boundary blurring, aggregation occlusion, weak expansion, and perturbation recovery processes, while enhancing the modeling ability for changes in metabolic load and the long-term effects of control actions. This provides more discriminative and stable input features for subsequent state representation, trend prediction, and closed-loop regulation.
[0042] The state characterization unit is used to construct a first characteristic curve characterizing cell proliferation activity, a second characteristic curve characterizing metabolic load, and a third characteristic curve characterizing environmental disturbance response. Based on the slope difference, inflection point spacing, and phase offset of the three characteristic curves, the culture stage migration index and abnormal risk index are generated. In one embodiment, the state characterization unit is located in a data processing server and is communicatively connected to the feature extraction unit and the fusion prediction unit. It receives temporal features of images, temporal features of parameters, and intermediate statistics generated during feature extraction. These features are further organized into three characteristic curves that reflect the evolution of the culture state. The three characteristic curves are: a first characteristic curve characterizing cell proliferation activity, a second characteristic curve characterizing metabolic load, and a third characteristic curve characterizing the response to environmental disturbances. After constructing the three characteristic curves, the state characterization unit calculates the differences in the rate of change, the order of transitions, and the overall lag relationship among the three curves to generate a culture stage migration index and an abnormal risk index, which are then output to the subsequent fusion prediction unit.
[0043] In this embodiment, the state characterization unit includes a curve construction subunit, a curve smoothing and normalization subunit, a curve relationship analysis subunit, and an exponential generation subunit. These subunits are connected sequentially, with the output of one subunit serving as the input of the next. This embodiment does not rely solely on a single detection value at a given moment; instead, it organizes multiple intermediate features reflecting cell state into a continuously changing curve on a unified time axis. The purpose of this is to transform originally discrete, localized, and noise-sensitive observations into comparable state trajectories. By comparing the relative speed, sequence, and synchronicity of different state trajectories, changes and abnormal trends during the culture phase can be identified more accurately. Specifically, the first characteristic curve primarily reflects changes in cell self-expansion capacity and morphological activity; the second characteristic curve primarily reflects nutrient consumption, accumulation of metabolic byproducts, and environmental compensation burden; and the third characteristic curve primarily reflects the system's recovery speed, recovery amplitude, and disturbance tolerance after external control actions. These three curves change together on the same time axis, constituting a dynamic characterization of the cell culture process.
[0044] The specific construction embodiment of the first characteristic curve is as follows. The first characteristic curve is used to characterize cell proliferation activity. In one embodiment, the state characterization unit receives the following intermediate statistics from the image temporal feature extraction subunit: cell confluence increment, cell count growth rate, clonal area expansion rate, cell motility, and change in aggregation compactness. These indicators reflect the cell expansion state from different perspectives. The cell confluence increment reflects the increase in cell coverage on the culture surface; the cell count growth rate reflects the change in the number of cells per unit time; the clonal area expansion rate reflects the trend of cell community expansion; cell motility reflects cell boundary advancement and local dynamics; and the change in aggregation compactness reflects whether the aggregation region tends to expand, compress, or reconstruct. When constructing the curve, the state characterization unit first uses a unified time window as the basic unit to statistically analyze and align the above indicators within each time window. For example, if the unified time window is one minute, a set of proliferation-related feature values is formed from the image analysis results within each minute. Subsequently, the features are scaled uniformly so that features of different dimensions can be fused in the same curve. Scale unification can be achieved by normalizing based on historical ranges or by mapping intervals based on empirical upper and lower limits corresponding to cell types. After scale unification, the state characterization unit performs weighted fusion of various proliferation-related features according to preset weights to obtain the proliferation activity value corresponding to each time window. Preferably, for adherent amplifying cells, the weights of cell confluence increment and clonal area expansion rate can be increased; for cells in suspension culture or growing in clusters, the weights of cell count growth rate and change in cluster compactness can be appropriately increased; for cells that are more sensitive to migration, the weight of cell motility activity can be increased. After obtaining the proliferation activity values for each time window, the curve smoothing and normalization subunit further performs moving average, exponential smoothing, or local polynomial smoothing on the value sequence to reduce short-term jitter caused by single-frame recognition errors, focal plane fluctuations, and illumination fluctuations, ultimately forming a continuous and smooth first feature curve. The higher the value of this curve, the stronger the cell proliferation activity and the better the amplification capacity; a plateau or decline in the curve indicates that cell amplification is slowing down or activity is weakening.
[0045] A specific embodiment of the construction of the second characteristic curve, used to characterize metabolic load. In one embodiment, the state characterization unit receives the following intermediate statistics from the parameter temporal feature extraction subunit: glucose consumption rate, lactate production rate, glutamine consumption rate, pH correction frequency, and dissolved oxygen compensation. Among these indicators, glucose consumption rate and glutamine consumption rate reflect the level of nutrient consumption, lactate production rate reflects the accumulation rate of metabolic byproducts, pH correction frequency reflects the regulatory burden borne by the system to maintain acid-base balance, and dissolved oxygen compensation reflects the degree of additional compensation required by the system to maintain the target oxygen supply. The state characterization unit also integrates the above indicators according to a unified time window. For metabolic indicators with low sampling frequency, the window interpolation results output by the temporal reconstruction unit can be directly used; for control compensation-related indicators, the number of compensations, total compensation, or compensation magnitude within the corresponding time window is statistically analyzed. After completing the time window mapping, the above indicators are scaled uniformly so that they can reflect the degree of metabolic load within a unified range. Then, the state characterization unit merges these metabolic-related indicators according to preset weights to obtain the metabolic load value corresponding to each time window. Preferably, in scenarios requiring increased sensitivity to metabolic imbalances, the weights of lactate production rate and pH correction frequency can be increased; in oxygen-sensitive cell cultures, the weight of dissolved oxygen compensation can be appropriately increased. By concatenating metabolic load values over a continuous time window, a second characteristic curve is formed. Further, a curve smoothing and normalization subunit smooths the second characteristic curve to reduce spikes caused by single-detection bias or individual sampling errors. A higher value on the smoothed second characteristic curve indicates greater metabolic pressure, faster nutrient consumption, and a heavier environmental compensation burden during cell culture. If this curve continues to rise and its rate of increase is faster than that of the first characteristic curve, it typically indicates that the system is transitioning from high-efficiency amplification to a high-load operating state.
[0046] A specific embodiment of the construction of the third characteristic curve, which is used to characterize the response to environmental disturbances. In one embodiment, the state characterization unit receives intermediate statistics related to control actions and system recovery processes, including: pH deviation recovery time after a control action, dissolved oxygen overshoot, temperature fluctuation amplitude, cell morphology recovery delay after agitation, and image activity recovery rate. Specifically, pH deviation recovery time represents the time required for pH to recover from a deviation state to the target stable range after a certain control action is triggered; dissolved oxygen overshoot represents the degree to which the dissolved oxygen value exceeds the target value after gas supply or agitation adjustment; temperature fluctuation amplitude represents the temperature fluctuation range before and after the control action; morphology recovery delay represents the time required for cell image features to recover to the stable state before the disturbance after disturbances such as liquid replacement, feeding, or agitation; and image activity recovery rate represents the proportion of cell boundary advancement, local movement, or texture activity that recovers to the baseline level after disturbance. Since the goal of the third characteristic curve is to reflect the "disturbance burden," during its construction, indicators such as recovery time, overshoot, fluctuation amplitude, and recovery delay—where "larger values indicate more unfavorable conditions"—are directly mapped forward. Conversely, indicators such as image activity recovery rate—where "larger values indicate better recovery"—are mapped backward, so that a higher recovery rate in the curve represents a smaller contribution to the disturbance burden. The state representation unit organizes these indicators according to a unified time window, performs scale unification, and then fuses the environmental disturbance response values corresponding to each time window according to preset weights. Subsequently, the third characteristic curve is formed through time series concatenation and smoothing. The higher the curve, the slower the system's recovery from control actions or environmental disturbances, the greater the fluctuations, and the worse the cell's tolerance to external stimuli. If the curve suddenly rises in a short period, it usually indicates a decrease in system stability or an aggravation of the disturbance's impact.
[0047] To facilitate subsequent comparisons, the three feature curves, after their individual construction and smoothing, require a unified processing step. In one embodiment, the state characterization unit maps the three curves to intervals based on the historical minimum and maximum values of the current batch, or to the empirical range of the corresponding cell type, ensuring that the three curves ultimately fall within the same numerical range. After unified processing, although the three curves have different physical meanings, they can be compared in terms of position, trend, and sequence on the same time axis. In a specific implementation of slope difference calculation, the slope difference is used to characterize the difference in the rate of change of two curves within the same time period. In one embodiment, the curve relationship analysis subunit sets a local analysis window for each curve, for example, taking the current moment and three or five consecutive time windows prior to it. For each curve, within this local analysis window, the average rate of change of the curve value over time is calculated, and this average rate of change is used as the local slope at the current time point. After obtaining the local slopes of the first, second, and third feature curves, the slope difference between the first and second, the slope difference between the first and third, and the slope difference between the second and third are calculated respectively. The specific method is as follows: Subtract the local slope values of the two curves and take the absolute value to obtain the corresponding slope difference. If the slope of the first characteristic curve is significantly greater than the slope of the second characteristic curve, it indicates that cell activity increases faster than metabolic burden increases, usually corresponding to a good expansion state. If the slope of the second characteristic curve is significantly greater than the slope of the first characteristic curve, it indicates that metabolic pressure increases faster than cell activity increases, potentially indicating a risk of metabolic imbalance. If the slope of the third characteristic curve suddenly increases in a short period and forms a large slope difference with the first characteristic curve, it usually indicates that the system disturbance response is intensifying and beginning to affect the cell state. Continuous monitoring of the slope difference can identify trend reversal points and potential precursors of abnormalities during the culture process.
[0048] In a specific implementation of the inflection point spacing calculation, the inflection point spacing is used to characterize the sequential relationship of trend reversals among different curves. In one embodiment, the curve relationship analysis subunit performs local trend analysis on each curve to determine whether the curve has undergone trend changes such as changing from rising to falling, from falling to rising, or from rapid rising to plateauing within the current observation window. When the direction of curve change changes before and after a certain time point, and the magnitude of the change exceeds a preset threshold, that time point is determined to be an inflection point of the curve. After obtaining the most recent valid inflection point of the three curves, the state characterization unit calculates the time intervals between the first and second, first and third, and second and third sets of inflection points, respectively. This time interval is the inflection point spacing. If the inflection point of the second characteristic curve appears earlier than that of the first characteristic curve, it indicates that changes in metabolic stress precede changes in cell proliferation, potentially suggesting that metabolic problems have manifested earlier. If the inflection point of the third characteristic curve appears immediately after the second characteristic curve, it indicates that control compensation or environmental disturbances are beginning to have a significant impact on the system. If the first characteristic curve shows a downward inflection point first, while the second and third curves are still in the upward phase, it indicates that cell activity is beginning to decline, but the system burden is still increasing, typically corresponding to a high-risk transition phase. By analyzing the intervals between inflection points, the temporal relationship of "which changes first and which changes later" in the culture system can be identified, rather than simply looking at the current values of each curve. Therefore, this approach is more suitable for determining phase transitions and abnormal propagation paths.
[0049] In a specific implementation of phase offset calculation, the phase offset is used to characterize the time lag relationship between the overall changing trends of two curves. In one embodiment, the curve relationship analysis subunit shifts a curve forward or backward along the time axis within a preset time range, and calculates the similarity between the two curves at each shift position. The similarity can be represented by the degree of curve overlap, the degree of synchronous change, or the correlation score. When the similarity between the two curves reaches its maximum at a certain shift amount, that shift amount is taken as the phase offset between the two curves. If the first characteristic curve needs to be shifted forward relative to the second characteristic curve to achieve optimal overlap, it indicates that the change in metabolic load lags behind the change in cell activity; if the third characteristic curve changes earlier relative to the second characteristic curve, it indicates that the response to environmental disturbances precedes the change in metabolic load. The larger the phase offset, the higher the degree of asynchronous change between the two curves, the weaker the coupling between the internal states of the system, or the more obvious the anomaly transmission. Through phase offset analysis, it is possible not only to determine whether the two curves change synchronously, but also to determine how long a state change will lag before being transmitted to another state, which is very important for early prediction of anomalies. For example, if the metabolic load curve shows a continuously widening lag offset relative to the proliferation activity curve, it indicates that a situation may have occurred in the culture system where "the surface is still expanding, but metabolic risks have already accumulated."
[0050] In one embodiment, the generation of the cultivation stage migration index involves the index generation subunit using three sets of slope differences, three sets of inflection point intervals, three sets of phase offsets, and the current absolute levels of the three curves as input to generate the cultivation stage migration index. Specifically, the various relationship indicators are first mapped to a unified range to allow for comparison, and then weighted and fused according to preset weights to obtain a continuously changing stage migration index value. This stage migration index reflects whether the current cultivation process is closer to the adaptation phase, amplification phase, steady-state phase, or harvest preparation phase. If the first characteristic curve is in an upward phase, the second characteristic curve is rising slowly from a low level, and the third characteristic curve remains low and stable, the stage migration index tends towards the amplification phase; if the first characteristic curve's rise slows and gradually plateaus, the second characteristic curve remains high, and the third characteristic curve shows moderate fluctuations, the stage migration index tends towards the steady-state maintenance phase; if the first characteristic curve declines and the second and third characteristic curves rise simultaneously, the stage migration index tends towards the harvest preparation phase or the decay transition phase. Preferably, the stage migration index can be output as a continuous value for direct use by the subsequent neural network, or it can be further mapped as a discrete stage label for use in control strategy selection.
[0051] The anomaly risk index is used to quantify the likelihood of metabolic imbalance, abnormal disturbances, activity decay, or contamination risks during the current culture process. In one embodiment, the anomaly risk index is mainly determined by the following factors: first, whether the slope difference between the second and first characteristic curves continuously increases; second, whether the third characteristic curve experiences an abnormal surge within a short period; third, whether the inflection point of the second or third curve is significantly earlier than that of the first curve; and fourth, whether the phase offset continuously expands and exceeds the normal batch range. Specifically, during generation, the state characterization unit first performs a unified scale mapping on the above risk-related indicators, and then performs weighted fusion according to preset weights to obtain the anomaly risk index. If the proliferation activity increases slowly while the metabolic load increases rapidly, and the disturbance response curve also continuously rises, the anomaly risk index increases; if the environmental disturbance response curve recovers well after the control action, and the metabolic load and proliferation activity remain well synchronized, the anomaly risk index decreases. Preferably, the anomaly risk index can be further divided into three levels: low risk, medium risk, and high risk. When the anomaly risk index exceeds a preset threshold, the state characterization unit sends the result to the fusion prediction unit and the decision control unit to trigger a more conservative prediction and control strategy.
[0052] In one specific embodiment, the workflow of the state characterization unit is as follows: Step 1, receiving the image temporal features, parameter temporal features, and corresponding intermediate statistics output by the feature extraction unit; Step 2, performing time window alignment, scale unification, and weighted fusion on the changes in cell confluence increment, cell count growth rate, clonal area expansion rate, cell motility, and aggregation compactness to construct a first feature curve; Step 3, performing time window alignment, scale unification, and weighted fusion on the glucose consumption rate, lactate production rate, glutamine consumption rate, pH correction frequency, and dissolved oxygen compensation to construct a second feature curve; Step 4, performing time window alignment on the pH recovery time, dissolved oxygen overshoot, temperature fluctuation amplitude, morphological recovery delay, and image activity recovery rate after the control action. Step 5: Window alignment, orientation unification processing, and weighted fusion are performed to construct the third feature curve. Step 6: Smoothing and unified range mapping are applied to the three feature curves respectively. Step 7: The rate of change of the three feature curves is calculated within a local time window, and three sets of slope differences are formed. Step 8: Trend inflection points of the three curves are detected and the spacing between the three inflection points is calculated. Step 9: The slope difference, inflection point spacing, phase shift, and the current level of the three curves are input into the index generation subunit to generate the cultivation stage migration index and the anomaly risk index. Step 10: The cultivation stage migration index and the anomaly risk index are output to the fusion prediction unit as input to the stage-gated cross-modal attention network.
[0053] Through the aforementioned state representation method, this embodiment transforms the originally scattered image features, metabolic features, and perturbation response features into three state curves with clear biological and technological significance. Then, by analyzing the slope difference, inflection point spacing, and phase offset, the relative relationship between the three curves is analyzed. This allows the system to not only see "what the current value is," but also "whether the rate of change is unbalanced," "which trend reversal occurs first," and "which state change lags behind another." Therefore, compared to traditional schemes relying on single-point thresholds or single-modal features, this embodiment can detect stage switching and abnormal precursors earlier, improve the accuracy of characterizing the dynamic state of cell culture, and provide a more stable and discriminative input basis for subsequent fusion prediction and closed-loop control.
[0054] The fusion prediction unit is used to input the culture stage migration index and abnormal risk index into the stage-gated cross-modal attention network to predict cell state and future culture trends. In one embodiment, the fusion prediction unit is located in a data processing server and is communicatively connected to the state representation unit and the decision control unit. It receives temporal features of images, temporal features of parameters, culture stage migration index, and anomaly risk index, and jointly predicts the current cell state and future culture trends based on a stage-gated cross-modal attention network. The fusion prediction unit includes an input mapping subunit, a stage-gated cross-modal attention network, and a result output subunit. The input mapping subunit performs unified dimensional mapping on features from different pre-level modules; the stage-gated cross-modal attention network performs adaptive fusion of image modalities and parameter modalities under stage information constraints; and the result output subunit decodes the fused high-level state features into an estimate of the current cell state and predictions of culture trends in multiple future prediction time domains.
[0055] In one specific implementation, the overall working principle of the fusion prediction unit is as follows: instead of simply concatenating the features of each modality and feeding them into the prediction network, the fusion prediction unit first generates a gating vector based on the culture stage migration index and anomaly risk index output by the state characterization unit. This gating vector is then used to weight and modulate image modal features, parametric modal features, and cross-modal attention heads, allowing the network to adaptively adjust its focus on various types of information at different culture stages and risk levels. For example, during the adaptation or rapid expansion phase, the network prioritizes cell boundary expansion, increased confluence, and clonal expansion features in the image modality; during the homeostasis maintenance or metabolic load increase phase, the network prioritizes glucose consumption, lactate accumulation, pH correction frequency, and dissolved oxygen compensation in the parametric modality; and when the anomaly risk increases, the network increases the weight of perturbation response features and historical control quantities, and reduces its dependence on low-confidence image frames, thereby improving prediction stability and robustness. In this way, the fusion prediction unit outputs not only a single indicator, but also the current cell viability, current cell density, current state category, and the trends of cell viability, cell density, lactate accumulation, stage migration probability, and abnormal risk level changes in at least two future prediction time domains.
[0056] In one embodiment, the stage-gated cross-modal attention network comprises: an image temporal branch, a parameter temporal branch, a stage-gated branch, a cross-modal attention fusion layer, a residual update layer, and a trend decoding layer. The image temporal branch receives the cell morphology evolution feature sequence output by the feature extraction unit. Preferably, this feature sequence is jointly extracted by the aforementioned convolutional neural network, boundary-enhanced temporal attention layer, and ConvLSTM, with a sequence length of 8 or 12 consecutive standard time windows. Each time window corresponds to a fixed-dimensional image hidden feature vector, such as 128 dimensions. The image temporal branch includes: an image feature mapping layer; an image position encoding layer; and an image self-attention encoding layer. The image feature mapping layer uses a fully connected layer or a 1×1 convolutional mapping to uniformly map the input image features to the network's internal hidden dimensions, such as 128 or 256 dimensions. The image position encoding layer preserves the temporal order of image features. The image self-attention encoding layer further extracts the correlation between different time slices within the image sequence, highlighting temporal patterns such as amplified boundary advancement, cluster reconstruction, and local recovery. The temporal branch of the image ultimately outputs the first hidden feature sequence.
[0057] The parameter temporal branch receives long- and short-term dependency features of environmental parameters, metabolic parameters, and historical control variables. This input comes from the output of the aforementioned dual-scale Transformer and can be a sequence of parameter hidden features for 20 or 30 consecutive standard time windows, each time window corresponding to a fixed-dimensional parameter hidden vector, such as 128 dimensions. The parameter temporal branch includes: a parameter feature mapping layer; a parameter location encoding layer; and a parameter self-attention encoding layer. The parameter feature mapping layer maps environmental parameters, metabolic parameters, and control variable features to the same hidden dimension as the image branch; the parameter location encoding layer preserves the temporal positional relationships in the parameter sequence; and the parameter self-attention encoding layer extracts the internal dependencies between short-term perturbation responses and long-term metabolic trends. The parameter temporal branch ultimately outputs a second hidden feature sequence. The stage-gated branch is one of the key improved structures in this application, used to dynamically adjust the weights of each modality and each attention head based on the cultivation stage migration index and anomaly risk index.
[0058] In one embodiment, the stage-gated branch receives the training stage migration index and anomaly risk index output by the state representation unit, and can also receive the normalized values of the current three feature curves as auxiliary inputs. The stage-gated branch includes: an index embedding layer; a two-layer gated perception network; and a gated vector generation layer. The index embedding layer maps the training stage migration index and anomaly risk index to a fixed-dimensional stage embedding vector; the two-layer gated perception network consists of two fully connected layers and an activation layer, used to learn the nonlinear mapping relationship between stage information and modal attention intensity; the gated vector generation layer outputs multiple gated vectors, corresponding to the image branch gated vector, the parameter branch gated vector, and the cross-modal attention head gated vector, respectively. Preferably, the image branch gated vector is used to weight each channel of the first hidden feature sequence; the parameter branch gated vector is used to weight each channel of the second hidden feature sequence; and the cross-modal attention head gated vector is used to assign different weights to multiple cross-modal attention heads. For example, during the amplification phase, the image branch gating vector increases the weight of channels related to boundary advancement and convergence expansion; during the metabolically high load phase, the parametric branch gating vector increases the weight of channels related to lactate accumulation, pH correction frequency, and dissolved oxygen compensation; and when the risk of anomalies increases, the cross-modal attention head gating vector increases the weight of attention heads related to perturbation response.
[0059] A cross-modal attention fusion layer is used to achieve deep association between image modalities and parametric modalities. The cross-modal attention fusion layer includes multiple cross-modal attention heads, each performing a set of cross-modal attention calculations. In one embodiment, the cross-modal attention fusion layer uses the first hidden feature sequence output by the image temporal branch as the query sequence and the second hidden feature sequence output by the parametric temporal branch as the key and value sequences to calculate the attention distribution of image features under different parameter states. Simultaneously, a reverse cross-modal attention sub-layer can be set, using parametric features as the query and image features as the key and value, thereby extracting the reverse dependency of the parametric modality on the image modality. The outputs of both are concatenated to form a bidirectional cross-modal fused feature. To reflect "stage gating," attention head gating weights are introduced into each cross-modal attention head. Specifically, the attention head gating vectors generated by the stage gating branch act on the outputs of different attention heads, causing some attention heads to be enhanced and others suppressed at different training stages. This avoids the problem of ordinary multi-head attention using a fixed attention pattern in all stages. Preferably, the cross-modal attention fusion layer has 4, 6, or 8 attention heads. Each attention head focuses on learning different types of associations, for example: the first attention head focuses on learning the association between cell morphology changes and glucose consumption; the second attention head focuses on learning the association between cell boundary expansion and lactate accumulation; the third attention head focuses on learning the association between cell recovery state and perturbation compensation amount; and the fourth attention head focuses on learning the association between control actions and subsequent image activity changes.
[0060] To preserve effective information in the original modal features and improve network training stability, a residual update layer is set after the cross-modal attention fusion layer. The residual update layer performs residual summation and layer normalization on the gated image features, gated parameter features, and cross-modal fusion features to form a unified fused state feature. The role of this layer is to prevent useful single-modal information from being excessively diluted during cross-modal fusion, while ensuring that the subsequent trend decoding layer receives a high-quality state representation that combines both original modal information and fused correlation information.
[0061] The trend decoding layer converts fused state features into current cell state estimates and future culture trend predictions. This layer includes: a current state decoding branch; a future trend prediction branch; and a multi-task output head. The current state decoding branch outputs the current cell state estimate, such as current cell viability, current cell density, current stage category, or current state category. The future trend prediction branch outputs trend results over multiple prediction time domains, such as cell viability trends, cell density trends, lactate accumulation trends, stage transition probabilities, and abnormal risk level trends over the next 30, 60, and 120 minutes. The multi-task output head uses multiple parallel fully connected layers corresponding to different prediction tasks, with each task outputting a continuous value, probability value, or level value.
[0062] The specific workflow of the stage-gated cross-modal attention network is as follows. In a specific embodiment, the fusion prediction unit operates according to the following process: Step S51, receiving image temporal features, parameter temporal features, training stage migration index, and anomaly risk index; Step S52, performing unified dimensional mapping on the image modality and parameter modality through the image feature mapping layer and parameter feature mapping layer respectively, and superimposing time position encoding; Step S53, the image temporal branch performs self-attention encoding on the image feature sequence and outputs the first hidden feature sequence; the parameter temporal branch performs self-attention encoding on the parameter feature sequence and outputs the second hidden feature sequence; Step S54, the stage-gated branch embeds the training stage migration index and anomaly risk index, and generates image branch gating vector, parameter branch gating vector, and attention head gating vector through a two-layer gating perception network; Step S55, using the image branch gating vector to perform gating on the first hidden feature sequence... The hidden feature sequence is weighted channel by channel, and the second hidden feature sequence is weighted channel by channel using the parameter branch gating vector; in step S56, in the cross-modal attention fusion layer, the first hidden feature sequence after gating is used as the query, and the second hidden feature sequence after gating is used as the key and value, multi-head cross-modal attention calculation is performed, and the output of each cross-modal attention head is weighted using the attention head gating vector; in step S57, the outputs of each cross-modal attention head are concatenated, and after linear transformation, the cross-modal fusion feature is obtained, and it is input into the residual update layer together with the gating original modal feature to generate the fusion state feature; in step S58, the trend decoding layer outputs the current cell state estimate and the culture trend prediction results in multiple future prediction time domains according to the fusion state feature; in step S59, the prediction results are sent to the decision control unit to provide a basis for the subsequent model prediction control algorithm to generate feeding, liquid replacement, gas supply and stirring adjustment commands.
[0063] In one specific embodiment, taking adherent mammalian cell culture as an example, the fusion prediction unit receives updated image temporal features and parameter temporal features every minute. The image temporal branch receives cell morphological evolution features from the most recent 8 minutes, while the parameter temporal branch receives environmental parameters, metabolic parameters, and long- and short-term dependence features of historical control variables from the most recent 30 minutes. The state characterization unit synchronously outputs the culture stage migration index and anomaly risk index at the current time point. During the first 12 hours of culture, the phase-gated branch increases the gating weight of the image branch based on the low anomaly risk index and the phase migration index, which is in the transition from the adaptation phase to the expansion phase. This allows the network to focus on cell adhesion stability, clonal margin expansion, and increased confluence. After 24 hours of culture, as the lactate production rate and pH correction frequency gradually increase, the phase-gated branch increases the weight of the parameter branch and attention heads related to metabolic load. This allows the network to focus on the balance between metabolic stress and cell proliferation. If the anomaly risk index suddenly increases during a certain period, the network automatically increases the weight of attention heads related to perturbation recovery and control compensation to predict the anomaly risk level in a certain time domain. The results are then sent to the decision control unit to implement conservative control strategies in advance.
[0064] Through the structural design of the stage-gated cross-modal attention network described above, this embodiment has at least the following technical advantages compared to traditional simple feature splicing networks or fixed-weight fusion networks: First, by introducing gating branches using the culture stage migration index and the abnormal risk index, the network can dynamically adjust its focus on image modalities and parametric modalities according to different culture stages and risk levels, thereby enhancing stage adaptability. Second, by weighting multiple cross-modal attention heads separately, rather than uniformly weighting the overall features, different types of modal relationships can be selectively strengthened or suppressed at different stages, improving the accuracy of cross-modal fusion. Third, by adopting a hierarchical structure of image branch, parametric branch, stage-gated branch, cross-modal attention layer, and trend decoding layer, the current state estimation and future trend prediction can share the underlying fusion state features, while retaining the independent output capabilities of each task, improving the consistency and stability of multi-task prediction. Fourth, by retaining effective information in the original modality through the residual update layer, excessive smoothing of information occurs during deep fusion, thereby improving the accuracy of cell state prediction and abnormal risk identification, providing a more reliable prediction basis for subsequent closed-loop dynamic regulation.
[0065] The decision control unit is used to generate feeding, liquid replacement, gas regulation and stirring regulation commands based on the prediction results through a model predictive control algorithm with safety barrier constraints, and output them to the execution equipment to achieve closed-loop dynamic control.
[0066] In one embodiment, the decision control unit is located in a central controller or independent control server and is communicatively connected to the fusion prediction unit, the execution device, and the data acquisition unit. It is used to receive the current cell state estimate, the culture trend prediction results in multiple future prediction time domains, the culture stage migration index, and the abnormal risk index output by the fusion prediction unit, and generate feeding, liquid replacement, gas regulation, and stirring regulation commands based on the model predictive control algorithm with safety barrier constraints, and output them to the execution device to realize closed-loop dynamic control of the cell culture process.
[0067] The decision control unit includes: a state reading subunit, a target setting subunit, a candidate control generation subunit, a safety barrier constraint subunit, a rolling optimization subunit, an instruction issuance subunit, and a feedback correction subunit. In this embodiment, the decision control unit does not trigger a control action only after detecting that a parameter exceeds a threshold. Instead, based on the future trend results provided by the fusion prediction unit, it pre-calculates the optimal sequence of control actions to be taken within a future control time domain. In other words, the decision control unit considers not only "what the current state is" but also "how the future state will change after executing a certain action," thereby selecting the scheme that is more beneficial to cell culture and meets safety constraints from multiple candidate control schemes.
[0068] Specifically, at the beginning of each control cycle, the status reading subunit receives the following information: current cell viability, current cell density, current lactate accumulation level, current culture stage label, and current anomaly risk level; trends in cell viability, cell density, lactate accumulation, stage transition probability, and anomaly risk level over multiple predicted time domains; the migration index and anomaly risk index of the current culture stage; and real-time environmental parameters, metabolic parameters, and execution equipment status fed back by the data acquisition unit. Subsequently, the target setting subunit generates control targets based on the current culture stage and process objectives. For example, during the expansion phase, the control targets are more focused on increasing cell viability and cell density; during the steady-state maintenance phase, the control targets are more focused on controlling lactate accumulation, maintaining dissolved oxygen stability, and reducing disturbances; and during the harvest preparation phase, the targets are more focused on maintaining a stable state, preventing sudden anomalies, and maintaining the expression conditions of the target product. Based on this, the candidate control generation subunit constructs multiple candidate action sequences for the future control time domain, focusing on four types of control variables: feeding, liquid replacement, gas regulation, and stirring regulation. The safety barrier constraint subunit performs safety screening on each of these candidate action sequences. The rolling optimization subunit selects the action sequence with the highest overall benefit from the candidate actions that meet the safety conditions and outputs the first control instruction to be executed at the current moment. The instruction issuing subunit sends the control instruction to the execution devices such as the feeding pump, liquid replacement valve group, gas mixer, and stirring actuator. After execution, the feedback correction subunit corrects the subsequent control weights and model parameters based on the new round of collected data and enters the next control cycle.
[0069] The specific structure and function of the decision control unit include a state reading subunit that receives the prediction results output by the fusion prediction unit and simultaneously reads the current environmental parameters, metabolic parameters, and feedback from the execution device. In one embodiment, the state reading subunit reads at least the following information: current cell viability estimate; current cell density estimate; current culture stage category; current abnormal risk level; predicted cell viability for the next 30, 60, and 120 minutes; predicted cell density for the next 30, 60, and 120 minutes; predicted lactate accumulation trend for the next 30, 60, and 120 minutes; stage migration probability for the next 30, 60, and 120 minutes; current temperature, pH, dissolved oxygen, liquid level, and carbon dioxide concentration; current glucose concentration, lactate concentration, and glutamine concentration; current feed pump status, valve group opening / closing status, gas supply device status, and stirring device status. The state reading subunit encapsulates the above information into a current control state vector and provides it to subsequent control subunits.
[0070] The target setting subunit is used to generate optimized targets for the current control cycle based on the culture stage and process requirements. In one embodiment, the target setting subunit presets multiple stage target templates, including target templates for the adaptation period, expansion period, steady-state maintenance period, and harvest preparation period. For example, during the adaptation period, the target setting subunit prioritizes ensuring a mild environment and low disturbance, with the main control targets being maintaining stable temperature, pH, and dissolved oxygen, reducing stirring intensity, and avoiding premature large-dose feeding; during the expansion period, the main control targets are improving cell viability, increasing cell density, maintaining glucose supply, and inhibiting rapid lactate rise; during the steady-state maintenance period, the main control targets are reducing the rate of lactate accumulation, reducing pH fluctuations, and controlling gas supply and stirring fluctuations; during the harvest preparation period, the main control targets are preventing state deterioration, controlling abnormal risks, and maintaining the final culture quality. Preferably, when the abnormal risk index increases, the target setting subunit automatically increases the weight of "reducing disturbances" and "controlling risks," and decreases the weight of "rapidly increasing cell density," thereby shifting the control strategy from aggressive to conservative. The candidate control generation subunit is used to generate multiple candidate action sequences around four types of controllable variables. These four types of controllable variables include: feed control; liquid change control; gas regulation control; and stirring regulation control.
[0071] In one embodiment, the feed control parameters include the feed start time, feed duration, feed flow rate, and feed volume; the liquid exchange control parameters include the liquid exchange start time, liquid exchange ratio, drain rate, and inlet rate; the gas regulation control parameters include oxygen flow rate, carbon dioxide flow rate, gas mixing ratio, and ventilation duration; and the stirring regulation control parameters include the target stirring speed, adjustment range, and holding time. The candidate control generation subunit generates multiple candidate combination schemes within each control cycle, for example: Scheme A: low-dose, high-frequency feed, maintaining moderate oxygen supply, with a slight increase in stirring; Scheme B: moderate-dose feed, supplemented by a low-ratio liquid exchange, enhancing oxygen supply; Scheme C: maintaining the feed rate unchanged, only increasing oxygen supply and decreasing stirring; Scheme D: performing partial liquid exchange and simultaneously decreasing stirring to maintain stable feed. Each scheme covers multiple future control steps, but only the first step is executed in the current cycle, with the remaining part used as a rolling prediction reference.
[0072] The safety barrier constraint subunit is used to ensure that candidate control actions do not lead the culture process into a dangerous state. In one embodiment, the safety barrier constraint subunit presets safety domains, including: pH safety domain; dissolved oxygen safety domain; temperature safety domain; shear strength safety domain; feed rate per unit time safety domain; liquid exchange ratio per unit time safety domain; liquid level change safety domain; and stirring speed safety domain. The safety barrier constraint subunit checks each candidate action sequence step by step to determine whether the action will cause any of the above variables to approach or exceed the boundary of the safety domain in the future prediction time domain. For example, if a candidate feed scheme may cause excessive feed rate per unit time, resulting in a rapid rise in liquid level or a sudden change in osmotic pressure, then the scheme is determined to violate the safety barrier constraint; if a gas supply scheme may cause excessive dissolved oxygen overshoot, then the scheme is determined to not meet the safety requirements; if a stirring adjustment scheme may cause excessive shear strength, then the scheme is determined to be unselectable. Preferably, the safety barrier constraint subunit not only checks whether the boundary is exceeded, but also checks whether the boundary is approached too quickly. That is, even if the predicted value has not yet exceeded the boundary, if it has entered the warning buffer zone, the candidate scheme will be downweighted or eliminated. This prevents the controlled movements from becoming too aggressive when approaching the limit.
[0073] The rolling optimization subunit is the core control component of this embodiment, used to select the optimal comprehensive solution from candidate action sequences that have passed safety screening. In one embodiment, the rolling optimization subunit employs a model predictive control algorithm to perform rolling optimization of the future culture process using preset prediction and control time domains. For example, the prediction time domain can be set to the next 60 minutes, the control time domain to the next 15 minutes, and the control step size to 5 minutes. During rolling optimization, the rolling optimization subunit comprehensively considers the following factors: whether cell viability increases or remains stable in the future time domain; whether cell density approaches the desired target in the future time domain; whether the lactate accumulation trend is suppressed; whether pH and dissolved oxygen fluctuations are sufficiently stable; whether the changes in control actions are too drastic; whether the abnormal risk level increases; and whether the current stage objective is met. Preferably, the rolling optimization subunit assigns different weights to the above factors and dynamically adjusts them according to the culture stage. For example, during the expansion phase, the weights of cell viability and cell density-related items are increased; during the steady-state maintenance phase, the weights of inhibiting lactate accumulation and reducing fluctuations are increased; and when the abnormal risk is high, the weights of risk penalty items are increased. After completing the multi-objective comparison, the rolling optimization sub-unit selects the candidate action sequence with the highest comprehensive score, but only uses the first step of this action sequence as the current actual control command. The command is then recalculated based on the latest state at the start of the next control cycle; this method is called rolling optimization.
[0074] The instruction issuing subunit is used to send the optimal control action determined in the current cycle to the specific execution device. In one embodiment, the instruction issuing subunit outputs one or more of the following instructions: sending instructions to the feed pump for feed start time, feed flow rate, and feed volume; sending instructions to the liquid exchange valve group for discharge duration, liquid inlet duration, and liquid exchange ratio; sending instructions to the gas mixer for oxygen flow rate, carbon dioxide flow rate, and gas mixing ratio; and sending instructions to the agitator driver for target speed, acceleration / deceleration gradient, and hold duration. To ensure execution reliability, the instruction issuing subunit may also include an action confirmation mechanism. That is, after the instruction is issued, the status signal returned by the execution device is read to confirm that the action has been executed as required; if it has not been executed or the execution deviation is too large, the deviation information is fed back to the feedback correction subunit.
[0075] The feedback correction subunit is used to update the control weights and subsequent control schemes based on the newly collected data after the control action is executed. In one embodiment, the feedback correction subunit performs the following operations: comparing the deviations between the predicted changes in pH, dissolved oxygen, lactic acid, and cell state and the actual collected values; if the deviation is small, the current optimized weights remain unchanged; if the deviation remains large, the dynamic response parameters corresponding to feeding, liquid replacement, gas supply, or stirring are corrected; if the abnormal risk index remains higher than the threshold, the conservative control weights are increased, and high-amplitude actions are restricted; if there is lag or insufficient response in the feedback from the executing equipment, the equipment compensation parameters are adjusted. In this way, the decision control unit can gradually adapt to different batches, different cell types, and different equipment operating conditions during continuous operation.
[0076] In one specific embodiment, the decision control unit operates according to the following process: Step S61, read the current state estimate, future trend prediction result, cultivation stage migration index, and abnormal risk index output by the fusion prediction unit; Step S62, determine the optimization objective and weight of each objective within the current control cycle based on the current cultivation stage and process task; Step S63, generate multiple candidate control action sequences around four types of control variables: feeding, liquid replacement, gas regulation, and stirring regulation; Step S64, input each candidate control action sequence into the safety barrier constraint subunit, check the constraints such as pH, dissolved oxygen, temperature, shear strength, feeding amount per unit time, and liquid replacement ratio per unit time one by one, and eliminate schemes that exceed or are close to exceeding the limits; Step S65: Perform model predictive control rolling optimization on the candidate solutions that have passed the safety screening, comprehensively compare the effects of cell viability improvement, cell density improvement, lactate inhibition, control stability, and abnormal risk changes to determine the comprehensive optimal action sequence; Step S66: Output only the first step control command in the comprehensive optimal action sequence and send it to the feed pump, liquid exchange valve group, gas mixer, and agitator driver through the command sending subunit; Step S67: Receive feedback from the execution equipment and the environmental parameters, metabolic parameters, and image status collected at the next moment, and compare the deviation between the prediction and the actual situation; Step S68: Adjust the optimization weights, dynamic response parameters, and control constraints based on the deviation information, and enter the next control cycle.
[0077] In one specific embodiment, taking a 96-hour adherent mammalian cell culture as an example, the decision control unit executes a control decision every 5 minutes, with the prediction time domain set to the next 60 minutes and the control time domain set to the next 15 minutes. During the first 24 hours of culture, the fusion prediction unit determines that the system is transitioning from the adaptation phase to the expansion phase, with high cell viability and low abnormal risk. The decision control unit prioritizes increasing cell density and maintaining nutrient supply. At this time, the rolling optimization results preferentially select a low-dose, high-frequency feeding scheme, and maintain the stirring speed in a low to medium range to reduce shear disturbance. During the 24- to 60-hour culture period, if the prediction results show an increase in lactate accumulation and an increase in pH correction frequency within the next 60 minutes, the decision control unit increases the weight of "inhibiting lactate accumulation" and "reducing metabolic burden." At this time, the candidate control scheme may preferentially select a combination of "medium-dose feeding + small-ratio medium replacement + enhanced oxygen supply," and appropriately reduce the stirring amplitude to mitigate the trend of metabolic imbalance. If the abnormal risk index increases and the environmental disturbance response curve remains high after 60 hours of cultivation, the decision control unit enters a conservative control mode, increases the weight of the risk penalty item, limits the amount of feed per unit time and the range of stirring adjustment, and only allows small-step correction actions to prevent drastic fluctuations in the cultivation system.
[0078] Through the above-described decision-making and control methods, this embodiment achieves at least the following technical effects: First, by making control decisions based on future trend predictions rather than single-point threshold judgments, the system can intervene in advance rather than passively adjusting after the state deteriorates, thereby improving the forward-looking nature of control. Second, by using a model predictive control algorithm with safety barrier constraints, while optimizing cell viability, cell density, and metabolic load, variables such as pH, dissolved oxygen, temperature, shear strength, feed volume, and medium replacement ratio are strictly limited from entering dangerous ranges, improving the safety of the culture process. Third, by using a rolling optimization method, only the current optimal action is executed in each control cycle, and subsequent actions are recalculated based on the latest feedback, which can adapt to dynamic changes in the culture state and enhance the flexibility of closed-loop regulation. Fourth, by continuously comparing the deviation between the predicted results and the actual execution results through the feedback correction subunit and adjusting the control weights and dynamic response parameters online, the robustness and generalization ability of the system under different cell types, different culture batches, and different equipment operating conditions can be improved.
[0079] Preferably, the culture environment parameter sequence includes temperature, pH, dissolved oxygen, carbon dioxide concentration, stirring speed, and liquid level; the metabolic parameter sequence includes glucose concentration, lactate concentration, glutamine concentration, ammonium ion concentration, and osmotic pressure; and the historical control sequence includes feed volume, liquid replacement ratio, gas supply volume, and stirring adjustment records. The time-series reconstruction unit resamples each sequence according to a preset sliding time window and generates modal confidence scores based on sampling completeness, noise variance, drift bias, and time delay to compensate for missing data.
[0080] Preferably, the first feature curve is a proliferation activity curve, which is generated by weighting the cell confluence increment, cell count growth rate, clonal area expansion rate, cell motility activity, and aggregation compactness obtained by extracting spatial features from a microscopic image sequence through a convolutional neural network and extracting temporal features through a ConvLSTM. The proliferation activity curve represents the trend of cell expansion capacity per unit time.
[0081] Preferably, the second characteristic curve is a metabolic load curve, which is generated by weighting the glucose consumption rate, lactate production rate, glutamine consumption rate, pH correction frequency, and dissolved oxygen compensation amount obtained by encoding the metabolic parameter sequence, some culture environment parameter sequence, and historical control sequence through a dual-scale Transformer. The short-scale Transformer is used to extract local fluctuation features, and the long-scale Transformer is used to extract the stage change trend.
[0082] Preferably, the third characteristic curve is an environmental disturbance response curve, which takes the trigger time of the control action in the historical control sequence as the anchor point and is generated by weighting the pH deviation recovery time, dissolved oxygen overshoot, temperature fluctuation amplitude, cell morphology recovery delay after stirring disturbance, and image activity recovery rate after the control action. It is used to characterize the dynamic response capability of the cell culture system to external control disturbances.
[0083] Preferably, the state characterization unit generates a cultivation stage migration index and an anomaly risk index based on the slope difference, adjacent inflection point spacing, peak-valley sequence, and phase shift of the first characteristic curve, the second characteristic curve, and the third characteristic curve; wherein, the cultivation stage migration index is used to characterize the stage transition trend between the adaptation period, the expansion period, the steady state period, and the harvest preparation period, and the anomaly risk index is used to characterize the probability of occurrence of metabolic imbalance, excessive disturbance, activity decay, or contamination anomalies.
[0084] Preferably, the stage-gated cross-modal attention network includes an image temporal branch, a parameter temporal branch, a stage-gated branch, a cross-modal attention layer, and a trend decoding layer; the cross-modal attention layer includes multiple cross-modal attention heads; the image temporal branch receives the cell morphological evolution features and outputs a first hidden feature; the parameter temporal branch receives the long-term and short-term dependence features of environmental parameters, metabolic parameters, and historical control quantities and outputs a second hidden feature; the stage-gated branch receives the culture stage migration index and abnormal risk index and generates a gating vector, which weights the first hidden feature, the second hidden feature, and the cross-modal attention heads; the cross-modal attention layer uses the first hidden feature as a query vector and the second hidden feature as a key vector and value vector to perform cross-modal association calculation and outputs a fused state feature.
[0085] Preferably, the trend decoding layer uses a multi-task prediction head to simultaneously output the current cell state estimate and the culture trend prediction results for at least two future prediction time domains. The culture trend prediction results include cell viability, cell density, lactate accumulation trend, stage transition probability, and abnormal risk level.
[0086] Preferably, the model predictive control algorithm in the decision control unit adopts a rolling optimization approach, with the joint optimization objectives of improving cell viability, increasing target cell density or target product expression level, inhibiting lactic acid accumulation, and reducing control fluctuations. A safety barrier function is constructed to limit pH, dissolved oxygen, temperature, shear intensity, feed rate per unit time, and liquid exchange rate per unit time within a preset safety range in the prediction time domain, so as to generate feed, liquid exchange, gas supply, and stirring adjustment commands.
[0087] Preferably, after each control cycle, the decision control unit recalculates the first characteristic curve, the second characteristic curve, and the third characteristic curve based on the latest acquired microscopic image sequence, culture environment parameter sequence, and metabolic parameter sequence, and corrects the modal confidence, the gating vector of the stage gating branch, and the optimization weight of the model predictive control algorithm online according to the deviation between the predicted curve and the actual curve; when the abnormal risk index exceeds a preset threshold, it switches to a conservative control mode.
[0088] In some embodiments, a boundary-enhanced temporal attention layer is inserted between the convolutional neural network and the ConvLSTM. This boundary-enhanced temporal attention layer includes a multi-scale spatial feature extraction branch, an adjacent frame difference branch, and a channel-space joint attention branch. The multi-scale spatial feature extraction branch processes the current frame feature map output by the convolutional neural network in parallel using 3×3 convolution and 3×3 dilated convolution with a dilation rate of 2, respectively, to extract multi-scale boundary features of cell contours and clustered regions. The adjacent frame difference branch performs a difference operation between the current frame feature map and the previous frame feature map, and then maps them using a 1×1 convolution to obtain temporal response features characterizing minute changes in cell morphology. The channel-space joint attention branch generates attention weights based on the multi-scale boundary features and the temporal response features, and then weights and enhances the feature map input to the ConvLSTM to highlight cell boundaries, clustering and separation regions, and regions of state change.
[0089] This invention inserts a boundary-enhanced temporal attention layer between the convolutional neural network and the ConvLSTM, and simultaneously extracts cell contours, clumping region edges, and local texture changes through a multi-scale spatial feature extraction branch. This effectively enhances the representation of blurred cell boundaries, interconnected regions, and clumped occlusion areas in microscopic images. Compared to directly inputting the output features of the convolutional neural network into the ConvLSTM, this invention can highlight boundary features and separation region features that are more sensitive to proliferation status before temporal modeling, thereby improving the accuracy of extracting features such as cell confluence, clonal expansion, and clumping compactness, and providing a more reliable input for subsequent proliferation activity curve construction.
[0090] This invention employs an adjacent-frame differential branch to perform differential processing on the feature map of the current frame and the feature map of the previous frame. This highlights subtle morphological changes during cell culture, such as proliferation and expansion, local contraction, aggregation and reconstruction, and disturbance recovery. Because cell culture images often exhibit small changes between adjacent time points, the traditional convolutional neural network and ConvLSTM cascade structure is easily affected by background noise, lighting fluctuations, or culture medium disturbances, leading to the obscuring of crucial change information. This invention introduces time-difference response features and incorporates attention weighting, enabling the temporal input received by the ConvLSTM to focus more on regions of real-world change, thereby enhancing the ability to characterize cell morphological evolution and identify early signs of stage transitions.
[0091] This invention generates attention weights for multi-scale boundary features and temporal response features through a channel-space joint attention branch, adaptively enhancing the feature map of the input ConvLSTM. This suppresses the interference of irrelevant background, random noise, and pseudo-change regions on the temporal modeling results. The resulting cell morphology evolution features have higher discriminativeness and stability, making the subsequently constructed proliferation activity curve smoother and more accurate. It also helps the stage-gated cross-modal attention network improve the accuracy of cell state prediction and abnormal risk identification, further enhancing the reliability and closed-loop control effect of the model predictive control algorithm when generating feeding, liquid replacement, gas supply, and stirring adjustment commands.
[0092] The cell culture dynamic control system based on multimodal time-series data of the present invention, such as Figure 2 As shown, the system mainly includes a data acquisition unit, a time-series reconstruction unit, a feature extraction unit, a state characterization unit, a fusion prediction unit, a decision control unit, an execution device, and a central controller / server. The data acquisition unit is connected to the cell culture device and is used to acquire microscopic image sequences, environmental parameter sequences, metabolic parameter sequences, and control history sequences. It then sends the acquired multimodal raw data to the time-series reconstruction unit and the decision control unit. The time-series reconstruction unit performs unified time window alignment, missing data compensation, and standardization on multimodal data with different sampling frequencies and time delays, and sends the processed standardized time-series data to the feature extraction unit.
[0093] The feature extraction unit extracts cell morphological evolution features, as well as long-term and short-term dependence features of environmental parameters, metabolic parameters, and historical control quantities from standardized time-series data, and sends the extraction results to the state characterization unit and the execution device correlation analysis path. The state characterization unit constructs proliferation activity curves, metabolic load curves, and environmental disturbance response curves, and generates a culture stage migration index and anomaly risk index based on these three feature curves. The fusion prediction unit receives the state characterization results and combines them with multimodal features to predict cell state and future culture trends. The decision control unit is connected to the fusion prediction unit, the state characterization unit, and the central controller / server, and executes model predictive control based on the prediction results, generating commands for feeding, liquid replacement, gas regulation, and stirring. The execution devices include a feeding pump, a liquid replacement valve assembly, a gas mixer, and a stirring device, each used to execute the corresponding culture regulation actions. The central controller / server enables data interaction, computation scheduling, control command issuance, and execution result feedback among the functional units, thereby forming a closed-loop dynamic control of the cell culture process.
[0094] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products, and therefore this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0095] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A cell culture dynamic control system based on multimodal time-series data, characterized in that, include: The data acquisition unit is used to acquire cell microscopic image sequences, culture environment parameter sequences, metabolic parameter sequences, and historical control sequences. The time series reconstruction unit is used to align each sequence according to a unified time window and perform missing data compensation based on modal confidence to generate standardized time series data. The feature extraction unit is used to extract cell morphological evolution features using convolutional neural networks and ConvLSTM, and to extract long-term and short-term dependence features of environmental parameters, metabolic parameters and historical control quantities using dual-scale Transformer. The state characterization unit is used to construct a first characteristic curve characterizing cell proliferation activity, a second characteristic curve characterizing metabolic load, and a third characteristic curve characterizing environmental disturbance response. Based on the slope difference, inflection point spacing, and phase offset of the three characteristic curves, the culture stage migration index and abnormal risk index are generated. The fusion prediction unit is used to input the culture stage migration index and abnormal risk index into the stage-gated cross-modal attention network to predict cell state and future culture trends. The decision control unit is used to generate feeding, liquid replacement, gas regulation and stirring regulation commands based on the prediction results through a model predictive control algorithm with safety barrier constraints, and output them to the execution equipment to achieve closed-loop dynamic control.
2. The cell culture dynamic control system based on multimodal time-series data according to claim 1, characterized in that, The culture environment parameter sequence includes temperature, pH, dissolved oxygen, carbon dioxide concentration, stirring speed, and liquid level; the metabolic parameter sequence includes glucose concentration, lactate concentration, glutamine concentration, ammonium ion concentration, and osmotic pressure; and the historical control sequence includes feed volume, liquid replacement ratio, gas supply volume, and stirring adjustment records. The time-series reconstruction unit resamples each sequence according to a preset sliding time window and generates modal confidence scores based on sampling completeness, noise variance, drift bias, and time delay to compensate for missing data.
3. The cell culture dynamic control system based on multimodal time-series data according to claim 2, characterized in that, The first feature curve is the proliferation activity curve, which is generated by weighting the cell confluence increment, cell count growth rate, clonal area expansion rate, cell motility activity, and aggregation compactness obtained by extracting spatial features from the microscopic image sequence through a convolutional neural network and extracting temporal features through ConvLSTM. The proliferation activity curve represents the changing trend of cell expansion capacity per unit time.
4. The cell culture dynamic control system based on multimodal time-series data according to claim 2, characterized in that, The second characteristic curve is the metabolic load curve, which is generated by weighting the glucose consumption rate, lactate production rate, glutamine consumption rate, pH correction frequency, and dissolved oxygen compensation amount obtained by encoding the metabolic parameter sequence, some culture environment parameter sequence, and historical control sequence through a dual-scale Transformer. The short-scale Transformer is used to extract local fluctuation features, and the long-scale Transformer is used to extract the stage change trend.
5. The cell culture dynamic control system based on multimodal time-series data according to claim 1, characterized in that, The third characteristic curve is the environmental disturbance response curve. It takes the trigger time of the control action in the historical control sequence as the anchor point and is generated by weighting the pH deviation recovery time, dissolved oxygen overshoot, temperature fluctuation amplitude, cell morphology recovery delay after stirring disturbance, and image activity recovery rate after the control action. It is used to characterize the dynamic response capability of the cell culture system to external control disturbances.
6. The cell culture dynamic control system based on multimodal time-series data according to claim 5, characterized in that, The state characterization unit generates a cultivation stage migration index and an anomaly risk index based on the slope difference, adjacent inflection point spacing, peak-valley sequence, and phase shift of the first, second, and third characteristic curves. The cultivation stage migration index is used to characterize the stage transition trend between the adaptation period, amplification period, steady state period, and harvest preparation period, while the anomaly risk index is used to characterize the probability of metabolic imbalance, excessive disturbance, activity decay, or contamination anomalies.
7. The cell culture dynamic control system based on multimodal time-series data according to claim 1, characterized in that, The stage-gated cross-modal attention network includes an image temporal branch, a parameter temporal branch, a stage-gated branch, a cross-modal attention layer, and a trend decoding layer; the cross-modal attention layer includes multiple cross-modal attention heads, and the image temporal branch receives the cell morphology evolution features and outputs the first hidden feature; The parameter time-series branch receives the long- and short-term dependency features of environmental parameters, metabolic parameters, and historical control quantities, and outputs the second hidden feature. The stage-gated branch receives the cultivation stage migration index and the abnormal risk index and generates a gating vector, which weights the first hidden feature, the second hidden feature, and the cross-modal attention head. The cross-modal attention layer uses the first hidden feature as the query vector and the second hidden feature as the key vector and value vector to perform cross-modal association calculation and outputs the fused state feature.
8. The cell culture dynamic control system based on multimodal time-series data according to claim 7, characterized in that, The trend decoding layer uses a multi-task prediction head to simultaneously output the current cell state estimate and the culture trend prediction results for at least two future prediction time domains. The culture trend prediction results include cell viability, cell density, lactate accumulation trend, stage transition probability, and abnormal risk level.
9. The cell culture dynamic control system based on multimodal time-series data according to claim 1, characterized in that, The model predictive control algorithm in the decision control unit adopts a rolling optimization approach, with the joint optimization objectives of improving cell viability, increasing target cell density or target product expression level, inhibiting lactic acid accumulation, and reducing control fluctuations. It also constructs a safety barrier function to limit pH, dissolved oxygen, temperature, shear intensity, feed rate per unit time, and liquid exchange rate per unit time within a preset safety range in the prediction time domain, so as to generate feed, liquid exchange, gas supply, and stirring adjustment commands.
10. A cell culture dynamic control system based on multimodal time-series data according to claim 9, characterized in that, After each control cycle, the decision control unit recalculates the first, second, and third characteristic curves based on the latest acquired microscopic image sequence, culture environment parameter sequence, and metabolic parameter sequence. It also adjusts the modal confidence, the gating vector of the stage gating branch, and the optimization weights of the model predictive control algorithm online according to the deviation between the predicted and actual curves. When the abnormal risk index exceeds a preset threshold, it switches to a conservative control mode.