Biological reaction cross-task model training method and cross-scene parameter generating method based on dynamic context

By employing a dynamic context-based cross-task model training method for bioreactors, and utilizing conditional diffusion models and denoising networks, the problems of model adaptability and data dependence during bioreactor process scale-up were solved, achieving high-quality training and accurate evaluation during the cold start phase of new processes.

CN121789807APending Publication Date: 2026-04-03SHAANXI HUANYAN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies cannot effectively adapt to changes in the physical environment such as fluid mass transfer, heat transfer, mixing, and shearing when scaling up or migrating bioreactor processes. This leads to a decrease in the accuracy of model predictions and a strong dependence on data, making it impossible to accurately assess new processes during the cold start phase.

Method used

A cross-task model training method for bioreaction based on dynamic context is adopted. By collecting reaction data from bioreactors of different sizes with the same process, a support set is constructed. The model is trained using a conditional diffusion model and a denoising network. By combining global dynamic latent variables and time-resolved state latent vectors, cross-scenario parameters are generated, achieving high-quality training with limited data.

Benefits of technology

During the cold start phase of the new process, model training can be completed using a small amount of initial data, enabling accurate assessment of the biological response state. This solves the model training problem caused by insufficient data and improves the accuracy and adaptability of the assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789807A_ABST
    Figure CN121789807A_ABST
Patent Text Reader

Abstract

The invention discloses a biological reaction cross-task model training method and a cross-scene parameter generation method based on dynamic context, and relates to the field of biological reactions.The method comprises the steps that reaction data of all time steps in the complete reaction process of two different bioreactor sizes under the same technology are collected; intercepting data of B time steps before the initial reaction stage to form a support set; obtaining a mean value of posterior Gaussian distribution of the global dynamic potential variables through all reaction data; performing K-step Gaussian noise addition on the global dynamic potential variable according to a forward noise addition network of the conditional diffusion model to obtain a noise addition sample; obtaining a time-resolved state potential vector sequence through the support set; inputting the time-resolved state potential vector sequence into a task encoder to obtain a dynamic context vector; predicting noise added in each step according to a denoising network of a conditional diffusion model, a noise adding sample and a dynamic context vector; and training a denoising network and a task encoder according to the real noise added in each step and the predicted noise added in each step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biological reactions, and in particular to a method for training cross-task models of biological reactions based on dynamic context and a method for generating parameters across scenarios. Background Technology

[0002] The research and development of bioproducts typically involves three stages: laboratory, pilot-scale, and industrial-scale. Although the biological reactions occurring in the bioreactor are the same at each stage, there are often differences in the mixing, mass transfer, heat transfer, shear, and apparent gas velocity of the reaction solution.

[0003] To assess the state of biological reactions in bioreactors of different sizes, current methods employ mathematical estimation models trained on specific experimental data, and then use these trained mathematical estimation models to evaluate the state of biological reactions.

[0004] However, this method typically requires a large amount of comprehensive historical data for training to ensure accuracy. In the cold start phase of a new process, the amount of data is often limited, resulting in inaccurate assessments of biological response states when using the trained mathematical estimation model. Summary of the Invention

[0005] The purpose of this application is to provide a method for training a cross-task model of biological response based on dynamic context and a method for generating parameters across scenarios, which can improve the accuracy of evaluation results.

[0006] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for training a cross-task model of biological responses based on dynamic context, the training method comprising: Reaction data at each time step were collected during the complete reaction process of two different bioreactor sizes under the same process. The support set is formed by extracting data from the reaction data at each time step, specifically the data from the first B time steps before the initial stage of the reaction. The mean of the posterior Gaussian distribution of the global kinetic latent variables was obtained from all reaction data; the global kinetic latent variables represent the properties of the bioreactor that do not change with reaction time during the reaction process. Based on the forward noise-adding network of the conditional diffusion model, K-step Gaussian noise addition is performed on the global dynamic latent variables to finally obtain the noisy samples; The time-resolved state latent vector sequence is obtained through the support set; the time-resolved state latent variable sequence represents the properties of the bioreactor that change with reaction time during the reaction process. Input the time-resolved state latent vector sequence into the task encoder to obtain the dynamic context vector; The denoising network based on the conditional diffusion model predicts the noise added at each step based on the noisy sample and the dynamic context vector. The denoising network and task encoder are trained based on the real noise added at each step and the predicted noise added at each step, resulting in a trained conditional diffusion model and a task encoder. The trained biological response cross-task model includes a trained conditional diffusion model and a trained task encoder.

[0007] Secondly, this application provides a method for generating parameters across different scenarios, the method comprising: Collect real data in the initial stage of the reaction; Obtain the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence based on real data; Using a trained task encoder, the real dynamic context vector is obtained based on the real time-resolved state latent vector sequence; Based on static metadata, calculate the true dimensionless physical criterion number, and construct the true physical feature vector based on the true dimensionless physical criterion number; The real physical feature vector is concatenated with the real type information to obtain the real joint input vector; The real joint input vector is input into the MLP network to obtain the real static context vector; the static context vector includes: physical environment and biological basis.

[0008] By using a trained conditional diffusion model, multiple global dynamic latent variables are obtained based on the real dynamic context vector; Based on the real-time-resolved state potential prediction vector corresponding to each global dynamic potential variable; The real static context vector, each real time-resolved state potential vector and the corresponding real global dynamic potential variable are decoded into predicted observation data, and multiple sets of predicted observation data constitute a predicted observation data sequence. The mean of the data at each time step in the predicted observation data sequence is obtained to get the final predicted observation data; Obtain the confidence intervals for the data at each time step; Output the final predicted observation data and confidence interval.

[0009] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method for training a cross-task model of biological reactions based on dynamic context and a method for generating parameters across scenarios. This disclosure integrates the conditional diffusion model into the cross-task model architecture of biological reactions. Leveraging the inherent advantages of the conditional diffusion model in small-sample learning scenarios, it only requires collecting a small amount of data from the first B time steps of the initial reaction to construct a support set, thus completing effective model training. Unlike traditional mathematical estimation models that rely on a large amount of complete long-term reaction data to fit the reaction patterns, the conditional diffusion model performs K-step Gaussian noise addition on the global dynamic latent variables through a forward noise addition process, achieving implicit data augmentation for limited data. This is equivalent to expanding the information dimension and diversity of effective samples during the training phase. Combined with the accurate prediction and learning of noise by a denoising network, the model can fully extract the core dynamic characteristics and initial state patterns of the reaction from a small amount of initial data, completing high-quality training without relying on a large amount of complete reaction data. This perfectly solves the model training problem caused by insufficient data in the cold start phase of new processes. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a cross-task model training method for biological responses based on a dynamic context, according to an exemplary embodiment. Figure 2 This is a schematic diagram of the functional modules of a cross-task model training device for biological responses based on dynamic context, according to an exemplary embodiment. Figure 3 This is a flowchart illustrating a cross-scene parameter generation method according to an exemplary embodiment; Figure 4 This is a schematic diagram of the functional modules of a cross-scene parameter generation device according to an exemplary embodiment; Figure 5 This is a schematic diagram of an experimental vessel according to an exemplary embodiment; Figure 6 This is illustrated according to an exemplary embodiment. Figure 5 Design data sheet for the middle tank; Figure 7 This is illustrated according to an exemplary embodiment. Figure 5 Design data sheet for the middle tank (Table 2); Figure 8 This is illustrated according to an exemplary embodiment. Figure 5 Design data for the intermediate tank (Table 3); Figure 9 This is illustrated according to an exemplary embodiment. Figure 5 Table 4 shows the design data for the middle tank. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] The research and development of any bioproduct generally involves three stages: laboratory stage, pilot-scale stage, and industrial-scale stage. Although the biological reactions occurring in the bioreactor are the same at each stage, the mixing, mass transfer, heat transfer, shear, and apparent gas velocity of the reaction solution often differ. How to assess the state of biological reactions in bioreactors of different scales, especially during bioreactor scale-up, is a key bottleneck in the fields of biopharmaceuticals and industrial biotechnology. With limited experimental and simulation data, accurately controlling the reactor to maintain cell growth rates similar to those of the biological reaction during scale-up is an urgent problem to be solved.

[0015] The disadvantages of existing technologies are as follows: 1. Existing technologies mostly optimize reaction processes, such as optimizing reaction processes based on historical data to establish corresponding relationships, or predicting the control processes of bioreactors through neural networks. However, there are fewer technologies that expand the scale or capacity of bioreactors by improving existing bioreactor processes or simulating bioreactor processes.

[0016] 2. When facing reactor process scale-up or migration, existing technologies become ineffective due to fundamental changes in key physical environments such as fluid mass transfer, heat transfer, mixing, shearing, and apparent gas velocity. For example, models using neural networks to predict control actions learn an "input-output" mapping within the reactor data distribution at a specific scale and process. However, when facing process scale-up, the changing physical environment leads to a catastrophic decrease in model prediction accuracy.

[0017] 3. Existing technologies fail to link apparent phenomena at different scales with underlying biological principles. The prediction process is a deterministic point estimation, which cannot quantify uncertainty. They lack the ability to extrapolate the patterns of bioreactors from a low scale (e.g., 100L) to a high scale (e.g., 10000L).

[0018] 4. Strong data dependence and limited extrapolation capabilities make them ineffective in addressing the cold start problem of new processes. Traditional methods, especially data-driven methods, typically require a large amount of comprehensive historical data for training to ensure accuracy. These methods have not effectively solved the problem of how to perform relatively accurate full-process extrapolation predictions using only a small amount of data from the initial stage of operation of a new process (such as a new cell line, virus, new culture medium, or reactor).

[0019] 5. Existing technologies are mostly optimizations of single-cell culture processes and cannot be extrapolated in terms of process transfer. However, under the same reaction vessel, various reaction processes should have many commonalities in parameters, which can be used for potential learning and mutual reference by the model.

[0020] To address the aforementioned technical issues, this disclosure provides a method for training a cross-task model of biological responses based on dynamic context and a method for generating parameters across scenarios.

[0021] Figure 1 This is a flowchart illustrating a cross-task model training method for biological responses based on a dynamic context, according to an exemplary embodiment. Figure 1 As shown, the method includes the following steps S101-S108: In step S101, reaction data at each time step are collected during the complete reaction process of two different bioreactor sizes under the same process.

[0022] The same process means that the cell line, culture medium formula, reaction target (such as product generation, cell expansion), and core operation logic (such as stirring / aeration mode) are completely consistent, and only the nominal volume (size) of the bioreactor is different (e.g., 10L and 1000L stirred tank reactors).

[0023] The complete reaction process covers the entire cycle from reaction initiation (cell seeding) to reaction termination (cell decline / product achievement), rather than fragmented data.

[0024] To address the challenges of adapting bioreactors to multiple factors during volume changes and cell migration processes, the core design logic of this data acquisition strategy revolves around "process migration feasibility" and "volume scaling regularity." Using key, directly controllable operating parameters during bioreactor operation as the central axis, and linking five dimensions—reaction environment characteristics, mass transfer efficiency, fluid dynamics, cell metabolic state, and inherent equipment properties—a dataset is constructed that accurately reflects the correlation between "volume change - process parameters - reaction effect," providing data support for process migration between bioreactors of different volumes.

[0025] In collecting data for each impact category, we closely followed the requirements of process migration and volume change: Culture medium composition: Focus on collecting key parameters related to cell compatibility to ensure that the basic environment of cells in bioreactors of different volumes is consistent, and avoid interference with the judgment of process migration effect due to differences in culture medium characteristics. At the same time, pay attention to the component properties that affect mass transfer (such as dissolved oxygen) to provide a reference for adjusting mass transfer efficiency after volume changes. Mass and heat transfer: Focusing on the transfer patterns of oxygen, nutrients, and metabolites, this study quantifies the core indicators of mass transfer efficiency in bioreactors of different volumes, clarifies the impact of volume changes on mass and heat transfer effects, provides a basis for adapting and adjusting mass transfer-related parameters during process migration, and ensures stable reaction temperatures in bioreactors of different volumes to avoid interference from temperature fluctuations on cell growth and process stability. Fluid properties: Collect basic physical property data of the bioreactor. These properties change with the volume of the bioreactor and the cultivation process, which in turn affects the stirring efficiency and mass transfer effect. The data can support the optimization of fluid-related parameters during process migration and ensure that the influence of fluid behavior on the reaction remains consistent under different volumes. Mixing and shearing: The focus is on analyzing the shear distribution and mixing efficiency in bioreactors of different volumes, clarifying the impact of volume changes on shear intensity and mixing uniformity, especially the influence of shearing on cell activity. This is a key aspect that needs to be controlled during process migration, as it can prevent excessive shearing or uneven mixing caused by volume scaling from affecting cell growth and product generation. Metabolites: Track the consumption of key substrates for cell growth and the generation of metabolites, establish the correlation between cell metabolic state and process parameters in bioreactors of different volumes, verify the stability of cell growth and product generation after process migration through metabolic data, and ensure that cell metabolism can still be maintained in the optimal state after volume changes.

[0026] Bioreactor attribute parameters: Structural characteristic data of bioreactors with different nominal volumes are collected in a targeted manner, including core structural parameters closely related to mixing and mass transfer. The inherent attribute differences of bioreactors of different volumes are clarified, providing a basis for parameter adjustment based on equipment characteristics during process migration, and ensuring that the process can be adapted to the structural characteristics of bioreactors of different volumes. Bioreactor control parameters: Focusing on three core control parameters—stirring, aeration, and working volume—the regulation patterns within bioreactors of different volumes are collected to clarify the impact of volume changes on the optimal range of control parameters, especially the linkage between key control parameters and mass transfer efficiency, as well as the changes in control parameter thresholds caused by volume changes. This provides data support for precise scaling of control parameters during process migration, ensuring that bioreactors of different volumes can achieve optimal reaction results through parameter regulation. The core value of this dataset lies in its integration of online real-time monitored process parameters, offline analyzed biochemical parameters, spectral data for real-time prediction of key indicators, and inherent structural parameters of the equipment, based on measured data from multi-scale bioreactors. By comparing data from bioreactors of different volumes, it uncovers the intrinsic patterns between volume changes and process parameters, as well as reaction effects. Each bioreactor volume underwent multiple repeated experiments to ensure data repeatability and reliability. The resulting multi-dimensional correlated dataset not only provides ample samples for subsequent process scale-up model training but also directly supports process transfer between bioreactors of different volumes. It helps accurately determine key process parameters that need adjustment after volume changes, ensuring stable and efficient operation of the process in bioreactors of varying volumes. The specific data types are shown in Table 1: Table 1

[0027] Based on the above description, if there are multiple samples, the response data at each time step in this disclosure includes: the time-series multimodal observation data of the i-th sample. and the static metadata of the i-th sample The time-series multimodal observation data of the i-th sample includes data at time t. The rich information collected includes: culture media, metabolomics, and physiological microscopic images of the observed subjects. (This microscopic image captures and quantifies visual information such as cell morphology, density, and aggregation state) and sensor readings of the biological response parameters of the i-th sample. (Sensor reading vectors containing key process parameters (KPPs) such as pH, dissolved oxygen (DO), temperature, substrate concentration, and metabolite concentration); static metadata of the i-th sample. A vector describing fixed attributes, including: type information of the basic bioreactor components for the i-th sample. Numerical information including (e.g., cell line ID, culture medium formulation number, vector type) and the bioreactor configuration for the i-th sample. (e.g., nominal reactor volume, height-to-diameter ratio, impeller type and diameter, baffle configuration, maximum aeration rate, etc.); among which, time-series multimodal observation data are data that change with reaction time, while static metadata is data that does not change with reaction time.

[0028] This disclosure uses the training of a set of samples as an example for illustration. It is worth noting that in the following embodiments, if the parameter has i, it represents the parameter corresponding to the i-th sample; if i is not present, it represents the current sample, and the parameter corresponding to the current sample is used.

[0029] In step S102, data from the first B time steps before the initial stage of the reaction are extracted from the reaction data at each time step to form a support set.

[0030] The initial stage of the reaction refers to the period after the reaction has started, when cells are in the adaptation phase / early logarithmic growth phase and have not yet entered a stable state (e.g., data from the first 24 hours, corresponding to a cell density of 1.0 × 10⁻⁶). 6 / ml increased to 6.0×10 6 / ml stage).

[0031] B needs to meet the requirement of "capturing initial dynamic features" but "having a small amount of data" (usually B=5~20 time steps) to simulate the pain point of "cold start of new process" in actual industrial scenarios, where only a small amount of initial data can be obtained and it is impossible to wait for complete data for optimization.

[0032] The support set is a "limited set of clues" used by the model to infer the laws of new processes, rather than complete training data. Its core value is that it contains "process-specific dynamic signals" (such as the initial DO decline rate and glucose consumption pattern of a cell line in a 1000L reactor).

[0033] In step S103, the mean of the posterior Gaussian distribution of the global kinetic latent variable is obtained from all reaction data; the global kinetic latent variable represents the property of the bioreactor that does not change with reaction time during the reaction process.

[0034] Global kinetic latent variables are static vectors that represent the intrinsic kinetic "genes" or "fingerprints" of the entire reaction batch. They are designed to capture the unique, time-invariant macroscopic properties of that batch, such as the maximum specific growth rate of a particular cell line in a particular culture medium, substrate consumption kinetics, and the tendency to generate metabolic byproducts.

[0035] In one embodiment, before performing step S103, the method further includes the following steps A1-A6: A1. Obtain the first global kinetic latent variables corresponding to all reaction data based on the initial HAVE encoder. The posterior Gaussian distribution parameters and the first time-resolved state latent vectors corresponding to each time step. The posterior Gaussian distribution parameters; the posterior Gaussian distribution parameters include: mean.

[0036] It is worth noting that all reaction data in step A1 can be the same as the reaction data in S101, or it can be new reaction data specifically used to train the HAVE encoder and decoder.

[0037] This disclosure presents a HAVE encoder, which is a Hierarchical Variational Autoencoder (HVAE). This HAVE encoder employs deep regularization by incorporating prior physical knowledge to ensure that the learned latent space organization aligns with underlying engineering principles. The model process begins by mapping heterogeneous, high-dimensional reaction data streams to a unified, structured, and physically meaningful low-dimensional latent space.

[0038] In one embodiment, step A1 obtains the first global kinetic latent variable corresponding to all reaction data based on the initial HAVE encoder. The posterior Gaussian distribution parameters are obtained, including the following sub-steps B1-B6: B1. A convolutional neural network using an initial HAVE encoder is used to process the microscopic images at each time step. Feature extraction is performed, converting image pixel information into fixed-dimensional morphological feature vectors. The morphological feature vectors of all time steps are arranged in chronological order to form a morphological feature sequence.

[0039] The EfficientNet-B0 convolutional neural network, pre-trained on ImageNet, is used to extract features from the microscopic images at each time step. Specifically, it extracts a high-dimensional morphological feature vector from the original microscopic images, condensing the image pixel information into a fixed-dimensional morphological feature vector. The morphological feature vectors of all time steps are arranged in chronological order to form a morphological feature sequence.

[0040] B2. Using the first MLP network of the initial HAVE encoder, the sensor readings at each time step are... By performing nonlinear transformation and dimension matching, biochemical feature vectors are obtained. Arrange the biochemical feature vectors of all time steps in chronological order to form a biochemical feature sequence; Sensor readings After nonlinear transformation and dimension matching by the first multilayer perceptron (MLP), a fixed-dimensional biochemical feature vector is obtained.

[0041] B3. Concatenate the feature vectors at the same time step in the morphological feature sequence and the biochemical feature sequence to obtain the time-series comprehensive feature sequence. ; The two feature sequences obtained from B1 and B2 are concatenated and combined with visual morphological information and biochemical parameter information to obtain a temporal integrated feature sequence.

[0042] B4. Input the comprehensive feature sequence into the Bi-GRU model of the initial HAVE encoder to obtain the complete hidden state at each time step. The complete hidden state at each time step includes the reaction data at the current time step, the reaction trend before the current time step, and the reaction result after the current time step. Arrange the complete hidden states of all time steps in chronological order to form a complete hidden state sequence. .

[0043] A bidirectional gated recurrent unit (Bi-GRU) is employed to process the comprehensive feature sequence, fully capturing the "past-present-future" correlation information of the time series data to generate hidden states containing complete temporal context. The comprehensive feature sequence is input into the Bi-GRU model, which captures the "reaction trend at the current time step and before" through a "forward GRU (from the first time step to the last time step)," and captures the "reaction result at the current time step and after" through a "backward GRU (from the last time step to the first time step). Finally, the forward and backward hidden states of each time step are concatenated to obtain the complete hidden state of each time step (containing condensed information of current observations, historical trends, and future results). The complete hidden states of all time steps are arranged chronologically to form a complete hidden state sequence. .

[0044] B5. Hidden state sequence Input the self-attention weighted pooling network of the initial HAVE encoder to obtain the global information vector.

[0045] A self-attention weighted pooling mechanism is applied to the hidden state sequence. Based on the importance of each time step to the "definition of batch global dynamics" (such as higher weights for key time steps like rapid cell growth and metabolic transition), weights are dynamically allocated and the hidden state sequence is aggregated to output a global information vector (rather than a "pooling information sequence" to avoid misleading; the essence is to condense the time sequence into a single vector).

[0046] B6. Input the global information vector into the second MLP network of the initial HAVE encoder to obtain the first global dynamic latent variables. The posterior Gaussian distribution parameters.

[0047] The global information vector is input into the second MLP network, which outputs the posterior Gaussian distribution parameters of the first global dynamic latent variable (including the mean and logarithmic variance; the logarithmic variance can avoid negative variance in the calculation).

[0048] To obtain global information representing the entire sequence, a self-attention pooling mechanism is applied to all hidden states of the Bi-GRU, allowing the model to dynamically focus on the time points most important for defining the global dynamics. The pooled vectors are then fed into an MLP network, which outputs the first global dynamic latent variable. The parameters of the posterior Gaussian distribution, i.e., the first global dynamic latent mean. and the first global dynamic potential log-variance Sample an instance from this distribution: .

[0049] In one embodiment, step A1 obtains the first time-resolved state latent vectors for each time step corresponding to all reaction data based on the initial HAVE encoder. The posterior Gaussian distribution parameters are obtained, including the following sub-steps C1-C2: C1. Concatenate the hidden state and the global dynamic latent variable at each time step in the hidden state sequence to obtain the concatenated sequence.

[0050] The hidden state at each time step in the hidden state sequence is concatenated with the first global dynamic potential vector to make the local information at each time step subject to global dynamic constraints, thus ensuring that the dynamic state conforms to the underlying physical laws.

[0051] C2. Input the concatenated sequence into the third MLP network of the initial HAVE encoder to obtain the first time-resolved state latent vector at each time step. The posterior Gaussian distribution parameters.

[0052] The concatenated sequence (each time step is a fused vector of "hidden state + global variable") is input into a separate third MLP network, and the model outputs the first-time discriminative state latent variables step by step. Posterior Gaussian distribution parameters: mean of dynamic latent state and the log-variance of the latent state Sample an instance from this distribution: .

[0053] A2. Based on static metadata corresponding to all reaction data Calculate the dimensionless physical criterion number, and construct a physical feature vector based on the dimensionless physical criterion number. .

[0054] From static metadata In this process, a set of dimensionless physical criterion numbers with cross-scale invariance or comparability are calculated using engineering formulas, forming a physical characteristic vector. These criteria include: Reynolds number. (characterizing the fluid state), where Where is the fluid density, N is the agitator speed, and D is the maximum rotating diameter of the agitator blades. μ is the fluid dynamic viscosity. Power rating (Characterizing the power consumption of stirring), where P is the effective power transmitted from the motor to the stirring paddle, i.e., the stirring input power, ρ is the fluid density, N is the stirring speed, and D is the stirring paddle diameter. Airflow rate. (Characterizing ventilation efficiency), where Na is the ventilation number, representing the ratio of ventilation rate to impeller discharge rate, reflecting the gas dispersion effect in the liquid; Q is the ventilation rate (volume flow rate), representing the volume of gas introduced into the liquid per unit time. N is the agitator rotation speed, and D is the impeller diameter. Also included is the volumetric oxygen transfer coefficient kLa, estimated based on a semi-empirical formula.

[0055] Reynolds number (fluid state, such as how fast the liquid rotates during stirring, whether there is turbulence), power number (stirring power consumption, such as how much electricity the motor consumes, indirectly reflecting the stirring intensity), aeration number (aeration efficiency, such as whether oxygen can enter the culture medium efficiently), kLa (volume oxygen transfer coefficient, directly related to whether cells can obtain enough oxygen). These are all core indicators of the reactor's physical environment, representing the "physical environment in which the reaction takes place."

[0056] A3. Concatenate the physical feature vector with the type information to obtain the joint input vector.

[0057] Type information (e.g., cell line ID) is converted into a dense vector through an embedding layer. , physical feature vector and The vectors are concatenated to obtain the joint input vector.

[0058] A4. Obtain the static context vector from the joint input vector. .

[0059] The joint input vector is passed through an MLP network to ultimately generate a high-information-density static context vector. This vector is a refined description of the physical environment of the bioreactor.

[0060] A5. The first global dynamic latent variable The posterior Gaussian distribution parameters and the first time-resolved state latent vectors corresponding to each time step. The posterior Gaussian distribution parameters and static context vector are input into the initial decoder to obtain the reconstructed response data.

[0061] After encoding, a generator network (i.e., a decoder) is used. The goal is to reconstruct the original response data with high quality from the first global dynamic latent variables and the first time-resolved state latent vectors. The core breakthrough here lies in preventing the latent space containing the first global dynamic latent variables and the first time-resolved state latent vectors from developing freely. Instead, it uses prior physical knowledge to strongly regularize and guide the formation of their geometric structure, which is the key to achieving cross-scale generalization.

[0062] decoder The core is an autoregressive model (such as a multilayer GRU). When generating reconstructed feature response data at each time step, its input is not only the output of the previous time step and the current time-resolved state latent vector. To achieve a fine-grained response to complex conditions, a feature-level linear modulation (FiLM) mechanism is employed. The first global dynamic latent variable and the static context vector are respectively input into two small MLPs to generate a pair of affine transformation parameters (scale vector) for each gated unit (or its linear transformation layer) within the decoder GRU. and bias vector The decoder's computation is thus dynamically modified to... Similar to related technologies, this disclosure utilizes the GRU mechanism to perform linear mapping of features; a linear transformation layer can be used instead of GRU. This mechanism enables the decoding process to simultaneously understand the intrinsic biological characteristics of the batch (encoded by the first global dynamic latent variable) and its external physical environment (encoded by the static context vector), thereby generating data that highly matches specific conditions.

[0063] A6. Based on the reconstructed reaction data and the collected reaction data, the initial HAVE encoder and the initial decoder are trained using the first loss function to obtain the trained HAVE encoder and the trained decoder; the trained biological response cross-task model also includes: the trained HAVE encoder and the trained decoder.

[0064] To enable the model to "understand" physics, the global dynamic latent variables learned by the HVAE encoder must be able to predict their physical environment. Therefore, an auxiliary prediction network is introduced. And define a physical regularization loss term: ; in, This represents the physical regularization loss value. It is a distance or divergence metric (such as mean squared error MSE or Wasserstein distance). This represents the static context vector of the i-th sample (which includes the physical environment and biological basis). Indicates an auxiliary prediction network, It is the mean of the first global dynamic latent variable inferred from the i-th sample. Used to predict the physical environment vector from the mean of the first global dynamic latent variable of the i-th sample.

[0065] This loss term forces the global dynamic potential variables to be closely related to the macroscopic, interpretable physical environment (such as fluid state, shear force level, mass transfer efficiency, etc.), rather than an uninterpretable black box code.

[0066] The first loss number mentioned above can be expressed as:

[0067] ; The components of this loss function are explained below: First item: This is the reconstruction loss. This is the main part of ELBO, which requires that the data decoded from the latent variables be as consistent as possible with the original input data, ensuring information fidelity. Among these... This represents the first global dynamic latent variable output by the HAVE encoder. This indicates that the potential state vector is distinguished at the first moment. Represents time-series multimodal observation data, Represents the static context vector. Indicates a completely hidden state; Indicates to and Seeking expectations; Indicates that from via HAVE encoder Get from and and through the decoder according to Restore The expected logarithmic probability.

[0068] Second item: It is the global variable KL divergence. It constrains... posterior distribution It approaches a simple prior distribution (Standard normal distribution) This plays a regularization role, making the latent space more regular and smooth. It is its preset weight.

[0069] Third item: It is the dynamic variable KL divergence. It constrains the KL divergence at each time step. posterior distribution Approaching a Prior distribution with condition This achieves the goal of the hierarchical Bayesian model, which is that dynamic changes are based on global characteristics. It is its preset weight.

[0070] Fourth item: This is the physical regularization loss. As mentioned earlier, this is key to injecting engineering knowledge into the model, and its weights are determined by... control.

[0071] The first loss number mentioned above is for the current sample, so no label i is written, but the first loss number is the same for each sample during training.

[0072] By analyzing large datasets covering different scales, cell lines, and operational parameters... Through optimization, a high-quality latent space was finally obtained. This space not only accurately represents the original reaction data, but more importantly, its internal organization (especially the distribution of the first global dynamic latent variable) reflects real physical and biological laws, laying the foundation for subsequent dynamic modeling and cross-scale generalization.

[0073] At this point, the step S103, which involves obtaining the mean of the posterior Gaussian distribution of the global dynamic latent variables through all response data, includes: encoding all response data using the trained HAVE encoder to obtain the mean of the posterior Gaussian distribution of the global dynamic latent variables.

[0074] In step S104, the forward noise network of the conditional diffusion model performs K-step Gaussian noise addition on the global dynamic latent variables, and finally obtains the noise-added samples.

[0075] This disclosure also introduces a conditional diffusion model operating on the space of the global dynamic latent variables (i.e., the "dynamic parameter manifold"). The goal of this model is no longer to predict a single global dynamic latent variable, but rather to learn and sample the entire posterior probability distribution. This allows for the generation of an ensemble of dynamic models, providing a complete probabilistic description of future predictions.

[0076] Treating the global dynamic latent variable as a data point, the conditional diffusion model includes two processes: a forward noise-adding process (a fixed process that does not require training) and a reverse noise-reducing process (a process that requires learning).

[0077] The forward noise addition process is a Markov chain that starts with a "clean" global dynamic latent variable and iteratively adds Gaussian noise to it, for a total of [number of steps]. Step. The noise addition process of the step is defined as follows: ; in, Represents the noisy transition distribution (representing the global dynamic latent variable after adding noise from step k-1). The global dynamic latent variables are transferred to the noisy global dynamic latent variables at the k-th step. The probability distribution is the core transition rule of the forward process in the conditional diffusion model. This represents the global dynamic latent variable at step k after adding noise. This represents the global dynamic latent variable at the (k-1)th step after adding noise. It is a pre-set, gradually increasing noise variance scheduling table. After... After stepping, The distribution will be very close to a standard normal distribution.

[0078] In step S105, a time-resolved state latent vector sequence is obtained through the support set. The time-resolved state latent variable sequence represents the properties of the bioreactor that change with reaction time during the reaction process.

[0079] This time-resolved state latent vector sequence captures the early dynamic behavior of the training task.

[0080] In one embodiment, before performing step S105, the method further includes the following sub-steps D1-D3: D1. Based on the trained HAVE encoder, obtain the sequence of the first mean in the posterior Gaussian distribution parameters of the second global dynamic latent variable and the second mean in the posterior Gaussian distribution parameters of the second time-resolved state latent vector corresponding to all reaction data.

[0081] D2. Input the first mean into the initial supernetwork to obtain the predicted mean sequence of the second time-resolved state latent vector.

[0082] D3. The initial supernetwork is trained based on the second mean sequence and the predicted mean sequence of the second time-discriminated state latent vector to obtain the trained supernetwork; the trained biological response cross-task model further includes the trained supernetwork. Specifically, mean squared error is used to quantify the difference between the second mean sequence and the predicted mean sequence of the second time-resolved state latent vector. The initial supernetwork is then trained through backpropagation to obtain the trained supernetwork.

[0083] After successfully constructing a structured and physically rich latent space using the HVAE encoder, the next core task is to accurately model the temporal evolution of the system within this space. This disclosure abandons discrete-time, RNN-based transition models because they struggle to guarantee prediction consistency across different time steps and lack physical interpretability. Instead, it assumes that the system's evolution follows a continuous-time latent neural ordinary differential equation (L-NODE). Furthermore, it is argued that this differential equation itself is not universal but uniquely determined by the global dynamics of the reaction. It is assumed that the evolutionary flow field of the time-resolved state latent vector of the reaction data in the latent space is parameterized by its global dynamic latent variables. Its mathematical expression is: ; in, Let t represent the second time-resolved state latent vector, and t represent the time step. Represents the static context vector. Represents a vector field function. The parameter mapping function represents the second global dynamic latent variable, indicating the mapping of the second global dynamic latent variable. (High-dimensional / low-dimensional vectors) are mapped to model parameters of the vector field function f.

[0084] The key idea here is: It is a vector field function parameterized by a neural network, which defines the direction and velocity of motion of the second time-resolved state latent vector at any point in the latent space. The weight parameters of this neural network... The set is not a fixed, globally constant that needs to be learned. Instead, it is a function of the global dynamic latent variables. This means that different reaction batches (with different global dynamic latent variables) will have different "physical laws" (i.e., different differential equations) for their latent states. The evolutionary process is also explicitly conditional on a static context vector, meaning that the macroscopic physical environment of the bioreactor can directly influence the instantaneous dynamics.

[0085] To model the weight parameters from abstract global dynamic latent variables to specific differential equations. This mapping relationship introduces a hypernetwork. It is itself a neural network (again using MLP), which learns a higher-order function, whose input is the parameters (or representations) of another network, and whose output is the parameters of the target network. Within this framework: ; This hypernetwork (its parameters are) It takes a second global dynamic latent variable (also known as a dynamic "gene") as input and outputs an L-NODE vector field that controls the dynamic evolution of this batch. All weight parameters This mechanism endows the model with great flexibility and generalization ability because it learns not a single dynamic model, but the meta-ability of "how to generate a suitable dynamic model based on global characteristics".

[0086] In order to train this complex dynamic system (i.e., train the supernetwork) parameters This requires a supervisory signal to ensure the correctness of the generated dynamic model. This supervisory signal comes from the second time-resolved state latent vector (also known as the "true" latent trajectory) output by the trained HAVE encoder. A Dynamics Consistency Loss is defined for this purpose. The specific training process is as follows: 1. First, based on the trained HAVE encoder, obtain the sequence of first mean values ​​in the posterior Gaussian distribution parameters of the second global dynamic latent variable corresponding to the reaction data and the second mean value sequence in the posterior Gaussian distribution parameters of the second time-resolved state latent vector.

[0087] 2. Then, the first mean is input into the initial supernetwork. In the process, L-NODE parameters tailored to the reaction data are generated in real time: , Let represent the first mean among the posterior Gaussian distribution parameters of the second global dynamic latent variable.

[0088] 3. Using the second mean of the first time step in the second mean sequence as the initial condition, a differentiable numerical ODE solver (implemented using the adjoint method in Dopri5) is used to solve the ODE. Integrating the parameterized L-NODE, the time interval is from arrive ,in, This represents the total number of time steps (time series length). This will generate a predicted potential trajectory driven entirely by the inferred dynamic model, i.e., a predicted sequence of second time-resolved state potential vectors, as shown in the following formula: ; in, This represents the predicted sequence of the second time-resolved state latent vectors. This represents the second mean of the second mean sequence at the first time step. Indicates the start time of integration (initial moment). This represents a parameterized vector field function. Representing the integral variable The second time-resolved state latent vector at time step (the "intermediate state variable" during the integration process).

[0089] 4. Finally, the dynamic consistency loss is defined as the difference between the predicted sequence of the second time-resolved state latent vector generated by the L-NODE integral and the second mean sequence of the posterior Gaussian distribution parameters of the second time-resolved state latent vector inferred directly from the reaction data by the HAVE encoder. This difference is typically measured using mean squared error (MSE). ; in, This represents the dynamic consistency loss value. This represents the total number of time steps (time series length) for the i-th sample. This represents the sequence of second mean values ​​in the posterior Gaussian distribution parameters of the second time-resolved state latent vector of the i-th sample output by the trained HAVE encoder. This represents the predicted sequence of the second time-resolved state latent vector of the i-th sample output by the initial supernetwork.

[0090] this The loss term plays a crucial bridging role between representation learning and dynamics modeling. It optimizes the parameters of both the HVAE encoder and the hypernetwork simultaneously through backpropagation. This forces the latent space structure learned by the HVAE encoder to be consistent with a continuous, hypernetwork-parameterized dynamic system. In other words, points in the latent space cannot be arbitrarily distributed; they must be connected by a smooth flow field. By minimizing this loss, the model learns how to infer complete "physical laws" from static global dynamic latent variables that accurately describe all future dynamic behaviors of the system. This is the core mechanism for achieving long-term, physically plausible extrapolation predictions of new processes.

[0091] At this point, obtaining the time-resolved state latent vector sequence through the support set includes: encoding the support set through the trained hypernetwork to obtain the time-resolved state latent vector sequence.

[0092] In step S106, the time-resolved state potential vector sequence is input into the task encoder to obtain the dynamic context vector.

[0093] To refine the time-resolved state latent vector sequence into a compact, conditionally usable representation, it is fed into a specialized task encoder. (its parameters are) This task encoder employs a Transformer architecture with a self-attention mechanism, and its final output is a fixed-dimensional dynamic context vector. This vector encapsulates all the key clues about the dynamic characteristics of the new task.

[0094] In step S107, the denoising network of the conditional diffusion model predicts the noise added at each step based on the noisy sample and the dynamic context vector.

[0095] This step is the reverse denoising process, which is the core learning task of the model. A denoising network is trained, the goal of which is to predict the denoising value at the _th ... The noise added step. This prediction must be based on the current noisy global dynamic latent variables. Noise level And the most important information provided by the supporting set is the condition.

[0096] Deep conditional injection and meta-learning context enhance the performance of denoising networks, which directly depends on how effectively they utilize conditional information from the support set. To address this, a deep conditional injection mechanism based on meta-learning principles was designed.

[0097] The architecture of the denoising network (e.g., a multi-layered MLP or a Transformer for high-dimensional global dynamic latent variables) is designed to accept multiple inputs. At each step of the denoising process, the input includes not only the noisy global dynamic latent variables but also a representation of the noise level. The temporal embedding also includes the extracted dynamic context vector. That is, the network computes... . The dynamic context vector is deeply injected into each layer of the denoising network through a cross-attention mechanism. This ensures that each step of the denoising process is precisely guided by the unique dynamic characteristics of the new task.

[0098] In step S108, the denoising network and task encoder are trained based on the real noise added at each step and the predicted noise added at each step, to obtain the trained conditional diffusion model and task encoder. The trained biological response cross-task model includes the trained conditional diffusion model and the trained task encoder.

[0099] Specifically, an L2 loss function is constructed based on the real noise added at each step and the predicted noise added at each step. The denoising network and the task encoder are simultaneously optimized through backpropagation of the loss, resulting in the trained conditional diffusion model and task encoder.

[0100] The training of the conditional diffusion model follows an episoded meta-learning paradigm, teaching it how to utilize the support set. In each training iteration: a. Collect reaction data at each time step in the complete reaction process of two different bioreactor sizes under the same cell technology.

[0101] b. Extract the data from the reaction data at each time step, specifically the data from the first B time steps before the initial stage of the reaction, to form the support set. .

[0102] c. Use the trained HVAE encoder to infer global dynamic latent variables from the complete reaction data.

[0103] d. From the support set Extracting dynamic context vectors .

[0104] e. Perform the standard training steps for the conditional diffusion model, i.e., sample a random noise level and a random noise source. Constructing noisy samples ,in, This represents the global dynamic latent variable at step k after adding noise. Represents the initial global dynamical latent variables. This is the cumulative product parameter in the conditional diffusion model, used to describe the "cleanliness" of the data at step k. The closer it is to 1, the cleaner the data; the closer it is to 0, the closer the data is to pure noise. , The noise scheduling parameter (also called "step noise figure") is a manually set sequence of small values ​​(usually starting from a very small number and gradually increasing) that controls the intensity of noise added at each step. (This represents real Gaussian noise).

[0105] By progressively removing noise from the noise data features, the model is fitted to the optimal reaction data features of a large-scale bioreactor. The reaction data of a small-scale bioreactor is used as a condition to guide the diffusion process of the model, thus moving closer to the optimal large-scale data.

[0106] f. Minimize the error between the predicted noise and the actual noise of the denoising network using L2 loss: ; Among them, L diff (ω,χ,ϕ) represents the loss function of the conditional diffusion model; ω is the model parameter of the denoising network in the conditional diffusion model; χ is the scheduling parameter of the diffusion process (noise variance table). For the model parameters of the HAVE encoder; E task,k,ϵ [⋅] represents the expectation operator (expectation of the task, number of noise addition steps, and noise); task is the task index (sample / scene identifier); k is the number of noise addition steps (diffusion step index); This is real Gaussian noise (original noise with added noise); The predicted noise for the denoising network; This represents the global dynamic potential variable at the k-th step after adding noise; This is a dynamic context vector.

[0107] By repeating this process on a large number of different tasks, the conditional diffusion model learns how to react to any given initial dynamic cues ( ), to perform accurate, conditional probability density estimation on the manifold of dynamic parameters.

[0108] In summary, this disclosure integrates the conditional diffusion model into a cross-task model architecture for biological reactions. Leveraging the inherent advantages of the conditional diffusion model in small-sample learning scenarios, it requires only a small amount of data from the first T time steps of the initial reaction to construct a support set, enabling effective model training. Unlike traditional mathematical estimation models that rely on large amounts of complete, long-term reaction data to fit the reaction patterns, the conditional diffusion model performs k-step Gaussian noise addition on the global dynamic latent variables through a forward noise addition process, achieving implicit data augmentation for limited data. This effectively expands the information dimension and diversity of the effective samples during the training phase. Combined with a denoising network for accurate noise prediction learning, the model can fully extract the core dynamic characteristics and initial state patterns of the reaction from a small amount of initial data, achieving high-quality training without relying on large amounts of complete reaction data. This perfectly solves the model training problem caused by insufficient data during the cold start phase of new processes. The cross-task model for biological reactions trained based on the conditional diffusion model requires only a small amount of real initial reaction data in practical applications to achieve complete prediction of the entire biological reaction trajectory. On the one hand, during the training phase, the model has deeply learned the intrinsic relationship between the initial reaction data and the global reaction law through the noise addition and denoising mechanism of the conditional diffusion model. It can extract the global dynamic characteristics of the reaction (represented by global dynamic latent variables) and the temporal evolution logic (represented by time-resolved state latent vector sequences) from the limited initial data. On the other hand, the task encoder deeply integrates the dynamic context vector transformed from the initial data with the diffusion denoising network, enabling the model to accurately infer the reaction state changes at subsequent time steps based on the dynamic characteristics of the initial data during the prediction process. This changes the limitation of traditional models that "require a large amount of data input to ensure prediction accuracy," and can still provide accurate basis for reaction state assessment and product yield prediction in data-scarce scenarios such as new process pilot production and small-batch experiments.

[0109] It is worth noting that in this disclosure, the HAVE encoder, decoder, supernetwork, and denoising network in the conditional diffusion model of the biological response cross-task model training based on dynamic context (also known as the HVDMD model) can not only be trained independently in their respective training stages, but also adaptively retrain the previous networks when training other networks later.

[0110] Furthermore, in this disclosure, the networks can be trained independently, but rather as an organic whole, undergoing end-to-end collaborative optimization through a unified, weighted joint loss function. This ensures that information flows between modules in a lossless and efficient manner.

[0111] The complete set of learnable parameters for the HVDMD model is These correspond to the HVAE encoder, decoder, hypernetwork, conditional diffusion model, and task encoder, respectively. The overall loss function is a weighted sum of the losses of each module: ; in, These are hyperparameters used to balance the importance of representation learning, dynamic consistency, and probabilistic generation. The training process strictly follows the aforementioned episodic meta-learning paradigm, sampling multiple tasks in each batch and partitioning the support set / query set within each task, then calculating the total loss. And update all parameters through backpropagation. .

[0112] The HVDMD model disclosed herein aims to shift from predicting "phenomena" to generating "laws." Reactors of different sizes are viewed as different projections or realizations within a unified, high-dimensional "bioprocess kinetic space." Each point in this abstract space represents a complete set of kinetic laws describing the evolution of a specific reaction process. Therefore, the fundamental task of the model is to learn the intrinsic geometry of this unified space and generate a high-level meta-ability: when faced with a scaled-up reaction task of a novel biological reaction process, the model should be able to locate the region corresponding to the new task within this kinetic space using only limited initial observational data, and make probabilistic and evidence-based inferences about the kinetic characteristics of that region.

[0113] The HVDMD framework achieves this goal through three deeply coupled and collaboratively optimized core modules: (1) a hierarchical variational autoencoder for mapping data to the latent space of physical meaning; (2) a neural ODE and supernetwork system for inferring continuous-time dynamic models from latent representations; and (3) a dynamic parameter spatial conditional diffusion model for uncertainty quantification and regularity generation under new tasks. These three steps together constitute a complete and coherent information processing chain from data representation and dynamic identification to probabilistic generation.

[0114] This disclosure also provides deployment strategies: Given the sensitivity and distributed nature of biopharmaceutical data, the HVDMD framework is naturally suited for deployment and training within a federated learning framework. Individual companies or laboratories can train the global model locally using their private data without sharing the original data. They only need to securely upload the model's parameter updates (not the data itself) to a central server. The server then uses aggregation algorithms such as weighted averaging (e.g., FedAvg) to merge the "wisdom" from multiple sources, update the global model, and then distribute it back to each party for the next round of local training. Through this decentralized, co-evolutionary model, the HVDMD model can learn from diverse data globally, ultimately evolving into a truly intelligent biological process prediction and optimization engine with industry-wide applicability, thus solving the challenge of universality in bioreactor algorithms.

[0115] Core deployment process: The federated learning training of the HVDMD framework adopts a closed-loop process of "global initialization → local training → parameter aggregation → model distribution": Phase 1: Global HVDMD model initialization (server side).

[0116] Prior knowledge embedding: Combine biological process mechanisms (such as the Monod equation and product synthesis kinetics model) to initialize the core data of HVDMD, avoiding slow convergence caused by the model learning from scratch; Process layering and adaptation: Different initial model structures are designed for different biopharmaceutical scenarios (such as upstream fermentation and downstream purification). For example, the fermentation scenario strengthens the dynamic correlation modeling of "temperature-pH-dissolved oxygen", while the purification scenario focuses on capturing the time-series dependence of "flow rate-pressure-purity". Initial model distribution: Encrypt and distribute the global model parameters (non-data) to all local nodes (such as enterprise production workshops and scientific research laboratories) participating in federated training.

[0117] Phase 2: Local node training (data does not leave the local machine).

[0118] Data preprocessing (compliance): Local data is cleaned according to biopharmaceutical data standards (such as FDA 21 CFR Part 11) to remove invalid time-series data such as abnormal reactor shutdowns and sensor malfunctions; Sensitive information (such as enterprise process parameter ranges and batch numbers) is anonymized, and only the "feature-label" pairs required for model training (such as "temperature / pH sequence → product concentration prediction value") are retained. HVDMD local training: Based on local private data (such as 1000 time steps of a certain batch of fermentation), the HVDMD model is trained, with a focus on optimizing the "dynamic mode extraction" module (to capture the process fluctuation patterns of the local reactor). Only the parameter update amount (such as the gradient change of model weights and the adjustment value of modal decomposition coefficients) is calculated, without uploading any raw data, thus reducing the risk of privacy leakage; Parameters are uploaded with encryption: The encrypted parameter update is sent to the central server through a secure communication layer, along with the local data volume and process type label (for subsequent weighted aggregation).

[0119] Phase 3: Global parameter aggregation (server side).

[0120] Weighted aggregation algorithm optimization: In response to the highly heterogeneous nature of biopharmaceutical data, FedProx (Federated Proximal Aggregation) is adopted to replace the traditional FedAvg. By introducing proximal term penalties, it alleviates the "model shift" caused by differences in data distribution at different nodes (such as different bioreactor types) and ensures that the global model is adaptable to multiple scenarios. Aggregation weights are assigned based on the amount of data and the representativeness of the process at the local node (e.g., whether it covers high / low yield batches) to avoid excessive influence of small sample nodes on the global model. Post-aggregation verification: The aggregated global HVDMD model was tested using a server-side "public validation set" (such as anonymized industry standard process data) to verify whether its prediction errors (such as the product concentration prediction MAE and the fermentation cycle prediction accuracy) met the threshold. If the verification fails, return to adjust the aggregate weights or trigger the node to retrain (e.g., check for nodes with abnormal parameter updates).

[0121] Phase 4: Global model distribution and iteration.

[0122] Model version update: The validated new global model is encrypted and distributed to all local nodes, overwriting the old model; Multiple rounds of iterative optimization: Repeat the process of "local training → parameter aggregation → model distribution" (usually 10-50 rounds) until the global HVDMD model achieves stable accuracy on the local test set of each node (e.g., product concentration prediction error <5%). Incremental training support: When adding a new bioreactor node or updating the process, there is no need to restart the full training. The parameters of the new node are integrated into the global model through "local incremental aggregation", which adapts to the process iteration needs of biopharmaceuticals.

[0123] Key technology adaptation: Co-optimization of HVDMD and federated learning.

[0124] To address the specific needs of the biopharmaceutical scenario, federated learning adaptation is required to suit the characteristics of the HVDMD framework. The core optimization points are as follows: 1. Parameter update volume compression: adapted to narrow bandwidth scenarios.

[0125] Biopharmaceutical manufacturing bases are often located in suburban areas with limited network bandwidth. The "dynamic mode decomposition" feature of the HVDMD framework can naturally compress parameter size—only "modal coefficient updates" (rather than the full weights) need to be transmitted, reducing parameter volume by 60%-80% compared to traditional deep learning models and reducing transmission latency.

[0126] 2. Heterogeneous data adaptation: mitigating model bias caused by "process differences".

[0127] The significant differences in reactor types (e.g., stirred tank, airlift) and process parameter ranges (e.g., fermentation temperature 28℃ vs 32℃) among different companies can easily lead to "model fragmentation" in federated training. HVDMD addresses this by: Dynamic modal alignment: During the local training phase, time series data from different processes are mapped to a unified "modal space" (e.g., by standardizing modal frequencies) to reduce the distribution differences of heterogeneous data; Federated distillation assistance: "Knowledge distillation" is introduced on the server side to distill the "general process knowledge" of the global model (such as the key time window for product synthesis) to each local node, thereby improving the model's versatility under different processes.

[0128] 3. Small sample training and adaptation: to cope with the scarcity of laboratory data.

[0129] In research laboratories or newly commissioned production facilities with limited local data (e.g., only 10-20 fermentation batches), traditional federated learning is prone to overfitting. HVDMD's "mechanism-data hybrid driven" characteristic can alleviate this problem: During local training, some modal parameters are initialized by combining biological process mechanism models (such as known cell growth curves) to reduce dependence on sample size; The server provides a "pre-trained modality library" (such as the fermentation fluctuation modality commonly used in the industry), and local nodes can fine-tune it based on small samples to quickly adapt to their own processes.

[0130] Key points of this disclosure: 1. This disclosure, through a unique and deeply coupled multi-module design, fundamentally solves the core pain points of traditional bioreaction models, namely poor versatility and low prediction accuracy in process scale-up and migration. The significant technical effects it produces can all be rigorously derived from the technical solutions disclosed herein, rather than being simple statements of conclusions.

[0131] 2. This disclosure proposes a novel bioreactor scale-up scheme, achieving a leap from "fitting phenomena" to "generation laws," and constructs an industry-wide universal bioreactor cross-scale prediction engine, solving the problem of bioreactor cross-scale generalization.

[0132] 3. For distributed scenarios in biopharmaceutical manufacturing, HVDMD federated learning deployment adopts a three-layer architecture of "central server - local nodes - secure communication layer," balancing training efficiency, data privacy, and process adaptability. It can address industry pain points such as "insufficient versatility" and "data silos" in bioreactor models while protecting data privacy. 4. This disclosure can be well generalized in both reactor scale and process dimensions, and has high practical and commercial application value.

[0133] Traditional models fail after scaling or transfer because they learn the apparent trajectory of observed data under specific conditions (such as a 10L reactor), rather than the intrinsic dynamics of the control system's evolution. This patent, through the following design, achieves the generation of dynamic laws, thus possessing unprecedented generalization capabilities: 1. High-precision, physically self-consistent long-term extrapolation predictions have been achieved.

[0134] Traditional models based on recurrent neural networks (RNNs) suffer from rapid error accumulation and may produce predictions that contradict physical reality during long-term forecasting. This patent significantly improves the long-term stability and physical accuracy of predictions by modeling continuous-time dynamics.

[0135] Derivation process: The core dynamic model disclosed herein is the Latent Neural Differential Equation (L-NODE). Unlike discrete models such as RNNs, L-NODE learns a continuous vector field defined over the entire latent space. Predicting future states is achieved by numerically integrating this vector field. This approach naturally guarantees the smoothness of the predicted trajectory and the consistency of predictions at arbitrary time steps.

[0136] The dynamic consistency loss plays a crucial role in training. It requires that the trajectory generated by the L-NODE integral must closely match the trajectory encoded by the HVAE encoder from real data. This forces the latent spatial structure learned by the HVAE encoder to be "integrable," meaning that points in space must be connectable by a smooth, continuous dynamic manifold.

[0137] Therefore, when the model performs extrapolation predictions, it does not make discrete, error-prone jumps, but rather "slides" on a learned, physically self-consistent continuous manifold. This greatly reduces the rate of error accumulation and ensures the accuracy and stability of long-term predictions.

[0138] Technical Results: By employing continuous-time L-NODE and constraining with dynamic consistency loss, this disclosure can generate long-term, stable, and physically self-consistent predicted trajectories, overcoming the shortcomings of poor extrapolation ability of traditional discrete models.

[0139] 2. It enables quantitative uncertainty assessment of new process predictions, providing a quantitative basis for risk control and decision-making.

[0140] When faced with a new process with only a limited amount of initial data, any deterministic point prediction carries extremely high risk. This disclosure innovatively introduces a probabilistic generative model that can quantify the uncertainty of prediction.

[0141] Derivation process: This disclosure recognizes that a novel task's unique global dynamics cannot be accurately inferred from only a finite support set. Therefore, this disclosure introduces a conditional diffusion model in the third step. The core task of this model is not to predict a single variable, but to learn the entire posterior probability distribution.

[0142] During the inference phase, the model uses the dynamic context vector extracted from the initial data of the new task as a condition, and samples from this posterior distribution multiple times to generate an ensemble of dynamic variables.

[0143] Each sample in this ensemble represents a possible hypothesis about the true "genes" of the new task. By feeding this ensemble sequentially into the subsequent supernetwork, L-ODE solver, and decoder, the model ultimately outputs an ensemble of future observation data trajectories.

[0144] The distribution range of this trajectory ensemble (e.g., calculating its 5% and 95% quantiles at each time point) directly constitutes the confidence interval for the prediction.

[0145] Technical Impact Conclusion: By performing conditional probability generation on the manifold of dynamic parameters, this disclosure elevates prediction from a single deterministic value to a probabilistic distribution with confidence intervals. This enables engineers to quantify the risk of predictions, assess process robustness, and thus gain unprecedented quantitative support for critical decisions such as feeding strategies and process optimization.

[0146] 3. Significantly improved data utilization efficiency and model interpretability.

[0147] The architecture design disclosed herein makes it more efficient and transparent in practical applications.

[0148] Derivation Process (Data Efficiency): The training paradigm disclosed herein is essentially an episodic meta-learning approach. By training the model to "infer the entire dynamic law from a small amount of support set data," it acquires the ability to quickly adapt to new tasks. When facing new processes, there is no need to train from scratch or perform large-scale fine-tuning; only a small amount of initial data is required to make high-quality probabilistic predictions, greatly saving experimental costs and time.

[0149] Derivation Process (Interpretability): As described in Effect 1, physical regularization establishes a clear connection between global dynamic variables and interpretable engineering parameters (static context vectors). Researchers can analyze the learned dynamic parameter manifolds to explore how different physical environments (such as different stirring intensities and aeration rates) affect the intrinsic dynamic characteristics of biological processes, potentially discovering new directions for process optimization and overcoming the "black box" dilemma of traditional deep learning models.

[0150] Technical results: This disclosure enables rapid adaptation to small samples, reducing the need for new process data; at the same time, its potential space of physical regularization provides a window for understanding the process mechanism, enhancing the credibility and practical value of the model.

[0151] The HVDMD framework disclosed herein is highly modular and extensible. Its core technology—namely, decoupling static laws from dynamic states through hierarchical variables, regularization using physical priors, and dynamically generating specific evolutionary rules through hypernetworks and neural ODEs—is indispensable for achieving cross-scale generalization. However, the specific implementation methods are flexible: for example, the conditional diffusion model used for probabilistic inference can be replaced by a more efficient conditional VAE or normalized flow, and specific network components (such as GRU, FiLM) can also be replaced according to task requirements.

[0152] This framework can be widely extended to other complex dynamic systems, such as predicting the full life decay curve of a battery using a small amount of early data, assessing the long-term performance degradation of engineered materials, or predicting the unique response patterns of different patients to drugs in personalized medicine, providing a new, probabilistic paradigm for prediction and analysis for all systems with inherent "dynamic personalities".

[0153] Based on the same inventive concept, this application also provides a dynamic context-based biological response cross-task model training device for implementing the above-described dynamic context-based biological response cross-task model training method. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the dynamic context-based biological response cross-task model training device provided below can be found in the limitations of the dynamic context-based biological response cross-task model training method described above, and will not be repeated here.

[0154] In one exemplary embodiment, such as Figure 2 As shown, a training device for a cross-task model of biological response based on dynamic context is provided. The training device includes: The first acquisition module is used to acquire reaction data at each time step in the complete reaction process of two different bioreactor sizes under the same process; The extraction module is used to extract data from the reaction data of each time step, specifically the data from the first T time steps before the initial stage of the reaction, to form a support set. The first acquisition module is used to acquire the mean of the posterior Gaussian distribution of the global kinetic latent variables through all reaction data; the global kinetic latent variables represent the properties of the bioreactor that do not change with reaction time during the reaction process; The noise-adding module is used to perform K-step Gaussian noise addition on the global dynamic latent variables according to the forward noise-adding network of the conditional diffusion model, and finally obtain the noise-adding sample. The second acquisition module is used to acquire the time-resolved state latent vector sequence through the support set; the time-resolved state latent variable sequence represents the properties of the bioreactor that change with reaction time during the reaction process; The third acquisition module is used to input the time-resolved state potential vector sequence into the task encoder to obtain the dynamic context vector. The first prediction module is used to predict the noise added at each step according to the denoising network of the conditional diffusion model, based on the noisy sample and the dynamic context vector. The first training module is used to train the denoising network and the task encoder based on the real noise added at each step and the predicted noise added at each step, so as to obtain the trained conditional diffusion model and the task encoder. The trained biological response cross-task model includes the trained conditional diffusion model and the trained task encoder.

[0155] In one embodiment, the data at each time step includes: time-series multimodal observation data and static metadata. The time-series multimodal observation data includes: sensor readings of culture medium, metabolomics, physiological microscopic images of the observed objects, and biological reaction parameters. The static metadata includes: type information of the basic components of the bioreactor and numerical information of the bioreactor configuration. The time-series multimodal observation data is data that changes with reaction time, while the static metadata is data that does not change with reaction time.

[0156] In one embodiment, the apparatus further includes: The fourth acquisition module is used to acquire the posterior Gaussian distribution parameters of the first global dynamic latent variables corresponding to all reaction data and the posterior Gaussian distribution parameters of the first time-resolved state latent vectors at each time step based on the initial HAVE encoder; the posterior Gaussian distribution parameters include: mean; The first construction module is used to calculate the dimensionless physical criterion number based on the static metadata corresponding to all reaction data, and to construct the physical feature vector based on the dimensionless physical criterion number. The first concatenation module is used to concatenate the physical feature vector with the type information to obtain a joint input vector; The fifth acquisition module is used to obtain the static context vector based on the joint input vector; The sixth acquisition module is used to input the posterior Gaussian distribution parameters of the first global dynamic latent variable, the posterior Gaussian distribution parameters of the first time-resolved state latent vector corresponding to each time step, and the static context vector into the initial decoder to obtain the reconstructed response data. The second training module is used to train the initial HAVE encoder and the initial decoder using the first loss function based on the reconstructed reaction data and the collected reaction data, so as to obtain the trained HAVE encoder and the trained decoder; the trained biological response cross-task model also includes: the trained HAVE encoder and the trained decoder. In obtaining the mean of the posterior Gaussian distribution of the global dynamic latent variables from all reaction data, the first acquisition module is specifically used for: The trained HAVE encoder encodes all response data to obtain the mean of the posterior Gaussian distribution of the global dynamic latent variables.

[0157] In one embodiment, the fourth acquisition module is specifically used for: A convolutional neural network with an initial HAVE encoder is used to extract features from the microscopic image at each time step, converting the image pixel information into a fixed-dimensional morphological feature vector. The morphological feature vectors of all time steps are arranged in chronological order to form a morphological feature sequence. The first MLP network of the initial HAVE encoder is used to perform nonlinear transformation and dimension matching on the sensor readings at each time step to obtain biochemical feature vectors; the biochemical feature vectors of all time steps are arranged in chronological order to form a biochemical feature sequence. The feature vectors at the same time step in the morphological feature sequence and the biochemical feature sequence are concatenated to obtain the time-series comprehensive feature sequence. The comprehensive feature sequence is input into the Bi-GRU model of the initial HAVE encoder to obtain the complete hidden state at each time step. The complete hidden state at each time step includes the reaction data at the current time step, the reaction trend before the current time step, and the reaction result after the current time step. The complete hidden states of all time steps are arranged in chronological order to form a hidden state sequence.

[0158] The hidden state sequence is input into the self-attention weighted pooling network of the initial HAVE encoder to obtain the global information vector; The global information vector is input into the second MLP network of the initial HAVE encoder to obtain the posterior Gaussian distribution parameters of the first global dynamic latent variable.

[0159] In one embodiment, the fourth acquisition module is specifically used for: The complete hidden state at each time step in the hidden state sequence and the first global dynamic latent variable are concatenated into a vector to obtain the concatenated sequence. The spliced ​​sequence is input into the third MLP network of the initial HAVE encoder to obtain the posterior Gaussian distribution parameters of the first time-resolved state latent vector at each time step.

[0160] In one embodiment, the apparatus further includes: The seventh acquisition module is used to acquire, based on the trained HAVE encoder, the sequence of the first mean in the posterior Gaussian distribution parameters of the second global dynamic latent variable corresponding to all reaction data and the second mean in the posterior Gaussian distribution parameters of the second time-resolved state latent vector. The eighth acquisition module is used to input the first mean into the initial supernetwork to obtain the predicted mean sequence of the second time-resolved state latent vector; The third training module is used to train the initial supernetwork based on the second mean sequence and the predicted mean sequence of the second time-discriminated state potential vector to obtain the trained supernetwork; the trained biological response cross-task model also includes the trained supernetwork. The second acquisition module is specifically used for: The trained hypernetwork encodes the support set to obtain a time-resolved state latent vector sequence.

[0161] In one embodiment, the third training module is specifically used for: The difference between the second mean sequence and the predicted mean sequence of the second time-resolved state latent vector is quantified using mean squared error. The initial supernetwork is then trained by backpropagation to obtain the trained supernetwork.

[0162] In one embodiment, the first training module is specifically used for: An L2 loss function is constructed based on the real noise added at each step and the predicted noise added at each step. The denoising network and the task encoder are simultaneously optimized through backpropagation of the loss, resulting in the trained conditional diffusion model and task encoder.

[0163] Figure 3 This is a flowchart illustrating a cross-scene parameter generation method according to an exemplary embodiment, such as... Figure 3 As shown, the method includes the following steps S201-S209: S201. Collect real data during the initial stage of the reaction. The real data includes: real time-series multimodal observation data and real static metadata; the real static metadata includes: real type information of the basic components of the bioreactor and real numerical information of the bioreactor configuration.

[0164] The data type of the real data here is similar to the reaction data in S101.

[0165] S202. Obtain the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence based on real data.

[0166] Specifically, a pre-trained HAVE encoder is used to encode real data to obtain the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence.

[0167] S203. Using the trained task encoder, obtain the real dynamic context vector based on the real time-resolved state latent vector sequence.

[0168] S204. Based on real static metadata, calculate the real dimensionless physical criterion number, and construct the real physical feature vector based on the real dimensionless physical criterion number.

[0169] S205. Concatenate the real physical feature vector with the real type information to obtain the real joint input vector.

[0170] S206. Obtain the real static context vector based on the real joint input vector; the real static context vector includes: physical environment and biological basis.

[0171] S207. Using the trained conditional diffusion model, based on the real dynamic context vector and each real global dynamic latent variable, diffuse each real global dynamic latent variable to obtain multiple diffused real global dynamic latent variables corresponding to each real global dynamic latent variable.

[0172] S208. Based on the real time-resolved state potential vector sequence and each diffusion real global dynamic potential variable, obtain the real time-resolved state potential prediction vector corresponding to each diffusion real global dynamic potential variable.

[0173] Specifically, the trained hypernetwork is used to obtain the real time-resolved state potential prediction vector corresponding to each diffusion real global dynamic potential variable.

[0174] For each diffusion real global dynamic latent variable, generate its corresponding L-NODE parameter. For each L-NODE parameter, obtain the corresponding real time-resolved state latent prediction vector from the real time-resolved state latent vector sequence of the last time step in the support set.

[0175] S209. Decode the real static context vector, each real time-resolved state potential prediction vector and the corresponding diffusion real global dynamic potential variable into predicted response data, and multiple sets of predicted response data constitute a predicted response data sequence.

[0176] The predicted response data sequence here is the cross-scenario parameter, which can also be called the predicted biological response trajectory.

[0177] Specifically, a pre-trained decoder decodes the real static context vector, each real time-resolved state latent vector, and the corresponding real global dynamic latent variables into predicted response data.

[0178] In one embodiment, the above method further includes the following steps S2010-S2012: S2010. Obtain the mean value of the data at each time step in the predicted reaction data sequence to obtain the final predicted reaction data.

[0179] S2011. Obtain the confidence interval of the data at each time step.

[0180] S2012, Output the final predicted response data and confidence interval.

[0181] Once the model is trained, when faced with a new task requiring scaled-up prediction (e.g., from 10L to 1000L), the complete inference process is as follows: 1. Online Adaptation and Context Extraction: First, collect real data from the initial reaction phase of the new task (e.g., a 1000L reactor) as the support set. Then, use the trained HVAE encoder and task encoder to calculate its specific true dynamic context vector from the support set. Simultaneously, calculate its true static context vector based on the static metadata of the new task (geometric parameters of the 1000L reactor, etc.). 2. Probabilistic Sampling of Kinetic Parameters: Using the dynamic context vector as a condition, run the inverse denoising process of the trained conditional diffusion model. From... Starting with different standard normal distributed random noise, they are executed in parallel. This completes one full denoising iteration. This will ultimately generate multiple diffused real global dynamic latent variables for each real global dynamic latent variable, forming an ensemble of global dynamic variables. This ensemble represents a set of Monte Carlo samples of the posterior probability distribution of the model's intrinsic dynamic characteristics for the new task. 3. Determination of the dynamic manifold ensemble: For each sampled diffused real global dynamic latent variable, it is input into the trained supernetwork to generate its corresponding L-NODE parameters. This yields an ensemble containing multiple different dynamic models, each representing a possible hypothesis about the future evolution of the new task. 4. Generation of the latent trajectory ensemble: For each determined dynamic model, using the real time-resolved state latent prediction vector at the last moment of the support set as initial conditions, its corresponding L-NODE is numerically solved, and then integrated forward to generate future real time-resolved state latent prediction vectors. Each real time-resolved state latent prediction vector can also be considered as generating a future latent trajectory. This process is repeated. This yields an ensemble of potential future trajectories. 5. Decoding to the Observable Space: Finally, using the decoder, each potential trajectory, along with its corresponding diffusion real global dynamics variables and real static context vectors, is decoded back into a high-dimensional, observable data space. This produces an ensemble of future observed data trajectories (including cell images and sensor readings). 6. Probabilistic Output and Decision Support: This final ensemble is the model's complete probabilistic prediction of the future. Its time-wise mean can be calculated as the most probable deterministic prediction of the future. More importantly, its time-wise quantiles (e.g., 5% and 95%) can be calculated to construct the confidence interval of the prediction. This prediction with confidence assessment provides unprecedented, quantitative decision support for process development scientists and engineers to conduct risk assessments, determine process robustness, design feeding strategies, and optimize control schemes.

[0182] This disclosure achieves the generation of dynamic laws through the following design, thereby possessing unprecedented generalization ability: Step 1: Constructing a "Regularity Space" Anchored to the Physical World. The Hierarchical Variational Autoencoder (HVAE) disclosed herein does not construct a black-box latent space. By introducing a physical regularization loss, it mandates that the global kinetic latent variables representing the batch's "kinetic genes" must be strongly correlated with a static context vector composed of dimensionless physical criterion numbers (such as Reynolds number and power number) calculated from reactor geometry and operating parameters. This design ensures that the organization of the latent space follows underlying engineering science principles. Therefore, for reaction processes at different scales (10L vs 1000L), as long as their core fluid environments and mass transfer characteristics are similar, the corresponding global kinetic latent variables will be located close to each other in this "regularity space." This lays a solid foundation for cross-scale comparisons and generalization.

[0183] Step Two: Learning the Generative Rules of the "Laws". This disclosure does not learn a single, fixed kinetic equation for all reaction processes. Instead, it learns a higher-order mapping relationship through a hyperNetwork: This means that the model learns "how to generate its own, continuous-time latent neural differential equations (L-NODEs) based on the intrinsic genes (global dynamic latent variables) of a process." When faced with a new 1000L task, the model only needs to infer its corresponding global dynamic latent variables to instantly generate a dynamic model "tailor-made" for the 1000L operating condition through the hypernetwork, rather than rigidly applying the 10L model.

[0184] Technical results: This paradigm of "global dynamic latent variables → generation of specific laws" ensures from a mechanistic perspective that the model can adapt to previously unseen and significantly different boundary conditions, thus fundamentally solving the bottleneck of model universality in the amplification and migration of biological reaction processes.

[0185] Based on the same inventive concept, this application also provides a cross-scene parameter generation apparatus for implementing the cross-scene parameter generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more cross-scene parameter generation apparatus embodiments provided below can be found in the limitations of the cross-scene parameter generation method described above, and will not be repeated here.

[0186] In one exemplary embodiment, such as Figure 4 As shown, a cross-scene parameter generation device is provided, the device comprising: The second acquisition module is used to acquire real data in the initial stage of the reaction; the real data includes: real time-series multimodal observation data and real static metadata; the real static metadata includes: real type information of the basic components of the bioreactor and real numerical information of the bioreactor configuration.

[0187] The ninth acquisition module is used to acquire the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence based on real data.

[0188] In one embodiment, the ninth acquisition module is specifically used to encode real data using a trained HAVE encoder to obtain the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence.

[0189] The tenth acquisition module is used to obtain the real dynamic context vector based on the real time-resolved state latent vector sequence using the trained task encoder.

[0190] The eleventh acquisition module is used to calculate the true dimensionless physical criterion number based on the true static metadata, and to construct the true physical feature vector based on the true dimensionless physical criterion number.

[0191] The concatenation module is used to concatenate the real physical feature vector with the real type information to obtain the real joint input vector.

[0192] The twelfth acquisition module is used to acquire the real static context vector based on the real joint input vector; the static context vector includes: physical environment and biological basis.

[0193] The thirteenth acquisition module is used to acquire multiple real global dynamic latent variables based on the real dynamic context vector through a trained conditional diffusion model.

[0194] The fourteenth acquisition module is used to acquire the real time-resolved state potential prediction vector corresponding to each real global dynamic potential variable.

[0195] The fourteenth acquisition module is specifically used to input each real global dynamic latent variable into the trained supernetwork to obtain the corresponding real time-resolved state latent prediction vector.

[0196] The decoding module is used to decode the real static context vector, each real time-resolved state potential prediction vector and the corresponding real global dynamic potential variable into predicted response data. Multiple sets of predicted response data constitute a predicted response data sequence.

[0197] In one embodiment, the decoding module is specifically used to decode the real static context vector, each real time-resolved state latent vector, and the corresponding real global dynamic latent variable into predicted response data using a trained decoder.

[0198] In one embodiment, the apparatus further includes: The fifteenth acquisition module is used to acquire the mean value of the data at each time step in the predicted reaction data sequence to obtain the final predicted reaction data. The sixteenth acquisition module is used to acquire the confidence interval of the data at each time step; The output module is used to output the final predicted response data and confidence intervals.

[0199] Figure 5 This is a schematic diagram of an experimental vessel according to an exemplary embodiment. Figures 6-9 for Figure 5 Design data sheet for the middle tank.

[0200] (a) Basic experimental information from 10L to 1000L.

[0201] Experimental subject: Cells (HEK-293).

[0202] Culture medium: suspension culture medium.

[0203] Culture information: passage density .

[0204] (ii) The optimal cell state control parameters for 10L are shown in Table 2: Table 2

[0205] (III) The cell state detection results at different time points after model simulation control were extended to 1000L, as shown in Table 3: Table 3

[0206] (iv) Other records.

[0207] Samples were taken and sent for testing at 24, 48, and 72 hours, with 3ml of sample retained per bottle.

[0208] The data clearly show that expanding the 10L data to 1000L allows the biological reaction to proceed effectively and ensures the survival rate.

[0209] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0210] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for training a cross-task model of biological responses based on dynamic context, characterized in that, The training method includes: Reaction data at each time step were collected during the complete reaction process of two different bioreactor sizes under the same process. The support set is formed by extracting data from the reaction data at each time step, specifically the data from the first B time steps before the initial stage of the reaction. The mean of the posterior Gaussian distribution of the global kinetic latent variables was obtained from all reaction data; the global kinetic latent variables represent the properties of the bioreactor that do not change with reaction time during the reaction process. Based on the forward noise-adding network of the conditional diffusion model, K-step Gaussian noise addition is performed on the global dynamic latent variables to finally obtain the noisy samples; The time-resolved state latent vector sequence is obtained through the support set; the time-resolved state latent variable sequence represents the properties of the bioreactor that change with reaction time during the reaction process. Input the time-resolved state latent vector sequence into the task encoder to obtain the dynamic context vector; The denoising network based on the conditional diffusion model predicts the noise added at each step based on the noisy sample and the dynamic context vector. The denoising network and task encoder are trained based on the real noise added at each step and the predicted noise added at each step, resulting in a trained conditional diffusion model and a task encoder. The trained biological response cross-task model includes a trained conditional diffusion model and a trained task encoder.

2. The training method according to claim 1, characterized in that, The data at each time step includes: time-series multimodal observation data and static metadata. The time-series multimodal observation data includes: sensor readings of culture medium, metabolomics, physiological microscopic images of the observed objects, and biological reaction parameters. The static metadata includes: type information of the basic components of the bioreactor and numerical information of the bioreactor configuration. Among them, the time-series multimodal observation data is data that changes with reaction time, while the static metadata is data that does not change with reaction time.

3. The training method according to claim 2, characterized in that, The method further includes: The posterior Gaussian distribution parameters of the first global dynamic latent variable corresponding to all reaction data and the posterior Gaussian distribution parameters of the first time-resolved state latent vector at each time step are obtained based on the initial HAVE encoder; the posterior Gaussian distribution parameters include: mean; Based on the static metadata corresponding to all reaction data, the dimensionless physical criterion number is calculated, and the physical feature vector is constructed based on the dimensionless physical criterion number. The physical feature vector is concatenated with the type information to obtain the joint input vector; Obtain the static context vector from the joint input vector; The posterior Gaussian distribution parameters of the first global dynamic latent variable, the posterior Gaussian distribution parameters of the first time-resolved state latent vector corresponding to each time step, and the static context vector are input into the initial decoder to obtain the reconstructed response data. Based on the reconstructed reaction data and the collected reaction data, the initial HAVE encoder and the initial decoder are trained using the first loss function to obtain the trained HAVE encoder and the trained decoder; the trained biological response cross-task model also includes: the trained HAVE encoder and the trained decoder. The process of obtaining the mean of the posterior Gaussian distribution of the global dynamic latent variables from all reaction data includes: The trained HAVE encoder encodes all response data to obtain the mean of the posterior Gaussian distribution of the global dynamic latent variables.

4. The training method according to claim 3, characterized in that, The step of obtaining the posterior Gaussian distribution parameters of the first global dynamic latent variables corresponding to all reaction data based on the initial HAVE encoder includes: A convolutional neural network with an initial HAVE encoder is used to extract features from the microscopic image at each time step, converting the image pixel information into a fixed-dimensional morphological feature vector. The morphological feature vectors of all time steps are arranged in chronological order to form a morphological feature sequence. The first MLP network of the initial HAVE encoder is used to perform nonlinear transformation and dimension matching on the sensor readings at each time step to obtain biochemical feature vectors; the biochemical feature vectors of all time steps are arranged in chronological order to form a biochemical feature sequence. The feature vectors at the same time step in the morphological feature sequence and the biochemical feature sequence are concatenated to obtain the time-series comprehensive feature sequence. The comprehensive feature sequence is input into the Bi-GRU model of the initial HAVE encoder to obtain the complete hidden state at each time step. The complete hidden state at each time step includes the reaction data at the current time step, the reaction trend before the current time step, and the reaction result after the current time step. The complete hidden states of all time steps are arranged in chronological order to form a hidden state sequence. The hidden state sequence is input into the self-attention weighted pooling network of the initial HAVE encoder to obtain the global information vector; The global information vector is input into the second MLP network of the initial HAVE encoder to obtain the posterior Gaussian distribution parameters of the first global dynamic latent variable.

5. The training method according to claim 4, characterized in that, The step of obtaining the posterior Gaussian distribution parameters of the first time-resolved state latent vectors for each time step corresponding to all reaction data based on the initial HAVE encoder includes: The complete hidden state at each time step in the hidden state sequence and the first global dynamic latent variable are concatenated into a vector to obtain the concatenated sequence. The spliced ​​sequence is input into the third MLP network of the initial HAVE encoder to obtain the posterior Gaussian distribution parameters of the first time-resolved state latent vector at each time step.

6. The training method according to claim 5, characterized in that, The method further includes: Based on the trained HAVE encoder, obtain the first mean sequence of the posterior Gaussian distribution parameters of the second global dynamic latent variable and the second mean sequence of the posterior Gaussian distribution parameters of the second time-resolved state latent vector corresponding to all reaction data. The first mean is input into the initial supernetwork to obtain the predicted mean sequence of the second time-resolved state latent vector; The initial supernetwork is trained based on the second mean sequence and the predicted mean sequence of the second time-discriminated state latent vector to obtain the trained supernetwork; the trained biological response cross-task model further includes the trained supernetwork. The step of obtaining the time-resolved state latent vector sequence through the support set includes: The trained hypernetwork encodes the support set to obtain a time-resolved state latent vector sequence.

7. The training method according to claim 6, characterized in that, The step of training the initial supernetwork based on the prediction sequence of the second mean sequence and the second time-discriminated state latent vector to obtain the trained supernetwork includes: The difference between the second mean sequence and the predicted mean sequence of the second time-resolved state latent vector is quantified using mean squared error. The initial supernetwork is then trained by backpropagation to obtain the trained supernetwork.

8. The training method according to claim 7, characterized in that, The process of training the denoising network and task encoder based on the actual noise added at each step and the predicted noise added at each step to obtain the trained conditional diffusion model and task encoder includes: An L2 loss function is constructed based on the real noise added at each step and the predicted noise added at each step. The denoising network and the task encoder are simultaneously optimized through backpropagation of the loss, resulting in the trained conditional diffusion model and task encoder.

9. A method for generating parameters across different scenarios, characterized in that, The method includes: Collect real data during the initial stage of the reaction; the real data includes: real time-series multimodal observation data and real static metadata; the real static metadata includes: real type information of the basic components of the bioreactor and real numerical information of the bioreactor configuration; Obtain the mean of the posterior Gaussian distribution of the real global dynamic latent variables and the real time-resolved state latent vector sequence based on real data; Using a trained task encoder, the real dynamic context vector is obtained based on the real time-resolved state latent vector sequence; Based on real static metadata, calculate the real dimensionless physical criterion number, and construct the real physical feature vector based on the real dimensionless physical criterion number; The real physical feature vector is concatenated with the real type information to obtain the real joint input vector; The true static context vector is obtained based on the true joint input vector; the static context vector includes: physical environment and biological basis; The trained conditional diffusion model is used to diffuse each real global dynamic latent variable based on the real dynamic context vector and each real global dynamic latent variable, resulting in multiple diffused real global dynamic latent variables corresponding to each real global dynamic latent variable. Based on the real time-resolved state potential vector sequence and each diffusion real global dynamic potential variable, obtain the real time-resolved state potential prediction vector corresponding to each diffusion real global dynamic potential variable; The real static context vector, each real time-resolved state potential prediction vector, and the corresponding diffusion real global dynamic potential variables are decoded into predicted response data, and multiple sets of predicted response data constitute a predicted response data sequence.

10. The method according to claim 9, characterized in that, The method further includes: The mean of the data at each time step in the predicted reaction data sequence is obtained to get the final predicted reaction data; Obtain the confidence intervals for the data at each time step; Output the final predicted response data and confidence interval.