Cross-process process parameter collaborative optimization method and system for intelligent manufacturing of traditional Chinese medicine formula
By combining cross-process multimodal datasets and digital twin generative adversarial networks, the problems of high trial-and-error costs and insufficient simulation fidelity in cross-process parameter optimization in traditional Chinese medicine production are solved, thereby improving the quality stability and process efficiency of traditional Chinese medicine formulation production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 陈金龙
- Filing Date
- 2026-04-02
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies for optimizing process parameters across processes in traditional Chinese medicine production face challenges such as high trial-and-error costs and insufficient simulation fidelity, making it difficult to achieve low-cost and high-fidelity collaborative optimization across processes.
By acquiring multimodal datasets across processes, combining mechanistic skeleton models and digital twin generative adversarial networks, we optimize the physical consistency constraint loss function and utilize a multi-agent collaborative strategy network to collaboratively optimize process parameters across processes, thereby generating efficient process optimization parameters.
A collaborative optimization environment with both low-cost trial-and-error capabilities and high fidelity was constructed, which improved the reliability of transferring optimization strategies to actual production and enhanced the quality stability and process efficiency of traditional Chinese medicine formulation production.
Smart Images

Figure CN122334816A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent pharmaceutical manufacturing, and in particular relates to a method and system for cross-process parameter collaborative optimization in intelligent manufacturing of traditional Chinese medicine formulas. Background Technology
[0002] With the in-depth development of intelligent manufacturing technology for traditional Chinese medicine, the collaborative optimization of process parameters across multiple processes has become a core requirement for improving product quality stability. This technology aims to solve the control challenges caused by the high coupling of parameters between multiple processes such as extraction, concentration, drying, and granulation, requiring the optimization method to have the ability to handle high-dimensional nonlinear problems and quickly adapt to batch fluctuations in medicinal materials.
[0003] Traditional process optimization methods mainly rely on offline experimental design, which involves conducting a large number of parameter combination experiments on the production line to establish response surfaces, or on static parameter optimization based on empirical formulas.
[0004] However, current process optimization methods face a core contradiction when dealing with the complex systems of traditional Chinese medicine: the trial-and-error costs in real production environments are extremely high, making large-scale parameter exploration unfeasible both in terms of safety and economy; while existing simulation models suffer from insufficient fidelity due to their inability to accurately capture multi-component reaction kinetics and nonlinear mass and heat transfer processes, making it difficult to directly transfer simulation-optimized strategies to actual production. Therefore, constructing a training environment that possesses both low-cost trial-and-error capabilities and high fidelity in reflecting the laws of real processes has become a critical technical bottleneck that urgently needs to be overcome to achieve cross-process collaborative optimization. Summary of the Invention
[0005] Therefore, it is necessary to provide a method and system for cross-process parameter collaborative optimization in intelligent manufacturing of traditional Chinese medicine formulas that can build a high-fidelity training environment at low cost and achieve efficient policy transfer and generalization, in order to address the above-mentioned technical problems.
[0006] Firstly, this application provides a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas, including:
[0007] S1. Obtain the cross-process multimodal dataset of traditional Chinese medicine formula in the production process, input the cross-process multimodal dataset into the mechanism skeleton model, solve the differential equations, and obtain the mechanism prediction value; wherein, the mechanism skeleton model is constructed by the material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the traditional Chinese medicine production process.
[0008] S2. Update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the mechanism prediction value, and iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment in this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is the process parameter sequence, and the output is the production process state sequence;
[0009] S3. Obtain historical real features in the real production environment, input the cross-process multimodal dataset into the multi-agent cooperative strategy network in the optimized digital twin generative adversarial network environment, generate historical simulation features, and calculate the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference metric.
[0010] S4. Based on the distribution difference metric, combine the kernel function with the optimal mapping function to solve for the mapping relationship from the real data feature space to the simulated data feature space. Based on the mapping relationship, adjust the process parameters of the cross-process multimodal dataset to obtain the corrected cross-process multimodal dataset.
[0011] S5. Input the corrected cross-process multimodal dataset into the multi-agent cooperative strategy network to perform multi-agent cooperative decision-making and generate cross-process optimization parameters.
[0012] In one embodiment, the physical consistency constraint loss function of the digital twin generative adversarial network environment is updated based on the mechanism prediction value, and the digital twin generative adversarial network environment is iteratively optimized based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment for this collaborative optimization round, including:
[0013] S11. Based on the cross-process multimodal dataset, obtain condition variable and real state parameter data;
[0014] S12. Input the condition variables into the generator in the digital twin generative adversarial network environment, and use the generator to perform forward propagation calculation on the condition variables to obtain the production process state sequence.
[0015] S13. Input the production process state sequence and mechanism prediction value into the physical consistency loss function, calculate the mean square error, and obtain the physical consistency loss. Input the production process state sequence into the time smoothness loss function, calculate the sum of squares of the Euclidean distances between adjacent time state vectors, and obtain the time smoothness loss.
[0016] S14. Input the production process state sequence into the discriminator in the digital twin generative adversarial network environment, perform forward propagation, and calculate the probability value of the real data distribution from the real data.
[0017] S15. Based on the probability values of the real data distribution, the adversarial loss of the generator is calculated using the adversarial loss function. The total loss function value of the generator is calculated by weighted summation of the adversarial loss, physical consistency loss, and temporal smoothness loss.
[0018] S16. Calculate the gradient value of the total loss function, input the gradient value into the learning rate function, calculate the generator update parameters, and update the generator according to the generator update parameters to obtain the updated generator.
[0019] S17. Input the condition variables into the update generator to obtain the updated production process state sequence, and calculate the residual sequence between the updated production process state sequence and the actual state parameter data.
[0020] S18. Input the residual sequence into the Gaussian mixture model and perform probability distribution fitting to obtain the residual distribution model;
[0021] S19. Repeat S12 to S18 until the update generator and residual distribution model reach the convergence condition. Based on the update generator and residual distribution model that have reached the convergence condition, construct the optimized digital twin generative adversarial network environment for this round of collaborative optimization.
[0022] Based on the above embodiments, the expression for the total loss function corresponding to the total loss function value is as follows:
[0023]
[0024] in, Let the total loss function of the generator be . For the generator, use a random noise vector and condition variables The production process state sequence generated as input; These are the probability values of the true data distribution. The first output of the generator Time-state vector For the first Mechanism prediction values at time points. The length of the time series. For the mechanism-constrained loss weighting coefficient, The weighting coefficients for time-series smoothing loss are denoted as .
[0025] In one embodiment, the multi-agent cooperative policy network is constructed through the following steps:
[0026] S21. In the optimized digital twin generative adversarial network environment, each process is defined as a corresponding process agent, and an upper-level coordinating agent is set up to load and obtain the initial multi-agent network.
[0027] S22. To initialize each process agent in the multi-agent network, design the observation space, action space, and reward function respectively to obtain the defined multi-agent network. The observation space includes the process parameters, state parameters, and upstream material attributes of the current process. The action space includes the adjustment amount of the process parameters of the current process. The reward function includes the quality index reward item and the energy consumption penalty item.
[0028] S23. Add a policy network based on the Actor-Critic architecture to the defined multi-agent network to obtain a multi-agent policy network.
[0029] S24. Obtain the process state parameters and process parameters at the current moment, input the process state parameters and process parameters at the current moment into the multi-agent policy network, and perform mapping calculation after splicing the process state parameters and process parameters at the current moment to obtain the target instruction;
[0030] S25. Input the target instruction into each process agent, and perform forward propagation calculation through the Actor network in each process agent to obtain the process parameter adjustment amount;
[0031] S26. Input the process parameter adjustment amount into the optimized digital twin generative adversarial network environment, perform state transition calculation, obtain the state parameters of the next moment, the output actions of each process agent and the reward value of each process agent, and then combine the current process state parameters, process parameters, next moment state parameters, output actions of each process agent and reward value of each process agent into an experience tuple and store it in the experience replay pool to obtain the experience replay pool.
[0032] S27. Sample a small batch of experience tuples from the experience replay pool, calculate the Critic network loss and Actor network policy gradient of each agent based on the experience tuples, update the network parameters of each agent through the backpropagation algorithm, and obtain the updated multi-agent policy network.
[0033] S28. Repeat S24 to S27 until the agents in each process converge, and obtain the multi-agent cooperative strategy network.
[0034] In one embodiment of the present invention, the distribution difference between historical simulation features and historical real features is calculated to obtain a distribution difference metric, including:
[0035] S31. Input the historical simulation features and historical real features into the maximum mean difference calculation function model to calculate the distribution difference measure. The expression of the maximum mean difference calculation function model is as follows:
[0036]
[0037] in, This is a measure of distributional difference. As a historical simulation feature, As a historical fact, For kernel function mapping, the expression for kernel function mapping is: ; These are historical simulation samples within the historical simulation features. A true historical sample with authentic historical characteristics; This is the bandwidth parameter of the kernel function, which controls the smoothness of the kernel function; For the regenerating nucleus Hilbert space norm, The first of the historical simulation features One historical simulation sample; The first characteristic of historical authenticity A true historical sample; This represents the total number of historical simulation samples included in the historical simulation features. The total number of historical real samples included in the historical real features.
[0038] Secondly, this application also provides a cross-process parameter collaborative optimization system for intelligent manufacturing of traditional Chinese medicine formulas, used to implement the method described in the first aspect, including:
[0039] The multimodal data and mechanism prediction module is used to acquire cross-process multimodal datasets of traditional Chinese medicine formulas in the production process. The cross-process multimodal datasets are input into the mechanism skeleton model, and differential equations are solved to obtain mechanism prediction values. The mechanism skeleton model is constructed from the material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the traditional Chinese medicine production process.
[0040] The twin environment physical constraint optimization module is used to update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the mechanism prediction value, and iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment in this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is the process parameter sequence, and the output is the production process state sequence;
[0041] The distribution difference measurement module is used to obtain historical real features in the real production environment. It inputs the cross-process multimodal dataset into the multi-agent cooperative strategy network in the optimized digital twin generative adversarial network environment to generate historical simulation features, and calculates the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference measurement value.
[0042] The feature space mapping and parameter correction module is used to divide the kernel function according to the distribution difference metric and the optimal mapping function, solve for the mapping relationship from the real data feature space to the simulated data feature space, and adjust the process parameters of the cross-process multimodal dataset based on the mapping relationship to obtain the corrected cross-process multimodal dataset.
[0043] The multi-agent collaborative optimization decision module is used to input the modified cross-process multimodal dataset into the multi-agent collaborative strategy network, perform multi-agent collaborative decision-making, and generate cross-process optimization parameters.
[0044] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods in the first aspect of this application.
[0045] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods in the first aspect of this application.
[0046] The aforementioned method for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas acquires a multimodal dataset across multiple processes and integrates a mechanistic framework model to solve differential equations. It embeds physicochemical mechanisms into the training process of a digital twin generative adversarial network environment, enabling the simulation environment to maintain data-driven flexibility while possessing physical consistency constraints, significantly improving the simulation environment's high-fidelity representation of the real production process. Furthermore, by calculating the distribution differences between historical simulation features and historical real features, and combining the optimal mapping function with kernel function partitioning to solve the mapping relationship from the real data feature space to the simulation data feature space, the multimodal dataset across processes is corrected, effectively reducing the distribution offset between simulation and reality. This allows the optimization process to make decisions based on data that more closely resembles the real state. On this basis, a multi-agent collaborative strategy network is used for cross-process collaborative decision-making, generating cross-process optimization parameters and achieving efficient linkage and global optimization of process parameters across multiple processes. This method constructs a collaborative optimization environment that combines low-cost trial-and-error capability with high fidelity, breaking through the technical bottlenecks of high trial-and-error costs and insufficient simulation fidelity in real production, improving the reliability of transferring optimization strategies to actual production, and thus achieving synergistic improvement in the quality stability and process efficiency of traditional Chinese medicine formulation production. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 A schematic diagram of an implementation environment provided for one embodiment of the present invention;
[0049] Figure 2 This is a flowchart of a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas according to one embodiment of the present invention.
[0050] Figure 3 This is a schematic diagram of the cross-process parameter collaborative optimization system for intelligent manufacturing of traditional Chinese medicine formulas in one embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0052] The cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the parameter optimization terminal 100 communicates with the sensor 101 via a network. The parameter optimization terminal 100 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. The sensor 101 can be, but is not limited to, temperature sensors, pressure sensors, flow meters, mass flow meters, and tachometers.
[0053] In one exemplary embodiment, such as Figure 2 As shown, a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas is provided, which is then applied to... Figure 1 Taking the parameter optimization terminal 100 as an example, the method includes:
[0054] S1. Obtain the cross-process multimodal dataset of Chinese medicine formula in the production process, input the cross-process multimodal dataset into the mechanism skeleton model, solve the differential equations, and obtain the mechanism prediction value; wherein, the mechanism skeleton model is constructed by the material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the Chinese medicine production process.
[0055] Specifically, the cross-process multimodal dataset can include process parameter data, process status data, and quality inspection data. Process parameter data consists of the adjustable parameters for each production process; process status data comprises the real-time operating status parameters of the materials and equipment for each process; and quality inspection data includes the quality indicators of intermediate and final products for each process. The mechanistic framework model can be constructed from the material and energy conservation differential equations for each process based on the physicochemical mechanisms of traditional Chinese medicine production. The expression for material conservation for each process can be: In the formula, The instantaneous concentration of the effective components in the liquid feed. For extraction time, This is the dissolution rate constant. The equilibrium concentration of the effective component; the expression for the energy conservation differential equation can be: In the formula, For the density of the liquid, For the volume of liquid feed, The specific heat capacity of the liquid under constant pressure. The temperature of the liquid feed. For heating input power, This is due to heat loss from the equipment.
[0056] For example, the parameter optimization terminal 100 can acquire a multi-modal dataset of traditional Chinese medicine formulas across processes during production via sensor 101. The terminal can then use the Grubbs criterion to remove outliers from the dataset, supplement missing data using linear interpolation, and finally standardize the dataset to a preset range using a min-max normalization method, resulting in a standardized multi-modal dataset. The terminal can then input this standardized dataset into a mechanistic framework model and solve the differential equations using the fourth-order Runge-Kutta method to obtain mechanistic prediction values reflecting the process states of each step.
[0057] S2. Update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the mechanism prediction value, and iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment of this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is the process parameter sequence, and the output is the production process state sequence.
[0058] Specifically, a digital twin generative adversarial network environment can include a generator and a discriminator. The generator can receive the process parameter sequence, simulate the production process of each step, and output the corresponding production process state sequence. The discriminator can distinguish the production process state sequence output by the generator from the real production state sequence. The realism of the simulation data is improved through adversarial training between the two.
[0059] For example, the parameter optimization terminal 100 can iteratively optimize the digital twin generative adversarial network environment by updating the physical consistency constraint loss function based on the mechanism prediction value. The expressions for the original physical consistency constraint loss function and the optimized physical consistency constraint loss function can be:
[0060]
[0061]
[0062] in, The loss function is the original physical consistency constraint. To optimize the physical consistency constraint loss function, The loss weight is the difference between the mechanism prediction and the simulated state sequence. The loss weights are the differences between the simulated state sequence and the actual production state sequence. This is a mechanistic prediction value. It is a sequence of production process states. This represents a real production state sequence. The parameter optimization terminal 100 can select multiple different sets of parameters using cross-validation. and The combinations are then substituted into the optimized digital twin generative adversarial network environment for training. The simulation error of each group after training is calculated, and the combination with the smallest simulation error is selected as the final weight coefficient combination. The parameter optimization terminal 100 can use a preset optimizer to adjust the parameters of the generator and discriminator. The generator loss function can be:
[0063]
[0064] In the formula, Let be the generator loss function. To generate adversarial losses, The loss weights are used for the discriminator; the discriminator loss function can be:
[0065]
[0066] In the formula, Let the discriminator loss function be... As expected, To determine the output probability of the discriminator, the parameter optimization terminal 100 can iterate and calculate the error according to a preset batch until the error reaches the standard or the preset number of iterations is reached, thus obtaining the optimized digital twin generative adversarial network environment for this round of collaborative optimization.
[0067] S3. Obtain historical real features in the real production environment, input the cross-process multimodal dataset into the multi-agent cooperative strategy network in the digital twin generative adversarial network environment to generate historical simulation features, and calculate the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference metric.
[0068] Specifically, the multi-agent cooperative strategy network can be the core decision-making unit for optimizing the digital twin generative adversarial network environment. The multi-agent cooperative strategy network can be composed of independent agents in each production process such as extraction, concentration, drying and granulation of traditional Chinese medicine. Each agent can adopt the DQN structure in deep reinforcement learning. The structure of each agent can include an input layer, a hidden layer and an output layer.
[0069] For example, the parameter optimization terminal 100 can access the historical production database of the traditional Chinese medicine production workshop, and filter out historical production batch data that are consistent with the current traditional Chinese medicine formula to be optimized, have similar production process parameters, and similar production environment. During the filtering, production data within the past year is selected first to obtain historical real data. The parameter optimization terminal 100 can preprocess the historical real data to obtain standard historical real data. The parameter optimization terminal 100 can use principal component analysis to extract key features from the standard historical real data. Combining the core quality requirements and process characteristics of traditional Chinese medicine production, a feature contribution rate threshold is set to filter out key indicators that can reflect the stability of the production process, the content of effective ingredients in the product, and production efficiency. Features; the parameter optimization terminal 100 can organize these key features by production batch to form a historical real feature matrix. Rows in the historical real feature matrix can represent different historical production batches, and columns can represent different feature dimensions. The parameter optimization terminal 100 can input the historical real feature matrix into the multi-agent cooperative policy network in the optimized digital twin generative adversarial network environment to obtain historical simulation features. The parameter optimization terminal 100 can input the historical simulation features and historical real features into the KL divergence and combine it with the Wasserstein distance calculation function to quantify the distribution differences, obtaining the KL divergence and Wasserstein distance. The KL divergence calculation expression can be:
[0070]
[0071] in, As a historical fact, As a historical simulation feature, This refers to the KL divergence, which is used to measure... and Information differences between them for exist The probability density at that location, for exist The probability density at point A, where the Wasserstein distance can be calculated using the function:
[0072]
[0073] in, For Wasserstein distance, For all distributions To distribution The set of joint distributions, For joint distribution variables, The value represents the true characteristics of history. The values for historical simulation features, for and The Euclidean distance between them; the parameter optimization terminal 100 can sum the KL divergence and Wasserstein distance with weights to obtain a measure of distribution difference.
[0074] S4. Based on the distribution difference metric, combine the kernel function with the optimal mapping function to solve for the mapping relationship from the real data feature space to the simulated data feature space. Based on the mapping relationship, adjust the process parameters of the cross-process multimodal dataset to obtain the corrected cross-process multimodal dataset.
[0075] For example, the parameter optimization terminal 100 can compare the distribution difference metric with a preset threshold. When the distribution difference metric is greater than the preset threshold, a kernel function is applied; when the distribution difference metric is less than the preset threshold, the original dataset is used. The parameter optimization terminal 100 can combine the distribution difference features to apply the kernel function, and based on all historical true feature vectors, calculate the radial basis function (RBF) kernel function value corresponding to any two feature vectors. The expression for calculating the RBF kernel function value is:
[0076]
[0077] in, The calculated value is the radial basis function kernel function. This represents the historical feature vector corresponding to the historical true features. This represents the historical feature vector corresponding to the historical true features. The kernel width parameter controls the smoothness of the kernel function. The distance between the historical real feature vectors and the historical simulated feature vectors is the Euclidean distance. The parameter optimization terminal 100 can construct a kernel matrix based on the radial basis function kernel function. Based on the kernel matrix, the parameter optimization terminal 100 uses a least squares support vector machine to solve the mapping relationship between the real data feature space and the simulated data feature space. The expression for the mapping relationship can be:
[0078]
[0079] in, These are the mapped simulation feature values. This is the mapping function from real data features to simulated data features. For historical true feature vectors, The weight vector represents the mapping relationship. It is a high-dimensional feature vector obtained by mapping the historical true feature vector through a kernel function. As a bias term; the parameter optimization terminal 100 can adjust the process parameters in the cross-process multimodal dataset based on the mapping relationship to obtain the corrected cross-process multimodal dataset.
[0080] S5. Input the corrected cross-process multimodal dataset into the multi-agent cooperative strategy network to perform multi-agent cooperative decision-making and generate cross-process optimization parameters.
[0081] For example, the parameter optimization terminal 100 can determine the optimization target weights and parameter constraint ranges using the analytic hierarchy process (AHP); the parameter optimization terminal 100 can input the modified cross-process multimodal dataset into a multi-agent cooperative strategy network, where the agents of each process in the multi-agent cooperative strategy network collaboratively calculate the preliminary cross-process optimization parameters; the parameter optimization terminal 100 can calculate the fitness index based on the preliminary cross-process optimization parameters, where the calculation expression for the fitness index can be:
[0082]
[0083] in, The adaptability index is used for verification. Adaptability For the initial cross-process optimization parameters, for The average value, Standard parameters for each process, for The average value; the parameter optimization terminal 100 can determine whether the adaptability index is within the preset adaptability range. When the adaptability index is within the preset adaptability range, the preliminary cross-process optimization parameters are used as the cross-process optimization parameters. When the adaptability index is not within the preset adaptability range, the preliminary cross-process optimization parameters are recalculated.
[0084] The aforementioned method for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas acquires a multimodal dataset across multiple processes and integrates a mechanistic framework model to solve differential equations. It embeds physicochemical mechanisms into the training process of a digital twin generative adversarial network environment, enabling the simulation environment to maintain data-driven flexibility while possessing physical consistency constraints, significantly improving the simulation environment's high-fidelity representation of the real production process. Furthermore, by calculating the distribution differences between historical simulation features and historical real features, and combining the optimal mapping function with kernel function partitioning to solve the mapping relationship from the real data feature space to the simulation data feature space, the multimodal dataset across processes is corrected, effectively reducing the distribution offset between simulation and reality. This allows the optimization process to make decisions based on data that more closely resembles the real state. On this basis, a multi-agent collaborative strategy network is used for cross-process collaborative decision-making, generating cross-process optimization parameters and achieving efficient linkage and global optimization of process parameters across multiple processes. This method constructs a collaborative optimization environment that combines low-cost trial-and-error capability with high fidelity, breaking through the technical bottlenecks of high trial-and-error costs and insufficient simulation fidelity in real production, improving the reliability of transferring optimization strategies to actual production, and thus achieving synergistic improvement in the quality stability and process efficiency of traditional Chinese medicine formulation production.
[0085] In one embodiment of the present invention, the physical consistency constraint loss function of the digital twin generative adversarial network environment is updated based on the mechanism prediction value, and the digital twin generative adversarial network environment is iteratively optimized based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment for this collaborative optimization round, including:
[0086] S11. Based on the cross-process multimodal dataset, obtain condition variable and real state parameter data.
[0087] Specifically, conditional variables can be adjustable process parameters for each production process, and can cover key adjustable indicators such as extraction time, extraction temperature, concentration pressure, and drying rate; real-state parameter data can be actual operating indicators synchronously acquired by data acquisition equipment at key nodes of each process during production, and can include the concentration of effective components in the feed liquid, equipment operating temperature, energy consumption data, etc.
[0088] For example, the parameter optimization terminal can preprocess and standardize the cross-process multimodal dataset to obtain a standardized cross-process multimodal dataset; the parameter optimization terminal can then extract conditional variables and real state parameter data based on the standardized cross-process multimodal dataset.
[0089] S12. Input the condition variables into the generator in the digital twin generative adversarial network environment, and use the generator to perform forward propagation calculation on the condition variables to obtain the production process state sequence.
[0090] Specifically, the generator in a digital twin generative adversarial network environment can adopt a multi-layer neural network structure, which can consist of an input layer, a hidden layer, and an output layer. The input layer can be used to receive conditional variables, and the hidden layer can include convolutional layers and fully connected layers. The convolutional layer can be used to extract deep features of the conditional variables, and the fully connected layer can be used to realize linear mapping and non-linear transformation of features. The output layer can use an activation function to map the features into state data that conforms to the actual production range.
[0091] For example, the parameter optimization terminal can input condition variables into the generator in the digital twin generative adversarial network environment to obtain a production process state sequence. The input layer of the generator can receive condition variables and input them into the hidden layer. The hidden layer can perform linear mapping and nonlinear transformation on the condition variables to obtain production process state features. The hidden layer can input the production process state features into the output layer, and the output layer outputs the production process state sequence.
[0092] S13. Input the production process state sequence and mechanism prediction value into the physical consistency loss function, calculate the mean square error, and obtain the physical consistency loss. Input the production process state sequence into the time smoothness loss function, calculate the sum of squares of the Euclidean distance between adjacent time state vectors, and obtain the time smoothness loss.
[0093] For example, the parameter optimization terminal can input the production process state sequence and mechanism prediction values into the physical consistency loss function, calculate the mean squared error, and obtain the physical consistency loss. The expression for calculating the physical consistency loss can be:
[0094]
[0095] in, This represents the physical consistency loss value. A smaller physical consistency loss value indicates a more reasonable production process state sequence. The total number of moments in the production process state sequence. for The numerical sequence of production process states generated at any given time. for The mechanism prediction value corresponding to each time step; the parameter optimization terminal can input the production process state sequence into the time smoothness loss function, calculate the sum of squares of the Euclidean distances between state vectors at adjacent time steps, and obtain the time smoothness loss. The expression for the time smoothness loss function can be:
[0096]
[0097] in, This represents the time smoothness loss value. for Time and The Euclidean distance between the state vectors at each time step.
[0098] S14. Input the production process state sequence into the discriminator in the digital twin generative adversarial network environment, perform forward propagation, and calculate the probability value of the real data distribution from the real data.
[0099] Specifically, the discriminator in a digital twin generative adversarial network environment can adopt a multi-layer neural network structure that matches the generator. The discriminator can consist of an input layer, a hidden layer, and an output layer. The input layer can be used to receive the state sequence of the production process. The hidden layer can include multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layers can be used to extract features from the state sequence. The pooling layers can be used to reduce the feature dimensionality and reduce redundant information. The fully connected layers can be used to classify and discriminate the extracted features. The output layer can use the sigmoid activation function to map the discrimination result to a probability value in the interval [0,1].
[0100] For example, the parameter optimization terminal can input the production process state sequence into the discriminator in the digital twin generative adversarial network environment. The production process state sequence passes through the input layer, hidden layer and output layer of the discriminator in sequence to calculate the real data distribution probability value from the real data.
[0101] S15. Based on the probability values of the real data distribution, the adversarial loss of the generator is calculated using the adversarial loss function. The total loss function value of the generator is calculated by weighting and summing the adversarial loss, physical consistency loss, and temporal smoothness loss.
[0102] For example, the parameter optimization terminal can input the probability value of the real data distribution into the adversarial loss function and calculate the adversarial loss through cross-entropy; the parameter optimization terminal can perform weighted summation of the adversarial loss, physical consistency loss and temporal smoothness loss according to preset weight coefficients, and the weight coefficients are determined through cross-validation to obtain the total loss function value.
[0103] S16. Calculate the gradient value of the total loss function, input the gradient value into the learning rate function, calculate the generator update parameters, and update the generator according to the generator update parameters to obtain the updated generator.
[0104] For example, the parameter optimization terminal can employ the backpropagation algorithm to calculate the gradient of the total loss function value, obtaining the gradient value. This gradient value reflects the trend of the total loss function value as the generator parameters change, and is used to determine the direction of generator parameter adjustment. The parameter optimization terminal can input the gradient value into a preset learning rate function. The learning rate function automatically adjusts the learning rate based on the magnitude of the gradient value, calculating the generator update parameters. These update parameters can include adjustments to core parameters such as network weights and biases. Based on the updated generator parameters, the parameter optimization terminal can iteratively adjust the generator's original parameters to obtain an updated generator.
[0105] S17. Input the condition variables into the update generator to obtain the updated production process state sequence, and calculate the residual sequence of the updated production process state sequence and the actual state parameter data.
[0106] For example, the parameter optimization terminal can input conditional variables into the update generator, calculate the updated production process state sequence through forward propagation, and then align the updated production process state sequence and the actual state parameter data according to the corresponding time points to calculate the difference and obtain the residual sequence.
[0107] S18. Input the residual sequence into the Gaussian mixture model and perform probability distribution fitting to obtain the residual distribution model.
[0108] For example, the parameter optimization terminal can input the residual sequence into the Gaussian mixture model to obtain the initial Gaussian mixture model parameters; the parameter optimization terminal can use the maximum likelihood estimation method to iteratively optimize the initial Gaussian mixture model parameters using the residual sequence as samples; the parameter optimization terminal can integrate the initial Gaussian mixture model parameters that have reached a preset threshold to obtain the residual distribution model.
[0109] S19. Repeat S12 to S18 until the update generator and residual distribution model reach the convergence condition. Based on the update generator and residual distribution model that have reached the convergence condition, construct the optimized digital twin generative adversarial network environment for this round of collaborative optimization.
[0110] Specifically, the convergence condition can be that the mean residual between the updated state sequence output by the update generator and the real state parameter data is less than a preset threshold, and the parameters of the residual distribution model tend to be stable and no longer change significantly.
[0111] For example, the parameter optimization terminal can repeatedly execute S12 to S18 to iteratively optimize the update generator and the residual distribution model. In each iteration, the generator parameters are adjusted based on the feedback from the residual distribution model in the previous round. When the update generator and the residual distribution model reach the convergence condition, the parameter optimization terminal can integrate and construct the optimized digital twin generative adversarial network environment for this round of collaborative optimization based on the update generator and the residual distribution model that have reached the convergence condition.
[0112] This application provides a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas. By accurately processing the input data, adapting the structure of the generator and discriminator, and iteratively optimizing the parameters, combined with loss calculation and residual distribution fitting verification, it achieves accurate optimization of the digital twin generative adversarial network environment, effectively improving the fidelity and stability of network simulation, ensuring that the generated production state sequence conforms to the actual production law, providing highly reliable and accurate simulation support for subsequent cross-process parameter collaborative optimization, while reducing the trial and error cost of network optimization and improving optimization efficiency.
[0113] Based on the above embodiments, the expression for the total loss function corresponding to the total loss function value can be:
[0114]
[0115] in, Let the total loss function of the generator be . For the generator, use a random noise vector and condition variables The production process state sequence generated as input; These are the probability values of the true data distribution. The first output of the generator Time-state vector For the first Mechanism prediction values at time points. The length of the time series. For the mechanism-constrained loss weighting coefficient, The weighting coefficients for time-series smoothing loss are denoted as .
[0116] This application provides a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas. By calculating the total loss function, it achieves multi-dimensional constraints on the generator output: adversarial loss ensures the generated state sequence closely matches real production data; mechanistic constraint loss guarantees the sequence conforms to the physicochemical mechanisms of traditional Chinese medicine production; and temporal smoothing loss avoids abrupt changes in the state sequence. These three factors synergistically improve the authenticity, rationality, and continuity of the generated sequence. Simultaneously, through flexible adjustment of weight coefficients, different constraint directions can be emphasized according to optimization needs, effectively improving the simulation fidelity of the digital twin generative adversarial network. This provides precise guidance for subsequent iterative optimization of generator parameters, ensuring that the optimized digital twin environment can accurately simulate actual production, laying a reliable foundation for cross-process parameter optimization.
[0117] In one embodiment of the present invention, a multi-agent cooperative strategy network is constructed through the following steps:
[0118] S21. In the optimized digital twin generative adversarial network environment, each process is defined as a corresponding process agent, and an upper-level coordinating agent is set up to load and obtain the initial multi-agent network.
[0119] Specifically, the upper-level coordinating agent can be used to coordinate the decision-making of agents in various processes and avoid parameter conflicts between processes.
[0120] For example, in an optimized digital twin generative adversarial network environment, the parameter optimization terminal can define each process as a corresponding process agent and set an upper-level coordinating agent. The parameter optimization terminal can load preset network initialization parameters to obtain an initialized multi-agent network. The network initialization parameters may include the initial weights and decision thresholds of each agent. The number of upper-level coordinating agents can be one.
[0121] S22. To initialize the agent in each process of the multi-agent network, design the observation space, action space, and reward function respectively to obtain the defined multi-agent network. The observation space includes the process parameters, state parameters, and upstream material attributes of the current process. The action space includes the adjustment amount of the process parameters of the current process. The reward function includes the quality index reward item and the energy consumption penalty item.
[0122] For example, the parameter optimization terminal can design the observation space, action space, and reward function for each agent in the initialization of the multi-agent network, thus obtaining a defined multi-agent network. The expression for the reward function can be:
[0123]
[0124] in, This represents the reward value for a single-process intelligent agent. As a weight for quality rewards, Points are awarded for meeting quality standards. As energy consumption penalty weight, This represents the actual energy consumption value of the process.
[0125] S23. Add a policy network based on the Actor-Critic architecture to the defined multi-agent network to obtain a multi-agent policy network.
[0126] Specifically, the Actor-Critic architecture consists of two parts: an Actor network and a Critic network. The Actor network, as the decision-making and execution unit, is mainly responsible for outputting specific action instructions, i.e., process parameter adjustment schemes for each process, based on the production data observed by the agents in each process. The core objective of the Actor network is to maximize long-term reward value. The Critic network, as the value evaluation unit, is mainly responsible for evaluating the rationality of the actions output by the Actor network. By calculating the value function corresponding to the action, it provides feedback on the quality of the action and provides a basis for updating the parameters of the Actor network.
[0127] For example, the parameter optimization terminal can configure a dedicated Actor network and Critic network associated with the upper-level coordinating agent for each process agent; the parameter optimization terminal can deeply fuse the observation space, action space, and reward function of the configured Actor network, Critic network, and the defined multi-agent network to obtain a multi-agent policy network.
[0128] S24. Obtain the process state parameters and process parameters at the current moment, input the process state parameters and process parameters at the current moment into the multi-agent policy network, and perform mapping calculation after splicing the process state parameters and process parameters at the current moment to obtain the target instruction.
[0129] Specifically, the parameter optimization terminal can acquire the process status parameters and process parameters of each process in real time through sensors. The process status parameters can include real-time operating indicators such as the operating speed of the equipment in each process, the concentration of the feed solution, and the content of the effective ingredients, while the process parameters can include adjustable indicators such as the currently set extraction temperature, concentration pressure, and drying time for each process.
[0130] For example, the parameter optimization terminal can perform min-max normalization on the current process state parameters and process parameters to eliminate the influence of different dimensions; the parameter optimization terminal can splice the normalized current process state parameters and process parameters and classify them by process to obtain comprehensive process parameters; the parameter optimization terminal can input the comprehensive process parameters into a multi-agent policy network, and perform feature mapping calculation through the fully connected layer inside the network to transform the comprehensive data into target instructions that can be recognized by the agents of each process. The target instructions can be used to set the parameter optimization direction and basic adjustment range of each process.
[0131] S25. Input the target instruction into each process agent, and perform forward propagation calculation through the Actor network in each process agent to obtain the process parameter adjustment amount.
[0132] For example, the parameter optimization terminal can categorize target instructions by process and input them into the agents of each process. The target instructions are then processed through the input layer, hidden layer, fully connected layer, and output layer of the Actor network in each process agent for forward propagation calculation. This process parses the target instructions, transforming the abstract instructions into specific process parameter adjustment quantities. Specifically, the input layer can standardize the target instructions before inputting them into the hidden layer. The hidden layer can then use the ReLU activation function and a fully connected layer to perform feature extraction, linear mapping, and nonlinear transformation on the standardized target instructions to obtain a process parameter feature vector. The output layer can then map the process parameter feature vector into process parameter adjustment quantities.
[0133] S26. Input the process parameter adjustment amount into the optimized digital twin generative adversarial network environment, perform state transition calculation, obtain the state parameters of the next moment, the output actions of each process agent and the reward value of each process agent, and then combine the current process state parameters, process parameters, next moment state parameters, output actions of each process agent and reward value of each process agent into an experience tuple and store it in the experience replay pool to obtain the experience replay pool.
[0134] For example, the parameter optimization terminal can perform compliance verification on the input process parameter adjustment amount. After the verification is passed, the parameter optimization terminal can control the digital twin environment to simulate the production state transition according to the adjusted process parameters, calculate and output the production state parameters of the next moment, the actual action performance of each process agent, energy consumption data and quality inspection results in real time; the parameter optimization terminal can calculate the reward value through the reward function; the parameter optimization terminal integrates the current process state parameters, process parameter adjustment amount, next moment state parameters, output actions of each agent and corresponding reward value into experience tuples according to the preset standard format, verifies the integrity and accuracy of the tuple data one by one, removes abnormal data, and stores qualified experience tuples in batches into the experience playback pool.
[0135] S27. Sample a small batch of experience tuples from the experience replay pool. Calculate the Critic network loss and Actor network policy gradient of each agent in each process based on the experience tuples. Update the network parameters of each agent in each process through the backpropagation algorithm to obtain the updated multi-agent policy network.
[0136] For example, the parameter optimization terminal can randomly sample a small batch of experience tuples from the experience replay pool, and calculate the Critic network loss and Actor network policy gradient for each agent in each process based on the experience tuples. The expression for calculating the Critic network loss can be:
[0137]
[0138] in, For Critic network loss, As a reward value, For the predictive value of the Critic network, The expected value is used; before calculating the gradient of the Actor network policy, the Actor network loss must be calculated first. The formula for calculating the Actor network loss is as follows:
[0139]
[0140] in, For Actor network loss, Let the action value function be... For an Actor network in a given state At that time, output action The probability, This indicates the calculation of the loss function relative to the Actor network parameters. gradient, Let be the logarithmic probability of the policy; the parameter optimization terminal can obtain the Actor network policy gradient based on the Actor network loss; the parameter optimization terminal can take the partial derivatives of the Actor network policy gradient and the Critic network loss respectively to obtain the parameter gradient value; the parameter optimization terminal can update the network parameters of each agent according to the preset learning rate based on the parameter gradient value to obtain the updated multi-agent policy network.
[0141] S28. Repeat S24 to S27 until the agents in each process converge, and obtain the multi-agent cooperative strategy network.
[0142] For example, the parameter optimization terminal can repeatedly execute all steps from S24 to S27, monitor the decision-making status of each process agent and the network convergence in real time, until the process parameter adjustment amount output by each process agent tends to stabilize and the reward value fluctuation is within the preset range, at which point the parameter optimization terminal can stop iterating and obtain the multi-agent collaborative strategy network.
[0143] This application provides a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas. By constructing a multi-agent network and incorporating an Actor-Critic architecture, the method clarifies the division of labor between each process agent and the upper-level coordinating agent, configures reasonable observation, action space, and reward mechanisms, verifies the rationality of process parameter adjustments in an optimized digital twin environment, and completes iterative optimization of network parameters based on an experience replay pool and backpropagation algorithm. Finally, a convergent and stable multi-agent collaborative strategy network is obtained, enabling precise decision-making and collaborative work among process agents, ensuring that process parameter adjustments are scientific and compliant, and adapting to the needs of the entire traditional Chinese medicine production process. This provides high-fidelity and high-reliability technical support for subsequent cross-process parameter optimization.
[0144] In one embodiment of the present invention, the distribution difference between historical simulation features and historical real features is calculated to obtain a distribution difference metric, including:
[0145] S31. Input the historical simulation features and historical real features into the maximum mean difference calculation function model to calculate the distribution difference measure.
[0146] Optionally, the expression for the maximum mean difference calculation function model can be:
[0147]
[0148] in, This is a measure of distributional difference. As a historical simulation feature, As a historical fact, For kernel function mapping, the expression for kernel function mapping is: ; These are historical simulation samples within the historical simulation features. A true historical sample with authentic historical characteristics; This is the bandwidth parameter of the kernel function, which controls the smoothness of the kernel function; For the regenerating nucleus Hilbert space norm, The first of the historical simulation features One historical simulation sample; The first characteristic of historical authenticity A true historical sample; This represents the total number of historical simulation samples included in the historical simulation features. The total number of historical real samples included in the historical real features.
[0149] For example, the parameter optimization terminal can unify the data format and align the dimensions of historical simulation features and historical real features respectively. The terminal can also use a maximum mean difference calculation function model to map the historical simulation feature set and the historical real feature set to a high-dimensional reproducing kernel Hilbert space using kernel functions. The terminal can calculate the mean of the historical simulation feature mapping vector and the mean of the historical real feature mapping vector respectively, and then calculate the squared norm of their difference in the reproducing kernel Hilbert space to obtain the distribution difference metric. Finally, the terminal can use a grid search combined with cross-validation to select the bandwidth parameter that maximizes the discriminative power of the distribution difference metric between the simulation and real distributions, thus calculating the distribution difference metric.
[0150] This application provides a cross-process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formulas. By introducing a maximum mean difference calculation function model, it maps historical simulation features and historical real features to a high-dimensional regenerative kernel Hilbert space and accurately quantifies the distribution difference between the two. This provides a unified and measurable quantitative basis for the degree of deviation between the simulation and the real environment. This step effectively makes up for the shortcomings of traditional methods in directly assessing simulation fidelity, enabling subsequent feature space mapping and parameter correction to be precisely adjusted based on clear difference indicators. This significantly improves the credibility and transfer reliability of the simulation environment's representation of the real production process during cross-process collaborative optimization.
[0151] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0152] Based on the same inventive concept, this application also provides a system for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas, used to implement the aforementioned method for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the system for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas provided below can be found in the limitations of the method for collaborative optimization of cross-process parameters in intelligent manufacturing of traditional Chinese medicine formulas described above, and will not be repeated here.
[0153] In one exemplary embodiment, such as Figure 3 As shown, a cross-process parameter collaborative optimization system 400 for intelligent manufacturing of traditional Chinese medicine formulas is provided, which can be used to implement the methods in the above-mentioned method embodiments, including:
[0154] The multimodal data and mechanism prediction module 401 can be used to acquire cross-process multimodal datasets of traditional Chinese medicine formulas in the production process, input the cross-process multimodal datasets into the mechanism skeleton model, solve the differential equations, and obtain the mechanism prediction values; wherein, the mechanism skeleton model is constructed from the material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the traditional Chinese medicine production process.
[0155] The twin environment physical constraint optimization module 402 can be used to update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the mechanism prediction value, and iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment in this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is a sequence of process parameters, and the output is a sequence of production process states;
[0156] The distribution difference measurement module 403 can be used to obtain historical real features in the real production environment. It inputs the cross-process multimodal dataset into the multi-agent cooperative strategy network in the optimized digital twin generative adversarial network environment to generate historical simulation features, and calculates the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference measurement value.
[0157] The feature space mapping and parameter correction module 404 can be used to divide the kernel function according to the distribution difference metric value and the optimal mapping function, solve for the mapping relationship from the real data feature space to the simulation data feature space, and adjust the process parameters of the cross-process multimodal dataset based on the mapping relationship to obtain the corrected cross-process multimodal dataset.
[0158] The multi-agent collaborative optimization decision module 405 is used to input the modified cross-process multimodal dataset into the multi-agent collaborative strategy network, perform multi-agent collaborative decision-making, and generate cross-process optimization parameters.
[0159] In one embodiment of the present invention, the twin environment physical constraint optimization module 402 can also be used for:
[0160] S11. Based on the cross-process multimodal dataset, obtain condition variable and real state parameter data;
[0161] S12. Input the condition variables into the generator in the digital twin generative adversarial network environment, and use the generator to perform forward propagation calculation on the condition variables to obtain the production process state sequence.
[0162] S13. Input the production process state sequence and mechanism prediction value into the physical consistency loss function, calculate the mean square error, and obtain the physical consistency loss. Input the production process state sequence into the time smoothness loss function, calculate the sum of squares of the Euclidean distances between adjacent time state vectors, and obtain the time smoothness loss.
[0163] S14. Input the production process state sequence into the discriminator in the digital twin generative adversarial network environment, perform forward propagation, and calculate the probability value of the real data distribution from the real data.
[0164] S15. Based on the probability values of the real data distribution, the adversarial loss of the generator is calculated using the adversarial loss function. The total loss function value of the generator is calculated by weighted summation of the adversarial loss, physical consistency loss, and temporal smoothness loss.
[0165] S16. Calculate the gradient value of the total loss function, input the gradient value into the learning rate function, calculate the generator update parameters, and update the generator according to the generator update parameters to obtain the updated generator.
[0166] S17. Input the condition variables into the update generator to obtain the updated production process state sequence, and calculate the residual sequence between the updated production process state sequence and the actual state parameter data.
[0167] S18. Input the residual sequence into the Gaussian mixture model and perform probability distribution fitting to obtain the residual distribution model;
[0168] S19. Repeat S12 to S18 until the update generator and residual distribution model reach the convergence condition. Based on the update generator and residual distribution model that have reached the convergence condition, construct the optimized digital twin generative adversarial network environment for this round of collaborative optimization.
[0169] In one embodiment of the present invention, the cross-process parameter collaborative optimization system 400 for intelligent manufacturing of traditional Chinese medicine formulas can also be used for:
[0170] S21. In the optimized digital twin generative adversarial network environment, each process is defined as a corresponding process agent, and an upper-level coordinating agent is set up to load and obtain the initial multi-agent network.
[0171] S22. To initialize each process agent in the multi-agent network, design the observation space, action space, and reward function respectively to obtain the defined multi-agent network. The observation space includes the process parameters, state parameters, and upstream material attributes of the current process. The action space includes the adjustment amount of the process parameters of the current process. The reward function includes the quality index reward item and the energy consumption penalty item.
[0172] S23. Add a policy network based on the Actor-Critic architecture to the defined multi-agent network to obtain a multi-agent policy network.
[0173] S24. Obtain the process state parameters and process parameters at the current moment, input the process state parameters and process parameters at the current moment into the multi-agent policy network, and perform mapping calculation after splicing the process state parameters and process parameters at the current moment to obtain the target instruction;
[0174] S25. Input the target instruction into each process agent, and perform forward propagation calculation through the Actor network in each process agent to obtain the process parameter adjustment amount;
[0175] S26. Input the process parameter adjustment amount into the optimized digital twin generative adversarial network environment, perform state transition calculation, obtain the state parameters of the next moment, the output actions of each process agent and the reward value of each process agent, and then combine the current process state parameters, process parameters, next moment state parameters, output actions of each process agent and reward value of each process agent into an experience tuple and store it in the experience replay pool to obtain the experience replay pool.
[0176] S27. Sample a small batch of experience tuples from the experience replay pool, calculate the Critic network loss and Actor network policy gradient of each agent based on the experience tuples, update the network parameters of each agent through the backpropagation algorithm, and obtain the updated multi-agent policy network.
[0177] S28. Repeat S24 to S27 until the agents in each process converge, and obtain the multi-agent cooperative strategy network.
[0178] In one embodiment of the present invention, the distribution difference measurement module 403 can also be used for:
[0179] S31. Input the historical simulation features and historical real features into the maximum mean difference calculation function model to calculate the distribution difference measure. The expression of the maximum mean difference calculation function model is as follows:
[0180]
[0181] in, This is a measure of distributional difference. As a historical simulation feature, As a historical fact, For kernel function mapping, the expression for kernel function mapping is: ; These are historical simulation samples within the historical simulation features. A true historical sample with authentic historical characteristics; This is the bandwidth parameter of the kernel function, which controls the smoothness of the kernel function; For the regenerating nucleus Hilbert space norm, The first of the historical simulation features One historical simulation sample; The first characteristic of historical authenticity A true historical sample; This represents the total number of historical simulation samples included in the historical simulation features. The total number of historical real samples included in the historical real features.
[0182] In one embodiment, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0183] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0184] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0185] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A cross-process process parameter collaborative optimization method for intelligent manufacturing of traditional Chinese medicine formula, characterized in that, The method includes: S1. Obtain the cross-process multimodal dataset of traditional Chinese medicine formula in the production process, input the cross-process multimodal dataset into the mechanism skeleton model, solve the differential equation, and obtain the mechanism prediction value; wherein, the mechanism skeleton model is constructed by material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the traditional Chinese medicine production process. S2. Update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the predicted value of the mechanism, and iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment of this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is a sequence of process parameters, and the output is a sequence of production process states; S3. Obtain historical real features in the real production environment, input the cross-process multimodal dataset into the multi-agent cooperative strategy network in the optimized digital twin generative adversarial network environment, generate historical simulation features, and calculate the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference metric. S4. Based on the distribution difference metric, combine the kernel function with the optimal mapping function to solve for the mapping relationship from the real data feature space to the simulated data feature space, and based on the mapping relationship, adjust the process parameters of the cross-process multimodal dataset to obtain the corrected cross-process multimodal dataset. S5. Input the modified cross-process multimodal dataset into the multi-agent cooperative strategy network to perform multi-agent cooperative decision-making and generate cross-process optimization parameters.
2. The method of claim 1, wherein, The process of updating the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the predicted value of the mechanism, and iteratively optimizing the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment for this collaborative optimization round includes: S11. Based on the cross-process multimodal dataset, obtain condition variable and real state parameter data; S12. Input the condition variable into the generator in the digital twin generative adversarial network environment, and perform forward propagation calculation on the condition variable through the generator to obtain the production process state sequence. S13. Input the production process state sequence and the mechanism prediction value into the physical consistency loss function, calculate the mean square error, and obtain the physical consistency loss. Input the production process state sequence into the time smoothness loss function, calculate the sum of squares of the Euclidean distances between adjacent time state vectors, and obtain the time smoothness loss. S14. Input the production process state sequence into the discriminator in the digital twin generative adversarial network environment, perform forward propagation, and calculate the real data distribution probability value from the real data. S15. Based on the probability value of the real data distribution, the adversarial loss of the generator is calculated using the adversarial loss function, and the total loss function value of the generator is calculated by weighted summation of the adversarial loss, the physical consistency loss, and the time smoothness loss. S16. Calculate the gradient value of the total loss function, input the gradient value into the learning rate function, calculate the generator update parameters, and update the generator according to the generator update parameters to obtain the updated generator. S17. Input the condition variable into the update generator to obtain the updated production process state sequence, and calculate the residual sequence of the updated production process state sequence and the real state parameter data; S18. Input the residual sequence into the Gaussian mixture model and perform probability distribution fitting to obtain the residual distribution model; S19. Repeat S12 to S18 until the update generator and the residual distribution model reach the convergence condition, and based on the update generator and the residual distribution model that have reached the convergence condition, construct the optimized digital twin generative adversarial network environment for this round of collaborative optimization.
3. The method of claim 2, wherein, The expression for the total loss function corresponding to the total loss function value is: in, The total loss function of the generator. The generator is given a random noise vector and condition variables The production process state sequence generated as input; The probability value of the true data distribution. The first output of the generator Time-state vector For the first The predicted value of the mechanism at time [time]. The length of the time series. For the mechanism-constrained loss weighting coefficient, The weighting coefficients for time-series smoothing loss are denoted as .
4. The method of claim 1, wherein, The multi-agent cooperative strategy network is constructed through the following steps: S21. In the optimized digital twin generative adversarial network environment, each process is defined as a corresponding process agent, and an upper-level coordinating agent is set to load and obtain an initial multi-agent network. S22. Design observation space, action space and reward function for each process agent in the initialized multi-agent network to obtain a defined multi-agent network. The observation space includes process parameters, state parameters and upstream material attributes of the current process. The action space includes the adjustment amount of process parameters of the current process. The reward function includes quality index reward items and energy consumption penalty items. S23. Add a policy network based on the Actor-Critic architecture to the defined multi-agent network to obtain a multi-agent policy network; S24. Obtain the process state parameters and process parameters at the current time, input the process state parameters and process parameters at the current time into the multi-agent policy network, and perform mapping calculation after concatenating the process state parameters and process parameters at the current time to obtain the target instruction; S25. Input the target instruction into each process agent, and perform forward propagation calculation through the Actor network in each process agent to obtain the process parameter adjustment amount; S26. Input the process parameter adjustment amount into the optimized digital twin generative adversarial network environment, perform state transition calculation, obtain the state parameters at the next moment, the output actions of each process agent and the reward value of each process agent, and then store the process state parameters at the current moment, the process parameters, the state parameters at the next moment, the output actions of each process agent and the reward value of each process agent into an experience tuple and store it in the experience replay pool to obtain the experience replay pool; S27. Sample a small batch of experience tuples from the experience replay pool, calculate the Critic network loss and Actor network policy gradient of each process agent based on the experience tuples, and update the network parameters of each process agent through the backpropagation algorithm to obtain the updated multi-agent policy network. S28. Repeat S24 to S27 until the agents of each process converge to obtain the multi-agent cooperative strategy network.
5. The method of claim 1, wherein, The step of calculating the distribution difference between the historical simulation features and the historical real features to obtain a distribution difference metric includes: S31. Input the historical simulation features and the historical real features into the maximum mean difference calculation function model to calculate the distribution difference metric, wherein the expression of the maximum mean difference calculation function model is: in, The distribution difference measure, For the aforementioned historical simulation features, For the aforementioned historical characteristics, This is a kernel function mapping, and the expression for the kernel function mapping is: ; These are the historical simulation samples in the historical simulation features. These are historical real samples that represent the aforementioned historical real characteristics; The bandwidth parameter is used to control the smoothness of the kernel function; For the regenerating nucleus Hilbert space norm, The first of the historical simulation features One historical simulation sample; The first of the historical true features A true historical sample; This refers to the total number of historical simulation samples included in the historical simulation features. The total number of historical real samples included in the historical real features.
6. A cross-process parameter collaborative optimization system for intelligent manufacturing of traditional Chinese medicine formulas, used to implement the method according to any one of claims 1 to 5, characterized in that, The system includes: The multimodal data and mechanism prediction module is used to acquire cross-process multimodal datasets of traditional Chinese medicine formulas in the production process, input the cross-process multimodal datasets into the mechanism skeleton model, solve the differential equations, and obtain the mechanism prediction values; wherein, the mechanism skeleton model is constructed from the material conservation and energy conservation differential equations of each process based on the physicochemical mechanism of the traditional Chinese medicine production process. The twin environment physical constraint optimization module is used to update the physical consistency constraint loss function of the digital twin generative adversarial network environment based on the mechanism prediction value, and to iteratively optimize the digital twin generative adversarial network environment based on the physical consistency constraint loss function to obtain the optimized digital twin generative adversarial network environment for this collaborative optimization round; wherein, the input of the digital twin generative adversarial network environment is a sequence of process parameters, and the output is a sequence of production process states; The distribution difference measurement module is used to obtain historical real features in the real production environment. It inputs the cross-process multimodal dataset into the multi-agent cooperative strategy network in the optimized digital twin generative adversarial network environment to generate historical simulation features, and calculates the distribution difference between the historical simulation features and the historical real features to obtain the distribution difference measurement value. The feature space mapping and parameter correction module is used to divide the kernel function according to the distribution difference metric value and the optimal mapping function to obtain the mapping relationship from the real data feature space to the simulated data feature space, and adjust the process parameters of the cross-process multimodal dataset based on the mapping relationship to obtain the corrected cross-process multimodal dataset. The multi-agent collaborative optimization decision module is used to input the modified cross-process multimodal dataset into the multi-agent collaborative strategy network, perform multi-agent collaborative decision-making, and generate cross-process optimization parameters.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.