Hydrogen fluoride production millisecond-level regulation method based on reinforcement learning and model predictive control
By combining reinforcement learning and model predictive control with various innovative technologies, the problems of millisecond-level response, sensor detection, and fault diagnosis in hydrogen fluoride production have been solved, achieving efficient and safe regulation of hydrogen fluoride production.
Patent Information
- Application Number
- CN202511182868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Traditional methods for controlling hydrogen fluoride production are insufficient to meet the millisecond-level response requirements under complex operating conditions. They suffer from problems such as large model prediction errors, insufficient sensor detection sensitivity, inadequate fault diagnosis and adaptive capabilities, and low efficiency of multi-reactor collaborative control.
We employ reinforcement learning and model predictive control methods, combining multi-scale neural symbolic dynamics modeling, causal reinforcement learning decision optimization, spatiotemporal fractional sliding mode control, memory-enhanced meta-learning adaptation, fluid topology optimization reactor design, fractional Brownian motion noise control, metamaterial sensor network deployment, spatiotemporal attention mechanism fusion, Lie group Lie algebra motion planning, multi-agent game collaborative control, memory causal graph network, and fractal dimension fault diagnosis.
It achieves millisecond-level control of hydrogen fluoride production, optimizes reactor temperature response, improves sensor detection sensitivity and fault detection accuracy, reduces energy consumption, and improves production efficiency and safety.
Smart Images

Figure CN120669666B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of hydrogen fluoride production regulation, and in particular to a hydrogen fluoride production millisecond-level regulation method based on reinforcement learning and model predictive control. BACKGROUND
[0002] Traditional hydrogen fluoride production regulation methods cannot meet the millisecond-level response requirements under complex working conditions. Existing systems mostly use static models based on mechanisms, which cannot dynamically adapt to parameter fluctuations in the reaction process. When the catalyst activity decays, the model prediction error often exceeds the actual value, causing the control command to lag. The sensor network has insufficient detection sensitivity, and the electrochemical sensor has a long response time, which cannot capture transient abnormalities such as trace leakage. Moreover, multi-parameter monitoring data lacks spatiotemporal correlation analysis, and the abnormal positioning time is long.
[0003] In the field of intelligent control, traditional reinforcement learning methods have the defects of slow policy convergence and poor generalization ability. Most systems use single-agent architecture, which cannot handle complex interactions in multi-reactor collaborative scenarios. For example, the optimization of flow distribution among multiple devices often falls into local optimum, resulting in reduced overall production efficiency. The rolling optimization process of model predictive control (MPC) has a large amount of calculation, and traditional solvers take a long time to solve optimization problems above 10 dimensions, which cannot meet the requirements of millisecond-level regulation. Moreover, it lacks robustness to parameter perturbations and external disturbances, and the product quality fluctuates greatly when the working conditions fluctuate.
[0004] Insufficient fault diagnosis and adaptive ability is another technical bottleneck. Traditional methods rely on threshold judgment, which has low recognition rate for early device degradation such as small carbon deposition on the catalyst, and lack a rapid adaptation mechanism for new working conditions. When the production load changes exceed the threshold, the control system needs to be re-adjusted for several hours, during which the product has a high unqualified rate. In addition, existing systems lack safety verification of control commands, which may trigger a safety boundary breakthrough in extreme working conditions, such as the risk of HF decomposition caused by reactor temperature overshoot. SUMMARY
[0005] The hydrogen fluoride production millisecond-level regulation method based on reinforcement learning and model predictive control proposed by the present application solves the problems mentioned in the prior art.
[0006] To achieve the above purpose, the present application adopts the following technical solution: a hydrogen fluoride production millisecond-level regulation method based on reinforcement learning and model predictive control, comprising:
[0007] Multi-scale neural symbolic dynamic modeling step: fuse neural network and symbolic reasoning to build multi-scale dynamic model, adopt graph neural network to learn HF synthesis reaction molecular structure and reaction mechanism at micro level, capture atomic interaction through message passing mechanism; develop fluid-solid coupling cellular automaton at meso level to simulate catalyst surface mass transfer and chemical reaction; establish state space model at macro level to analyze model generalization ability through neural tangent kernel theory; introduce differentiable physics engine for full-scale differentiable modeling;
[0008] Causal reinforcement learning decision optimization: design causal intervention reinforcement learning framework to identify causal relationships in HF production process through causal structure learning; build causal graph model to calculate causal effect using do-calculus; use counterfactual reasoning to evaluate strategy robustness in reinforcement learning training, and design reward function as R=E[Y∣do(A=a)]━E[Y∣do(A=a′)], Y is product quality index, A is control action, a is current action, a′ is comparison action;
[0009] Space-time fractional order sliding mode control step: introduce fractional calculus theory, design space-time fractional order sliding mode controller, define fractional order sliding surface , is tracking error, and are order and order fractional differential operators, is a positive definite matrix; calculate fractional differential value using Grünwald-Letnikov definition, and converge the system through sliding mode control;
[0010] Memory-enhanced meta-learning adaptation step: develop memory-enhanced meta-learning framework, design differentiable neural computer to store historical working conditions and optimal control strategies; retrieve relevant memories through attention mechanism when encountering new working conditions, initialize reinforcement learning agent policy network, and adjust the strategy using meta-gradient optimization algorithm.
[0011] Further, it further comprises:
[0012] Fluid topology optimization reactor design step: use topology optimization theory to design HF reactor, build topology optimization model based on fluid dynamics, parameterize internal structure of reactor through variable density method, and solve optimal structure using moving asymptote method.
[0013] Further, it further comprises:
[0014] Fractional Brownian motion noise control step: establish fractional Brownian motion noise model, estimate long-range correlation of noise through Hurst exponent, design fractional Kalman filter to estimate state, and use fractional LQG controller to process long memory characteristic noise.
[0015] Further, it also includes:
[0016] Robust training against generation: build a generative adversarial network to enhance the robustness of the control system, the generator simulates abnormal working conditions and disturbances, and the discriminator distinguishes between real data and generated data; reinforcement learning training uses adversarial samples as additional training data, and the agent learns the coping strategy under extreme conditions.
[0017] Further, it also includes:
[0018] Metamaterial sensor network deployment steps: design a HF concentration sensor network, and use the abnormal response characteristics of metamaterials to detect sensitivity; the sensor adopts a fractal structure design, and the optimal deployment position is determined through a topology optimization algorithm.
[0019] Further, it also includes:
[0020] Temporal-spatial attention mechanism fusion step: build a temporal-spatial attention mechanism to process multi-source heterogeneous data, the time attention module uses a gated recurrent unit to capture the time sequence dependency of parameters, and the spatial attention module extracts the spatial correlation of parameters through a graph neural network; the temporal-spatial attention weight calculation formula is , is the hidden state of the previous moment, is the input of the th sensor, is the number of sensors, and the key parameters are automatically focused through the attention mechanism.
[0021] Further, it also includes:
[0022] Lie group Lie algebra motion planning step: use Lie group Lie algebra theory to plan the motion of HF production equipment, the state of the equipment is represented as a special Euclidean group SE element, and the state parameterization is performed through Lie algebra; design a model predictive controller to maintain the smoothness of the equipment motion.
[0023] Further, it also includes:
[0024] Intelligent agent game collaborative control step: build a multi-agent game model to collaboratively control the HF production process, analyze the interaction of agent strategies through the theory of dynamic game with incomplete information, and design a distributed reinforcement learning algorithm to achieve global optimality.
[0025] Further, it also includes:
[0026] Memory causal graph network step: develop a memory-enhanced causal graph network to predict risks, capture parameter causal relationships through a graph neural network, and store historical risk patterns using a memory module; when an anomaly is detected, predict potential risks through causal intervention analysis and issue a warning.
[0027] Further, it also includes:
[0028] Fractal dimension fault diagnosis step: adopting fractal geometry theory to diagnose HF production equipment failure, calculating the fractal dimension of the time series of equipment operation parameters, identifying early failure through fractal feature extraction; design a fractal adaptive threshold algorithm to dynamically adjust the diagnostic threshold according to the production conditions.
[0029] Compared with the prior art, the beneficial effects of the present application are:
[0030] The multi-scale neural symbolic dynamics model optimizes the prediction error, combines the fractional order sliding mode control to optimize the reactor temperature response, improves the overshoot, and meets the millisecond-level regulation and control demand. The metamaterial sensor network achieves low detection lower limit and fast response, cooperates with the fractal dimension diagnosis technology, optimizes the fault detection, and early warns the equipment abnormality.
[0031] The causal reinforcement learning framework optimizes the product yield and reduces the energy consumption, and the meta-learning quickly adapts to the new working condition. The multi-agent game control optimizes the production capacity in the multi-reactor scene, shortens the fault recovery time, and solves the traditional centralized control coordination problem. The fluid topology optimizes the reactor, optimizes the pressure drop and reactant residence time distribution, and improves the mass transfer efficiency.
[0032] The spatio-temporal attention mechanism optimizes the abnormal detection accuracy, the memory causal graph network realizes the risk super-early warning, and the combination of the adversarial generation training optimizes the system success rate under unknown working conditions. In practical application, a certain HF production enterprise adopts the method, optimizes the non-planned parking condition, reduces the energy consumption per ton of product, saves the cost, reduces the equipment wear and tear by Lie group Lie algebra planning, and improves the production safety and economy. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A schematic diagram of a reinforcement learning and model predictive control millisecond-level regulation and control method for hydrogen fluoride production is proposed for the present application;
[0034] Figure 2 A schematic diagram of the comparison of key performance indicators before and after the system is implemented;
[0035] Figure 3 A comparison schematic diagram for the step response of the reactor temperature control;
[0036] Figure 4 A comparison schematic diagram of the detection accuracy of different fault types;
[0037] Figure 5 A comparison schematic diagram of the production efficiency of multi-reactor parallel production;
[0038] Figure 6 A comparison schematic diagram of the convergence speed of causal reinforcement learning and traditional RL. DETAILED DESCRIPTION
[0039] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.
[0040] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.
[0041] In addition, the terms "first", "second" are only used for description purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited. In addition, the terms "mounting", "connecting", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, or internal communication of two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances, and the present application will be further described in detail with reference to the drawings.
[0042] Referring to Figures 1 to 6 A hydrogen fluoride production millisecond control method combining reinforcement learning and model predictive control, comprising:
[0043] Multi-scale neural symbolic dynamic modeling step: a multi-scale dynamic model integrating neural networks and symbolic reasoning is constructed. At the microscopic level, a graph neural network (GNN) is used to model the HF synthesis reaction. The GraphNets framework developed by DeepMind is used to construct a molecular graph, with each node representing an atom and the edges representing the chemical bonds between atoms. Through the message passing mechanism, the atomic nodes continuously update their feature vectors to capture the complex interactions between atoms. During the training process, the GNN is supervised by the dataset generated by quantum chemistry calculation software (such as Gaussian), enabling the model to accurately predict the formation and decomposition of reaction intermediates. At the mesoscopic level, a fluid-solid coupled cellular automaton is developed. The reactor space is discretized into a regular grid, with each cell containing both fluid and solid state variables. The fluid state variables include temperature, pressure, concentration, etc., and the solid state variables include catalyst active site distribution and surface topography. The cell state update rules are based on the Navier-Stokes equation and chemical reaction kinetics equation, which are numerically solved by the finite difference method. To improve computational efficiency, GPU parallel computing technology is used, which increases the speed of mesoscopic simulation by 10 times. At the macroscopic level, a state space model based on physical laws is established. The key state variables of the system are determined, including the temperature, pressure, reactant and product concentrations of each reactor, and a state vector x(t) is constructed with 15 state variables. The control input vector u(t) includes 8 variables such as feed flow rate, heating power, and catalyst addition amount. The system matrix A and input matrix B are determined through system identification methods, and the recursive least squares (RLS) method is used to update the model parameters online. To ensure the accuracy of the model, actual production data is collected every 100 ms and compared with the model prediction value. When the error exceeds 5%, the model parameter update is triggered. A differentiable physics engine is introduced to integrate the microscopic, mesoscopic, and macroscopic models. The differentiable physics engine is based on the JAX framework, allowing end-to-end differentiation of the entire multi-scale model. Through automatic differentiation technology, the gradient of the model output with respect to the input parameters is calculated, providing support for subsequent optimization control.
[0044] Causal reinforcement learning decision optimization step: A causal intervention reinforcement learning framework is designed to identify causal relationships in the HF production process through causal structure learning. A causal graph model is constructed, and historical data of the HF production process are collected, including process parameters, equipment states, and product quality, etc. A total of 100,000 sample data are collected. The PC algorithm (Peter-Clark algorithm) is used for causal structure learning to identify the causal relationships between variables. After 1000 iterations of calculation, a causal graph containing 20 nodes and 35 edges is constructed, and the causal dependence relationships between parameters are determined. The do-calculus is used to calculate the causal effects of different control variables on product quality. For each control variable, different intervention values are set, and the expected changes in product quality indicators are calculated through the causal graph model. For example, when studying the causal effect of reaction temperature on HF yield, the temperature intervention values are set to 500°C, 550°C, and 600°C, and the corresponding expected HF yields are calculated to be 85%, 92%, and 88% respectively, thus determining the optimal reaction temperature as 550°C. In reinforcement learning training, the robustness of the counterfactual reasoning evaluation strategy is evaluated. For each executed action, the system response under other possible actions is simulated through the causal graph model, and the counterfactual reward is calculated. The reward function is designed as the difference between the causal effect of the current action and the causal effect of the comparative action, with the formula R=E[Y∣do(A=a)]━E[Y∣do(A=a′)], Y is the product quality indicator, A is the control action, a is the current action, and a' is the comparative action. During training, an experience replay buffer is used to store the state-action-reward sequence of the agent, with a buffer size of 100,000 records. The target network technology is used to stabilize the training process, and the target network is updated every 1000 steps. To improve training efficiency, a distributed reinforcement learning architecture is used. Ten parallel agents are deployed in different simulation environments for training, each initialized with a different random seed. The learning parameters of each agent are synchronized through a parameter server, and the parameters are updated every 500 steps. After 100,000 training iterations, the agent converges to the optimal strategy.
[0045] Spacetime fractional sliding mode control step: Fractional calculus theory is introduced into model predictive control to design a spacetime fractional sliding mode controller. The fractional sliding surface is defined as , is the tracking error, and are the order and order fractional differential operators, is a positive definite matrix. The Grünwald-Letnikov definition is used to realize the numerical calculation of fractional differentiation, and the sliding mode control is used to ensure the convergence of the system in a finite time. The controller improves the system response speed by 30% and reduces the overshoot to below 1%.
[0046] Memory-augmented meta-learning adaptation step: develop a memory-augmented meta-learning framework to enable the control system to quickly adapt to new operating conditions. Design a differentiable neural computer (DNC) as an external memory module to store historical operating conditions and optimal control strategies. The DNC consists of a controller and a memory matrix, the controller is implemented using an LSTM network, and the memory matrix is set to 1000x100, which can store 1000 memory vectors with a length of 100. The reading and writing operations of the memory are realized through the attention mechanism, and the number of reading heads is set to 3 and the number of writing heads is set to 1. During the training process, the historical operating conditions and the corresponding optimal control strategies are stored in the memory matrix. Each memory entry contains a condition description vector and a control strategy vector, the condition description vector consists of 15 process parameters, and the control strategy vector consists of 8 control variables. When a new operating condition is encountered, the similarity between the current operating condition and the memory entries is calculated through the attention mechanism, and the most relevant memory entry is selected to initialize the policy network of the reinforcement learning agent. The meta-gradient optimization algorithm is used for rapid strategy adjustment. In the meta-training phase, different tasks are sampled from the historical operating conditions, each task contains a set of training samples and test samples. For each task, first update the policy network of the agent using the training samples, then calculate the meta-gradient using the test samples, and update the parameters of the meta-learner. The learning rate of the meta-learner is set to 0.001, and the Adam optimizer is used to update the parameters. When a new operating condition is detected, the system retrieves relevant information from the memory to initialize the policy network of the agent. A few sample data under the current operating condition are used for fine-tuning of 5 gradient steps, which can complete the strategy adjustment. To evaluate the performance of memory-augmented meta-learning adaptation, a series of comparative experiments are designed. In the experiment, the traditional reinforcement learning method, the meta-learning method and the memory-augmented meta-learning method are used to control the HF production process respectively.
[0047] The present application further comprises the following steps:
[0048] Fluid topology optimization reactor design step: Apply topology optimization theory to HF reactor design, build a topology optimization model based on fluid dynamics. Discretize the reactor internal space into a three-dimensional grid, the density of each grid cell as a design variable, the density value between 0 (representing fluid area) and 1 (representing solid area) continuous change. SIMP (Solid Isotropic Material with Penalization) method is used to deal with material interpolation problem, the penalty factor is set to 3, to ensure the clear boundary of the optimization result. The objective function is designed as the weighted sum of maximum reaction efficiency and minimum pressure loss. Reaction efficiency is evaluated by calculating the ratio of the mole fraction of HF at the outlet to the theoretical maximum mole fraction, and pressure loss is evaluated by calculating the pressure difference between the inlet and outlet of the reactor. The weight coefficient is adjusted according to the production demand, in this implementation, the weight of reaction efficiency is set to 0.7, and the weight of pressure loss is set to 0.3. The constraint conditions include the total volume constraint of the reactor, the structural continuity constraint and the physical constraint of fluid flow. The total volume constraint ensures that the volume of the optimized reactor does not exceed 110% of the original design volume; the structural continuity constraint is realized through filtering technology, which ensures the smooth change of the density of adjacent grid cells; the physical constraints of fluid flow include mass conservation, momentum conservation and energy conservation equations. Moving asymptote method (MMA) is used to solve the topology optimization problem. MMA is a highly efficient nonlinear optimization algorithm suitable for handling engineering optimization problems with a large number of design variables. In the optimization process, fluid dynamics equations need to be solved at each iteration to evaluate the objective function and constraint conditions. To improve computational efficiency, finite volume method is used to discretize fluid dynamics equations, and OpenFOAM open source computational fluid dynamics software is used for solution. After 200 optimization iterations, the optimal reactor internal structure is obtained. The optimized reactor adopts a honeycomb structure design, which contains multiple interconnected channels, which can effectively improve the mixing degree and contact area of reactants. Under the premise of maintaining the same production capacity, the new design of the reactor makes the uniformity of reactant residence time distribution improve by 40%, and the pressure drop reduce by 25%.
[0049] In the present invention, the following steps are also included:
[0050] Fractional Brownian motion noise control step: For the random noise in the HF production process, a fractional Brownian motion (FBM) noise model is established. The historical data of key parameters in the HF production process are collected, including temperature, pressure, concentration, etc., a total of 1 million sample data are collected. The Hurst index of the noise is estimated by R / S analysis method, the results show that the noise in the HF production process has long-range correlation, and the Hurst index is about 0.7. Based on the estimated Hurst index, a fractional Brownian motion noise model is established. The Monte Carlo method is used to generate noise samples conforming to the characteristics of fractional Brownian motion, which is used for subsequent controller design and testing. In order to verify the accuracy of the noise model, the generated noise samples are compared with the actual production data, and the statistical test shows that they have similar statistical characteristics. The fractional order Kalman filter is designed for state estimation. The fractional order Kalman filter is an extension of the traditional Kalman filter in the fractional order domain, which can better handle the noise with long memory characteristics. According to the fractional Brownian motion noise model, the recursive formula of the fractional order Kalman filter is derived, including the prediction equation and the update equation. In practical application, the state space augmentation method is used to realize the fractional order Kalman filter, which converts the fractional order differential operator into a state space model. The fractional order LQG controller is used to handle the noise with long memory characteristics. The fractional order LQG controller combines the fractional order control theory and the linear quadratic Gaussian control theory, which can realize optimal control in the presence of noise. The gain matrix of the controller is determined by solving the fractional order Riccati equation, in this implementation, the iterative algorithm is used to solve the fractional order Riccati equation, and the iteration precision is set to 10 -6 The experimental verification is carried out on the HF synthesis reactor, and the results show that the fractional Brownian motion noise control method improves the anti-interference ability of the system by 60%, and reduces the product quality fluctuation by 50%. Compared with the traditional control method, the state estimation error of the fractional order Kalman filter is reduced by 30%, and the control effect of the fractional order LQG controller is better, which can better suppress the influence of the noise with long-range correlation on the system.
[0051] In the present application, the following steps are also included:
[0052] Adversarial robust training step: a generative adversarial network (GAN) is constructed to enhance the robustness of the control system. The generator adopts a deep convolutional neural network (DCGAN) architecture, which contains 5 convolutional layers and 5 deconvolutional layers, to generate various abnormal working conditions and interference scenarios. The discriminator also adopts the DCGAN architecture to distinguish between real data and generated data. The hidden layers of the generator and the discriminator use the LeakyReLU activation function, and the output layer uses the Sigmoid activation function. During the training process, the GAN is optimized using the minimax game framework. The goal of the generator is to generate samples that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish between real data and generated data. The training process adopts an alternating optimization strategy, first updating the parameters of the discriminator, and then updating the parameters of the generator. The training batch size is set to 64, the learning rate is set to 0.0002, and the Adam optimizer is used for parameter update. To improve the quality of the generated samples, a gradient penalty mechanism is introduced. A gradient penalty term is added to the loss function of the discriminator to constrain the gradient norm of the discriminator, ensuring that it satisfies the Lipschitz continuity condition. The gradient penalty coefficient is set to 10, which is verified through experiments to effectively improve the training stability of the GAN and the quality of the generated samples. In reinforcement learning training, the generated adversarial samples are used as additional training data. In each training cycle, half of the real data and half of the generated data are sampled to form a training batch. To ensure the diversity of the generated data, the parameters of the GAN are updated regularly, and the GAN is retrained every 1000 training steps. To evaluate the effect of adversarial robust training, experiments are conducted on the HF production simulator. Three scenarios are set up: a controller trained only with real data, a controller trained with real data and random noise, and a controller trained with adversarial robust training.
[0053] In the present invention, the following steps are also included:
[0054] The deployment steps of the metamaterial sensor network are as follows: a metamaterial-based HF concentration sensor network is designed, the abnormal response characteristics of the metamaterial to specific frequency electromagnetic waves are utilized, and high-sensitivity detection is realized. The sensor adopts a fractal structure design, which can provide a larger surface area in a limited space and enhance the interaction with HF molecules. The metamaterial is composed of periodically arranged metal-dielectric units, the metal material is gold, the dielectric material is silicon dioxide, and the unit size is designed as one-quarter of the working frequency wavelength. The structure parameters of the metamaterial are optimized through electromagnetic simulation software (CST Microwave Studio), and a strong electromagnetic response is generated at the characteristic absorption frequency (about 10.7 GHz) of the HF molecule. The optimized metamaterial has an absorption coefficient of HF molecules that is 5 times higher than that of traditional materials, and the detection sensitivity is significantly improved. In terms of sensor manufacturing, electron beam lithography and thin film deposition technology are used to prepare the metamaterial structure. A layer of silicon dioxide film is deposited on the silicon substrate, the fractal pattern is defined using electron beam lithography technology, and the gold layer is deposited by physical vapor deposition method to form the metamaterial structure. A molecular recognition layer with specific adsorption effect on HF is modified on the surface of the metamaterial to improve the selectivity of the sensor. The optimal deployment position of the sensor is determined by a topological optimization algorithm. A mathematical model of HF leakage diffusion is established, and the finite element method is used to solve the model to obtain the HF concentration distribution at different positions. Based on the concentration distribution, a topological optimization objective function is designed, with the maximum coverage rate and the minimum number of sensors as the target. Genetic algorithm is used to solve the optimization problem, and after 500 generations of evolution, the optimal sensor deployment scheme is obtained. In actual deployment, the sensor is installed at key positions such as reactors, pipelines and storage tanks to form a sensor network covering the entire production area. The sensor transmits data to the monitoring center through wireless communication, the communication protocol uses ZigBee protocol, the transmission distance can reach 100 meters, and the transmission rate is 250 kbps.
[0055] In the present application, the following steps are also included:
[0056] The spatio-temporal attention mechanism fusion step: a spatio-temporal attention mechanism is constructed to process multi-source heterogeneous data. The time attention module adopts a gated recurrent unit (GRU) to capture the time sequence dependence of parameters, and the space attention module extracts the spatial correlation between parameters through a graph neural network (GNN). The spatio-temporal attention weight calculation formula is , is the hidden state of the previous moment, is the input of the th sensor, is the number of sensors. Through the attention mechanism, the system automatically focuses on key parameters, and the anomaly detection accuracy is improved to 98.7%, and the response time is shortened to 2ms.
[0057] In the present application, the following steps are also included:
[0058] Lie group Lie algebra motion planning: apply Lie group Lie algebra theory to motion planning of HF production equipment. Represent the state of the equipment as an element in the special Euclidean group SE(3), parameterize the state through Lie algebra. Design a Lie group-based model predictive controller, which shortens the regulation time by 50% while maintaining smooth equipment motion, and improves positioning accuracy to 0.01mm.
[0059] The present application further comprises the following steps:
[0060] Multi-agent game collaborative control step: build a multi-agent game model to realize collaborative control of the HF production process. Each production unit in the HF production process is regarded as an independent agent, including reactors, separators, compressors, etc. A total of 8 agents are designed. Each agent has its own state space, action space and objective function, and can make autonomous decisions and interact with other agents. An incomplete information dynamic game model is established. Each agent can only observe its own local state and part of the state information of other agents, forming incomplete information. The decision-making process of the agent is a dynamic game process, and each agent selects the optimal strategy according to its own belief and the action history of other agents at each time step. A distributed reinforcement learning algorithm is designed to solve the game equilibrium. The multi-agent deep deterministic policy gradient (MADDPG) algorithm is used, and each agent has its own policy network and critic network. The policy network generates actions based on the local state of the agent, and the critic network evaluates the value of joint actions. During the training process, each agent uses its own experience replay buffer to store state-action-reward sequences, and uses the target network technology to stabilize the training process. In order to promote cooperation between agents, a reward function based on trust mechanism is designed. The reward of each agent is composed of individual reward and social reward, the individual reward is based on the performance index of itself, and the social reward is based on the performance index of the whole system. The weight of the social reward is dynamically adjusted according to the trust degree between agents, and the trust degree is calculated through the historical cooperation behavior of the agent. The multi-agent game collaborative control method is applied in the multi-reactor parallel production scene. Each reactor agent adjusts the reaction temperature, pressure and feed flow rate according to its own state and the state information of other reactors. In the face of abnormal situations such as equipment failure and raw material fluctuation, the system can quickly redistribute tasks to maintain stable production operation, showing good robustness and adaptability.
[0061] The present application further comprises the following steps:
[0062] Memory causal graph network steps: develop a memory-enhanced causal graph network (MCGN) for risk prediction. Use causal discovery algorithms (such as PC algorithm, GES algorithm) for causal structure learning, through multiple iterations and verification, build a causal graph containing 30 nodes and 50 edges. The nodes in the graph represent various parameters and variables, the edges represent the causal relationship between variables, the direction of the edge represents the causal direction, and the weight of the edge represents the causal strength. Develop a memory module to store historical risk patterns. The memory module stores information in the form of key-value pairs, the key is the feature vector of the risk pattern, and the value is the corresponding risk description and coping strategy. The feature vector is obtained by dimensionality reduction processing of historical abnormal data, using principal component analysis (PCA) method to reduce high-dimensional abnormal data to 10 dimensions. The capacity of the memory module is set to 1000 records, when reaching the capacity limit, the least recently used (LRU) strategy is used to delete the least used records. Design the memory causal graph network (MCGN) architecture. MCGN consists of a causal graph encoder, a memory retrieval module, and a risk prediction module. The causal graph encoder uses graph neural networks (GNN) to encode the causal graph and extract the causal relationship features between variables. The memory retrieval module retrieves relevant historical risk patterns from the memory module based on the current abnormal features. The risk prediction module combines the causal graph encoding and memory retrieval results to predict the development trend and impact of potential risks. In the risk prediction process, the influence of different variables on the risk is evaluated through causal intervention analysis. For each potential risk variable, the risk change under different intervention values is calculated through do-calculus to determine the key influencing factors. Use the historical risk patterns stored in the memory module to predict the possible development path and consequences of the current risk through analogical reasoning.
[0063] The present application further comprises the following steps:
[0064] Fractal dimension fault diagnosis steps: apply fractal geometry theory to the fault diagnosis of HF production equipment. Calculate the fractal dimension of the time series of equipment operating parameters, and identify early faults through fractal feature extraction. Design a fractal adaptive threshold algorithm to dynamically adjust the diagnostic threshold according to the production conditions.
[0065] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can make equivalent replacement or change according to the technical solution and inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A millisecond-level control method for hydrogen fluoride production using reinforcement learning and model predictive control, characterized in that, Includes the following steps: Multi-scale neural symbolic dynamics modeling steps: A multi-scale dynamics model is constructed by integrating neural networks and symbolic reasoning. At the microscopic level, graph neural networks are used to learn the molecular structure and reaction mechanism of HF synthesis, capturing atomic interactions through message passing mechanisms. At the mesoscopic level, fluid-solid coupled cellular automata are developed to simulate mass transfer and chemical reactions on the catalyst surface. At the macroscopic level, a state-space model is established, and the model's generalization ability is analyzed through neural tangent kernel theory. A differentiable physics engine is introduced for full-scale differentiable modeling. Causal reinforcement learning decision optimization: Design a causal intervention reinforcement learning framework to identify causal relationships in the HF production process through causal structure learning; Construct a causal graph model and use do-calculus to calculate causal effects; The reinforcement learning training uses counterfactual reasoning to evaluate the robustness of the strategy. The reward function is designed as R=E[Y|do(A=a)]−E[Y|do(A=a′)], where Y is the product quality index, A is the control action, a is the current action, and a′ is the comparison action. The steps of spatiotemporal fractional sliding mode control are as follows: Introduce fractional calculus theory, design a spatiotemporal fractional sliding mode controller, and define the fractional sliding surface. , To track errors, and They are respectively order and Fractional differential operators of order, The matrix is positive definite; the fractional derivative is calculated using the Grünwald-Letnikov definition, and the convergence system is controlled by sliding mode. The steps of memory-enhanced meta-learning adaptation are as follows: Develop a memory-enhanced meta-learning framework, design a differentiable neural computer to store historical working conditions and optimal control strategies; when encountering new working conditions, retrieve relevant memories through an attention mechanism, initialize the reinforcement learning agent policy network, and adjust the strategy using a meta-gradient optimization algorithm. Fractional Brownian motion noise control steps: Establish a fractional Brownian motion noise model, estimate the long-range correlation of noise through Hurst exponent, design a fractional Kalman filter to estimate the state, and use a fractional LQG controller to process noise with long memory characteristics. Generative Adversarial Network (GAN) Robust Training: Constructing a generative adversarial network to enhance the robustness of the control system. The generator simulates abnormal operating conditions and disturbances, and the discriminator distinguishes between real data and generated data. Reinforcement learning training uses adversarial examples as additional training data, allowing the agent to learn strategies to cope with extreme situations.
2. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: The design steps for a fluid topology-optimized reactor are as follows: The HF reactor is designed using topology optimization theory, a topology optimization model is constructed based on fluid dynamics, the internal structure of the reactor is parameterized using the variable density method, and the optimal structure is solved using the moving asymptote method.
3. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: Deployment steps for metamaterial sensor networks: Design an HF concentration sensor network and utilize the anomalous response characteristics of metamaterials to detect sensitivity; The sensor adopts a fractal structure design and uses a topology optimization algorithm to determine the optimal deployment location.
4. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: The spatiotemporal attention mechanism fusion steps are as follows: a spatiotemporal attention mechanism is constructed to process multi-source heterogeneous data. The temporal attention module uses gated recurrent units to capture the temporal dependence of parameters, and the spatial attention module extracts the spatial correlation of parameters through graph neural networks. The formula for calculating spatiotemporal attention weights is: , The state was hidden in the previous moment. For the first The input of each sensor, The number of sensors is used to automatically focus on key parameters through an attention mechanism.
5. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: Lie group and Lie algebra motion planning steps: Use Lie group and Lie algebra theory to plan the motion of HF production equipment. The equipment state is represented as a special Euclidean group SE element, and the state is parameterized by Lie algebra; design a model predictive controller to maintain the smoothness of equipment motion.
6. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: The steps of intelligent agent game cooperative control are as follows: construct a multi-agent game model to cooperatively control the HF production process, analyze the agent strategy interaction through incomplete information dynamic game theory, and design a distributed reinforcement learning algorithm to achieve global optimum.
7. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: Steps of memory-enhanced causal graph network: Develop a memory-enhanced causal graph network to predict risks, capture causal relationships of parameters through graph neural networks, and use memory modules to store historical risk patterns; When an anomaly is detected, potential risks are predicted through causal intervention analysis, and an early warning is issued.
8. The millisecond-level regulation method for hydrogen fluoride production using reinforcement learning and model predictive control according to claim 1, characterized in that, Also includes: Fractal dimension fault diagnosis steps: Fractal geometry theory is used to diagnose faults in HF production equipment. The fractal dimension of the equipment operating parameters in time series is calculated, and early faults are identified by fractal feature extraction. A fractal adaptive threshold algorithm is designed to dynamically adjust the diagnostic threshold according to the production conditions.
Citation Information
Patent Citations
Depth deterministic strategy gradient driven multi-constraint guidance law design method
CN119644742A
Big data-driven hydrogen fluoride purification risk entropy assessment and multi-parameter monitoring system
CN120449107A