Millisecond regulation and control method for hydrogen fluoride production based on reinforcement learning and model prediction control
By combining reinforcement learning and model predictive control methods with a variety of innovative technical means, the problems of millisecond-level response and fault detection in hydrogen fluoride production have been solved, efficient and safe hydrogen fluoride production regulation has been achieved, and production efficiency and product quality have been improved.
Patent Information
- Application Number
- CN202511182868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Traditional hydrogen fluoride production control methods are unable to meet the millisecond-level response requirements under complex working conditions. They have problems such as large model prediction errors, insufficient sensor detection sensitivity, insufficient fault diagnosis and adaptive capabilities, and insufficient safety verification, resulting in low production efficiency and unstable product quality.
By adopting the methods of reinforcement learning and model predictive control, combined with multi-scale neural symbolic dynamics modeling, causal reinforcement learning decision optimization, spatiotemporal fractional-order sliding mode control, memory-enhanced meta-learning adaptation, fluid topology optimization reactor design, fractional Brownian motion noise control, metamaterial sensor network deployment, spatiotemporal attention mechanism fusion, Lie group and Lie algebra motion planning, multi-agent game collaborative control, memory causal graph network and fractal dimension fault diagnosis and other technical means, the real-time regulation and fault detection of the hydrogen fluoride production process are optimized.
It has achieved millisecond-level regulation of the hydrogen fluoride production process, improved sensor detection sensitivity and fault detection accuracy, optimized production efficiency and product quality, reduced energy consumption and equipment wear, and improved production safety and economy.
Smart Images

Figure CN120669666A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrogen fluoride production control, and in particular to a millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control. Background Art
[0002] Traditional hydrogen fluoride production control methods struggle to meet the millisecond-level response requirements under complex operating conditions. Existing systems often use static, mechanism-based models that are unable to dynamically adapt to parameter fluctuations during the reaction process. When catalyst activity decays, model prediction errors often exceed actual values, resulting in delayed control instructions. Sensor networks lack detection sensitivity, and electrochemical sensors have long response times, making them unable to capture transient anomalies such as trace leaks. Furthermore, multi-parameter monitoring data lacks temporal and spatial correlation analysis, making it difficult to locate anomalies.
[0003] In the field of intelligent control, traditional reinforcement learning methods suffer from slow policy convergence and poor generalization capabilities. Most systems use a single-agent architecture, which makes it difficult to handle complex interactions in multi-reactor collaborative scenarios. For example, flow distribution optimization between multiple devices often falls into local optimality, resulting in reduced overall production efficiency. The rolling optimization process of model predictive control (MPC) is computationally intensive. Traditional solvers take a long time to solve optimization problems with more than 10 dimensions, unable to meet millisecond-level control requirements. They also lack robustness to parameter perturbations and external interference, resulting in large fluctuations in product quality when operating conditions fluctuate.
[0004] Insufficient fault diagnosis and adaptive capabilities are another technical bottleneck. Traditional methods rely on threshold judgments, with low recognition rates for early equipment degradation, such as minor carbon deposits on catalysts, and lack mechanisms for rapid adaptation to new operating conditions. When production load changes exceed thresholds, the control system requires several hours to re-debug, during which time the product rejection rate is high. Furthermore, existing systems lack safety verification of control commands, potentially triggering safety boundary breaches under extreme operating conditions, such as the risk of HF decomposition caused by reactor temperature overshoot. Summary of the Invention
[0005] The present invention proposes a millisecond-level regulation method for hydrogen fluoride production based on reinforcement learning and model predictive control to solve the problems mentioned in the above-mentioned prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control, comprising: Multi-scale neural symbolic dynamics modeling steps: A multi-scale dynamics model is constructed by integrating neural networks and symbolic reasoning. At the microscopic level, a graph neural network is used to learn the molecular structure and reaction mechanism of the HF synthesis reaction, and atomic interactions are captured through a message passing mechanism. At the mesoscopic level, a fluid-solid coupled cellular automaton is developed to simulate mass transfer and chemical reactions on the catalyst surface. At the macroscopic level, a state space model is established, and the model's generalization capability is analyzed through neural tangent kernel theory. A differentiable physics engine is introduced for full-scale differentiable modeling. Causal reinforcement learning decision optimization: A causal intervention reinforcement learning framework is designed to identify causal relationships in the HF production process through causal structure learning. A causal graph model is constructed and the do-calculus is used to calculate causal effects. Reinforcement learning training uses counterfactual reasoning to evaluate the robustness of the strategy. The reward function is designed as R=E[Y|do(A=a)]━E[Y|do(A=a′)], where Y is the product quality indicator, A is the control action, a is the current action, and a′ is the comparison action. Steps of space-time fractional-order sliding mode control: Introduce fractional-order calculus theory, design space-time fractional-order sliding mode controller, and define fractional-order sliding mode surface , is the tracking error, and They are Order and order fractional differential operators, is a positive definite matrix; the Grünwald-Letnikov definition is used to calculate the fractional differential value, and the system is converged by sliding mode control; Memory-enhanced meta-learning adaptation steps: Develop a memory-enhanced meta-learning framework and design a differentiable neural computer to store historical operating conditions and optimal control strategies; when encountering new operating conditions, retrieve relevant memories through the attention mechanism, initialize the reinforcement learning agent policy network, and use the meta-gradient optimization algorithm to adjust the strategy.
[0007] Furthermore, it also includes: The design steps of the fluid topology optimization reactor are as follows: the HF reactor is designed using topology optimization theory, a topology optimization model is constructed based on fluid dynamics, the internal structure of the reactor is parameterized using the variable density method, and the moving asymptote method is used to solve the optimal structure.
[0008] Furthermore, it also includes: The steps of fractional Brownian motion noise control are as follows: establish a fractional Brownian motion noise model, estimate the long-range correlation of noise through Hurst exponent, design a fractional-order Kalman filter to estimate the state, and use a fractional-order LQG controller to process long-memory characteristic noise.
[0009] Furthermore, it also includes: Adversarial generative robust training: A generative adversarial network is constructed to enhance the robustness of the control system. The generator simulates abnormal operating conditions and interference, and the discriminator distinguishes between real data and generated data. Reinforcement learning training uses adversarial samples as additional training data, and the intelligent agent learns response strategies in extreme situations.
[0010] Furthermore, it also includes: Deployment steps for metamaterial sensor networks: Design an HF concentration sensor network and use the abnormal response characteristics of metamaterials to detect sensitivity; the sensors adopt a fractal structure design, and the optimal deployment location is determined through a topology optimization algorithm.
[0011] Furthermore, it also includes: Steps of spatiotemporal attention mechanism fusion: Construct a spatiotemporal attention mechanism to process multi-source heterogeneous data. The temporal attention module uses a gated recurrent unit to capture the temporal dependency of parameters, and the spatial attention module extracts the spatial correlation of parameters through a graph neural network. The spatiotemporal attention weight calculation formula is: , is the hidden state at the previous moment, For the The input of the sensor, For the number of sensors, the key parameters are automatically focused through the attention mechanism.
[0012] Furthermore, it also includes: Lie group and Lie algebra motion planning steps: Use Lie group and Lie algebra theory to plan the motion of HF production equipment. The equipment state is represented as a special Euclidean group SE element, and the state is parameterized through Lie algebra; design a model predictive controller to maintain the smoothness of equipment motion.
[0013] Furthermore, it also includes: Agent game collaborative control steps: Build a multi-agent game model to collaboratively control the HF production process, analyze the agent strategy interaction through incomplete information dynamic game theory, and design a distributed reinforcement learning algorithm to achieve global optimization.
[0014] Furthermore, it also includes: Memory Causal Graph Network Steps: Develop a memory-enhanced causal graph network to predict risks, capture parameter causal relationships through graph neural networks, and use memory modules to store historical risk patterns; when anomalies are detected, predict potential risks through causal intervention analysis and issue early warnings.
[0015] Furthermore, it also includes: Fractal dimension fault diagnosis steps: Use fractal geometry theory to diagnose HF production equipment faults, calculate the fractal dimension of the equipment operating parameter time series, and identify early faults through fractal feature extraction; design a fractal adaptive threshold algorithm and dynamically adjust the diagnostic threshold according to production conditions.
[0016] Compared with the existing technology, the beneficial effects of the present invention are: A multiscale neural symbolic dynamics model optimizes prediction errors, combined with fractional-order sliding mode control to optimize reactor temperature response and reduce overshoot, meeting millisecond-level control requirements. A metamaterial sensor network achieves low detection limits and rapid response, combined with fractal dimension diagnostic technology to optimize fault detection and provide early warning of equipment anomalies.
[0017] A causal reinforcement learning framework optimizes product yield, reduces energy consumption, and leverages meta-learning to rapidly adapt to new operating conditions. Multi-agent game control optimizes production capacity in multi-reactor scenarios, shortens fault recovery time, and addresses coordination issues inherent in traditional centralized control. Fluid topology optimization of reactors optimizes pressure drop and reactant residence time distribution, improving mass transfer efficiency.
[0018] The spatiotemporal attention mechanism optimizes anomaly detection accuracy, the memory causal graph network enables ultra-early risk warning, and adversarial generative training is combined to optimize system success rates under unknown operating conditions. In practical applications, a HF production company implemented this method to optimize unplanned downtime, reduce energy consumption per ton of product, and save costs. Using Lie group and Lie algebraic programming, it reduced equipment wear and improved production safety and economic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic block diagram of a millisecond-level regulation method for hydrogen fluoride production based on reinforcement learning and model predictive control proposed by the present invention; Figure 2 Schematic diagram comparing key performance indicators before and after system implementation; Figure 3 This is a schematic diagram for comparing the step response of reactor temperature control; Figure 4 This is a schematic diagram comparing the detection accuracy of different fault types; Figure 5 This is a schematic diagram comparing the parallel production efficiency of multiple reactors; Figure 6 Schematic diagram comparing the convergence speed of causal reinforcement learning and traditional RL. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0022] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.
[0023] Reference Figures 1 to 6 A millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control, comprising: Multiscale neural symbolic dynamics modeling steps: Construct a multiscale dynamics model that integrates neural networks and symbolic reasoning. At the microscopic level, a graph neural network (GNN) is used to model the HF synthesis reaction. The molecular graph is constructed using the GraphNets framework developed by DeepMind, where each node represents an atom and edges represent chemical bonds between atoms. Through a message-passing mechanism, atomic nodes continuously update their feature vectors, capturing the complex interactions between atoms. During training, the GNN is supervised using datasets generated by quantum chemical computational software (such as Gaussian), enabling the model to accurately predict the formation and decomposition of reaction intermediates. At the mesoscopic level, a fluid-solid coupled cellular automaton is developed. The reactor space is discretized into a regular grid, with each cell containing two state variables: fluid and solid. Fluid state variables include temperature, pressure, and concentration, while solid state variables include the distribution of catalyst active sites and surface morphology. The cellular state update rules are based on the Navier-Stokes equations and chemical reaction kinetics, and are numerically solved using the finite difference method. To improve computational efficiency, GPU parallel computing technology is used, increasing the mesoscopic simulation speed by 10 times. At the macro level, a state-space model based on physical laws is established. The system's key state variables, including the temperature, pressure, reactant and product concentrations of each reactor, are determined. Fifteen state variables are selected to form the state vector x(t). The control input vector u(t) includes eight variables, such as feed flow rate, heating power, and catalyst addition amount. System matrix A and input matrix B are determined through system identification methods, and the model parameters are updated online using recursive least squares (RLS). To ensure model accuracy, actual production data is collected every 100ms and compared with the model's predicted values. Model parameter updates are triggered when the error exceeds 5%. A differentiable physics engine is introduced, integrating microscopic, mesoscopic, and macroscopic models. Implemented based on the JAX framework, the differentiable physics engine allows for end-to-end differential calculations of the entire multiscale model. Automatic differentiation techniques are used to calculate the gradients of the model output with respect to the input parameters, providing support for subsequent optimization control.
[0024] Causal reinforcement learning decision optimization steps: A causal intervention reinforcement learning framework was designed to identify causal relationships in the HF production process through causal structure learning. A causal graph model was constructed, and historical data on the HF production process, including process parameters, equipment status, and product quality, totaling 100,000 samples, was collected. Causal structure learning was performed using the PC algorithm (Peter-Clark algorithm) to identify causal relationships between variables. After 1,000 iterations, a causal graph with 20 nodes and 35 edges was constructed, clarifying the causal dependencies between parameters. A do-calculus algorithm was used to calculate the causal effects of different control variables on product quality. For each control variable, different intervention values were set, and the expected changes in product quality indicators were calculated using the causal graph model. For example, when studying the causal effect of reaction temperature on HF yield, the temperature intervention values were set at 500°C, 550°C, and 600°C. The corresponding expected HF yields were 85%, 92%, and 88%, respectively, calculated using do-calculus. The optimal reaction temperature was determined to be 550°C. Counterfactual reasoning was used during reinforcement learning training to evaluate the robustness of the strategy. For each executed action, a causal graph model simulates the system response under other possible actions and calculates a counterfactual reward. The reward function is designed as the difference between the causal effect of the current action and the causal effect of the comparison action, using the formula R=E[Y|do(A=a)]━E[Y|do(A=a′)], where Y represents the product quality indicator, A represents the control action, a represents the current action, and a′ represents the comparison action. During training, an experience replay buffer is used to store the agent's state-action-reward sequence, with a buffer size of 100,000 records. A target network technique is used to stabilize the training process, with the target network updated every 1,000 steps. To improve training efficiency, a distributed reinforcement learning architecture is used. Ten parallel agents are deployed for training in different simulation environments, each initialized with a different random seed. The learning parameters of each agent are synchronized via a parameter server, with parameter updates performed every 500 steps. After 100,000 training iterations, the agents converge to the optimal policy.
[0025] Steps of space-time fractional sliding mode control: Introduce fractional-order calculus theory into model predictive control and design space-time fractional-order sliding mode controller. Define fractional-order sliding surface , is the tracking error, and They are Order and order fractional differential operators, is a positive definite matrix. The Grünwald-Letnikov definition is used to numerically compute fractional differentials, and sliding mode control is used to ensure system convergence within a finite time. This controller improves system response speed by 30% and reduces overshoot to less than 1%.
[0026] Memory-enhanced meta-learning adaptation step: A memory-enhanced meta-learning framework is developed to enable the control system to rapidly adapt to new operating conditions. A differentiable neural computer (DNC) is designed as an external memory module to store historical operating conditions and optimal control policies. The DNC consists of a controller and a memory matrix. The controller is implemented using an LSTM network. The memory matrix is sized at 1000×100 and can store 1000 memory vectors of length 100. Reading and writing to the memory are performed using an attention mechanism, with three read heads and one write head. During training, historical operating conditions and the corresponding optimal control policies are stored in the memory matrix. Each memory entry contains a condition description vector consisting of 15 process parameters and a control policy vector consisting of eight control variables. When encountering a new operating condition, the attention mechanism calculates the similarity between the current condition and the memory entry, selecting the most relevant memory entry to initialize the reinforcement learning agent's policy network. A meta-gradient optimization algorithm is used for rapid policy adjustment. During the meta-training phase, different tasks are sampled from the historical operating conditions, each consisting of a set of training and test examples. For each task, the agent's policy network is first updated using training samples, and then the meta-gradients are calculated using test samples to update the meta-learner's parameters. The meta-learner's learning rate is set to 0.001, and the Adam optimizer is used for parameter updates. When a new operating condition is detected, the system retrieves relevant information from memory and initializes the agent's policy network. Policy adjustment is completed by fine-tuning with 5 gradient steps using a small amount of sample data under the current operating condition. To evaluate the performance of memory-enhanced meta-learning adaptation, a series of comparative experiments were designed. In the experiments, traditional reinforcement learning methods, meta-learning methods, and memory-enhanced meta-learning methods were used to control the HF production process.
[0027] The present invention further comprises the following steps: Fluid topology optimization reactor design steps: Topology optimization theory is applied to the design of the HF reactor, and a fluid dynamics-based topology optimization model is constructed. The reactor interior is discretized into a three-dimensional grid. The density of each grid cell is used as a design variable, with the density value varying continuously between 0 (representing the fluid region) and 1 (representing the solid region). The SIMP (Solid Isotropic Material with Penalization) method is used to address material interpolation, with a penalty factor set to 3 to ensure clear boundaries in the optimization results. The objective function is designed as a weighted sum of maximizing reaction efficiency and minimizing pressure loss. Reaction efficiency is evaluated by calculating the ratio of the HF mole fraction at the outlet to the theoretical maximum mole fraction, while pressure loss is evaluated by calculating the pressure difference between the reactor inlet and outlet. The weighting coefficients are adjusted according to production requirements. In this implementation, the weight for reaction efficiency is set to 0.7, and the weight for pressure loss is set to 0.3. Constraints include the total reactor volume, structural continuity, and physical constraints on fluid flow. The total volume constraint ensured that the optimized reactor volume did not exceed 110% of the original design volume. Structural continuity constraints were implemented using filtering techniques to ensure smooth density changes between adjacent grid cells. Physical constraints for fluid flow included the conservation of mass, momentum, and energy equations. The moving asymptote method (MMA) was used to solve the topology optimization problem. MMA is an efficient nonlinear optimization algorithm suitable for engineering optimization problems with a large number of design variables. During the optimization process, each iteration required solving the fluid dynamics equations to evaluate the objective function and constraints. To improve computational efficiency, the fluid dynamics equations were discretized using the finite volume method and solved using the open-source computational fluid dynamics software OpenFOAM. After 200 optimization iterations, the optimal reactor internal structure was obtained. The optimized reactor adopts a honeycomb structure with multiple interconnected channels, which effectively improves the mixing of reactants and the contact area. While maintaining the same production capacity, the newly designed reactor improved the uniformity of the reactant residence time distribution by 40% and reduced the pressure drop by 25%.
[0028] The present invention further comprises the following steps: Fractional Brownian motion noise control steps: A fractional Brownian motion (FBM) noise model was developed to address the random noise in the HF production process. Historical data on key parameters in the HF production process, including temperature, pressure, and concentration, was collected, totaling 1 million samples. The Hurst exponent of the noise was estimated using the R / S analysis method. Results showed that the noise in the HF production process exhibited long-range correlation, with a Hurst exponent of approximately 0.7. Based on the estimated Hurst exponent, a fractional Brownian motion noise model was developed. A Monte Carlo method was used to generate noise samples that conformed to the characteristics of fractional Brownian motion for subsequent controller design and testing. To verify the accuracy of the noise model, the generated noise samples were compared with actual production data, and statistical tests showed that the two had similar statistical characteristics. A fractional-order Kalman filter was designed for state estimation. The fractional-order Kalman filter is an extension of the traditional Kalman filter in the fractional domain and is better able to handle noise with long memory characteristics. Based on the fractional Brownian motion noise model, the recursive formula for the fractional-order Kalman filter, including the prediction equation and the update equation, was derived. In practical applications, the state space augmentation method is used to implement the fractional-order Kalman filter, converting the fractional-order differential operator into a state space model. The fractional-order LQG controller is used to handle noise with long memory characteristics. The fractional-order LQG controller combines fractional-order control theory and linear quadratic Gaussian control theory and can achieve optimal control in the presence of noise. The controller gain matrix is determined by solving the fractional-order Riccati equation. In this implementation, an iterative algorithm is used to solve the fractional-order Riccati equation, and the iteration accuracy is set to 10 -6 Experimental verification on an HF synthesis reactor demonstrated that the fractional Brownian motion noise control method improved the system's anti-interference capability by 60% and reduced product quality fluctuations by 50%. Compared with traditional control methods, the state estimation error of the fractional-order Kalman filter was reduced by 30%. The fractional-order LQG controller achieved superior control performance and was able to better suppress the impact of long-range noise correlation on the system.
[0029] The present invention further comprises the following steps: Adversarial generative robust training steps: A generative adversarial network (GAN) is constructed to enhance the robustness of the control system. The generator uses a deep convolutional neural network (DCGAN) architecture, consisting of five convolutional layers and five deconvolutional layers, to generate various abnormal operating conditions and disturbance scenarios. The discriminator also uses the DCGAN architecture to distinguish between real and generated data. The hidden layers of both the generator and discriminator use the LeakyReLU activation function, and the output layer uses the Sigmoid activation function. During training, the GAN is optimized using a min-max game framework. The generator's goal is to generate samples that can deceive the discriminator, while the discriminator's goal is to accurately distinguish between real and generated data. The training process uses an alternating optimization strategy, first updating the discriminator parameters and then the generator parameters. The training batch size is set to 64, the learning rate is set to 0.0002, and the Adam optimizer is used for parameter updates. To improve the quality of generated samples, a gradient penalty mechanism is introduced. A gradient penalty term is added to the discriminator's loss function to constrain the discriminator's gradient norm to ensure that it meets the Lipschitz continuity condition. The gradient penalty coefficient is set to 10, and experiments have verified that this value effectively improves the training stability of the GAN and the quality of the generated samples. The generated adversarial examples are used as additional training data during reinforcement learning training. In each training cycle, half of the training batch is sampled from real data and half from generated data. To ensure the diversity of the generated data, the GAN parameters are regularly updated, and the GAN is retrained every 1000 training steps. To evaluate the effectiveness of robust training using adversarial generation, experiments were conducted on a HF production simulator. The experiments set up three scenarios: a controller trained only with real data, a controller trained with real data and random noise, and a controller trained using robust adversarial generation.
[0030] The present invention further comprises the following steps: Deployment steps for a metamaterial sensor network: Design a metamaterial-based HF concentration sensor network, leveraging the metamaterial's unusual response to electromagnetic waves of specific frequencies to achieve highly sensitive detection. The sensor utilizes a fractal structure, providing a large surface area within a limited space and enhancing interaction with HF molecules. The metamaterial consists of periodically arranged metal-dielectric units, with gold as the metal and silicon dioxide as the dielectric. The unit size is designed to be one-quarter the wavelength of the operating frequency. Using electromagnetic simulation software (CST Microwave Studio), the metamaterial's structural parameters are optimized to generate a strong electromagnetic response at the characteristic absorption frequency of HF molecules (approximately 10.7 GHz). The optimized metamaterial's absorption coefficient for HF molecules is five times higher than that of conventional materials, significantly improving detection sensitivity. Sensor fabrication utilizes electron beam lithography and thin-film deposition techniques to create the metamaterial structure. A silicon dioxide film is deposited on a silicon substrate, a fractal pattern is defined using electron beam lithography, and a gold layer is deposited using physical vapor deposition to form the metamaterial structure. A molecular recognition layer with specific adsorption for HF is modified on the metamaterial surface to enhance sensor selectivity. Topology optimization algorithms are used to determine the optimal sensor deployment location. A mathematical model of HF leakage and diffusion was established and solved using the finite element method to determine the HF concentration distribution at different locations. Based on this concentration distribution, a topology optimization objective function was designed to maximize detection coverage and minimize the number of sensors. A genetic algorithm was used to solve this optimization problem, and after 500 generations of evolution, the optimal sensor deployment solution was obtained. In actual deployment, sensors were installed in key locations such as reactors, pipelines, and storage tanks, forming a sensor network covering the entire production area. The sensors transmit data to the monitoring center via wireless communication using the ZigBee protocol, with a transmission range of up to 100 meters and a transmission rate of 250 kbps.
[0031] The present invention further comprises the following steps: Steps of integrating spatiotemporal attention mechanism: Construct spatiotemporal attention mechanism to process multi-source heterogeneous data. The temporal attention module uses gated recurrent unit (GRU) to capture the temporal dependency of parameters, and the spatial attention module uses graph neural network (GNN) to extract the spatial correlation between parameters. The spatiotemporal attention weight calculation formula is: , is the hidden state at the previous moment, For the The input of the sensor, The system automatically focuses on key parameters through the attention mechanism, improving anomaly detection accuracy to 98.7% and shortening response time to 2ms.
[0032] The present invention further comprises the following steps: Lie group and Lie algebra motion planning: Lie group and Lie algebra theory is applied to the motion planning of HF production equipment. The equipment state is represented as elements of the special Euclidean group SE(3) and parameterized using Lie algebra. A Lie group-based model predictive controller is designed. While maintaining the smoothness of the equipment motion, the adjustment time is reduced by 50% and the positioning accuracy is improved to 0.01mm.
[0033] The present invention further comprises the following steps: Multi-agent game collaborative control steps: A multi-agent game model is constructed to achieve collaborative control of the HF production process. Each production unit in the HF production process, including the reactor, separator, and compressor, is considered an independent agent, resulting in a total of eight agents. Each agent has its own state space, action space, and objective function, capable of making autonomous decisions and interacting with other agents. An incomplete information dynamic game model is established. Each agent only observes its own local state and partial state information of other agents, resulting in incomplete information. The agent's decision-making process is a dynamic game. At each time step, each agent selects the optimal strategy based on its own beliefs and the action history of other agents. A distributed reinforcement learning algorithm is designed to solve the game equilibrium. The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is used. Each agent has its own policy network and critic network. The policy network generates actions based on the agent's local state, and the critic network evaluates the value of joint actions. During training, each agent uses its own experience replay buffer to store state-action-reward sequences, and a target network technique is used to stabilize the training process. To promote cooperation between agents, a reward function based on a trust mechanism is designed. Each agent's reward consists of an individual reward based on its own performance metrics, and a social reward based on the performance metrics of the entire system. The weight of the social reward is dynamically adjusted based on the trust between agents, which is calculated from their historical cooperative behavior. A multi-agent game-based collaborative control approach is applied in a multi-reactor parallel production scenario. Each reactor agent autonomously adjusts parameters such as reaction temperature, pressure, and feed flow based on its own status and information about the status of other reactors. In the face of abnormal situations such as equipment failures and raw material fluctuations, the system can quickly reallocate tasks to maintain stable production operations, demonstrating excellent robustness and adaptability.
[0034] The present invention further comprises the following steps: Memory Causal Graph Network Steps: Develop a memory-enhanced causal graph network (MCGN) for risk prediction. Causal structure learning is performed using causal discovery algorithms (such as the PC algorithm and the GES algorithm). After multiple iterations and validation, a causal graph consisting of 30 nodes and 50 edges is constructed. Nodes in the graph represent various parameters and variables, edges represent causal relationships between variables, edge directions indicate causal direction, and edge weights indicate causal strength. A memory module is developed to store historical risk patterns. This memory module stores information in the form of key-value pairs, with the key being the feature vector of the risk pattern and the value being the corresponding risk description and response strategy. Feature vectors are obtained by dimensionality reduction of historical anomaly data. Principal component analysis (PCA) is used to reduce high-dimensional anomaly data to 10 dimensions. The memory module is set to a capacity of 1000 records. When the capacity limit is reached, a least recently used (LRU) strategy is used to delete the least recently used records. The Memory Causal Graph Network (MCGN) architecture is designed. The MCGN consists of a causal graph encoder, a memory retrieval module, and a risk prediction module. The causal graph encoder uses a graph neural network (GNN) to encode the causal graph and extract causal relationship features between variables. The memory retrieval module retrieves relevant historical risk patterns from the memory module based on the current anomaly features. The risk prediction module combines the causal graph encoding and memory retrieval results to predict the development trend and impact of potential risks. During the risk prediction process, causal intervention analysis is used to assess the impact of different variables on the risk. For each potential risk variable, a do-calculus is used to calculate the risk change under different intervention values and identify the key influencing factors. Using the historical risk patterns stored in the memory module, analogical reasoning is used to predict the possible development path and consequences of the current risk.
[0035] The present invention further comprises the following steps: Fractal dimension fault diagnosis steps: Apply fractal geometry theory to HF production equipment fault diagnosis. Calculate the fractal dimension of the equipment's operating parameter time series and identify early-stage faults through fractal feature extraction. Design a fractal adaptive threshold algorithm to dynamically adjust the diagnostic threshold based on production conditions.
[0036] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control, characterized in that: The following steps are involved: Multi-scale neural symbolic dynamics modeling steps: A multi-scale dynamics model is constructed by integrating neural networks and symbolic reasoning. At the microscopic level, a graph neural network is used to learn the molecular structure and reaction mechanism of the HF synthesis reaction, and atomic interactions are captured through a message passing mechanism. At the mesoscopic level, a fluid-solid coupled cellular automaton is developed to simulate mass transfer and chemical reactions on the catalyst surface. At the macroscopic level, a state space model is established, and the model's generalization capability is analyzed through neural tangent kernel theory. A differentiable physics engine is introduced for full-scale differentiable modeling. Causal reinforcement learning decision optimization: Design a causal intervention reinforcement learning framework to identify causal relationships in the HF production process through causal structure learning; Construct a causal graph model and use do-calculus to calculate causal effects; Reinforcement learning training uses counterfactual reasoning to evaluate the robustness of the strategy. The reward function is designed as R = E[Y | do(A = a)] - E[Y | do(A = a′)], where Y is the product quality indicator, A is the control action, a is the current action, and a′ is the comparison action. Steps of space-time fractional-order sliding mode control: Introduce fractional-order calculus theory, design space-time fractional-order sliding mode controller, and define fractional-order sliding mode surface , is the tracking error, and They are Order and order fractional differential operators, is a positive definite matrix; the Grünwald-Letnikov definition is used to calculate the fractional differential value, and the system is converged by sliding mode control; Memory-enhanced meta-learning adaptation steps: Develop a memory-enhanced meta-learning framework and design a differentiable neural computer to store historical operating conditions and optimal control strategies; when encountering new operating conditions, retrieve relevant memories through the attention mechanism, initialize the reinforcement learning agent policy network, and use the meta-gradient optimization algorithm to adjust the strategy.
2. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: The design steps of the fluid topology optimization reactor are as follows: the HF reactor is designed using topology optimization theory, a topology optimization model is constructed based on fluid dynamics, the internal structure of the reactor is parameterized using the variable density method, and the moving asymptote method is used to solve the optimal structure.
3. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: The steps of fractional Brownian motion noise control are as follows: establish a fractional Brownian motion noise model, estimate the long-range correlation of noise through Hurst exponent, design a fractional-order Kalman filter to estimate the state, and use a fractional-order LQG controller to process long-memory characteristic noise.
4. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Adversarial generative robust training: Build a generative adversarial network to enhance the robustness of control systems. The generator simulates abnormal operating conditions and interference, and the discriminator distinguishes between real data and generated data. Reinforcement learning training uses adversarial samples as additional training data, and the agent learns coping strategies in extreme situations.
5. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Steps for deploying a metamaterial sensor network: Design an HF concentration sensor network and use the abnormal response characteristics of metamaterials to detect sensitivity; The sensor adopts fractal structure design and the optimal deployment position is determined through topology optimization algorithm.
6. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Steps for integrating the spatiotemporal attention mechanism: Build a spatiotemporal attention mechanism to process multi-source heterogeneous data. The temporal attention module uses a gated recurrent unit to capture the temporal dependencies of parameters, and the spatial attention module uses a graph neural network to extract the spatial correlation of parameters. The calculation formula for spatiotemporal attention weight is: , is the hidden state at the previous moment, For the The input of the sensor, For the number of sensors, the key parameters are automatically focused through the attention mechanism.
7. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Lie group and Lie algebra motion planning steps: Use Lie group and Lie algebra theory to plan the motion of HF production equipment. The equipment state is represented as a special Euclidean group SE element, and the state is parameterized through Lie algebra; design a model predictive controller to maintain the smoothness of equipment motion.
8. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Agent game collaborative control steps: Build a multi-agent game model to collaboratively control the HF production process, analyze the agent strategy interaction through incomplete information dynamic game theory, and design a distributed reinforcement learning algorithm to achieve global optimization.
9. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Memory Causal Graph Network Steps: Develop a memory-enhanced causal graph network to predict risks, capture parameter causal relationships through graph neural networks, and use memory modules to store historical risk patterns; When an anomaly is detected, potential risks are predicted through causal intervention analysis and an early warning is issued.
10. The millisecond-level control method for hydrogen fluoride production based on reinforcement learning and model predictive control according to claim 1, characterized in that: Also includes: Fractal dimension fault diagnosis steps: Use fractal geometry theory to diagnose HF production equipment faults, calculate the fractal dimension of the equipment operating parameter time series, and identify early faults through fractal feature extraction; design a fractal adaptive threshold algorithm and dynamically adjust the diagnostic threshold according to production conditions.
Citation Information
Patent Citations
Active power filter based on fractional order sliding mode control and filtering method
CN116780542A
Private domain live broadcast peak hot spot prediction and content scheduling method based on deep learning
CN119450099A
Depth deterministic strategy gradient driven multi-constraint guidance law design method
CN119644742A
Track control method based on high-order finite time observer
CN120010273A
Intelligent control system and method for injection molding processing of gas-liquid photoelectric slip ring
CN120422434A
Cited By
Modeling method of MOS device
CN120974792A
A method of modeling a mos device
CN120974792B
Method and device for predicting bleeding risk in spine surgery based on machine learning
CN121460178A
Automatic operation monitoring method for high-altitude platform process system
CN122332883A