A dynamic optimization method for pickling process based on deep reinforcement learning
By combining deep reinforcement learning and mechanistic models, the problem of sensor failure during pickling was solved, achieving stable and optimized control of the pickling process and improving the robustness and economy of production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN HUALING LIANYUAN STEEL SPECIAL NEW MATERIAL CO LTD
- Filing Date
- 2026-03-11
- Publication Date
- 2026-07-21
AI Technical Summary
The existing pickling process control relies heavily on high-precision external sensors, which are susceptible to factors such as acid mist corrosion, media contamination, or mechanical vibration, leading to data distortion or failure, and affecting production stability and economy.
A deep reinforcement learning-based approach is adopted to build a reinforcement learning framework. By combining a mechanistic model with a data-driven residual compensation simulation model, an agent is trained using the TD3 algorithm to achieve dynamic optimization control that does not rely on easily failing sensors. A hierarchical deployment architecture is used for online control.
It improves the robustness and adaptability of the pickling process, enhances the stability and reliability of the system, realizes comprehensive optimized control of the pickling process, and reduces the dependence on high-precision sensors.
Smart Images

Figure CN122436052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pickling continuous rolling production technology, specifically to a dynamic optimization method for the pickling process based on deep reinforcement learning. Background Technology
[0002] In the pickling and rolling production process, pickling is a crucial preliminary step aimed at removing iron oxide scale from the surface of the strip steel. The core challenge lies in minimizing acid and energy consumption while ensuring pickling quality. Compared to traditional PID control that relies on fixed parameter settings and optimization control based on static mathematical models, data-driven intelligent optimization methods are better able to adapt to the strong nonlinearity and large time delay of the pickling process, exhibiting superior adaptability and robustness.
[0003] As a typical example of intelligent optimization methods, reinforcement learning, through continuous interaction and learning between an agent and its environment, does not rely on precise mechanistic models and is applicable to pickling industrial processes. Currently, pickling process control largely depends on high-precision online sensors and comprehensive environmental perception data. However, in actual industrial settings, sensor measurement errors can become large or even momentarily malfunction due to factors such as acid mist corrosion, media contamination, and mechanical vibration. This leads to a sharp decline in the performance of the optimization system, severely impacting production stability and economic efficiency.
[0004] Therefore, researching autonomous optimization methods for pickling processes under conditions where external sensor information is limited or unreliable, and exploring new intelligent control approaches that do not heavily rely on high-cost, high-precision sensor data, is of great significance for improving the robust operation capability of pickling production lines in complex industrial environments, and is also a key link in promoting the practical implementation of intelligent manufacturing. Summary of the Invention
[0005] (a) Technical problems to be solved The technical problem to be solved by this invention is to propose a dynamic optimization method for pickling process based on deep reinforcement learning. The existing pickling process control relies heavily on high-precision external sensors. However, in complex industrial environments, these sensors are prone to data distortion or failure due to uncertain factors such as acid mist corrosion, media contamination or mechanical vibration, which leads to a decline in control performance or even production interruption.
[0006] (II) Technical Solution To solve the above-mentioned technical problems, the technical solution provided by the present invention is: a dynamic optimization method for acid washing process based on deep reinforcement learning, characterized by comprising the following steps: S1. Construct a reinforcement learning framework model, define the state space, action space, and reward function. The deep reinforcement learning process in the framework model can be described as a Markov decision process, where the agent receives and executes actions. The subsequent environmental state and reward value At this time, the policy network Mapping outputs the action at the current moment. ; Next, the agent performs the action. And according to the reward function Calculate the corresponding reward value Meanwhile, the state transition function Calculate the next state ; Within a set period, the agent updates the policy network according to an optimization algorithm. By gradually understanding the long-term effects of taking different actions under different conditions, we can eventually arrive at the optimal strategy that maximizes cumulative benefits. ; The reward function is key to guiding the agent's learning; the total reward value is defined using a multi-objective weighted sum. ,in, , , , Weighting coefficients S2. Build a mechanism model-data-driven residual compensation simulation model to make up for the shortcomings of the pure mechanism model, and use deep reinforcement learning algorithm to train the agent offline in the simulation model to obtain the optimal policy model; ensure that the agent not only converges in the simulation environment, but also improves the robustness and reliability of the system.
[0007] S3. Deploy the trained optimal strategy model online in the pickling production line control system to achieve dynamic optimization control of the pickling process.
[0008] (III) Beneficial Effects The advantages of this invention compared to the prior art are: (1) By designing a state space that does not rely on easily failed external sensors, and combining a hybrid simulation environment of mechanism model and data-driven residual compensation, this invention effectively solves the process optimization control problem when sensor data is unreliable in complex industrial environments, and improves the robustness and adaptability of the system. (2) The agent is trained using the TD3 algorithm. Through mechanisms such as dual Critic network, delayed update and target policy smoothing, the problem of overestimation of value function is effectively avoided, and the stability and convergence of policy learning are improved. (3) A hierarchical deployment architecture is adopted, and the trained intelligent agent model is deployed on the L2 edge computing layer. Combined with the safety early warning module, the action limit verification is performed to ensure the safety and reliability of online control. (4) This invention achieves comprehensive optimization control of the pickling process by designing a multi-objective reward function, which comprehensively considers pickling quality, cost consumption, production efficiency and operational stability. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the process dynamic optimization based on deep reinforcement learning, as exemplified by the present invention. Figure 2 This is a schematic diagram of the mechanism model and data-driven residual compensation simulation model for an example of the present invention; Figure 3 This is the logic diagram of the TD3 algorithm used in the example of this invention; Figure 4 This is a diagram of the hierarchical deployment architecture model for reinforcement learning in an example of the present invention. Detailed Implementation
[0010] The invention will now be described in further detail with reference to specific embodiments, but this should not be construed as limiting the scope of the subject matter of the invention to the following embodiments.
[0011] A dynamic optimization method for the acid washing process based on deep reinforcement learning includes the following steps: S1. Build a reinforcement learning framework model; the deep reinforcement learning process can be described as a Markov decision process, such as... Figure 1 As shown. The intelligent agent receives and executes actions. The subsequent environmental state and reward value At this time, the policy network Mapping outputs the action at the current moment. Next, the agent performs the action. And according to the reward function Calculate the corresponding reward value Meanwhile, the state transition function Calculate the next state Within a set period, the agent updates the policy network according to an optimization algorithm. By gradually understanding the long-term effects of taking different actions under different conditions, we can eventually arrive at the optimal strategy that maximizes cumulative benefits. .
[0012] S2. Offline training of the agent; building a mechanism model-data-driven residual compensation simulation model to make up for the shortcomings of the pure mechanism model; at the same time, the TD3 algorithm is adopted to ensure that the agent not only converges in the simulation environment, but also improves the robustness and reliability of the system.
[0013] S3. Online Deployment and Application: A hierarchical deployment architecture is adopted, in which the trained intelligent agent model is deployed on the L2 control system of the actual pickling production line. The actions of the safety early warning module to correct and limit are sent to the underlying actuators to achieve dynamic optimization control of the pickling process.
[0014] Further, step S1 includes the following steps: S11, State Space: This patent's state space deliberately avoids direct reliance on easily failing external sensor data, focusing instead on basic automation data from the L1 control layer, state data from the L2 edge computing layer, and PDI information from the L3 production planning layer. Specifically, it includes: ;in, For steel type, For hot-rolled thickness, For cold-rolled thickness, For hot roll width, For cold roll width, For hot roll length, For strip speed; For acid temperature, For acid concentration, For acid flow rate, Cumulative steel throughput.
[0015] S12, Action Space: The action space is defined as the control commands for key actuators in the pickling process, including: changes in the opening degree of the fresh acid supply valve, changes in heating power, and changes in the frequency of the circulating pump. The action vector can be represented as: ;in, This represents the change in valve opening. This represents the change in heating power. This represents the change in the frequency of the circulating pump.
[0016] S13, Reward Function: The reward function is crucial for guiding the agent's learning. This patent defines the total reward value using a multi-objective weighted sum:
[0017] in, , , , These are the weighting coefficients. The definitions of each sub-function are as follows: Quality Awards:
[0018] in, It is the reward constant. It is a cleaning exponential function. These are quality thresholds related to steel grade and thickness.
[0019] Cost incentive:
[0020] in, For acid consumption, Energy consumption.
[0021] Efficiency Rewards:
[0022] Stability Bonus:
[0023] Further, step S2 includes the following steps: S21. Based on the physicochemical laws of the pickling process, a mechanism simulation model is constructed as the training environment for the reinforcement learning agent, specifically including: Acid concentration kinetic model:
[0024] in, The reaction rate constant is... To balance the concentration, For the reaction area of the strip steel, For strip speed, For acid tank volume, To replenish acid flow, To replenish acid concentration.
[0025] Acid temperature dynamic model:
[0026] in, For heater power, For circulating flow, For acid density, For the specific heat capacity of acid solution, For the overall heat transfer coefficient, For heat dissipation area, For acid replenishment temperature, The ambient temperature.
[0027] Quality prediction model:
[0028] in, For comprehensive proportional factors, It is the temperature sensitivity coefficient. It is the time that the strip steel stays in the acid bath.
[0029] S22, Data-driven residual compensation: This patent employs a hybrid modeling method in constructing a simulation environment for the pickling process. First, based on collected historical production data, including states, actions, and the corresponding state at the next moment, a mechanistic model predicts the state at the next moment. Then, using the previous state and actions as input, and the difference between the actual state and the mechanistic model's prediction as the learning objective, a residual model is trained based on a deep neural network. Finally, the state at the next moment of the simulation environment is constructed by superimposing the mechanistic model's prediction and the residual model's prediction, thereby compensating for unmodeled dynamics and system errors while ensuring physical correctness. The specific process is as follows: Figure 2 As shown.
[0030] S23. Reinforcement Learning Algorithm Selection and Configuration: The TD3 algorithm proposes three key improvements over DDPG: a dual-critic network, delayed updates, and target policy smoothing. These improvements effectively reduce overestimation bias and enhance the policy's robustness and stability against noise and target value fluctuations. The specific implementation logic is as follows: Figure 3 As shown.
[0031] Policy Network The input is a state vector. The output is an action vector. The system employs a 4-layer network structure: Input layer → Fully connected layer (512 neurons, ReLU) → Fully connected layer (256 neurons, ReLU) → Output layer (tanh activation function).
[0032] Value network Q: Input is a state vector and action vectors The splicing and output of the data are a single value index. The network structure consists of four layers: input layer → fully connected layer (512 neurons, ReLU) → fully connected layer (256 neurons, ReLU) → output layer (linear).
[0033] The key parameters are set as follows: discount factor γ = 0.99, experience replay pool capacity = 1,000,000, soft update parameter τ = 0.005, Actor network learning rate = 0.0001, Critic network learning rate = 0.001, policy update frequency = 2, and the exploration noise standard deviation is initially 0.2 and decreases linearly to 0.05 as the training episode increases to encourage policy utilization in later stages.
[0034] Further, step S3 includes the following steps: S31. An edge computing host computer is deployed in the electrical room as the L2 intelligent decision-making unit, and interconnected with the L3 server and the L1 PLC via Ethernet. The L2 host computer is selected as follows: Intel Core i7-12700TE CPU, NVIDIA T1000 GPU, 64GB DDR4*2 memory, 2TB NVMe M.2 SSD system disk, and dual Intel I210-AT gigabit Ethernet ports.
[0035] S32. Deploy the online service of the reinforcement learning agent on the L2 host computer and load the optimal policy model trained offline. Simultaneously, a safety early warning module is defined, with preset upper and lower safety thresholds for pickling process parameters.
[0036] S33, The hierarchical architecture executes closed-loop control processes at fixed cycles, such as Figure 4 As shown: Unit L2 collects the pickling process status from system L1 in real time via the OPC-UA protocol. It obtains the current and future PDI information of the strip from the L3 system; the L2 unit will construct the state vector. Input to the optimal policy model In the process, the calculated action Regarding the action Perform amplitude limiting and logic verification to obtain safe actions. The L2 unit issues commands via the OPC-UA protocol. The PLC system in unit L1 uses this setpoint as its input value to drive field actuators such as control valves and heaters.
[0037] Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents. In short, if those skilled in the art are inspired by these claims and design similar structural methods and embodiments without departing from the inventive spirit of the present invention, they should all fall within the protection scope of the present invention.
Claims
1. A dynamic optimization method for acid washing process based on deep reinforcement learning, characterized in that, Includes the following steps: S1. Construct a reinforcement learning framework model, define the state space, action space, and reward function. The deep reinforcement learning process in the framework model can be described as a Markov decision process, where the agent receives and executes actions. The subsequent environmental state and reward value At this time, the policy network Mapping outputs the action at the current moment. ; Next, the agent performs the action. And according to the reward function Calculate the corresponding reward value Meanwhile, the state transition function Calculate the next state ; Within a set period, the agent updates the policy network according to an optimization algorithm. By gradually understanding the long-term effects of taking different actions under different conditions, we can eventually arrive at the optimal strategy that maximizes cumulative benefits. ; The reward function is key to guiding the agent's learning; the total reward value is defined using a multi-objective weighted sum. ,in, , , , These are the weighting coefficients; S2. Build a mechanism model-data-driven residual compensation simulation model to make up for the shortcomings of the pure mechanism model, and use deep reinforcement learning algorithm to train the agent offline in the simulation model to obtain the optimal policy model; ensure that the agent not only converges in the simulation environment, but also improves the robustness and reliability of the system. S3. Deploy the trained optimal strategy model online in the pickling production line control system to achieve dynamic optimization control of the pickling process.
2. The dynamic optimization method for acid washing process based on deep reinforcement learning according to claim 1, characterized in that, The state space in step S1 focuses on the basic automation data of the L1 control layer, the state data of the L2 edge computing layer, and the PDI information of the L3 production planning layer. Specifically, it includes: ; in, For steel type, For hot-rolled thickness, For cold-rolled thickness, For hot roll width, For cold roll width, For hot roll length, For strip speed; For acid temperature, For acid concentration, For acid flow rate, Cumulative steel throughput.
3. The dynamic optimization method for acid washing process based on deep reinforcement learning according to claim 1, characterized in that, The action space mentioned in step S1 is defined as the control commands for the key actuators in the pickling process, including: the change in the opening degree of the fresh acid supply valve, the change in heating power, and the change in the frequency of the circulating pump. The action vector can be expressed as: ;in, This represents the change in valve opening. This represents the change in heating power. This represents the change in the frequency of the circulating pump.
4. The dynamic optimization method for the acid washing process based on deep reinforcement learning according to claim 1, characterized in that, The reward function mentioned in step S1 has the following definitions for each component: Quality Awards: , in, It is the reward constant. It is a cleaning exponential function. Quality thresholds related to steel grade and thickness; Cost incentive: , in, For acid consumption, Energy consumption; Efficiency Rewards: , Stability Bonus: 。 5. The dynamic optimization method for the acid washing process based on deep reinforcement learning according to claim 1, characterized in that, The mechanistic model constructed in step S2 includes: Acid concentration kinetic model: , in, The reaction rate constant is... To balance the concentration, For the reaction area of the strip steel, For strip speed, For acid tank volume, To replenish acid flow, To replenish acid concentration; Acid temperature dynamic model: , in, For heater power, For circulating flow, For acid density, For the specific heat capacity of acid solution, For the overall heat transfer coefficient, For heat dissipation area, For acid replenishment temperature, The ambient temperature; Quality prediction model: , in, For comprehensive proportional factors, It is the temperature sensitivity coefficient. It is the time that the strip steel stays in the acid bath.
6. The dynamic optimization method for acid washing process based on deep reinforcement learning according to claim 1, characterized in that, The data-driven residual compensation in step S2 includes a hybrid modeling method in the construction of the simulation environment for the pickling process. First, based on the collected historical production data, including the state, actions, and the corresponding state at the next moment, the state at the next moment is predicted through a mechanistic model. Then, using the state and actions at the previous moment as input, the residual model is trained based on a deep neural network with the difference between the actual state and the predicted value of the mechanistic model as the learning target. Finally, the state at the next moment of the simulation environment is composed of the predicted value of the mechanistic model and the predicted value of the residual model, thereby compensating for unmodeled dynamic and system errors while ensuring physical correctness.
7. The dynamic optimization method for the pickling process based on deep reinforcement learning according to claim 1, characterized in that, In step S2, the deep reinforcement learning algorithm used is the TD3 algorithm, which employs a dual Critic network, delayed policy update, and target policy smoothing mechanism, wherein: (1) Policy Network The input is a state vector. The output is an action vector. ; The network uses a 4-layer structure: Input layer → Fully connected layer (512 neurons, ReLU) → Fully connected layer (256 neurons, ReLU) → Output layer (tanh activation function). (2) Value network Q: The input is the state vector and action vectors The splicing and output are a single value indicator; A four-layer network structure is adopted: input layer → fully connected layer (512 neurons, ReLU) → fully connected layer (256 neurons, ReLU) → output layer (linear).
8. The dynamic optimization method for the acid washing process based on deep reinforcement learning according to claim 7, characterized in that, In the TD3 algorithm, the parameters of the policy network and the value network are set as follows: discount factor γ = 0.99, experience replay pool capacity = 1,000,000, soft update parameter τ = 0.005, Actor network learning rate = 0.0001, Critic network learning rate = 0.001, policy update frequency = 2; the exploration noise standard deviation is initially 0.2 and decreases linearly to 0.05 as the training episode increases to encourage policy utilization in later stages.
9. The dynamic optimization method for acid washing process based on deep reinforcement learning according to claim 1, characterized in that, Step S3 specifically includes: S31. Deploy an edge computing host computer in the electrical room as an L2 intelligent decision-making unit, and interconnect it with the L3 server and the L1 PLC via Ethernet; S32. Deploy the online service of the reinforcement learning agent on the L2 host computer and load the optimal policy model trained offline. Simultaneously, a safety early warning module is defined, with preset upper and lower safety thresholds for pickling process parameters; S33. The hierarchical architecture executes closed-loop control procedures at fixed intervals: The L2 unit collects the pickling process status from the L1 system in real time via the OPC-UA protocol. It obtains the current and future PDI information of the strip from the L3 system; the L2 unit will construct the state vector. Input to the optimal policy model In the process, the calculated action Regarding the action Perform amplitude limiting and logic verification to obtain safe actions. The L2 unit issues commands via the OPC-UA protocol. The PLC system in unit L1 uses this setpoint as a reference to drive field actuators such as control valves and heaters.
10. The dynamic optimization method for acid washing process based on deep reinforcement learning according to claim 9, characterized in that, The L2 host computer selected in the S32 is as follows: Intel Core i7-12700TE CPU, NVIDIA T1000 GPU, 64GB DDR4*2 memory, 2TB NVMe M.2 SSD system disk, and dual Intel I210-AT Gigabit Ethernet ports.