Top-blown furnace spray gun position control method based on safety reinforcement learning

By constructing a state space for multi-source data perception and using secure reinforcement learning, combined with hard constraints on action and state layers, the safety and efficiency issues of top-blown furnace spray gun position control were solved, achieving high-precision and high-safety automatic spray gun control.

CN121994037APending Publication Date: 2026-05-08CENT SOUTH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional methods for controlling the position of the spray guns in top-blown furnaces are difficult to cope with complex and ever-changing actual working conditions, which can easily lead to operational errors or safety accidents. Directly applying reinforcement learning may increase the risks.

Method used

A state space for multi-source data perception is constructed. By combining action output and hard constraints of the state layer, a safety reward function is designed, a constraint reinforcement learning agent is trained, and a two-layer safety check and offline update are implemented to ensure that the spray gun position is controlled within the safe operating boundary.

Benefits of technology

This achieves improved spray gun positioning accuracy and smelting efficiency, reduced accident risks, and enhanced system intelligence while ensuring safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121994037A_ABST
    Figure CN121994037A_ABST
Patent Text Reader

Abstract

The invention discloses a top-blown furnace spray gun position control method based on safety reinforcement learning. The top-blown furnace spray gun position control method comprises the following steps that S1, a state space based on multi-source data perception is constructed; s2, constructing security constraint conditions based on parameters in the state space, wherein the security constraint conditions comprise action output and state layer hard constraint; s3, designing a reward function containing security constraints; S4, training a constraint reinforcement learning agent based on S1-S3, and adjusting an instruction; s5, acquiring an adjustment instruction, and performing security verification; and S7, the verified safety action is sent to an execution mechanism, the spray gun is driven to complete position adjustment, and interaction data is stored for intelligent agent updating. Compared with the prior art, the top-blowing furnace spray gun position control method based on safety reinforcement learning has the advantage that the top-blowing furnace spray gun position control method based on multi-source data acquisition and state space limitation is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of top-blown furnace spray lance control technology, specifically to a top-blown furnace spray lance position control method based on safety reinforcement learning. Background Technology

[0002] As one of the most widely used pieces of equipment in the metallurgical industry, the position control of the top-blown furnace's spray guns is crucial to production efficiency and safety.

[0003] Traditional control methods rely on fixed rules or simple feedback mechanisms, which are difficult to cope with complex and ever-changing actual working conditions and can easily lead to operational errors or safety accidents.

[0004] In recent years, reinforcement learning, as an intelligent decision-making method, has begun to be applied in automated control systems.

[0005] However, directly applying reinforcement learning may increase risks due to violations of safety constraints during the exploration process.

[0006] Therefore, how to achieve efficient automatic control of the spray gun position while ensuring safety has become an urgent problem to be solved. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the above-mentioned technical defects and provide a top-blown furnace spray gun position control method based on security reinforcement learning, which is based on multi-source data acquisition and state space constraints.

[0008] To solve the above-mentioned technical problems, the technical solution provided by the present invention is: a method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning, comprising the following steps: S1: Construct a state space based on multi-source data perception; S2: Construct safety constraints based on parameters within the state space, including action outputs and hard constraints at the state layer; S3: Design a reward function that includes safety constraints: Positive rewards will be given when process targets are met; Negative penalties are imposed when safety constraints are violated or when the safe operating area is deviated from. S4: Training a constraint-based reinforcement learning agent based on S1-S3; S5: During the operation phase, real-time data on the current working condition is collected and input into the state space. The trained agent outputs the candidate spray gun adjustment actions to form adjustment commands. S6: Obtain the adjustment command and perform a safety check. If the action exceeds the amplitude limit or rate limit threshold of the action, return to S5 to readjust and output the adjustment command. S7: Send the verified safety action to the actuator, drive the spray gun to complete the position adjustment, and update the agent by storing the data of this interaction.

[0009] Preferably, the action output includes: The range of motion amplitude limits and the threshold of rate limits are combined, with parameters including the current position coordinates of the spray gun in the top-blown furnace, the temperature distribution in the furnace, the height of the molten pool, the furnace structure parameters, and historical safe operation data, which are integrated to form a geometric constraint. The hard constraints of the state layer include the minimum safe distance between the spray gun and the furnace wall, the maximum allowable height above the molten pool surface, the highest tolerable temperature of the target area, and the maximum load of the equipment drive motor.

[0010] Preferably, the action output condition in S2 is the position adjustment amount of the spray gun, including horizontal displacement, vertical displacement and rotation angle.

[0011] Preferably, the state layer hard constraints obtain constraint parameters through deployed sensor data; It includes a ranging unit, a thermal imaging unit, an angle detection unit, and a pressure detection unit.

[0012] Preferably, the security verification in S6 includes: Action layer verification: Check whether the displacement amplitude and rotation angle of the candidate action are within the limit range, and whether the action rate is lower than the threshold. State layer verification: Predict the position of the spray gun after the action is executed, and determine whether the distance from the furnace wall, the height of the molten pool, and the temperature of the target area meet the hard constraints of the state layer; It also includes emergency handling: guiding the agent to adjust its output in case of minor violations, and triggering emergency braking, switching to a preset safe posture and recording the forbidden state in case of serious violations.

[0013] Preferably, in step S7, the agent update is based on offline updates, periodically importing the stored interaction data into the simulation environment. After fine-tuning the agent parameters based on the new data and verifying its safety, it was deployed to the field.

[0014] Preferably, the reward function in S3 is designed such that the positive reward is positively correlated with the degree of achievement of the process target and is related to the positioning accuracy of the spray gun. The higher the positioning accuracy, the greater the positive reward. Negative penalties are based on a combination of risk levels, including collision risk, overheating risk, equipment overload risk, and the degree of violation of deviating from the safe operating area. The higher the risk level or the degree of deviation, the greater the negative penalty imposed.

[0015] The advantages of this invention compared with the prior art are: this invention deeply integrates constraint reinforcement learning technology with the safety control of top-blown furnace spray guns, effectively balancing process performance and operational safety; It overcomes the safety risks that may arise during the exploration process of ordinary reinforcement learning, thereby improving smelting efficiency, spray gun positioning accuracy and system intelligence level while ensuring the safety of equipment and personnel. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings.

[0018] Combined with appendix Figure 1 As shown, a top-blown furnace lance position control method based on safety reinforcement learning is used to ensure that the lance always operates within the safe operating boundary while ensuring the smelting process objectives.

[0019] In S1, a state space based on multi-source data perception is constructed. This state space integrates real-time sensing information from the ranging unit, thermal imaging unit, angle detection unit, and pressure detection unit, and integrates the current position coordinates of the spray gun, the temperature distribution inside the furnace, the height of the molten pool, the furnace structural parameters, and historical safe operation data to comprehensively characterize the current working condition. This design makes the state representation high-dimensional and physically interpretable, laying a reliable foundation for subsequent safety decisions. In S2, based on the above state space, the system establishes dual safety constraints: on the one hand, it defines hard limits on action output, including the amplitude range and rate threshold of the spray gun in horizontal displacement, vertical displacement, and rotation angle; on the other hand, it sets hard constraints at the state layer, such as the minimum safe distance between the spray gun and the furnace wall, the maximum allowable height above the molten pool, the highest tolerable temperature of the target area, and the maximum load capacity of the drive motor. These constraints are directly derived from the physical limits of the equipment and process safety specifications, effectively preventing the agent from exploring dangerous areas.

[0020] S3 incorporates a reward function with safety constraints: when the spray gun positioning accuracy is high and the process target is well achieved, a positive reward is given; however, if the action or state constraints are violated, or the known safe operating area is deviated from, an incremental negative penalty is applied according to the risk level (such as collision, overheating, equipment overload, etc.). This mechanism guides the agent to actively avoid high-risk behaviors while pursuing performance, significantly improving the safety and robustness of the strategy. The above state space, constraints, and reward function are used to train the constraint reinforcement learning agent, enabling it to learn a spray gun adjustment strategy that balances efficiency and safety in the simulation environment. During operation, current working condition data is collected and input into the state space. The trained agent outputs candidate spray gun adjustment actions and performs dual safety checks: action layer check ensures that the displacement amplitude, angle change and movement speed do not exceed preset limits; state layer check predicts the spray gun position after the action is executed and verifies whether it meets all state hard constraints. If a minor violation occurs, the system guides the agent to regenerate a compliant action; if a serious violation occurs (such as an imminent collision with the furnace wall), an emergency brake is immediately triggered, switching the spray gun to a preset safe position and recording the forbidden state to avoid repeating the error. This dual-layer verification mechanism constitutes a double safety guarantee of soft constraints and hard protection, which greatly reduces the risk of on-site accidents.

[0021] In addition, the verified safety actions in this invention are sent to the actuator to drive the spray gun to complete the adjustment, and the interaction data (including status, actions, rewards and constraint violations) is stored. The system periodically imports the accumulated data into the simulation environment for offline intelligent agent fine-tuning, and then deploys it to the field after verifying the safety of the updated strategy, so as to achieve continuous evolution and adaptive improvement of the model.

[0022] This invention achieves high-precision and high-safety autonomous control of spray guns without human intervention by deeply embedding physical safety constraints into a reinforcement learning framework and combining it with a runtime dual-layer verification and closed-loop update mechanism. It also has process optimization capabilities and intrinsic safety characteristics, making it suitable for intelligent equipment control in complex high-temperature metallurgical scenarios.

[0023] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

[0024] This invention constructs a multi-source sensing state space, enabling the system to comprehensively and in real-time grasp the furnace operating conditions. Based on this, a dual hard constraint mechanism of action layer and state layer is introduced, combined with a safety-oriented reward function design, so that the agent is always guided to a safe and feasible operating area during training and operation. By adopting a closed-loop control process of generation-verification-execution-feedback, the system performs two-layer safety verification on candidate actions online and has a graded emergency response capability, which significantly improves the robustness and fault tolerance of the system.

[0025] Furthermore, by combining offline simulation fine-tuning with on-site data-driven updates, continuous optimization and adaptive evolution of the strategy are achieved. The overall solution not only avoids the dependence on human experience and the rigidity of fixed rules in traditional control methods.

[0026] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0027] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning, characterized in that: Includes the following steps: S1: Construct a state space based on multi-source data perception; S2: Construct safety constraints based on parameters within the state space, including action outputs and hard constraints at the state layer; S3: Design a reward function that includes safety constraints: Positive rewards will be given when process targets are met; Negative penalties are imposed when safety constraints are violated or when the safe operating area is deviated from. S4: Training a constraint-based reinforcement learning agent based on S1-S3; S5: During the operation phase, real-time data on the current working condition is collected and input into the state space. The trained agent outputs the candidate spray gun adjustment actions to form adjustment commands. S6: Obtain the adjustment command and perform a safety check. If the action exceeds the amplitude limit or rate limit threshold of the action, return to S5 to readjust and output the adjustment command. S7: Send the verified safety action to the actuator, drive the spray gun to complete the position adjustment, and update the agent by storing the data of this interaction.

2. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 1, characterized in that: The action output includes: The range of motion amplitude limits and the threshold of rate limits are combined, with parameters including the current position coordinates of the spray gun in the top-blown furnace, the temperature distribution in the furnace, the height of the molten pool, the furnace structure parameters, and historical safe operation data, which are integrated to form a geometric constraint. The hard constraints of the state layer include the minimum safe distance between the spray gun and the furnace wall, the maximum allowable height above the molten pool surface, the highest tolerable temperature of the target area, and the maximum load of the equipment drive motor.

3. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 1, characterized in that: The action output condition in S2 is the position adjustment amount of the spray gun, including horizontal displacement, vertical displacement and rotation angle.

4. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 2, characterized in that: The state layer hard constraints obtain constraint parameters through data from deployed sensors; It includes a ranging unit, a thermal imaging unit, an angle detection unit, and a pressure detection unit.

5. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 1, characterized in that: The security verification in S6 includes: Action layer verification: Check whether the displacement amplitude and rotation angle of the candidate action are within the limit range, and whether the action rate is lower than the threshold. State layer verification: Predict the position of the spray gun after the action is executed, and determine whether the distance from the furnace wall, the height of the molten pool, and the temperature of the target area meet the hard constraints of the state layer; It also includes emergency handling: guiding the agent to adjust its output in case of minor violations, and triggering emergency braking, switching to a preset safe posture and recording the forbidden state in case of serious violations.

6. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 1, characterized in that: In S7, the agent update is based on offline updates, periodically importing the stored interaction data into the simulation environment. After fine-tuning the agent parameters based on the new data and verifying its safety, it was deployed to the field.

7. The method for controlling the position of a top-blown furnace spray gun based on safety reinforcement learning according to claim 1, characterized in that: The design of the reward function in S3 is as follows: the positive reward is positively correlated with the degree of achievement of the process target and is related to the positioning accuracy of the spray gun. The higher the positioning accuracy, the greater the positive reward. Negative penalties are based on a combination of risk levels, including collision risk, overheating risk, equipment overload risk, and the degree of violation of deviating from the safe operating area. The higher the risk level or the degree of deviation, the greater the negative penalty imposed.