A closed-loop intelligent control system for coke ovens
Through the deep Q network DQN and DDPG algorithms, the coke oven control system is optimized, and the closed-loop regulation of gas flow and flue suction is realized, which solves the problem of intelligent insufficient in traditional systems, improves production efficiency and energy utilization, and reduces the risk of human error.
Patent Information
- Application Number
- CN202410860638.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-06-28
AI Technical Summary
Traditional coke oven control systems lack intelligent learning and adaptability and cannot achieve global optimization, resulting in low production efficiency and energy utilization, and require manual intervention to deal with abnormal situations, increasing the risk of human error.
The coke oven closed-loop intelligent control system built using the deep Q network DQN and the depth deterministic strategy gradient algorithm DDPG, uses real-time monitoring of the coke oven temperature and gas flow, and combines the reward function optimization control strategy to achieve closed-loop adjustment of gas flow and flue suction, improving the system's adaptability and prediction capabilities.
It improves the stability and efficiency of the coke oven combustion process, reduces energy losses, optimizes production parameters, improves production efficiency and product quality, and reduces accident risks.
Smart Images

Figure CN118599555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to coke oven control, and more particularly to a closed-loop intelligent control system for coke ovens. Background Art
[0002] For coke oven control, traditional automated control systems and rule-based expert systems have obvious limitations. These systems generally do not have intelligent learning and adaptation capabilities, cannot utilize big data for deep learning and optimization, but rely on static adjustment and preset rules. Due to the lack of dynamics and real-time performance, existing systems cannot flexibly cope with complex production environment changes, nor do they have the ability to predict future trends and problems.
[0003] Traditional systems often require manual intervention to adjust parameters or handle abnormal situations, increasing the risk of human errors. Most importantly, these systems cannot achieve global optimization, but only local optimization, resulting in inefficient production efficiency and energy utilization rate. Therefore, existing systems have obvious deficiencies in terms of real-time performance, flexibility, prediction ability, and global optimization. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the above-mentioned drawbacks of the prior art, the present invention provides a closed-loop intelligent control system for coke ovens, which can effectively overcome the defects of the prior art that the coke oven cannot be effectively controlled to improve the combustion efficiency and energy utilization rate.
[0006] (2) Technical Solutions
[0007] To achieve the above object, the present invention is realized through the following technical solutions:
[0008] A closed-loop intelligent control system for coke ovens includes a closed-loop control unit for the flow rate of the main return gas pipe and a closed-loop control unit for the flue gas suction.
[0009] The closed-loop control unit for the flow rate of the main return gas pipe monitors the change of the temperature uniformity of the coke oven in real time, and according to the fluctuation of the coke oven temperature, makes a feedback increase or decrease compensation for the heat consumption of the coke oven, and comprehensively adjusts the main pipe flow rate by feedforward, adjusts the main pipe flow rate by feedback, and predicts the flow rate of the main return gas pipe to realize the adjustment of the gas flow rate. At the same time, according to the change of the actual gas flow rate, the pressure behind the main valve of the return gas pipe is stably regulated to ensure the stability of the gas flow rate.
[0010] The flue gas suction closed-loop control unit, in combination with the coupling optimization regulation model of the coke oven heat consumption and flue gas suction, automatically matches the flue gas suction value according to the actual gas flow rate and the gas main pipe pressure, comprehensively considers the fluctuation of the excess coefficient of the branch flue, and performs feedback regulation on the flue gas suction value according to the fluctuation of the air coefficient. At the same time, based on the closed-loop control of the return gas main pipe flow, a feed-forward regulation is performed on the flue gas suction value to stabilize the excess coefficient of the branch flue, and the flue gas suction value is sent to the flue gas suction regulating flap control mechanism, and the actual flue gas suction reaches the target value by adjusting the flap mechanism.
[0011] Preferably, the return gas main pipe flow closed-loop control unit performs the closed-loop control task of the return gas main pipe flow by constructing a Deep Q-Network (DQN), and constructs a first reward function to guide the Deep Q-Network (DQN) to learn the optimal control strategy;
[0012] Among them, the states involved in the Deep Q-Network (DQN) include the change in the temperature uniformity of the coke oven, the feedback increase and decrease compensation of the coke oven heat consumption, the feed-forward quantity, the feedback quantity, the predicted gas flow rate of the gas main pipe, and the actual gas flow rate. The actions involved in the Deep Q-Network (DQN) include adjusting the gas flow rate and regulating the pressure after the return gas main pipe valve.
[0013] Preferably, the Deep Q-Network (DQN) learns the mapping relationship between the state and the action through model training, takes the state as the input, and outputs the Q value of each possible action taken in this state;
[0014] Among them, the update process of the Q value is represented by the following formula:
[0015]
[0016] In the above formula, s t is the state at time t, a t is the action taken at time t, Q(s t , a t ) is the Q value of the action a t taken in the state s t at time t, Q’(s t , a t ) is the updated Q value of the action a t taken in the state s t at time t, α is the learning rate used to represent the speed at which new information covers old information, r is the immediate reward obtained by taking the action a t in the state s t at time t, δ is the discount factor used to weigh the importance of future rewards, α, δ ∈ [0, 1], is the maximum value of the Q values of all possible actions a taken in the state s t+1 at time t + 1.
[0017] Preferably, the Deep Q - Network (DQN) implements an experience replay mechanism during model training. By storing and randomly sampling historical experiences, the correlation between data is reduced to improve the efficiency and stability of model training. At the same time, the Deep Q - Network (DQN) continuously updates the network parameters using the interaction between the system and the environment, and gradually learns the optimal control strategy.
[0018] Preferably, the construction factors of the first reward function include:
[0019] Regarding the change in the temperature uniformity of the coke oven, a reward is given within a certain range to encourage the system to maintain the stability of the coke oven temperature;
[0020] Regarding the feedback increase and decrease compensation of the coke oven heat consumption, corresponding rewards or punishments are given according to the compensation effect to guide the system to learn an effective compensation strategy;
[0021] Regarding the regulation of the gas flow rate and the control of the pressure after the main valve of the return gas to the coke oven, it is considered to give a reward when the actual gas flow rate and the pressure after the valve reach the expected range to encourage the system to learn effective control actions;
[0022] Regarding the suppression of pressure fluctuations, a penalty is imposed on the situation with large pressure fluctuations to guide the system to learn a stable control strategy.
[0023] Preferably, the flue gas draft closed - loop control unit constructs a Deep Deterministic Policy Gradient (DDPG) algorithm - based Deep Deterministic Policy Network (Actor) and Q - value Network (Critic) to perform the flue gas draft closed - loop control task, and constructs a second reward function to guide the core component Agent to learn the optimal control strategy;
[0024] Among them, the core component Agent is the entity for the interaction between the system and the environment, including the Deep Deterministic Policy Network (Actor) and the Q - value Network (Critic);
[0025] The states involved in the Deep Deterministic Policy Network (Actor) include the actual gas flow rate, the gas main pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, and the air coefficient. The actions involved in the Deep Deterministic Policy Network (Actor) include outputting a control signal for regulating the flue gas draft to control the flue gas draft regulating flap control mechanism;
[0026] The states involved in the Q - value Network (Critic) include the actual gas flow rate, the gas main pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, the air coefficient, and the actions taken in this state. The actions involved in the Q - value Network (Critic) include outputting the corresponding Q - value by evaluating the pros and cons of the state - action combination.
[0027] Preferably, the process of the Deep Deterministic Policy Gradient (DDPG) algorithm includes:
[0028] S1. Initialize the network parameters of the Actor current network, the Critic current network, the Actor target network, and the Critic target network, and clear the experience playback set D;
[0029] S2, start iteration, initialize S to the first state in the current state sequence;
[0030] S3, obtain the feature vector φ(S) of state S, and obtain the action A=πθ(φ(S))+N taken in state S based on the Actor current network;
[0031] S4, execute action A, obtain the immediate reward R, the new state S', and whether the state is_end is terminated, and store the data set {φ(S), A, R, φ(S'), is_end} in the experience replay set D;
[0032] S5. Sample the experience playback set D to obtain m data group samples {φ(S j ),A j ,R j ,φ(S' j ),is_end j}, j = 1, 2, ..., m, calculate the current target Q value yj:
[0033]
[0034] Among them, R j is the instant reward in the jth data group in the data group sample, γ is the decay factor, Q'[φ(S' j ),π θ' (φ(S' j )),w'] is in state S' j The eigenvector φ(S' j ), the Actor target network for the state S'j feature vector φ (S' j ) is an estimated value of π θ' (φ(S' j )), and the network parameter w' of the Critic target network, the Q value obtained by the Critic target network;
[0035] S6. Update the parameter w of the Critic current network;
[0036] S7, update the parameters θ of the Actor's current network;
[0037] S8. If the parameter update frequency C of the Critic target network is C = 1, update the network parameters w' of the Critic target network and the network parameters θ' of the Actor target network:
[0038] w' ← τw + (1 - τ)w';
[0039] θ' ← τθ + (1 - τ)θ';
[0040] where τ is the soft update coefficient;
[0041] S9. If the new state S' is a terminal state, end the current round of iteration; otherwise, use the new state S' as the state S and jump to S3.
[0042] Preferably, the update of the parameters w of the current Critic network in S6 includes:
[0043] Based on the first loss function, use the backpropagation algorithm to update the parameters w of the current Critic network;
[0044] where the first loss function is expressed by the following formula:
[0045]
[0046] In the above formula, Q(φ(S j ), A j , w) is the Q value obtained by the current Critic network under the condition of the feature vector φ(S j ) of the state S j , the action A j taken in the state S j , and the parameters w of the current Critic network;
[0047] The update of the parameters θ of the current Actor network in S7 includes:
[0048] Based on the second loss function, use the backpropagation algorithm to update the parameters θ of the current Actor network;
[0049] where the second loss function is expressed by the following formula:
[0050]
[0051] In the above formula, Q(S j , A j , θ) is the Q value obtained by the current Critic network under the condition of the state S j , the action A j taken in the state S j , and the parameters θ of the current Actor network.
[0052] Preferably, during the interaction between the core component Agent and the environment, the Agent selects actions according to the current state, observes the feedback of the environment, stores the experience data of interacting with the environment in the experience replay buffer, randomly samples from the experience replay buffer to reduce the correlation between data, and adopts an offline learning method for model training to continuously optimize the network parameters of the deep deterministic policy network Actor and the Q-value network Critic, and gradually learns the optimal control strategy.
[0053] Preferably, the construction factors of the second reward function include:
[0054] Comprehensively consider the stability, combustion efficiency and energy loss of the coke oven combustion process, and balance these factors;
[0055] Regarding the suppression of flue gas suction fluctuations, in the case of large flue gas suction fluctuations, that is, when the actual flue gas suction deviates significantly from the target value, impose penalties to guide the core component Agent to learn a more stable flue gas suction regulation strategy and improve the stability of system operation;
[0056] Regarding combustion efficiency, reward the state and action combinations in which the system has a high combustion efficiency to encourage the core component Agent to learn a more effective flue gas suction regulation strategy and improve the combustion efficiency of the system;
[0057] Regarding energy loss, reward the state and action combinations in which the system has a high energy utilization rate to encourage the core component Agent to learn a more energy-saving flue gas suction regulation strategy and reduce the energy loss of the system.
[0058] (III) Beneficial effects
[0059] Compared with the prior art, a coke oven closed-loop intelligent control system provided by the present invention has the following beneficial effects:
[0060] 1) Ensure the stability of the gas flow during the coke oven combustion process to improve the horizontal and vertical uniformity of the furnace temperature, thereby improving the coke oven temperature distribution, increasing the gas utilization efficiency. By real-time monitoring the change of the coke oven temperature uniformity, comprehensively feedforward regulating the main pipe flow, feedback regulating the main pipe flow and predicting the return coke oven gas main pipe flow, realize the closed-loop intelligent regulation of the gas flow. At the same time, the system stably regulates the pressure after the main valve of the return coke oven gas, reduces the pressure fluctuation after the valve, to ensure the stability of the main pipe gas flow and reduce the energy loss caused by the pressure fluctuation;
[0061] 2) Combine the coke oven heat consumption and flue gas suction coupling optimization control model, automatically match the flue gas suction value according to various factors, and comprehensively consider the fluctuations of the excess coefficient and air coefficient of the branch flue for feedback adjustment, so that the actual flue gas suction reaches the target value, which helps to improve the combustion performance of the gas in the combustion chamber, improve the combustion efficiency, and ensure the longitudinal temperature uniformity of the combustion process;
[0062] 3) Improve production efficiency: Through real-time monitoring and intelligent control, the system can optimize operation parameters, maximize production efficiency, reduce downtime, and increase production;
[0063] 4) Reduce energy consumption: Through intelligent optimization control, the system can effectively manage energy consumption, reduce waste, and thus reduce production costs;
[0064] 5) Optimize product quality: The system can monitor key parameters and adjust the process in a timely manner to ensure the stability and consistency of product quality;
[0065] 6) Improve safety: Through real-time monitoring and intelligent control, the system can reduce human errors, improve production safety, and reduce the risk of accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0067] Figure 1 It is a working schematic diagram of the closed-loop control unit for the main flow of the return gas in the present invention;
[0068] Figure 2 It is a working schematic diagram of the closed-loop control unit for the flue gas suction in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0070] A coke oven closed-loop intelligent control system includes a closed-loop control unit for the main flow of the return gas and a closed-loop control unit for the flue gas suction;
[0071] As shown Figure 1 in the figure, the closed-loop control unit for the main flow of the recycled gas monitors the change in the uniformity of the coke oven temperature in real time. According to the fluctuation of the coke oven temperature, it makes feedback increase or decrease compensation for the heat consumption of the coke oven, and comprehensively adjusts the main flow through feedforward, adjusts the main flow through feedback, and predicts the main flow of the recycled gas main pipe to realize the adjustment of the gas flow. At the same time, according to the change of the actual gas flow, it stably controls the pressure after the main valve of the recycled gas main pipe to ensure the stability of the gas flow.
[0072] The closed-loop control unit for the main flow of the recycled gas main pipe executes the closed-loop control task of the main flow of the recycled gas main pipe by constructing a deep Q-network (DQN), and constructs a first reward function to guide the deep Q-network (DQN) to learn the optimal control strategy;
[0073] Among them, the states involved in the deep Q-network (DQN) include the change in the uniformity of the coke oven temperature, the feedback increase or decrease compensation of the coke oven heat consumption, the feedforward quantity, the feedback quantity, the predicted flow of the gas main pipe, and the actual gas flow. The actions involved in the deep Q-network (DQN) include adjusting the gas flow and controlling the pressure after the main valve of the recycled gas main pipe.
[0074] 1) The deep Q-network (DQN) learns the mapping relationship between states and actions through model training, takes the state as input, and outputs the Q value of each possible action taken in this state;
[0075] Among them, the update process of the Q value is represented by the following formula:
[0076]
[0077] In the above formula, s t is the state at time t, a t is the action taken at time t, Q(s t , a t ) is the Q value of the action a t taken in the state s t at time t, Q’(s t , a t ) is the updated Q value of the action a t taken in the state s t at time t, α is the learning rate used to represent the speed at which new information covers old information, r is the immediate reward obtained by taking the action a t in the state s t at time t, δ is the discount factor used to weigh the importance of future rewards, α, δ ∈ [0, 1], is the maximum value of the Q values of all possible actions a taken in the state s t+1 at time t + 1.
[0078] 2) During the model training process of the Deep Q-Network (DQN), the experience replay mechanism is implemented. By storing and randomly sampling historical experiences, the correlation between data is reduced to improve the efficiency and stability of model training. At the same time, the DQN continuously updates the network parameters using the interaction between the system and the environment, gradually learning the optimal control strategy.
[0079] 3) The construction factors of the first reward function include:
[0080] Regarding the change in the temperature uniformity of the coke oven, a reward is given within a certain range to encourage the system to maintain the stability of the coke oven temperature;
[0081] Regarding the feedback increase and decrease compensation of the coke oven heat consumption, corresponding rewards or punishments are given according to the compensation effect to guide the system to learn an effective compensation strategy;
[0082] Regarding the regulation of the gas flow rate and the control of the pressure after the main valve of the return gas to the oven, it is considered to give a reward when the actual gas flow rate and the pressure after the valve reach the expected range to encourage the system to learn effective control actions;
[0083] Regarding the suppression of pressure fluctuations, a penalty is imposed on the situation with large pressure fluctuations to guide the system to learn a stable control strategy.
[0084] In the technical solution of this application, for the problem of closed-loop control of the main return gas flow to the oven, the Deep Q-Network (DQN) in the deep reinforcement learning method of the reinforcement learning algorithm is used to perform the closed-loop control task of the main return gas flow to the oven.
[0085] The DQN combines the ideas of deep learning and reinforcement learning and can effectively handle problems involving high-dimensional state spaces and action spaces. In the closed-loop control of the main return gas flow to the oven, the system needs to make decisions based on multiple factors such as the change in the temperature uniformity of the coke oven, the feedback increase and decrease compensation of the coke oven heat consumption, the feedforward quantity, the feedback quantity, and the predicted gas flow rate of the main gas pipe. This requires an algorithm that can handle high-dimensional state spaces and multi-variable decisions.
[0086] The DQN algorithm learns the mapping relationship from states to actions by establishing a deep neural network and combines techniques such as experience replay and target networks to improve the stability and efficiency of learning. During the real-time monitoring of the change in the temperature uniformity of the coke oven, the system can use the DQN algorithm to continuously learn and optimize the control strategy to achieve the closed-loop intelligent regulation of the gas flow rate, and stably control the pressure after the main valve of the return gas to the oven according to the change in the actual gas flow rate, ultimately ensuring the smoothness of the coke oven combustion process, improving the combustion efficiency, and reducing energy loss.
[0087] Such as Figure 2As shown in the figure, the flue gas suction closed-loop control unit, combined with the coupling optimization regulation model of coke oven heat consumption and flue gas suction, automatically matches the flue gas suction value according to the actual gas flow rate and the main gas pipe pressure, comprehensively considers the fluctuation of the excess coefficient of the branch flue, and performs feedback regulation on the flue gas suction value according to the fluctuation of the air coefficient. At the same time, based on the closed-loop control of the main return gas pipe flow, a feed-forward regulation is performed on the flue gas suction value to stabilize the excess coefficient of the branch flue, and the flue gas suction value is sent to the flue gas suction regulating flap control mechanism, and the actual flue gas suction reaches the target value by adjusting the flap mechanism.
[0088] The flue gas suction closed-loop control unit constructs a deep deterministic policy network Actor and a Q-value network Critic based on the deep deterministic policy gradient algorithm DDPG to execute the flue gas suction closed-loop control task, and constructs a second reward function to guide the core component Agent to learn the optimal control strategy;
[0089] Among them, the core component Agent is the entity for the system to interact with the environment, including the deep deterministic policy network Actor and the Q-value network Critic;
[0090] The states involved in the deep deterministic policy network Actor include the actual gas flow rate, the main gas pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, and the air coefficient. The actions involved in the deep deterministic policy network Actor include outputting a control signal for regulating the flue gas suction to control the flue gas suction regulating flap control mechanism;
[0091] The states involved in the Q-value network Critic include the actual gas flow rate, the main gas pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, the air coefficient, and the actions taken in this state. The actions involved in the Q-value network Critic include outputting the corresponding Q-value by evaluating the pros and cons of the state-action combination.
[0092] 1) The process of the deep deterministic policy gradient algorithm DDPG includes:
[0093] S1. Initialize the network parameters of the current Actor network, the current Critic network, the target Actor network, and the target Critic network, and clear the experience replay set D;
[0094] S2. Start iteration, and initialize S as the first state in the current state sequence;
[0095] S3. Obtain the feature vector φ(S) of state S, and based on the current Actor network, obtain the action A = πθ(φ(S)) + N taken in state S;
[0096] S4. Execute action A to obtain the immediate reward R, the new state S', and whether it is a terminal state is_end obtained by taking action A in state S, and store the data set {φ(S), A, R, φ(S'), is_end} in the experience replay set D;
[0097] S5. Sample from the experience replay set D to obtain m data set samples {φ(S j ), A j , R j , φ(S' j ), is_end j}, where j = 1, 2,..., m, and calculate the current target Q value yj:
[0098]
[0099] where R j is the immediate reward in the j-th data set of the data set sample, γ is the decay factor, and Q'[φ(S' j ), π θ' (φ(S' j )), w'] is the Q value obtained by the Critic target network under the condition of the feature vector φ(S' j ) in state S', the estimated value π j (φ(S' j )) of the feature vector φ(S' θ' (φ(S' j )) of the Actor target network for state S'j, and the network parameters w' of the Critic target network;
[0100] S6. Update the parameters w of the Critic current network;
[0101] S7. Update the parameters θ of the Actor current network;
[0102] S8. If the parameter update frequency C of the Critic target network is 1, update the network parameters w' of the Critic target network and the network parameters θ' of the Actor target network:
[0103] w' ← τw + (1 - τ)w';
[0104] θ' ← τθ + (1 - τ)θ';
[0105] where τ is the soft update coefficient;
[0106] S9. If the new state S' is a terminal state, end the current round of iteration; otherwise, use the new state S' as state S and jump to S3.
[0107] Specifically, in S6, the parameters w of the current Critic network are updated, including:
[0108] Based on the first loss function, the parameters w of the current Critic network are updated using the backpropagation algorithm;
[0109] Among them, the first loss function is expressed by the following formula:
[0110]
[0111] In the above formula, Q(φ(S j ), A j , w) is the Q value obtained by the current Critic network under the condition of the feature vector φ(S j ) of the state Sj, the action A j taken in the state Sj, and the parameters w of the current Critic network;
[0112] In S7, the parameters θ of the current Actor network are updated, including:
[0113] Based on the second loss function, the parameters θ of the current Actor network are updated using the backpropagation algorithm;
[0114] Among them, the second loss function is expressed by the following formula:
[0115]
[0116] In the above formula, Q(S j , A j , θ) is the Q value obtained by the current Critic network under the condition of the state S j , the action A j taken in the state S j , and the parameters θ of the current Actor network.
[0117] 2) During the interaction between the core component Agent and the environment, it selects actions according to the current state, observes the feedback of the environment, stores the experience data of the interaction with the environment in the experience replay buffer, randomly samples from the experience replay buffer to reduce the correlation between data, and uses the offline learning method for model training to continuously optimize the network parameters of the deep deterministic policy network Actor and the Q-value network Critic, and gradually learns the optimal control strategy.
[0118] 3) The construction factors of the second reward function include:
[0119] Comprehensively consider the stability, combustion efficiency, and energy loss of the coke oven combustion process, and balance these factors;
[0120] For suppressing the fluctuations of the flue gas suction, in the case of large fluctuations in the flue gas suction, that is, when the actual flue gas suction deviates significantly from the target value, penalties are imposed to guide the core component Agent to learn a more stable flue gas suction adjustment strategy and improve the stability of the system operation;
[0121] Regarding the combustion efficiency, rewards are given to the system in the state and action combinations with high combustion efficiency to encourage the core component Agent to learn a more effective flue gas suction adjustment strategy and improve the combustion efficiency of the system;
[0122] Regarding the energy loss, rewards are given to the system in the state and action combinations with high energy utilization efficiency to encourage the core component Agent to learn a more energy-saving flue gas suction adjustment strategy and reduce the energy loss of the system.
[0123] In the technical solution of this application, in view of the complexity and diversity of the closed-loop control of the flue gas suction, the Deep Deterministic Policy Gradient algorithm DDPG is used as the reinforcement learning algorithm, and a deep deterministic policy network Actor and a Q-value network Critic are constructed to execute the closed-loop control task of the flue gas suction.
[0124] The DDPG algorithm has certain advantages in dealing with problems in continuous action spaces and continuous state spaces and is suitable for control scenarios that require fine adjustment and continuous actions. In the closed-loop control of the flue gas suction, the system needs to comprehensively consider multiple factors to adjust the flue gas suction value, and these factors may change continuously. Therefore, the DDPG algorithm can help the system learn complex control strategies.
[0125] The DDPG algorithm combines deep learning and deterministic policy gradient methods, can effectively handle continuous action spaces, and can stably learn the optimal control strategy during the training process. In addition, since the system needs to continuously adjust the flue gas suction value according to the feedback of the environment, which involves the processing of continuous state spaces, the DDPG algorithm can also provide effective learning capabilities in this regard. Through the DDPG algorithm, the system can learn the optimal flue gas suction adjustment strategy according to the actual flue gas suction combined with the changes of various factors to achieve the closed-loop control of the flue gas suction.
[0126] The output of the deep deterministic policy network Actor will affect the adjustment of the flue gas suction regulating flap control mechanism, so that the actual flue gas suction reaches the target value, ensuring the stable combustion of the gas in the combustion chamber and ensuring the longitudinal temperature uniformity of the combustion process. By continuously optimizing the learning processes of the deep deterministic policy network Actor and the Q-value network Critic, the system can achieve more efficient closed-loop control of the flue gas suction.
[0127] The construction of the second reward function comprehensively considers the stability, combustion efficiency, and energy loss of the coke oven combustion process, and converts these elements into reward signals to guide the core component Agent to learn the optimal control strategy. By continuously adjusting the design of the second reward function, it is possible to guide the core component Agent to more efficiently explore the state space during the learning process, learn a better flue gas suction adjustment strategy, and thus improve the overall performance of the system.
[0128] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A closed-loop intelligent control system for a coke oven, characterized in that: It includes a closed-loop control unit for the flow rate of the main return gas pipe and a closed-loop control unit for the flue gas suction force. The closed-loop control unit for the flow rate of the main return gas pipe monitors the change of the temperature uniformity of the coke oven in real time. According to the fluctuation of the coke oven temperature, it makes feedback increase and decrease compensation for the heat consumption of the coke oven, comprehensively adjusts the main pipe flow rate by feedforward, adjusts the main pipe flow rate by feedback, and predicts the flow rate of the main return gas pipe to realize the regulation of the gas flow rate. At the same time, according to the change of the actual gas flow rate, it stably controls the pressure after the main valve of the return gas pipe to ensure the stability of the gas flow rate. The closed-loop control unit for the flue gas suction force combines the coupling optimization control model of the coke oven heat consumption and the flue gas suction force, automatically matches the flue gas suction force value according to the actual gas flow rate and the gas main pipe pressure, comprehensively considers the fluctuation of the excess coefficient of the branch flue, and makes feedback adjustment to the flue gas suction force value according to the fluctuation of the air coefficient. At the same time, it makes feedforward adjustment to the flue gas suction force value based on the closed-loop control of the flow rate of the main return gas pipe to stabilize the excess coefficient of the branch flue, and sends the flue gas suction force value to the control mechanism of the flue gas suction adjustment flap, and makes the actual flue gas suction force reach the target value by adjusting the flap mechanism. The closed-loop control unit for the flue gas suction force constructs a deep deterministic policy network Actor and a Q-value network Critic based on the deep deterministic policy gradient algorithm DDPG to execute the closed-loop control task of the flue gas suction force, and guides the core component Agent to learn the optimal control strategy by constructing a second reward function. Among them, the core component Agent is an entity that interacts with the environment of the system, including a deep deterministic policy network Actor and a Q-value network Critic. The states involved in the deep deterministic policy network Actor include the actual gas flow rate, the gas main pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, and the air coefficient. The actions involved in the deep deterministic policy network Actor include outputting a control signal for adjusting the flue gas suction force to control the flue gas suction adjustment flap control mechanism. The states involved in the Q-value network Critic include the actual gas flow rate, the gas main pipe pressure, the coke oven heat consumption, the excess coefficient of the branch flue, the air coefficient, and the action taken in this state. The actions involved in the Q-value network Critic include outputting the corresponding Q-value by evaluating the pros and cons of the state-action combination. The process of the deep deterministic policy gradient algorithm DDPG includes: S1. Initialize the network parameters of the current Actor network, the current Critic network, the target Actor network, and the target Critic network, and clear the experience replay set D. S2. Start the iteration, and initialize S as the first state in the current state sequence. S3. Obtain the feature vector φ(S) of the state S, and based on the current Actor network, obtain the action A taken in the state S, A = πθ(φ(S)) + N. S4. Execute the action A, obtain the immediate reward R, the new state S', and whether it is a terminal state is_end obtained by taking the action A in the state S, and store the data group {φ(S), A, R, φ(S'), is_end} in the experience replay set D. S5. Sample the experience replay set D to obtain m data group samples {φ(S j ), A j , R j , φ(S' j ), is_end j}, where j = 1, 2,..., m, and calculate the current target Q value yj: Among them, R j is the immediate reward in the j-th data group of the data group sample, γ is the attenuation factor, Q'[φ(S' j ), π θ' (φ(S' j ))), w'] is the Q value obtained by the Critic target network under the condition of the feature vector φ(S' j ) of the state S’, the estimated value π j (φ(S' j )) of the feature vector φ(S' θ' (φ(S' j )) of the Actor target network for the state S’j, and the network parameters w' of the Critic target network; S6. Update the parameters w of the current Critic network; S7. Update the parameters θ of the current Actor network; S8. If the parameter update frequency C of the Critic target network is 1, update the network parameters w' of the Critic target network and θ' of the Actor target network: w' ← τw + (1 - τ)w'; θ' ← τθ + (1 - τ)θ'; where τ is the soft update coefficient; S9. If the new state S' is a terminal state, end the current round of iteration; otherwise, use the new state S' as the state S and jump to S3; During the interaction with the environment, the core component Agent selects actions according to the current state, observes the feedback of the environment, stores the experience data of the interaction with the environment in the experience replay buffer, randomly samples from the experience replay buffer to reduce the correlation between data, and uses the off-line learning method for model training, continuously optimizing the network parameters of the deep deterministic policy network Actor and the Q-value network Critic, and gradually learning the optimal control strategy; The construction factors of the second reward function include: Comprehensively consider the stability, combustion efficiency, and energy loss of the coke oven combustion process, and balance these factors; Regarding the suppression of flue gas suction fluctuations, penalize the situation where the flue gas suction fluctuates greatly, that is, the actual flue gas suction deviates significantly from the target value, so as to guide the core component Agent to learn a more stable flue gas suction regulation strategy and improve the stability of system operation; Regarding the combustion efficiency, give rewards to the state-action combinations with high combustion efficiency of the system, so as to encourage the core component Agent to learn a more effective flue gas suction regulation strategy and improve the combustion efficiency of the system; Regarding the energy loss, give rewards to the state-action combinations with high energy utilization efficiency of the system, so as to encourage the core component Agent to learn a more energy-saving flue gas suction regulation strategy and reduce the energy loss of the system.
2. The coke oven closed-loop intelligent control system according to claim 1, wherein: The return gas main pipe flow closed-loop control unit performs the return gas main pipe flow closed-loop control task by constructing a deep Q-network DQN, and constructs the first reward function to guide the deep Q-network DQN to learn the optimal control strategy; Among them, the states involved in the deep Q-network DQN include the change of coke oven temperature uniformity, the feedback increase and decrease compensation of coke oven heat consumption, the feedforward quantity, the feedback quantity, the predicted flow of the gas main pipe, and the actual gas flow, and the actions involved in the deep Q-network DQN include adjusting the gas flow and regulating the pressure after the return gas main pipe valve.
3. The coke oven closed-loop intelligent control system according to claim 2, wherein: The deep Q-network DQN learns the mapping relationship between states and actions through model training, takes the state as the input, and outputs the Q-values of taking each possible action in this state; Among them, the update process of the Q-value is expressed by the following formula: In the above formula, s t is the state at time t, a t is the action taken at time t, Q(s t , a t ) is the Q-value of taking action a t under the state s t at time t, Q’(s t , a t ) is the updated Q-value of taking action a t under the state s t at time t, α is the learning rate used to represent the speed at which new information covers old information, r is the immediate reward obtained by taking action a t under the state s t at time t, δ is the discount factor used to weigh the importance of future rewards, α, δ ∈ [0, 1], is the maximum value among the Q-values of taking all possible actions a under the state s t+1 at time t + 1.
4. The coke oven closed-loop intelligent control system according to claim 2, wherein: During the model training process of the deep Q-network DQN, the experience replay mechanism is implemented, and the correlation between data is reduced by storing and randomly sampling historical experiences to improve the efficiency and stability of model training. At the same time, the deep Q-network DQN continuously updates the network parameters by using the interaction between the system and the environment, and gradually learns the optimal control strategy.
5. The coke oven closed-loop intelligent control system according to claim 2, characterized in that: The construction factors of the first reward function include: For the change in the temperature uniformity of the coke oven, a reward is given within a certain range to encourage the system to maintain the stability of the coke oven temperature; For the feedback increase and decrease compensation of the coke oven heat consumption, corresponding rewards or punishments are given according to the compensation effect to guide the system to learn effective compensation strategies; For the regulation of the gas flow rate and the control of the pressure after the main valve of the return gas to the oven, a reward is considered when the actual gas flow rate and the pressure after the valve reach the expected range to encourage the system to learn effective control actions; For the suppression of pressure fluctuations, a penalty is imposed on the situation with large pressure fluctuations to guide the system to learn stable control strategies.
6. The coke oven closed-loop intelligent control system according to claim 1, characterized in that: In S6, the parameters w of the current Critic network are updated, including: Based on the first loss function, the parameters w of the current Critic network are updated using the backpropagation algorithm; Among them, the first loss function is expressed by the following formula: In the above formula, Q(φ(S j ), A j , w) is the Q value obtained by the current Critic network under the condition of the feature vector φ(S j ) in state Sj, the action A j taken in state Sj, and the parameters w of the current Critic network; In S7, the parameters θ of the current Actor network are updated, including: Based on the second loss function, the parameters θ of the current Actor network are updated using the backpropagation algorithm; Among them, the second loss function is expressed by the following formula: In the above formula, Q(S j , A j , θ) is the Q value obtained by the current Critic network under the conditions of the state Sj, the action Aj taken in the state Sj, and the parameters θ of the current Actor network.
Citation Information
Patent Citations
Coke oven system thermal state regulation and control method and device, electronic equipment and storage medium
CN118092550A