Dry quenching RTO accurate control system and method based on Transform model

Through the precision control system of dry quenching RTO based on the Transformer model, traditional control systems are solved, and efficient and safe dry quenching control is achieved, which significantly improves energy utilization efficiency and system stability.

CN120103767APending Publication Date: 2025-06-06SHANDONG QINGBO IND TECH CO LTD

Patent Information

Application Number
CN202510295589.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional dry-quenching control systems are difficult to cope with the nonlinear and multivariable coupling characteristics of the process, resulting in low operating efficiency, high energy consumption and increased safety risks.

Method used

The dry-extinguishing RTO precision control system based on the Transformer model is adopted, and the system's global collaborative control and dynamic adjustment strategies are realized through data acquisition, time series prediction, RTO real-time optimization algorithm, edge computing and abnormal detection and alarm modules.

Benefits of technology

It improves prediction accuracy, realizes global system optimization, significantly reduces energy consumption, improves safety, and ensures stable system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103767A_ABST
    Figure CN120103767A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of dry quenching, and particularly relates to a dry quenching RTO accurate control system and method based on a Transform model, and the system comprises a data collection module, a time sequence prediction module, an RTO real-time optimization algorithm module, an edge calculation module and an anomaly detection and alarm module. The method comprises the following steps: firstly, acquiring data in a dry quenching furnace in real time by a data acquisition module, and predicting a key parameter change trend in a dry quenching process by a time sequence prediction module in combination with historical data and real-time data; the RTO real-time optimization algorithm module dynamically adjusts system state parameters according to a prediction result; the anomaly detection module compares the prediction data with the actual data and gives out early warning in time; and an optimization result is issued to the PLC control system in real time through the edge calculation module. Through modular design, the coupling complexity of the dry quenching system and data acquisition, prediction, optimization, control and other modules is reduced, and the energy utilization efficiency is improved by dynamically adjusting the control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of dry coke quenching, and in particular relates to a Transformer model-based dry coke quenching (RTO) precise control system and method. Background Art

[0002] CDQ technology has been widely used in the steel industry to achieve heat recovery and environmental protection during the coke cooling process. However, traditional CDQ control systems rely on empirical rules or simple PID control, which is difficult to cope with the nonlinear and multivariable coupling characteristics of the process, resulting in low operating efficiency, high energy consumption, and increased safety risks.

[0003] The closest existing technologies include rule control systems, PID control systems, and predictive control based on neural networks. Rule control systems use preset logical rules to perform simple control of key points (such as air guides and circulating air). The disadvantage is that they lack real-time optimization capabilities and have slow responses. PID control systems perform closed-loop control of key parameters such as temperature and pressure. The disadvantage is that they are only suitable for single-variable control and have difficulty handling multi-variable coupling problems. Predictive control based on neural networks uses simple neural networks to predict and control some parameters. The disadvantage is that the prediction accuracy is limited and it is difficult to capture dependencies in long time series.

[0004] Existing technologies mostly focus on the optimization of a single module, cannot achieve global coordinated control of the system, lack dynamic adjustment strategies, and have low energy utilization. In addition, traditional algorithms are difficult to accurately capture the nonlinear dynamic characteristics of the CDQ process, and abnormal detection is delayed, resulting in safety hazards that cannot be warned and handled in advance. Summary of the invention

[0005] In view of the various shortcomings of the existing technology, the inventors have researched and designed a dry coke quenching RTO precise control system and method based on the Transformer model in long-term practice. Through modular design, the coupling complexity of the dry coke quenching system with data acquisition, prediction, optimization, control and other modules is reduced, and the energy utilization efficiency is improved by dynamically adjusting the control strategy.

[0006] The present invention provides a dry coke quenching (RTO) precise control system based on a Transformer model, comprising the following parts.

[0007] Data acquisition module: used to collect the status parameters of relevant equipment in the CDQ process, including coke temperature, circulating gas temperature, circulating gas flow, coke discharge, air introduction flow, material level, and boiler inlet temperature.

[0008] Time series prediction module: Based on the Transformer model, the self-attention mechanism is used to capture the long-term dependencies in the CDQ process, and predict the future states of key parameters such as coke loading temperature, coke discharge temperature, boiler inlet temperature, circulating gas temperature, boiler inlet pressure, pre-storage chamber pressure, discharge flow, circulating gas flow, air introduction flow, material level, gas content, etc. in the multivariable coupling system. Identify the dynamic changes of coke cooling rate, circulating gas efficiency, air-to-material ratio, and gas production rate at the beginning and end of coke cooling.

[0009] RTO real-time optimization algorithm module: using reinforcement learning and heuristic algorithms, combined with the characteristics of the dry coke quenching process, based on the prediction results, dynamically adjust the circulation fan speed, air intake, bypass valve opening, vent valve opening and other system parameters. In the RTO real-time optimization algorithm module, the reinforcement learning reward function is embedded, and the weights of energy recovery efficiency and coke quality maintenance are added to dynamically adjust the coke discharge speed.

[0010] Edge computing module: used to perform real-time calculation and closed-loop control of the dry coke quenching process on industrial edge devices to ensure rapid response and execution of control instructions.

[0011] Abnormal detection and alarm module: used to monitor the operating status of the CDQ system, compare the predicted data with the actual data, and provide early warning and alarm information when abnormal temperature rise, pressure imbalance, insufficient flow, and abnormal increase in gas content are found in the system.

[0012] The Transformer model includes the following parts.

[0013] Encoder: It is composed of multiple identical layers stacked together, each of which contains a multi-head self-attention mechanism and a feed-forward neural network. The encoder converts the input data into a potential representation.

[0014] Decoder: Also composed of the same multiple layers, but each layer contains a cross-attention mechanism in addition to a multi-head self-attention mechanism and a feedforward neural network. The cross-attention mechanism allows the decoder to refer to the output of the encoder when generating output. The decoder is used to combine the encoder output and the representation of the target sequence.

[0015] Multi-head self-attention mechanism unit: The self-attention mechanism captures global dependencies by weightedly combining each element of the sequence with other elements. The multi-head self-attention mechanism enables the Transformer model to capture different features by computing multiple different attention representations in parallel.

[0016] Positional encoding: used to introduce sequential information into the Transformer model.

[0017] Feedforward Neural Network: A feedforward network in each layer consists of two linear transformations and an activation function, which are applied independently at each sequence position.

[0018] The RTO real-time optimization algorithm module includes a reinforcement learning module, a heuristic algorithm module, and a joint optimization module.

[0019] Furthermore, the reinforcement learning module is provided with a reinforcement learning model, and the core of the reinforcement learning model is the reward function.

[0020] The reinforcement learning model defines a high-dimensional state space according to data collected by sensors. The high-dimensional state space includes multiple state parameters, and all state parameters together reflect the real-time operating state of the CDQ system.

[0021] The reinforcement learning module adjusts the circulating air volume, air introduction valve opening, coke discharge speed, bypass valve opening, and vent valve opening of the dry coke quenching system according to the real-time operating status.

[0022] The reinforcement learning module adopts a deep reinforcement learning algorithm, uses historical data and real-time data to train the model online, and obtains a preliminary control strategy through multiple interactions with the environment, providing a basis for the heuristic algorithm.

[0023] Furthermore, the state parameters include but are not limited to coke temperature distribution, furnace pressure, circulating air volume, air introduction ratio and coke discharge speed.

[0024] Furthermore, the heuristic algorithm module embeds a heuristic algorithm to globally optimize the key factors affecting system performance during the dry coke quenching process. The heuristic algorithm constructs a global objective function, comprehensively considering key factors of system performance such as coke cooling uniformity, energy consumption minimization, air-to-material ratio, gas production rate, burn rate, equipment safety, etc., and constructs the objective function through reasonable weight allocation and constraint setting.

[0025] Furthermore, the heuristic algorithm adopts a simulated annealing algorithm, which avoids local optimal solutions by gradually reducing the search range and searches for global optimal solutions. In the dry coke quenching process, the simulated annealing algorithm is used in combination with a genetic algorithm, a particle swarm algorithm, an ant colony algorithm, and a differential evolution algorithm.

[0026] Furthermore, the joint optimization module adopts an online collaborative mechanism to combine the online learning ability of the reinforcement learning model with the optimization results of the heuristic algorithm to jointly drive the iterative improvement of the control strategy.

[0027] A Transformer model-based dry coke quenching (RTO) precise control method is implemented using the above system and includes the following steps.

[0028] S1, the data acquisition module obtains the data in the CDQ oven, including coke temperature Tc(t), circulating gas temperature Tg(t), circulating gas flow Qg(t), coke discharge Mc(t), air introduction flow Qa(t), material level L(t), and boiler inlet temperature Tb(t).

[0029] S2, the time series prediction module uses the Transformer model to combine historical data and real-time data to predict the changing trends of key parameters in the CDQ process in the future, which specifically includes the following sub-steps.

[0030] S2.1, the collected data are stored in time series format, and each variable is timestamped, i.e. .

[0031] S2.2, using the Min-Max normalization method, Normalize the data to the interval [0,1] to prevent convergence difficulties due to numerical differences during model training.

[0032] S2.3, using the sliding window method, set the time window length to , forming an input-output pair: .

[0033] S2.4, train the Transformer model to generate future trend predictions.

[0034] S3, RTO real-time optimization algorithm module dynamically adjusts the circulation fan speed, air introduction volume, bypass valve opening, and vent valve opening according to the prediction results to optimize the coke cooling effect and energy recovery efficiency.

[0035] The S4 anomaly detection module compares predicted data with actual data, detects potential anomalies in the CDQ process, and issues timely warnings.

[0036] The S5 optimization results are sent to the PLC control system through the edge computing module.

[0037] Furthermore, in step S2.4, the specific method for training the Transformer model is as follows.

[0038] S2.4.1 Input data format: Transformer requires sequence input, and the input data format is The tensor structure of . Among them, Batch is the number of samples for one training, Sequence Length is the time window length, and Features is the feature dimension of multiple parameters.

[0039] S2.4.2 Generate position codes using sine and cosine functions: ; ; in, represents the time step, Represents the feature dimension.

[0040] S2.4.3 Design model structure: Transformer consists of multiple encoders, each of which contains: Multi-head attention: capturing the correlation of different time steps in a time series; Feedforward neural network: used for feature transformation; Layer normalization: speeds up convergence; The self-attention is calculated by the following formula: ; in is the transformation matrix of the input data, is the characteristic dimension.

[0041] S2.4.4 Training model: Using MSE as the loss function: ; The AdamW optimizer is used and the learning rate is set to , weight decay .

[0042] S2.4.5 Input the latest time window data, Transformer processes the data, calculates multi-head attention, outputs the predicted value, and generates future trend predictions.

[0043] Furthermore, in step S2.4.4, the training steps of the training model are specifically as follows: S2.4.4.1 Input historical time series data into Transformer; S2.4.4.2 Calculate predicted values ; S2.4.4.3 Calculate the loss function and backpropagate; S2.4.4.4 Update model weights; S2.4.4.5 Iterate training until convergence.

[0044] Furthermore, in step S3, the RTO real-time optimization algorithm module dynamically adjusts the circulation fan speed, air introduction amount, bypass valve opening, and vent valve opening according to the prediction results to optimize the coke cooling effect and energy recovery efficiency, which specifically includes the following sub-steps.

[0045] S3.1 In the CDQ process, confirm the core control parameters, including: Circulating air Qc: controls the circulation volume of hot gas in the CDQ furnace; Air intake Qa: used to adjust the oxygen content of the system; Bypass vent opening Bp: adjusts whether low-temperature gas enters the boiler directly; Relief valve opening Vd: used for system pressure regulation.

[0046] S3.2 Design the state space of reinforcement learning: Based on the core control parameters described in step S3.1, the reinforcement learning model defines a high-dimensional state space based on the data collected by the sensor. The state space contains variables that affect the CDQ energy consumption and quality, including the circulating air flow Qc(t), the air introduction amount Qa(t), the bypass air valve opening Bp(t), and the relief valve opening Vd(t).

[0047] S3.3 Design the action space of reinforcement learning: The action space is the actions that the system can take, which is the basis for the reinforcement learning model to optimize the control strategy by interacting with the environment. The reinforcement learning algorithm adjusts the circulating air flow Qc(t) to optimize the uniformity of coke cooling, adjusts the air introduction Qa(t) to maintain the appropriate oxygen content, adjusts the bypass air valve opening Bp(t) to control the proportion of low-temperature gas entering the boiler, and adjusts the vent valve opening Vd(t) to maintain pressure stability.

[0048] S3.4 Setting the reward function: The optimization goal is to improve the energy recovery rate and maintain the coke quality. Since the reward function is the result of weighting the dimensions of each variable, and the dimensions and magnitudes of each variable are different, they are normalized to the same scale range [0, 1] to obtain the weighted reward function R: ; in: : Energy recovery efficiency, which is mainly determined by the boiler steam output and exhaust gas temperature; : Coke quality, including cooling uniformity and final coke strength; : Fan energy consumption, which is related to the circulation wind speed and fan speed; : Boiler energy consumption, depends on exhaust gas temperature and heat exchange efficiency; : Over-temperature penalty to avoid safety problems caused by boiler overheating; Reward Weight It can be adjusted experimentally to ensure the optimal balance of objectives.

[0049] S3.5 Reinforcement learning online adjustment: Use deep reinforcement learning to dynamically adjust the core control parameters, calculate the reward value, and update the strategy based on real-time data.

[0050] The beneficial effects of the present invention are:

[0051] 1) Higher prediction accuracy: The Transformer model can accurately capture the dynamic changes of multivariable coupled systems.

[0052] 2) Global optimization: realize the coordinated optimization of each module within the system.

[0053] 3) Significant reduction in energy consumption: Improve energy efficiency by dynamically adjusting control strategies.

[0054] 4) Improved security: Detect potential anomalies in advance to ensure stable system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a system architecture diagram of the system of the present invention.

[0056] Figure 2 It is a design principle diagram of the Transformer model architecture of the present invention.

[0057] Figure 3 It is a design principle diagram of the RTO real-time optimization and dry quenching process optimization algorithm architecture of the present invention.

[0058] Figure 4 This is a diagram showing the effect of using the system of the present invention to predict the priority circulating air volume for coke removal in a dry coke quenching process. DETAILED DESCRIPTION

[0059] The present invention will be further described in detail below in conjunction with implementation examples. By describing these implementation examples in sufficient detail, it is possible for those skilled in the art to understand and practice the present invention. Without departing from the spirit and scope of the present invention, logical, achievable and other changes may be made to implementation. Therefore, the following detailed description should not be construed as limiting, and the scope of the present invention is limited only by the claims.

[0060] See also Figure 1 The present invention is specially designed for the dry coke quenching process and proposes a dry coke quenching RTO precise control system based on the Transformer model, which includes the following parts.

[0061] Data acquisition module: used to collect key parameters of relevant equipment status data during the CDQ process, including coke temperature, circulating gas temperature, circulating gas flow, coke discharge, air introduction flow, material level, and boiler inlet temperature.

[0062] Time series prediction module: Based on the Transformer model, the self-attention mechanism is used to capture long-term dependencies in the CDQ process, such as the cooling process of coke from high temperature to low temperature and its impact on environmental parameters, and predict the future states of key parameters such as coke loading temperature, coke discharge temperature, boiler inlet temperature, circulating gas temperature, boiler inlet pressure, pre-storage chamber pressure, discharge flow, circulating gas flow, air introduction flow, material level, gas content, etc. in the multivariable coupling system, especially the dynamic changes of coke cooling rate, circulating gas efficiency, air-to-material ratio, gas production rate, etc. in the early and late stages of coke cooling, so as to identify potential fluctuations in advance.

[0063] RTO real-time optimization algorithm module: using reinforcement learning and heuristic algorithms, combined with the characteristics of the dry coke quenching process, based on the prediction results, dynamically adjust the circulation fan speed, air intake, bypass valve opening, vent valve opening and other system parameters to optimize the coke cooling effect and energy recovery efficiency. At the same time, the module is embedded with a reinforcement learning reward function, and the weights of energy recovery efficiency and coke quality maintenance are specially added to dynamically adjust the coke discharge speed to ensure that the coke quality is not sacrificed while optimizing energy consumption.

[0064] Edge computing module: Real-time computing and closed-loop control of the CDQ process are implemented on industrial edge devices to ensure rapid response and execution of control instructions, especially when dealing with emergency situations such as emergency shutdown, abnormal temperature fluctuations, abnormal pressure fluctuations, and abnormal gas content fluctuations.

[0065] Abnormal detection and alarm module: Real-time monitoring of the operating status of the CDQ system. By comparing the predicted data with the actual data, it can promptly detect potential problems such as abnormal temperature increase, pressure imbalance, insufficient flow, abnormal increase in gas content, etc. in the system, and provide early warning and alarm information.

[0066] See also Figure 2 The present invention designs a Transformer model architecture as shown in the figure, which is suitable for the CDQ precise control scenario. The Transformer model includes:

[0067] Encoder: It is composed of multiple identical layers stacked together, each of which contains a multi-head self-attention mechanism and a feedforward neural network. The role of the encoder is to convert the input data into a potential representation.

[0068] Decoder: Also composed of multiple identical layers, but contains an additional criss-cross attention mechanism for combining the encoder output and the representation of the target sequence. The criss-cross attention mechanism allows the decoder to refer to the encoder output when generating outputs, i.e. it is able to "pay attention" to different parts of the input sequence, thereby generating outputs based on the context of the input sequence. This mechanism is unique to the decoder because it needs to use the encoder output to guide its own output generation process.

[0069] Multi-head self-attention mechanism unit: The self-attention mechanism captures global dependencies by weightedly combining each element of the sequence with other elements. The multi-head self-attention mechanism enables the model to better capture different features by computing multiple different attention representations in parallel.

[0070] Position encoding: Because the Transformer model has no built-in sequential processing capabilities, it requires position encoding to introduce sequential information.

[0071] Feedforward Neural Networks: A feedforward network in each layer consists of two linear transformations and an activation function (such as ReLU), which are applied independently at each sequence position.

[0072] The Transformer model is suitable for parallel training and has significant advantages over recurrent neural networks (such as LSTM and GRU) in long sequence modeling.

[0073] See also Figure 3 The present invention designs an RTO real-time optimization algorithm module architecture as shown in the figure, which is suitable for the dry coke quenching precise control scenario. The RTO real-time optimization algorithm module includes a reinforcement learning module, a heuristic algorithm module, and a joint optimization module.

[0074] The reinforcement learning module has a reinforcement learning model built in. The core of the reinforcement learning model is the reward function. The high-dimensional state space is defined based on the data collected by the sensor. These states include but are not limited to coke temperature distribution, furnace pressure, circulating air volume, air introduction ratio, and coke discharge speed. These parameters together reflect the real-time operating status of the CDQ system. The system parameters such as the circulating air volume, air introduction valve opening, coke discharge speed, bypass valve opening, and vent valve opening of the CDQ system are adjusted according to the real-time operating status.

[0075] The reward function is the core of reinforcement learning, which gives feedback based on the system status and control actions. In the CDQ process, the reward function needs to comprehensively consider multiple objectives such as energy consumption, coke cooling uniformity, process safety, and equipment wear, and construct a reward function that can reflect the overall performance of the system through reasonable weight allocation.

[0076] Using deep reinforcement learning algorithms (such as PPO or DDPG), the model is trained online using historical data and real-time data. Through multiple interactions with the environment, the reinforcement learning model can gradually learn preliminary control strategies to ensure that the system has basic operating capabilities under complex working conditions.

[0077] The reinforcement learning model obtains a preliminary control strategy through multiple interactions, which provides a basis for the heuristic algorithm. The heuristic algorithm is based on the RL strategy and starts optimization from the local part, which greatly shortens the convergence time and narrows the search space.

[0078] Although reinforcement learning can quickly generate preliminary control strategies, it may be insufficient in global optimization and complex multi-objective balance. Heuristic algorithms can further optimize strategies and improve control effects.

[0079] The heuristic algorithm module is embedded with a heuristic algorithm to globally optimize the key factors that affect system performance during the dry coke quenching process. The heuristic algorithm constructs a global objective function, comprehensively considering multiple objectives such as coke cooling uniformity, energy consumption minimization, air-to-material ratio, gas production rate, burn rate, equipment safety, etc., and constructs an objective function that can reflect the overall optimization goal of the system through reasonable weight allocation and constraint setting.

[0080] The simulated annealing algorithm (SA) is commonly used in heuristic algorithms. The simulated annealing algorithm avoids local optimal solutions by gradually reducing the search range and finds the global optimal solution. In the CDQ process, the simulated annealing algorithm is used in combination with the genetic algorithm, particle swarm algorithm, ant colony algorithm, and differential evolution algorithm to improve optimization efficiency and accuracy.

[0081] Based on the reinforcement learning strategy, the heuristic algorithm covers unexplored potential optimal areas, performs more sophisticated control parameter optimization for complex coupled working conditions, finds the global optimal solution through global search and optimization, and further improves system performance.

[0082] The joint optimization module adopts an online collaborative mechanism to combine the online learning capability of the reinforcement learning model with the optimization results of the heuristic algorithm to jointly drive the iterative improvement of the control strategy.

[0083] The reinforcement learning model continuously monitors the environmental status and adjusts the control strategy in real time to adapt to dynamic operating conditions such as coke temperature fluctuations and air volume changes. Its rapid response capability ensures the timeliness of the control strategy.

[0084] The heuristic algorithm runs periodically (such as every hour or every shift), performs global optimization based on the reinforcement learning strategy, uses the collected operation data to correct the objective function, optimizes the control effect and feeds back to the reinforcement learning model. This periodic optimization mechanism ensures the continuous improvement of system performance.

[0085] Reinforcement learning focuses on short-term dynamic adjustments, while heuristic algorithms focus on long-term parameter optimization. This division of tasks avoids resource waste and computational conflicts.

[0086] The present invention reduces the coupling complexity of the CDQ system with modules such as data acquisition, prediction, optimization, and control through modular design, supports independent optimization of each module, and realizes global coordination to ensure the stability and efficiency of the overall system.

[0087] The present invention also proposes a dry coke quenching (RTO) precise control method based on the Transformer model, which is implemented using the above-mentioned precise control system and specifically includes the following steps.

[0088] S1, the data acquisition module acquires the data in the CDQ oven in real time, such as coke temperature Tc(t), circulating gas temperature Tg(t), circulating gas flow Qg(t), coke discharge Mc(t), air inlet flow Qa(t), material level L(t), boiler inlet temperature Tb(t). These data are collected in real time through PLC (Programmable Logic Controller) and transmitted to the optimization system for processing through OPC protocol.

[0089] S2, the time series prediction module uses the Transformer model to combine historical data and real-time data to predict the changing trends of key parameters in the CDQ process in the future, which specifically includes the following sub-steps.

[0090] S2.1, the collected data are stored in time series format, and each variable is timestamped, i.e. .

[0091] S2.2, using the Min-Max normalization method, Normalize the data to the interval [0,1] to prevent convergence difficulties due to numerical differences during model training.

[0092] S2.3, using the sliding window method, set the time window length to , forming an input-output pair: .

[0093] S2.4, training the Transformer model, specifically includes the following sub-steps.

[0094] S2.4.1 Input data format: Transformer requires sequence input, and the input data format is The tensor structure of . Among them, Batch is the number of samples for one training, Sequence Length is the time window length, and Features is the feature dimension of multiple parameters.

[0095] S2.4.2 Positional Encoding: Since the Transformer structure itself does not have the order information of the time series, the positional encoding is generated using sine and cosine functions: ; ; in, represents the time step, Represents the feature dimension.

[0096] S2.4.3 Design model structure: Transformer consists of multiple encoders, each of which contains: Multi-Head Attention: Capture the correlation of different time steps in the time series; Feedforward Network: used for feature transformation; Layer Normalization: speeds up convergence; Self-Attention is calculated by the following formula: ; in is the transformation matrix of the input data, is the characteristic dimension.

[0097] S2.4.4 Training model: MSE (mean square error) is used as the loss function: ; The AdamW optimizer is used and the learning rate is set to , weight decay .

[0098] The specific training steps are: S2.4.4.1 Input historical time series data into Transformer; S2.4.4.2 Calculation of predicted values ; S2.4.4.3 Calculate the loss function and back propagate; S2.4.4.4 Update model weights; S2.4.4.5 Iterate training until convergence.

[0099] S2.4.5 Enter the latest time window data, such as the parameters for the past 24 hours: ; Transformer processes data, calculates multi-head attention, outputs predictions, and generates future trend forecasts, such as the trend of coke temperature in the next 6 hours: .

[0100] The S3 RTO real-time optimization algorithm module dynamically adjusts the circulating fan speed, air intake, bypass valve opening, and vent valve opening according to the prediction results to optimize the coke cooling effect and energy recovery efficiency. It specifically includes the following sub-steps.

[0101] S3.1 In the CDQ process, confirm the core control parameters: Circulating air Qc: controls the circulation volume of hot gas in the CDQ furnace, affecting the coke cooling uniformity and waste heat recovery efficiency; Air intake Qa: used to adjust the oxygen content of the system to ensure proper combustion and heat exchange; Bypass vent opening Bp: adjusts whether low-temperature gas enters the boiler directly; Relief valve opening Vd: used for system pressure regulation to avoid equipment damage caused by excessive pressure.

[0102] S3.2 Design of state space for reinforcement learning: During the CDQ process, the reinforcement learning model defines a high-dimensional state space based on the data collected by the sensors. The state space contains key variables that affect the energy consumption and quality of CDQ, including circulating air flow Qc(t), air introduction Qa(t), bypass air valve opening Bp(t), and relief valve opening Vd(t).

[0103] S3.3 Design the action space of reinforcement learning: The action space defines the actions that the system can take, which are the basis for the reinforcement learning model to optimize the control strategy by interacting with the environment. The reinforcement learning algorithm can dynamically adjust the control parameters, including adjusting the circulating air flow Qc(t) to optimize the uniformity of coke cooling, adjusting the air introduction Qa(t) to maintain the appropriate oxygen content, adjusting the bypass air valve opening Bp(t) to control the proportion of low-temperature gas entering the boiler, and adjusting the vent valve opening Vd(t) to maintain pressure stability.

[0104] S3.4 Set reward function: The optimization goal is to improve energy recovery while maintaining coke quality. Since the reward function is the result of weighting the dimensions of each variable, and the dimensions and magnitudes of each variable are different, they are normalized to the same scale range [0, 1] to obtain the weighted reward function R: ; in: : Energy recovery efficiency, which is mainly determined by the boiler steam output and exhaust gas temperature; : Coke quality, including cooling uniformity and final coke strength; : Fan energy consumption, which is related to the circulation wind speed and fan speed; : Boiler energy consumption, depends on exhaust gas temperature and heat exchange efficiency; : Over-temperature penalty to avoid safety problems caused by boiler overheating; Reward Weight It can be adjusted experimentally to ensure the optimal balance of objectives.

[0105] S3.5 Online adjustment of reinforcement learning: Use deep reinforcement learning (such as DQN, PPO) to dynamically adjust control parameters, calculate reward values, and update strategies based on real-time data.

[0106] The S4 anomaly detection module compares the predicted data with the actual data in real time, detects potential anomalies in the dry coke quenching process, such as temperature anomalies, pressure imbalance, etc., and issues early warnings in a timely manner.

[0107] The S5 optimization results are sent to the PLC control system in real time through the edge computing module, achieving rapid response and control execution, ensuring the stable and efficient operation of the CDQ process.

[0108] See also Figure 4 The figure shows the effect diagram of a steel group company using this system to predict the priority circulation air volume for coke discharge in the dry quenching process. In the figure, the blue line represents the recommended value of the priority circulation air volume for coke discharge, that is, the predicted value of the system; the area composed of yellow lines represents the real-time value of the priority circulation air volume for coke discharge. It can be seen from the figure that the error between the predicted value and the actual value of the priority circulation air volume for coke discharge in the dry quenching process using the system of the present invention is very small, the prediction accuracy is high, and the stable and efficient operation of the dry quenching process can be ensured.

[0109] It is not possible to describe all possible combinations of components or methods for the purpose of describing the above-described embodiments, but it should be recognized by those skilled in the art that the various embodiments may be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications and variations that fall within the scope of protection of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, the word is encompassed in a manner similar to the term "including," as explained by "including," used as a transitional word in the claims. In addition, any term "or" used in the specification of the claims is intended to mean "non-exclusive or."

Claims

1. A dry coke quenching RTO precise control system based on Transformer model, characterized in that: include: Data acquisition module: used to collect status parameters of related equipment during the dry coke quenching process; Time series prediction module: Based on the Transformer model, the self-attention mechanism is used to capture the long-term dependencies in the CDQ process and predict the future state of system parameters; RTO real-time optimization algorithm module: uses reinforcement learning and heuristic algorithms to dynamically adjust system parameters based on prediction results; Edge computing module: used to calculate and close-loop control the dry coke quenching process on industrial edge devices to ensure the response and execution of control instructions; Anomaly detection and alarm module: used to monitor the operating status of the system, compare predicted data with actual data, and provide early warning and alarm information; The Transformer model includes: Encoder and decoder: Both the encoder and decoder are stacked with multiple identical layers, each of which contains a multi-head self-attention mechanism and a feed-forward neural network; the decoder also contains a cross-attention mechanism; Multi-head self-attention mechanism unit: By computing multiple different attention representations in parallel, the Transformer model can capture different features; Position encoding: introduces sequence order information into the Transformer model; Feedforward neural network: consists of two linear transformations and an activation function.

2. The system according to claim 1, characterized in that The RTO real-time optimization algorithm module includes a reinforcement learning module, a heuristic algorithm module, and a joint optimization module; The reinforcement learning module has a reinforcement learning model, and the core of the reinforcement learning model is the reward function; The reinforcement learning model defines a high-dimensional state space according to the data collected by the sensor, and the high-dimensional state space includes multiple state parameters, and all the state parameters jointly reflect the real-time operating state of the CDQ system; The reinforcement learning module adopts a deep reinforcement learning algorithm, uses historical data and real-time data to train the model online, and obtains a preliminary control strategy through multiple interactions with the environment, providing a basis for the heuristic algorithm.

3. The system according to claim 2, characterized in that The state parameters include but are not limited to coke temperature distribution, furnace pressure, circulating air volume, air introduction ratio and coke discharge speed.

4. The system according to claim 2, characterized in that The heuristic algorithm module has an embedded heuristic algorithm to globally optimize the key factors affecting system performance during the dry coke quenching process; the heuristic algorithm constructs a global objective function, comprehensively considers the key factors of system performance, and constructs the objective function through reasonable weight allocation and constraint setting.

5. The system according to claim 4, characterized in that The heuristic algorithm adopts a simulated annealing algorithm, which avoids local optimal solutions and searches for global optimal solutions by gradually reducing the search range. In the dry coke quenching process, the simulated annealing algorithm is used in combination with a genetic algorithm, a particle swarm algorithm, an ant colony algorithm, and a differential evolution algorithm.

6. The system according to claim 2, characterized in that The joint optimization module adopts an online collaborative mechanism to combine the online learning ability of the reinforcement learning model with the optimization results of the heuristic algorithm to jointly drive the iterative improvement of the control strategy.

7. A dry coke quenching RTO precise control method based on Transformer model, characterized in that: The method is implemented by using the system described in any one of claims 1 to 6, comprising the following steps: S1, the data acquisition module obtains the data in the CDQ oven, including coke temperature Tc(t), circulating gas temperature Tg(t), circulating gas flow Qg(t), coke discharge Mc(t), air introduction flow Qa(t), material level L(t), boiler inlet temperature Tb(t); S2, the time series prediction module uses the Transformer model to combine historical data and real-time data to predict the changing trend of key parameters in the CDQ process in the future. It specifically includes the following sub-steps: S2.1, the collected data are stored in time series format, and each variable is timestamped, i.e. ; S2.2, using the Min-Max normalization method, Normalize the data to the interval [0,1] to prevent convergence difficulties due to numerical differences during model training; S2.3, using the sliding window method, setting the time window length to , forming an input-output pair: ; S2.4, train the Transformer model to generate future trend predictions; S3, RTO real-time optimization algorithm module dynamically adjusts the circulation fan speed, air introduction volume, bypass valve opening, and vent valve opening according to the prediction results to optimize the coke cooling effect and energy recovery efficiency; The S4 anomaly detection module compares the predicted data with the actual data, detects potential anomalies in the CDQ process, and issues early warnings in a timely manner; The S5 optimization results are sent to the PLC control system through the edge computing module.

8. The method according to claim 7, characterized in that In step S2.4, the specific method of training the Transformer model is as follows: S2.4.1 Input data format: Transformer requires sequence input, and the input data format is The tensor structure of ; where Batch is the number of samples for one training, Sequence Length is the time window length, and Features is the feature dimension of multiple parameters; S2.4.2 Generate position codes using sine and cosine functions: ; ; in, represents the time step, Represents feature dimension; S2.4.3 Design model structure: Transformer consists of multiple encoders, each of which contains: Multi-head attention: capturing the correlation of different time steps in a time series; Feedforward neural network: used for feature transformation; Layer normalization: speeds up convergence; The self-attention is calculated by the following formula: ; in is the transformation matrix of the input data, is the characteristic dimension; S2.4.4 Training model: Using MSE as the loss function: ; The AdamW optimizer is used and the learning rate is set to , weight decay ; S2.4.5 Input the latest time window data, Transformer processes the data, calculates multi-head attention, outputs the predicted value, and generates future trend predictions.

9. The method according to claim 8, characterized in that In step S2.4.4, the training steps of the training model are specifically as follows: S2.4.4.1 Input historical time series data into Transformer; S2.4.4.2 Calculate predicted values ; S2.4.4.3 Calculate the loss function and backpropagate; S2.4.4.4 Update model weights; S2.4.4.5 Iterate training until convergence.

10. The method according to claim 7, characterized in that In step S3, the RTO real-time optimization algorithm module dynamically adjusts the circulating fan speed, air introduction amount, bypass valve opening, and vent valve opening according to the prediction results to optimize the coke cooling effect and energy recovery efficiency, which specifically includes the following sub-steps: S3.1 In the CDQ process, confirm the core control parameters, including: Circulating air Qc: controls the circulation volume of hot gas in the CDQ furnace; Air intake Qa: used to adjust the oxygen content of the system; Bypass vent opening Bp: adjusts whether low-temperature gas enters the boiler directly; Release valve opening Vd: used for system pressure regulation; S3.2 Design the state space of reinforcement learning: According to the core control parameters described in step S3.1, the reinforcement learning model defines a high-dimensional state space based on the data collected by the sensor; the state space contains variables that affect the CDQ energy consumption and quality, including the circulating air flow Qc(t), the air introduction amount Qa(t), the bypass air valve opening Bp(t), and the relief valve opening Vd(t); S3.3 Design the action space of reinforcement learning: The action space is the actions that the system can take, which is the basis for the reinforcement learning model to optimize the control strategy by interacting with the environment; the reinforcement learning algorithm adjusts the circulating air flow Qc(t) to optimize the uniformity of coke cooling, adjusts the air introduction Qa(t) to maintain the oxygen content, adjusts the bypass air valve opening Bp(t) to control the proportion of low-temperature gas entering the boiler, and adjusts the vent valve opening Vd(t) to maintain pressure stability; S3.4 Setting the reward function: The optimization goal is to improve energy recovery and maintain coke quality. Since the reward function is the result of weighting the dimensions of each variable, and the dimensions and magnitudes of each variable are different, they are normalized to the same scale range [0, 1] to obtain the weighted reward function R: ; in: : Energy recovery efficiency, which is mainly determined by the boiler steam output and exhaust gas temperature; : Coke quality, including cooling uniformity and final coke strength; : Fan energy consumption, which is related to the circulation wind speed and fan speed; : Boiler energy consumption, depends on exhaust gas temperature and heat exchange efficiency; : Over-temperature penalty to avoid safety problems caused by boiler overheating; Reward Weight Can be adjusted experimentally to ensure optimal target balance; S3.5 Reinforcement learning online adjustment: Use deep reinforcement learning to dynamically adjust the core control parameters, calculate the reward value, and update the strategy based on real-time data.

Citation Information

Patent Citations

  • Speech recognition network and method based on local information fusion of Transform model, and terminal

    CN114333824A

  • Reservoir prediction method based on cross attention diffusion model

    CN117236390A

  • Dry quenching burn-out rate optimization method based on particle swarm optimization

    CN118709538A

  • Multi-unmanned battle vessel cooperative control optimization method based on machine learning and improved genetic algorithm

    CN119088008A

  • New material production process parameter optimization method and system based on artificial intelligence

    CN119397914A

Cited By

  • Kelp drying energy consumption optimization method and system combined with big data comprehensive analysis

    CN120337783A

  • Coke oven small flue intelligent temperature control system based on Internet of Things

    CN121254947A