Plasma gas heating-reforming intelligent regulation and control system and method based on deep reinforcement learning and storage medium

The intelligent control system for plasma gas heating and reforming based on deep reinforcement learning solves the problems of high energy consumption and complex control in existing technologies, realizes efficient and stable gas heating and reforming processes, adapts to fluctuations in gas composition and flow, and reduces CO2 emissions.

CN121578633APending Publication Date: 2026-02-27UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511451777.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing coal gas heating technologies are energy-intensive, accompanied by secondary CO2 emissions, and have limited temperature increases, making it difficult to meet the requirements of low-carbon processes for high temperature and high reduction potential. Traditional reforming control is complex and difficult to adapt to fluctuations in coal gas composition and flow rate.

Method used

A plasma gas heating-reforming intelligent control system based on deep reinforcement learning is adopted. By integrating artificial intelligence control technology with AI self-learning algorithms, gas distribution, plasma heating, sensing and monitoring, and data acquisition and communication modules are constructed to achieve dynamic optimization of process parameters.

Benefits of technology

It achieves full-process adaptive and intelligent control under complex working conditions, reduces energy consumption, improves system stability and control efficiency, reduces dependence on mechanism models, and enhances system generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121578633A_ABST
    Figure CN121578633A_ABST
Patent Text Reader

Abstract

The invention discloses a plasma gas heating-reforming intelligent regulation and control system and method based on deep reinforcement learning and a storage medium, and relates to the technical field of metallurgical gas treatment. The system comprises a gas distribution module, a plasma heating module, a sensing and monitoring module, a data acquisition and communication module and an AI intelligent control module, and dynamic optimization of process parameters is realized through a plasma heating-reforming integrated system integrating an AI self-learning algorithm and an artificial intelligence control technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of metallurgical gas treatment technology, specifically designing an intelligent control system, method, and storage medium for plasma gas heating and reforming based on deep reinforcement learning. Background Technology

[0002] Currently, mainstream low-carbon blast furnace technologies such as top gas recirculation (TGR-BF), hydrogen-rich smelting, and hydrogen-based direct reduction processes all rely on high-temperature (typically ≥1100℃) and high-reduction-potential (high H2 / CO ratio) reducing gas as a key support. The sources of high-temperature, high-reduction-potential gas are mainly twofold: one is the purification and reheating of metallurgical top gas (such as BFG and COG); the other is the catalytic reforming and reheating of multiple gases (such as COG and converter gas LDG). However, existing gas heating technologies largely rely on traditional combustion methods, which suffer from bottlenecks such as high energy consumption, secondary CO2 emissions, and limited temperature increases (typically ≤900℃), making it difficult to meet the stringent requirements of low-carbon processes for high-temperature and high-reduction-potential gas. For reforming processes, to protect the catalyst, traditional reforming temperatures are typically low (800-900℃), and secondary heating of the syngas is still required after reforming, significantly increasing system complexity and carbon emissions. Although high-temperature non-catalytic reforming (≥1300℃) can achieve high conversion rates, its control is extremely complex and the operating conditions fluctuate greatly, making it difficult for traditional PID control to achieve stable and optimized operation.

[0003] In existing technologies, control methods mostly rely on preset logic, which cannot adapt to changes in gas composition, flow rate fluctuations, and the state of the plasma torch itself. Therefore, there is an urgent need for a plasma heating reforming system that can intelligently sense changes in operating conditions, make autonomous decisions, and continuously learn and optimize. Summary of the Invention

[0004] To address the above issues, the present invention provides an intelligent control system, method, and storage medium for plasma gas heating and reforming based on deep reinforcement learning. This integrated plasma heating and reforming system, which integrates AI self-learning algorithms with artificial intelligence control technology, achieves dynamic optimization of process parameters.

[0005] According to a first aspect of the technical solution of the present invention, a plasma gas heating-reforming intelligent control system based on deep reinforcement learning is provided, comprising: a gas distribution module, a plasma heating module, a sensing and monitoring module, a data acquisition and communication module, and an AI intelligent control module, wherein,

[0006] The gas distribution module is used to receive various types of coal gas from the front end and distribute them proportionally to the working gas path and the temperature regulating gas path, including: an inlet pipe, a working gas mixing tank, a temperature regulating gas mixing tank, a nitrogen compression tank, a working gas path control valve, and a temperature regulating gas path control valve.

[0007] The plasma heating module, connected to the gas distribution module, is used to generate a high-temperature plasma arc and heat or reform the gas, and includes: a plasma power torch, a plasma sealed heating furnace, and a gas outlet pipe.

[0008] The sensing and monitoring module is used to collect data from the entire system process in real time, including: temperature sensor, gas composition analyzer, flow sensor and electrical parameter sensor;

[0009] The data acquisition and communication module is connected to the sensing and monitoring module and is used to receive real-time data and upload it to the AI ​​decision-making intelligent control module.

[0010] The AI ​​decision-making intelligent control module is connected to the gas distribution module and the plasma heating module. It is used to extract key feature vectors based on real-time data and calculate system state vectors based on a deep reinforcement learning (DRL) model. Then, based on the system state vectors and optimization objectives, it outputs control commands for the gas distribution module and the plasma heating module. The module includes: a state perception and evaluation unit, a decision-making unit based on a deep reinforcement learning model, and a process knowledge base.

[0011] Furthermore, the various gases from the front end include blast furnace gas (BFG), coke oven gas (COG), and converter gas (LDG).

[0012] Furthermore, the air inlet pipe is connected to the working gas mixing tank and the temperature regulating gas mixing tank, and the nitrogen compression tank and the working gas mixing tank are connected to the plasma sealing heating furnace, forming a working gas supply path; the nitrogen compression tank and the temperature regulating gas mixing tank are connected to the plasma sealing heating furnace, forming a temperature regulating gas supply path, and the working gas path control valve and the temperature regulating gas path control valve are respectively installed on the working gas supply path and the temperature regulating gas supply path.

[0013] Furthermore, the plasma sealed heating furnace is provided with a working gas inlet and a temperature regulating gas inlet, which respectively receive working gas and temperature regulating gas from the gas distribution module.

[0014] Furthermore, in the sensing and monitoring module:

[0015] The temperature sensor is located near the outlet of the plasma power torch and the exhaust pipe inside the plasma sealed heating furnace, and is used to monitor the temperature at key points.

[0016] The gas composition analyzer extracts gas samples from the plasma sealed heating furnace through a gas pump and a cooling pipe for online composition analysis.

[0017] The flow sensor is installed on the working gas supply path and the temperature regulating gas supply path to monitor the real-time flow.

[0018] The electrical parameter sensor is integrated into the plasma power cabinet and is used to monitor the voltage, current, and power of the plasma power cabinet.

[0019] Furthermore, the gas composition analyzer analyzes components including H2, CO, CO2, CH4, and N2.

[0020] Furthermore, after receiving real-time data, the data acquisition and communication module also performs cleaning, preprocessing, and temporary storage.

[0021] Furthermore, the data acquisition and communication module uploads real-time data to the AI ​​decision-making intelligent control module via industrial Ethernet.

[0022] Furthermore, the AI ​​decision-making intelligent control module is connected to the working gas path control valve and the temperature regulating gas path control valve of the gas distribution module, and is also connected to the plasma power torch of the plasma heating module.

[0023] Furthermore, the state vector is expressed as:

[0024] S(t) = [T, C, F, P, Δ];

[0025] Where T is the temperature vector, C is the composition vector (H2%, CO%, N2%, CO2%, CH4%), P is the electrical power vector, and Δ is the rate of change vector of the key parameters.

[0026] Furthermore, in the AI ​​decision-making intelligent control module,

[0027] The state perception and evaluation unit is used to receive the raw data stream, extract key feature vectors from the multi-dimensional data after data processing, and calculate the system state vector S(t) in real time.

[0028] The decision-making unit based on the deep reinforcement learning model adopts the TD3 (TwinDelayed Deep Deterministic Policy Gradient) model based on the Actor-Critic framework. In this model, the Actor network outputs the optimal control action A(t) based on the system state vector S(t), and the value (Q value) of the action is evaluated by two independent Critic networks. The minimum value of the two is taken as the target Q value to update the policy.

[0029] The process knowledge base is used to store historical optimal strategies, abnormal operating condition handling cases, and DRL network parameters. The process knowledge base is continuously expanded and updated over time to form system memory.

[0030] Furthermore, the AI ​​decision-making intelligent control module includes:

[0031] State space (S): Defined by eigenvectors, defined as state vector S(t), which covers all key parameters affecting the process;

[0032] Action space (A): Control commands output by the engine, A = [ΔF] work ,ΔF temper ,ΔP plasma This refers to the adjustment amount of the working airflow, the adjustment amount of the temperature-regulating airflow, and the adjustment amount of the plasma torch power; the action output is a continuous value.

[0033] Reward function (R): R = w1 * (T) actual -T target ) 2 +w2*(H2 / CO actual -H2 / CO target ) 2 -w3*P power -w4*|ΔAction|, where T actual T represents the actual temperature. target Target temperature; H2 / CO actual This represents the actual H2 / CO ratio. target The target H2 / CO ratio; P power The power is represented by w1, w2, w3, and w4, which are weighting coefficients whose values ​​are adjusted according to the actual process. |ΔAction| represents the magnitude of the change in action and is used to penalize excessively drastic control actions to ensure system stability.

[0034] According to a second aspect of the present invention, a method for intelligent control of plasma gas heating and reforming based on deep reinforcement learning is provided. The method operates based on a system according to any one of the above aspects, and the method includes:

[0035] S1: System initialization: Start the device and load the initial deep reinforcement learning model and process knowledge base of the AI ​​intelligent control module;

[0036] S2: Target setting: Input or receive the target process parameters for this operation from the upstream system, including target temperature T, target composition C (or target H2 / CO ratio), and target total flow rate F;

[0037] S3: Data Acquisition: Acquire real-time data, preprocess it, and then upload it to the AI ​​intelligent control module;

[0038] S4: Intelligent Decision Making: Extract key feature vectors from real-time data and calculate the system state vector S(t).

[0039] Then, based on the system state vector S(t), the optimal control action A(t) is output;

[0040] S5: Action Execution: Convert the optimal control action A(t) into control commands and send them to each high-precision flow control valve and plasma power supply cabinet for execution;

[0041] S6: Knowledge Base Update: Update the stable and efficient optimal process points and their corresponding states discovered during operation.

[0042] - Action pairs (S,A) are stored in the process knowledge base. When a state similar to that in the knowledge base is encountered, the best historical action can be retrieved from the knowledge base first, thus accelerating decision-making efficiency.

[0043] S7: Iterative Optimization. Repeat steps S3-S7 to form a closed loop of "perception-decision-execution-learning".

[0044] This allows the system to continuously approach and maintain its optimal operating point.

[0045] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored thereon, and when the computer program is executed by a processor, it implements the intelligent control method for plasma gas heating-reforming based on deep reinforcement learning as described in any of the preceding aspects.

[0046] Compared with the prior art, the present invention has the following technical effects:

[0047] 1. Application of TD3 Algorithm Based on Actor-Critic Framework: This invention introduces a dual-delay deep deterministic policy gradient (TD3) model based on the Actor-Critic architecture, enabling the system to possess stronger policy stability and anti-interference capabilities. This model alleviates the overestimation problem through a dual value network and, combined with a policy delay update mechanism, allows the agent to more robustly learn the optimal control policy in the continuous action space. This achieves full-process adaptive and intelligent control under complex operating conditions such as fluctuating gas composition and changes in flow rate.

[0048] 2. Coordinated design of state space, action space and reward function: This invention constructs a multi-dimensional state space (covering multi-source sensor information such as temperature, composition, and flow rate) and a continuous action space (corresponding to the adjustment commands of the actuator), and designs a reward function that integrates multiple objectives (such as temperature stability, composition compliance, energy consumption reduction, and operational stability). This enables the reinforcement learning agent to optimize multiple key indicators simultaneously during training, thereby achieving a comprehensive optimal solution rather than single-objective optimization, significantly improving the overall system performance and quality control level.

[0049] 3. Structural improvements to the system architecture: This invention constructs a data-driven and model-coupled intelligent control architecture, integrates the TD3 algorithm with the actual process system in a closed loop, and embeds a real-time sensing and decision-making module. This allows the system to sense environmental dynamics and output control commands without relying on precise physical or chemical reaction models, thereby significantly reducing the dependence on mechanistic models, improving the system's generalization ability, and reducing development and application costs.

[0050] 4. Full-process intelligent decision-making mechanism: Through an end-to-end training and deployment framework, this invention enables the system to autonomously learn control strategies from historical data and real-time interactions, and respond to disturbances and changes online, thereby quickly approaching the optimal operating point and suppressing production fluctuations. This achieves a dual improvement in control efficiency and stability, while reducing reliance on human experience, providing a feasible technical path for intelligent manufacturing and "lights-out factories". Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0052] Figure 1 This is a schematic diagram of the system structure of an AI self-learning plasma intelligent gas heating device provided in an embodiment of the present invention;

[0053] Figure 2 This is a comparison chart showing the load disturbance resistance effects of the embodiments of the present invention and existing technical solutions;

[0054] Figure 3 This is a comparison chart showing the anti-component disturbance effect of the embodiments of the present invention and the prior art;

[0055] Figure 4 This is a flowchart comparing the embodiments of the present invention with existing technical solutions.

[0056] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0058] The terms "first," "second," etc., used in this disclosure are for distinguishing similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein.

[0059] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0060] Multiple, including two or more.

[0061] And / or, it should be understood that, for the purposes of this disclosure, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0062] The technical solution of this invention first provides an intelligent control system for plasma gas heating and reforming based on deep reinforcement learning. Figure 1 The diagram showcases its overall hardware architecture. It illustrates the physical connections between the gas distribution module (including mixing tank, valves, and flow meters), the plasma heating module (plasma torch and heating furnace), the multi-sensor monitoring module (temperature, composition, flow, and electrical parameter sensors), the edge computing gateway, and the AI ​​decision control center, clearly demonstrating the closed loop of data flow and control flow.

[0063] As shown in Figure 1, it includes: a gas distribution module, a plasma heating module, a sensing and monitoring module, a data acquisition and communication module, and an AI intelligent control module.

[0064] The gas distribution module is used to receive various types of coal gas (such as blast furnace gas BFG, coke oven gas COG, and converter gas LDG) from the front end and distribute them proportionally to the working gas path and the temperature regulating gas path; it includes an inlet pipe, a working gas mixing tank, a temperature regulating gas mixing tank, a nitrogen compression tank, and high-precision flow control valves located on the working gas path and the temperature regulating gas path respectively.

[0065] The plasma heating module, connected to the gas distribution module, is used to generate a high-temperature plasma arc and heat or reform the gas; it includes a plasma torch, a sealed heating furnace, and a gas outlet pipe; the plasma heating furnace is provided with a working gas inlet and a temperature-regulating gas inlet, which respectively receive working gas and temperature-regulating gas from the gas distribution module.

[0066] The sensing and monitoring module is used to collect data from the entire system process in real time. It includes: a temperature sensor located at the plasma torch outlet, the sealed heating furnace, and the gas outlet pipe, used to monitor the temperature at key points; a gas composition analyzer that extracts gas samples from the sealed heating furnace via a pump and cooling pipe for online component analysis (H2, CO, CO2, CH4, N2, etc.); a flow sensor installed in each gas path to monitor real-time flow; and an electrical parameter sensor integrated into the plasma power cabinet to monitor the voltage, current, and power of the plasma torch.

[0067] The data acquisition and communication module is connected to all sensors of the multi-sensor monitoring module. It is used to receive, clean, preprocess and temporarily store real-time data, and upload the processed data to the AI ​​decision control center via industrial Ethernet.

[0068] The AI ​​decision-making intelligent control module, as the core of the system, includes:

[0069] The state awareness and evaluation unit receives the raw data stream, performs data cleaning (outlier removal), filtering, smoothing, and normalization, and extracts key feature vectors from multi-dimensional data. Based on multi-sensor data, it calculates the system state vector S(t) in real time. This state vector S(t) = [T, C, F, P, Δ], where T is the temperature vector, C is the composition vector (H2%, CO%, N2%, CO2%, CH4%), F is the flow rate vector, P is the electrical power vector, and Δ is the rate of change vector of the key parameters. These features together constitute the state vector describing the current state of the system.

[0070] The decision-making unit based on a deep reinforcement learning model, employing the TD3 (TwinDelayed Deep Deterministic Policy Gradient) model based on the Actor-Critic framework, is the core of intelligent decision-making. It receives a state vector S(t) and outputs a control action A(t).

[0071] State space (S): Defined by eigenvectors, defined as state vector S(t), which encompasses all key parameters affecting the process;

[0072] Action space (A): Control commands output by the engine, A = [ΔF] work ,ΔF temper ,ΔP plasmaThis refers to the adjustment amounts of the working airflow, the temperature-regulating airflow, and the plasma torch power. The output is a continuous value.

[0073] Reward function (R): Designed as a multi-objective optimization function, it combines with the aforementioned state vector and control commands, simultaneously pursuing high efficiency, high quality, low energy consumption, and stable operation. R = w1*(T) actual -T target ) 2 +w2*(H2 / CO actual -H2 / CO target ) 2 -w3*P power -w4*|ΔAction|. Where, T actual T represents the actual temperature. target Target temperature; H2 / CO actual This represents the actual H2 / CO ratio. target The target H2 / CO ratio; P power The power is represented by w1, w2, w3, and w4, which are weighting coefficients whose values ​​are adjusted according to the actual process. |ΔAction| represents the magnitude of the change in action and is used to penalize excessively drastic control actions to ensure system stability.

[0074] Based on the system state provided by S(t), the TD3 agent generates a corresponding action A(t) to adjust the process operation. This action calculates an immediate reward value through the reward function R(t), thereby guiding the policy model to learn how to coordinate multiple control variables and optimize multiple competing objectives under complex operating conditions. Thus, the system not only eliminates its reliance on precise mechanistic models but also achieves highly autonomous and intelligent control in a variable environment, maintaining an ideal H2 / CO ratio, stabilizing temperature distribution, and reducing energy consumption and operational fluctuations. This effectively solves the core control challenges in plasma gas reforming processes, such as multivariate coupling, nonlinearity, and disturbance sensitivity.

[0075] The above TD3 (Twin Delayed Deep Deterministic Policy Gradient) model based on the Actor-Critic framework outputs the optimal control action A(t) based on the system state vector S(t) through the Actor network, and evaluates the value (Q value) of the action through two independent Critic networks. The minimum of the two values ​​is taken as the target Q value to update the policy.

[0076] Process knowledge base: Used to store historical best strategies, abnormal operating condition handling cases, and DRL network parameters. The knowledge base is continuously expanded and updated over time, forming system memory.

[0077] Therefore, it can be seen that the gas distribution module is connected to the front end of the plasma heating module; the sensors of the sensing and monitoring module are distributed at key nodes of the plasma heating module; the data acquisition and communication module is connected to all sensing and monitoring modules; and the AI ​​decision control center is connected to the data acquisition and communication module and all actuators (flow control valves, power cabinets) to form a closed-loop control.

[0078] The present invention also provides a method for intelligent control of plasma gas heating and reforming based on deep reinforcement learning. The method operates based on a system according to any of the above aspects, and includes:

[0079] S1: System initialization: Start the device and load the initial deep reinforcement learning model and process knowledge base of the AI ​​intelligent control module;

[0080] S2: Target setting: Input or receive the target process parameters for this operation from the upstream system, including target temperature T, target composition C (or target H2 / CO ratio), and target total flow rate F;

[0081] S3: Data Acquisition: Acquire real-time data, preprocess it, and then upload it to the AI ​​intelligent control module;

[0082] S4: Intelligent Decision Making: Extract key feature vectors from real-time data and calculate the system state vector S(t).

[0083] Then, based on the system state vector S(t), the optimal control action A(t) is output;

[0084] S5: Action Execution: Convert the optimal control action A(t) into control commands and send them to each high-precision flow control valve and plasma power supply cabinet for execution;

[0085] S6: Knowledge Base Update: Update the stable and efficient optimal process points and their corresponding states discovered during operation.

[0086] - Action pairs (S,A) are stored in the process knowledge base. When a state similar to that in the knowledge base is encountered, the best historical action can be retrieved from the knowledge base first, thus accelerating decision-making efficiency.

[0087] S7: Iterative Optimization. Repeat steps S3-S7 to form a closed loop of "perception-decision-execution-learning".

[0088] This allows the system to continuously approach and maintain its optimal operating point.

[0089] The present invention also provides a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, it implements the intelligent control method for plasma gas heating-reforming based on deep reinforcement learning as described in any of the above aspects.

[0090] Example 1: Intelligent Control of Pure Heating of Blast Furnace Gas Based on AI Self-Learning (To Cope with Load Disturbances)

[0091] Application Scenario and Purpose: A steel plant needs to heat hydrogen-rich blast furnace gas (BFG, composition shown in Table 1) from room temperature to 1100℃ before injecting it back into the blast furnace. The core objective of this embodiment is to address step disturbances in the upstream gas main flow rate by using AI intelligent control to quickly stabilize the outlet temperature at the target value (1100℃) while maintaining a constant gas composition, and to optimize system power consumption during this process.

[0092] Table 1. Composition (volume fraction) of hydrogen-rich blast furnace gas (BFG)

[0093]

[0094] Traditional control bottlenecks: Traditional PID control relies on fixed PID parameters. When the flow rate changes drastically, the adjustment process is slow, the temperature overshoot or undershoot is serious, the stability is poor, and energy-saving optimization cannot be achieved.

[0095] The intelligent control process of this invention:

[0096] Initialization and Target Setting: Upon system startup, the AI ​​decision control center loads the initial DRL model (Actor-Critic network parameters) and process knowledge base. Target parameters are set: T target =1100℃, total flow rate F target Initially 50000 Nm 3 / h (specified by the upstream).

[0097] Steady-state operation: The system operates smoothly. The state vector S(t) includes: [T 出口 =1100,T 炬口 =3500,...,C H2 =15,C CO =53,...,F 工作气 =25000,F 调温气 =25000,P 功率 =P1,ΔF=0,ΔT=0,...]. The AI ​​decision engine outputs action A(t)=[0,0,0], which means maintaining the current action.

[0098] Disturbance occurred: Upstream blast furnace operation caused a sudden 20% increase in the total flow rate of the BFG. target It becomes 60000Nm 3 / h. The higher gas flow rate disrupts the system's thermal equilibrium, and the outlet temperature begins to drop.

[0099] AI intelligent response:

[0100] State awareness: The data acquisition and communication module captures sudden changes in flow meter readings in real time. Multiple temperature sensors show that the outlet temperature decreases at a rate of ΔT / Δt = -2.5℃ / s. The state vector S(t) is rapidly updated, with significant changes in the flow rate F and the rate of change Δ vector.

[0101] Intelligent Decision Making: The Actor network in the AI ​​decision engine receives a new S(t). A query of the process knowledge base did not find a flow rate of 60000 Nm. 3 A complete match record of / h. The DRL model performs online calculations based on the current strategy, outputting the optimal control action A(t) = [+7000, +3000, +15%] within 1 second. That is: increase the working gas flow rate by 7000 Nm. 3 / h (up to 32000), temperature regulating gas flow rate increased by 3000 Nm 3 / h (up to 28000), plasma torch power increased by 15%. This action ensures the total flow requirement while quickly replenishing energy by adjusting the gas ratio and power.

[0102] Action execution: Control commands are immediately issued to the high-precision flow control valve and plasma power supply cabinet for execution.

[0103] Effectiveness evaluation and learning:

[0104] Effect Evaluation: After implementing the new action, the system outlet temperature experienced a brief, slight dip (to 1093℃) before rapidly recovering and eventually stabilizing within the range of 1100℃±2℃. The reward function R(t) was calculated: Due to the rapid reduction and eventual elimination of the temperature deviation, -w1*(T) actual -T target ) 2 This item contributed the main positive reward; although the increase in power resulted in -w3*P power The item is negative, but the overall reward value is positive.

[0105] Online learning: This experience tuple (S(t), A(t), R(t), S(t+1)) is stored in the experience replay pool. The policy learning and update unit uses this experience data to train and update the Critic and Actor networks, strengthening the policy to take effective actions under such perturbations.

[0106] Knowledge base update: The state S and action A under this stable operating condition are stored as a new optimal process point in the process knowledge base for future use in similar operating conditions. See Table 2 for details. Figure 2 As shown.

[0107] Table 2 Comparison of the effects of traditional PID control and the intelligent control method of this invention.

[0108]

[0109] The results show that, compared with traditional PID control which has significant overshoot (+15℃), undershoot (-30℃) and multiple oscillations, the intelligent control method of this invention has significant advantages in dealing with flow step disturbances: the system responds quickly (stabilizes within 60 seconds, speed increase of about 67%), temperature fluctuations are minimal (maximum drop of only 7℃, no overshoot), and the control accuracy is high after stabilization (±1℃). It also has online learning and multi-objective optimization capabilities, realizing a fundamental leap from "passive response" to "active perception-decision-optimization", and greatly improving the dynamic quality and overall energy efficiency of the system.

[0110] Example 2: Intelligent Control of Coke Oven Gas-Converter Gas Reforming Based on AI Self-Learning (To Cope with Composition Disturbances)

[0111] Application Scenario and Purpose: To mix coke oven gas (COG, rich in CH4 and H2) and converter gas (LDG, rich in CO and CO2) in a certain proportion and reform them at ≥1300℃ to produce reducing gas for direct reduction processes (target H2 / CO≈2.5). The core objective is to simultaneously and stably control the outlet temperature and syngas composition (H2 / CO ratio) when the coke oven gas composition fluctuates drastically, achieving multi-objective synergistic optimization.

[0112] Traditional control bottleneck: Composition fluctuations are complex disturbances that cannot be handled by traditional PID control. Manual adjustments are blind and lagging, which can easily lead to runaway reaction temperature and catalyst deactivation (if present) or a sharp drop in conversion rate, resulting in a high risk of production interruption.

[0113] The intelligent control process of this invention:

[0114] Steady-state operation: The system operates stably. COG flow rate: 10000 Nm³ 3 / h (CH4 = 24%), LDG flow rate 33000 Nm 3 / h, temperature 1300℃, H2 / CO ratio = 2.5, CH4 conversion rate >96%.

[0115] Disturbance occurs: The CH4 concentration in the coke oven gas suddenly drops to 18% vol (with a corresponding increase in H2 concentration). This directly alters the reactant concentration, the calorific value of the mixed gas, and the heat endothermic reaction.

[0116] AI intelligent response:

[0117] State awareness: The laser gas analyzer detected a sharp drop in CH4 concentration and a sharp rise in H2 concentration within seconds. This drastic change was immediately reflected in the composition vector C and the rate of change vector Δ of the state vector S(t). The outlet temperature sensor also detected a decreasing trend in temperature due to the increase in endothermic reaction (ΔT / Δt is negative).

[0118] Intelligent Decision Making: The AI ​​decision engine receives S(t) containing a sudden change in composition. A query of the process knowledge base reveals no perfectly matching cases. The Actor network performs emergency inference based on the learned strategy, calculating a multivariable decoupling control action:

[0119] A(t)=[ΔF work =+1500,ΔF temper =-1000,ΔP plasma =+12%]

[0120] Meaning: Increase the flow rate of hydrogen-rich COG working gas (to 11500 Nm). 3 / h) to compensate for calorific value and utilize the high enthalpy of H2; appropriately reduce the LDG temperature regulating gas flow rate (to 32000 Nm³). 3 / h) to fine-tune CO2 supply and total material balance; simultaneously increase plasma torch power to provide additional reaction energy.

[0121] Action execution: When instructions are issued, each executing mechanism responds quickly.

[0122] Effectiveness evaluation and learning:

[0123] Performance Evaluation: System status changes were as follows: the outlet temperature quickly returned to 1300℃ after a brief fluctuation of ±20℃; gas analysis showed that the H2 / CO ratio remained stable at 2.6, and the CH4 conversion rate remained at 95.8%. The reward function R(t) was calculated as: temperature and target component deviation terms (-w1*(T)). actual -T target ) 2 -w2*(H2 / CO actual -H2 / CO target ) 2 The value rapidly approaches zero from a negative value; the power term (-w3*P) power The value of -w4*|ΔAction| is negative; the motion smoothing term (-w4*|ΔAction|) is negative due to the large motion amplitude. However, the overall reward value is positive because it successfully resolved the crisis.

[0124] Online learning and knowledge base updates: This successful case of handling extreme component disturbances is considered valuable experience by the system. Complete data is stored in the experience replay pool for network training, and this "strategy for handling low CH4 concentration conditions" is stored as a classic case in the process knowledge base, forming part of the system's memory. See below for details. Figure 3 And Table 3.

[0125] Table 3 Comparison of the effects of traditional PID control and the intelligent control method of this invention.

[0126]

[0127] Compared to the passive situation of traditional control systems, which suffer from severe instability of the H2 / CO ratio (as low as 1.75) and drastic temperature fluctuations (up to -55℃) under component disturbances and require manual intervention, the intelligent control system of this invention, through its multivariate decoupling and autonomous decision-making capabilities, completes strategy response within 1 minute and achieves full parameter stabilization within 15 minutes. It precisely maintains the H2 / CO ratio in the range of 2.56–2.60 (reducing the deviation by 85%), controls temperature fluctuations within 10℃, and maintains a CH4 conversion rate of over 95% without any manual intervention throughout the process. This achieves highly adaptive control of complex component disturbances and fully autonomous optimization of the production process.

[0128] Figure 4 A flowchart comparing the present invention with traditional control methods is presented. Compared to the open-loop model of traditional control schemes ("disturbance → manual monitoring → experience-based adjustment → passive response"), the present invention constructs an intelligent closed loop of "perception → decision-making → execution → learning": real-time state perception drives the TD3 algorithm for decision-making, relying on a process knowledge base to achieve microsecond-level strategy invocation, forming a continuously optimized system memory. This architectural innovation transforms the system from a lagging control reliant on human experience into an intelligent agent with multi-objective autonomous optimization, strong anti-disturbance capabilities, and continuous evolution, fundamentally solving the industrial bottlenecks of traditional control, such as slow response, large fluctuations, and the inability to balance energy efficiency and quality.

[0129] Compared with existing intelligent control methods, the TD3 algorithm based on the Actor-Critic framework adopted in this invention not only effectively overcomes the problem of overestimation of Q-value in algorithms such as DDPG through a dual Critic network, target policy smoothing, and delayed update mechanism, significantly improving policy stability and convergence efficiency—reducing control fluctuation by about 40% and increasing convergence speed by 25% within the same training period—but also demonstrates comprehensive advantages in multi-objective optimization, system generalization, and real-time decision-making: the system can reduce energy consumption by 12% to 18% while ensuring that the H2 / CO ratio deviation is ≤ ±0.05, and still keep the control effect fluctuation within ±3% when the gas composition fluctuates by ±15%. At the same time, it realizes microsecond-level real-time policy invocation with the help of the process knowledge base mechanism, which greatly improves the control accuracy, adaptability, and operating efficiency in complex industrial environments.

[0130] In summary, Examples 1 and 2 clearly demonstrate how the present invention utilizes its advanced AI self-learning core to perform intelligent, rapid, and multi-objective optimized closed-loop control in two modes—pure heating and heating-reforming—when facing two typical industrial scenarios: flow disturbance and composition disturbance. This fully reflects the innovation and practicality of the invention.

[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0132] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that the above implementation methods can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0134] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A deep reinforcement learning-based intelligent control system for plasma coal gas heating-reforming, characterized in that, include: The module includes a gas distribution module, a plasma heating module, a sensing and monitoring module, a data acquisition and communication module, and an AI intelligent control module. The gas distribution module is used to receive various types of coal gas from the front end and distribute them proportionally to the working gas path and the temperature regulating gas path, including: an inlet pipe, a working gas mixing tank, a temperature regulating gas mixing tank, a nitrogen compression tank, a working gas path control valve, and a temperature regulating gas path control valve. The plasma heating module, connected to the gas distribution module, is used to generate a high-temperature plasma arc and heat or reform the gas, and includes: a plasma power torch, a plasma sealed heating furnace, and a gas outlet pipe. The sensing and monitoring module is used to collect data from the entire system process in real time, including: temperature sensor, gas composition analyzer, flow sensor and electrical parameter sensor; The data acquisition and communication module is connected to the sensing and monitoring module and is used to receive real-time data and upload it to the AI ​​decision-making intelligent control module. The AI ​​decision-making intelligent control module is connected to the gas distribution module and the plasma heating module. It is used to extract key feature vectors based on real-time data using a deep reinforcement learning model and calculate the system state vector. Then, based on the system state vector and the optimization objective, it outputs control commands for the gas distribution module and the plasma heating module. The module includes: a state perception and evaluation unit, a decision-making unit based on a deep reinforcement learning model, and a process knowledge base.

2. The intelligent control system for plasma gas heating and reforming according to claim 1, characterized in that, The air inlet pipe is connected to the working gas mixing tank and the temperature regulating gas mixing tank. The nitrogen compression tank and the working gas mixing tank are connected to the plasma sealing heating furnace, forming a working gas supply path. The nitrogen compression tank and the temperature regulating gas mixing tank are connected to the plasma sealing heating furnace, forming a temperature regulating gas supply path. The working gas path control valve and the temperature regulating gas path control valve are respectively installed on the working gas supply path and the temperature regulating gas supply path.

3. The intelligent control system for plasma gas heating and reforming according to claim 1, characterized in that, The plasma power torch is equipped with a working gas inlet and a temperature regulating gas inlet, which respectively receive working gas and temperature regulating gas from the gas distribution module.

4. The intelligent control system for plasma gas heating and reforming according to claim 1, characterized in that, In the sensing and monitoring module: The temperature sensor is located at the outlet of the plasma power torch, outside the plasma sealed heating furnace, and near the gas outlet pipe, and is used to monitor the temperature at key points. The gas composition analyzer extracts gas samples from the plasma sealed heating furnace through a gas pump and a cooling pipe for online composition analysis. The flow sensor is installed on the working gas supply path and the temperature regulating gas supply path to monitor the real-time flow. The electrical parameter sensor is integrated into the plasma power cabinet and is used to monitor the voltage, current, and power of the plasma power cabinet.

5. The intelligent control system for plasma gas heating and reforming according to claim 1, characterized in that, The AI ​​decision-making intelligent control module is connected to the working gas path control valve and the temperature regulating gas path control valve of the gas distribution module, and is also connected to the plasma power torch of the plasma heating module.

6. The intelligent control system for plasma gas heating and reforming according to claim 1, characterized in that, The state vector is expressed as follows: S(t) = [T, C, F, P, Δ]; Where T is the temperature vector, C is the composition vector, F is the flow vector, P is the electrical power vector, and Δ is the rate of change vector of the key parameters.

7. The intelligent control system for plasma gas heating and reforming according to claim 6, characterized in that, In the AI ​​decision-making intelligent control module The state perception and evaluation unit is used to receive the raw data stream, extract key feature vectors from the multi-dimensional data after data processing, and calculate the system state vector S(t) in real time. The decision-making unit based on the deep reinforcement learning model adopts the Twin DelayedDeep Deterministic Policy Gradient model based on the Actor-Critic framework. In this model, the Actor network outputs the optimal control action A(t) based on the system state vector S(t), and the value of the action is evaluated by two independent Critic networks. The minimum value of the two is taken as the target Q value to update the policy. The process knowledge base is used to store historical optimal strategies, abnormal operating condition handling cases, and DRL network parameters. The process knowledge base is continuously expanded and updated over time to form system memory.

8. The intelligent control system for plasma gas heating and reforming according to claim 6, characterized in that, The AI ​​decision-making intelligent control module includes: State space S: defined by eigenvectors, defined as state vector S(t), which covers all key parameters affecting the process; Action space A: control instruction output by the engine, A = [ΔF work , ΔF temper , ΔP plasma ], that is, adjustment amount of working gas flow, adjustment amount of temperature-adjusting gas flow, adjustment amount of plasma torch power, and action output is a continuous value; Reward function R: R = w1*(T actual -T target ) 2 +w2*(H2 / CO actual -H2 / CO target ) 2 -w3*P power -w4*|ΔAction|, wherein, T actual is actual temperature, T target is target temperature; H2 / CO actual is actual H2 / CO ratio, H2 / CO target is target H2 / CO ratio; P power is power; w1, w2, w3, w4 are weight coefficients, whose values are adjusted according to actual process; |ΔAction| represents change amplitude of action, used to punish too drastic control action.

9. A method for intelligent control of plasma gas heating and reforming based on deep reinforcement learning, characterized in that, The method operates based on the system according to any one of claims 1 to 8, the method comprising: S1: System initialization: Start the device and load the initial deep reinforcement learning model and process knowledge base of the AI ​​intelligent control module; S2: Target setting: Input or receive the target process parameters for this operation from the upstream system, including target temperature T, target composition C or target H / CO ratio, and target total flow rate F; S3: Data Acquisition: Acquire real-time data, preprocess it, and then upload it to the AI ​​intelligent control module; S4: Intelligent Decision Making: Extract key feature vectors from real-time data and calculate the system state vector S(t). Then, based on the system state vector S(t), the optimal control action A(t) is output; S5: Action Execution: Convert the optimal control action A(t) into control commands and send them to each high-precision flow control valve and plasma power supply cabinet for execution; S6: Knowledge base update: Store the stable and efficient optimal process points and their corresponding state-action pairs (S,A) discovered during operation into the process knowledge base. When encountering a state similar to that in the knowledge base, prioritize calling the historical optimal action from the knowledge base to accelerate decision-making efficiency. S7: Iterative Optimization. Repeat steps S3-S7 to form a closed loop of "perception-decision-execution-learning". This allows the system to continuously approach and maintain its optimal operating point.

10. A computer-readable storage medium, wherein, It stores a computer program, characterized in that, when the computer program is executed by a processor, it implements the intelligent control method for plasma gas heating and reforming based on deep reinforcement learning as described in claim 9.