Expert strategy constrained blast furnace smelting safety reinforcement learning decision optimization method

By employing a safety reinforcement learning framework constrained by expert strategies, combined with a conditional generation adversarial mechanism and a state-action safety evaluation network, the safety and efficiency issues in blast furnace ironmaking operations are resolved, enabling intelligent operation and precise control of the blast furnace.

CN121167488APending Publication Date: 2025-12-19CENT SOUTH UNIV

Patent Information

Application Number
CN202511349624.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing optimization methods for blast furnace ironmaking operations are difficult to make safe and efficient decisions in the face of complex decision-making environments. Traditional methods rely on expert experience and have cognitive limitations. Furthermore, reinforcement learning algorithms cannot guarantee reliability and safety in blast furnace operations and lack explicit safety constraints.

Method used

A safe reinforcement learning framework with expert policy constraints is adopted. The policy distribution is aligned with expert decision-making through a conditional generation adversarial mechanism. A state-action safety evaluation network and a memory enhancement network are introduced. A triple guidance mechanism is designed for decision optimization, including distribution alignment, reward optimization and safety constraints.

Benefits of technology

It enables the learning of safe and effective operating strategies without online exploration, improving the safety and efficiency of blast furnace operation, providing refined control guidance, and ensuring the stable operation of the blast furnace and the quality of molten iron.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167488A_ABST
    Figure CN121167488A_ABST
Patent Text Reader

Abstract

The invention discloses a blast furnace smelting safety reinforcement learning decision optimization method based on expert strategy constraint. According to the method, an optimization strategy can be learned from an off-line expert track on the premise that operation safety is ensured. Specifically, a conditional generative adversarial mechanism is introduced to realize the alignment of strategy distribution and expert decision, and the exploration ability in a security domain is retained while the expert experience is inherited. And an independent state-action safety evaluation network is designed and a discount factor is introduced, so that the long-term accumulation risk caused by the large hysteresis characteristic of the blast furnace is effectively dealt with. In addition, memory is adopted to enhance network coding historical information, incomplete state observation is supplemented, and information deviation in decision is reduced. According to the method, effective integration of triple guidance mechanisms is realized, knowledge inheritance is realized through distributed alignment, performance improvement is driven through reward optimization, and risk prevention and control are ensured through explicit security constraints.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of blast furnace ironmaking safety, in particular to a blast furnace smelting safety reinforcement learning decision optimization method with expert strategy constraint. BACKGROUND

[0002] Making safe and efficient decisions in the process of blast furnace ironmaking is the key to the sustainable development and economic benefit improvement of the steel industry. The inherent challenges such as high-dimensional continuous decision space, limited observable state, and conflicting optimization objectives make it difficult for traditional control methods to meet the requirements. In the face of complex decision-making environment, on-site smelting control mainly relies on expert experience, but this control method is highly dependent on expert experience and cognitive level, and has cognitive limitations under the conditions of decision variable coupling and multi-objective constraints, and tends to use conservative strategies rather than seeking optimal strategies. Under the background of the "double carbon strategy", the traditional experience-driven decision-making method cannot meet the requirements of fine and low-carbon production, and it is urgent to introduce intelligent methods to realize autonomous optimization of decision-making.

[0003] Blast furnace ironmaking is a strongly coupled production process with continuous blast and periodic charging and tapping, and its operation optimization is essentially a complex sequential decision-making problem with high-dimensional state space, continuous action space, and long time delay characteristics. Ironmaking process operation involves multiple key operation variables, and these variables have strong nonlinear coupling relationship. At the same time, the blast furnace system has significant large lag characteristics, and the influence of operation adjustment on molten iron quality often needs several hours to appear, which further increases the complexity of decision optimization. The industrial site has accumulated a large amount of historical data including furnace condition state variables, expert operation trajectories, molten iron quality indicators, and furnace abnormal events, which contain rich process knowledge and operation rules, providing valuable basic resources for intelligent decision-making based on data-driven. Reinforcement learning, as a machine learning method for sequential decision-making problems, has made breakthrough progress in automatic driving, robot control, medical recommendation and other fields, providing a new technical path for solving complex decision problems such as blast furnace operation. However, traditional reinforcement learning algorithms rely on the interaction between agents and the environment to learn optimal strategies through a large number of online exploration, which is fundamentally in conflict with the strict safety requirements of blast furnace operation. Therefore, the present application adopts an offline reinforcement learning paradigm, fully utilizes the decision-making knowledge contained in the historical expert operation data, and learns safe and effective operation strategies without any online exploration, which not only inherits the valuable experience of experts, but also realizes strategy optimization through intelligent algorithms, providing a safe and reliable technical solution for intelligent operation of blast furnace.

[0004] The Chinese invention patent with publication number CN116562127A discloses a blast furnace smelting operation optimization method and system based on offline reinforcement learning. By obtaining historical data of the blast furnace, an expert database is established. Based on the DDPG algorithm, a blast furnace smelting operation optimization model is established. The difference between the action output by the expert database and the action output by the policy network is used to construct a safety signal. According to the output of the safety signal and the evaluation network, the parameter update rule of the policy network is obtained, and the blast furnace smelting operation optimization model is trained based on the parameter update rule of the expert database and the policy network. Thus, the blast furnace smelting optimization operation is obtained, which solves the technical problem that reinforcement learning cannot guarantee reliability and safety when applied to blast furnace smelting operation optimization. Without any data model or mechanism model as support, the decision scheme provided by the policy network trained based on the expert operation trajectory can provide reasonable operation guidance and support for the furnace master to realize fine control of the blast furnace, ensuring the smooth operation of the blast furnace and improving the quality of molten iron.

[0005] The blast furnace smelting operation optimization method and system disclosed by the invention use the difference between expert actions and policy network actions to construct a safety signal. Based on the safety signal and the multi-element molten iron quality reward signal, the policy network is trained to maximize long-term decision-making benefits, achieving optimization of blast furnace smelting operation parameters.

[0006] However, the patent uses point estimation to construct a safety signal, which may overlook the multi-modal data distribution characteristics under multiple working conditions, and cannot effectively mimic the expert decision-making approach.

[0007] The Chinese invention patent with publication number CN118112922A discloses an indirect data-driven blast furnace molten iron quality optimal tracking control method, which includes: collecting data samples of the blast furnace ironmaking process, taking the silicon content and temperature of molten iron as the output variables of the blast furnace molten iron quality model, and applying typical correlation analysis and correlation analysis to determine the input variables of the blast furnace molten iron quality model; using nonlinear subspace identification technology to establish a blast furnace molten iron quality Hammerstein state space model; setting the output reference trajectory of the blast furnace molten iron quality index variable and the quadratic performance index of the blast furnace molten iron quality index tracking control, and establishing an augmented system related to the system state and reference trajectory; using the Krotov method to design a blast furnace molten iron quality optimal tracking controller. The blast furnace molten iron quality Hammerstein model designed by the invention can better depict the nonlinear dynamics of blast furnace ironmaking, and has high prediction accuracy; the designed blast furnace molten iron quality control algorithm has good tracking control performance and high anti-interference performance.

[0008] The application discloses an indirect data-driven blast furnace molten iron quality optimal tracking control method, adopts typical correlation analysis to determine model variables, establishes a Hammerstein state space model to describe a blast furnace ironmaking nonlinear dynamic process, designs an augmented system and a quadratic performance index, and uses a Krotov method to optimize a tracking controller, so that effective control and optimization of multivariate molten iron quality are realized.

[0009] However, the nonlinear model established by the patent has a large amount of calculation and is difficult to meet the real-time control requirement, online parameter updating is difficult, and the control strategy lacks safety boundary constraints, so that the optimal solution may violate the equipment operation regulation limit.

[0010] A blast furnace molten iron quality self-adaptive robust prediction control method based on lazy learning is disclosed in Chinese patent application No. CN109001979A, and relates to the technical field of blast furnace smelting automation control. The method comprises determining a controlled variable and a control variable; collecting blast furnace production history input and output measurement data to construct an initial database; constructing a query regression vector to determine abnormal data; querying a similar learning subset from the database, selecting an optimal learning subset, and processing the abnormal data; taking the optimal learning subset as a training set to establish a prediction model; calculating a molten iron quality index reference trajectory, constructing a prediction control performance index, and obtaining an optimal control vector; sending the optimal control vector to a bottom-layer PLC system and adjusting an actuator; collecting a new set of blast furnace measurement data, pre-processing the data, and updating the database. The method provided by the application can effectively suppress the influence of input and output interference and overcome the influence of abnormal data, so that the blast furnace molten iron quality is stabilized near the expected value, which is beneficial to stable operation and high-quality and high-yield of the blast furnace.

[0011] The application discloses a blast furnace molten iron quality self-adaptive robust prediction control method based on lazy learning, which queries similar samples from a database to form a learning sample set by using lazy learning, and establishes a local predictor by using a multiple-output least square support vector regression machine. A control performance index is constructed according to a future output expected value and a prediction value after multiple corrections, and an optimal control vector is obtained by sequential quadratic programming calculation.

[0012] However, the application needs to query similar samples to construct a data set for real-time training of a data model, and stable operation principles of the blast furnace make the samples under fluctuating furnace conditions less, so that the model established under the condition has poor precision, and the reliability of the decision under the condition is affected.

[0013] In summary, the existing blast furnace operation optimization is mostly based on the model predictive control framework, and the control effect of this method is directly related to the established prediction model. The uncertainty of the mine source and the dynamic change of the market order may cause the data-driven model to mismatch the actual blast furnace ironmaking process, resulting in loss of controller performance. The existing offline reinforcement learning method uses a single-point estimation of the expert strategy distribution with a supervisory signal, which easily ignores the diversity of decision combinations in multi-modal, limiting the exploration and generalization ability of the strategy network. In addition, this method lacks an explicit safety constraint mechanism, and in the face of uncertain states or operating boundary conditions, it is easy to cause policy failure or dangerous decisions. SUMMARY

[0014] The present application aims to propose an expert strategy constraint blast furnace smelting safety reinforcement learning decision optimization method, which learns the optimized strategy from the offline expert trajectory under the premise of ensuring operation safety through the safety reinforcement learning framework of expert strategy constraint. Specifically, the conditional generative adversarial mechanism is introduced to realize the alignment of the policy distribution and the expert decision, while inheriting the expert experience and retaining the exploration ability within the safety domain. An independent state-action safety evaluation network is designed and a discount factor is introduced to effectively deal with the long-term cumulative risk brought by the large lag characteristics of the blast furnace. In addition, a memory enhancement network is used to encode historical information, supplement incomplete state observations, and reduce information bias in decision making. The present application realizes the effective integration of the triple guidance mechanism: distribution alignment realizes knowledge inheritance, reward optimization drives performance improvement, and explicit safety constraints ensure risk prevention and control.

[0015] To achieve the above purpose, the technical scheme adopted by the present application is:

[0016] An expert strategy constraint blast furnace smelting safety reinforcement learning decision optimization method, characterized in that it comprises the following steps:

[0017] S1, acquiring field data for preprocessing, including outlier rejection, missing value filling, mean value processing and standardization processing, and constructing a reinforcement learning sample set based on the processed data;

[0018] S2, learning the distribution mode of expert data using the conditional generative adversarial idea, realizing distribution alignment by limiting the difference between the policy network and the expert decision;

[0019] S3, introducing reinforcement learning to optimize the strategy learned by the policy network, guiding the policy network to realize the improvement of molten iron quality in a weakly supervised manner through the reward signal;

[0020] S4, designing an independent state-action safety evaluation network to explicitly constrain the safety of the decision, avoiding the delayed outbreak of safety risks caused by short-sighted decisions;

[0021] S5, introduce memory network maintenance history information to enhance state description ability, guide policy network by combining distribution alignment, reward optimization and safety constraint triple mechanism, realize organic unification of knowledge inheritance, performance improvement and safety guarantee;

[0022] S6, randomly sample operation trajectory in experience replay pool to train safety reinforcement learning framework under expert policy constraint, save trained policy network structure and parameters, and use trained model to provide real-time decision assistance for furnace master.

[0023] As a preferred technical scheme of the present application: in step S1, the field data preprocessing and the construction of the reinforcement learning sample set are as follows:

[0024] S11, field data preprocessing:

[0025] S111, use the box plot method to identify and eliminate abnormal values caused by extreme working conditions or human input errors;

[0026] S112, for the problem of missing field data, fill in the key information by using the average value before and after;

[0027] S113, considering the inconsistency of the sampling frequency of process variables, the field data is averaged by hour and then time stamped;

[0028] S114, in order to eliminate the deviation caused by the non-uniformity of the data dimension to the model training, the process data is standardized;

[0029] S12, construction of reinforcement learning sample set:

[0030] For the processed data set, the corresponding sample set is prepared according to the training paradigm of reinforcement learning, which specifically includes:

[0031] Observation set: select process variables as observation variables when experts make decisions from sensor data, including four main categories of temperature, pressure, flow and quality indicators that experts focus on when making decisions;

[0032] Decision set: lower thermal parameters of blast furnace regulation and control;

[0033] Reward set: hot metal temperature and chemical composition are the core indicators of product quality, based on the operation experience and process requirements of field experts, a weighted comprehensive reward function is designed to calculate the decision return;

[0034] Safety set: based on the experience of field experts, the safety score is calculated based on the fitted silicon content trend;

[0035] Expert experience replay pool: select decision trajectory within a preset time to form a round, and use sliding window to intercept operation trajectory, and store the prepared operation trajectory in experience replay area for subsequent training.

[0036] As a preferred technical solution of the present application: in step S2, the expert policy imitation learning based on conditional generative adversarial network is as follows:

[0037] A conditional generator with a parameter is used to represent the decision network to be learned, and the purpose is to generate the action related to the state according to the given condition state ;

[0038] A conditional discriminator with a parameter is used to judge whether the state-action pair generated by the policy network obeys the joint distribution probability of the state-action pair in the expert data set ; The goal of the conditional generator

[0039] is to generate the closest expert decision based on the state condition to deceive the conditional discriminator, and the loss function formula of the conditional generator is expressed as:

[0040] (1);

[0041] Where , are the state and decision vectors at time , is the policy network, is random noise,

[0042] When the conditional discriminator is fixed, the parameters of the conditional generator are updated in the gradient descent manner, that is:

[0043] (2);

[0044] Where is the conditional generator parameter at time step , is the learning rate, is the gradient operator,

[0045] The goal of the conditional discriminator is to accurately judge whether the input smelting state-action pair is from the expert data or the policy network , and the loss function formula of the conditional discriminator is expressed as:

[0046] (3);

[0047] Where is the expert data, ​is a policy network,

[0048] When the condition generator is fixed, the parameters of the condition discriminator are updated in a gradient ascent manner, that is:

[0049] (4);

[0050] wherein is the condition discriminator parameter at time step ,

[0051] The conditional generative adversarial learning can effectively capture the multi-modal characteristics of expert decision and alleviate the distribution drift problem by simultaneously introducing state information into the generator and the discriminator to form a conditional constraint, combining the dynamic feedback mechanism of the adversarial learning framework, and realizing effective inheritance of expert decision knowledge by aligning the distribution difference.

[0052] As a preferred technical solution of the application: in step S3, the reinforcement learning decision optimization of the expert policy constraint is as follows:

[0053] The deep deterministic policy gradient algorithm is used to optimize the policy network, and the main network is composed of a parameterized condition generator and a parameterized state-action evaluation network The second objective of the condition generator is to generate a sequence decision trajectory with high return, and the long-term return of the decision is evaluated by the parameterized evaluation network The formulation of the loss function of the expected return maximized by the condition generator is:

[0054] (5);

[0055] wherein , are the state and decision vector at time , respectively, is random noise, is the policy network decision trajectory, is the evaluation network,

[0056] In order to achieve the goal of maximizing the expected return, the parameters of the condition generator are updated in a gradient ascent manner, that is:

[0057] (6);

[0058] wherein is the condition generator parameter at time step , is the learning rate, is the gradient operator,

[0059] The evaluation network Co-learned with the condition generator to estimate the action value associated with the policy network and guide the direction of policy learning, therefore, accurate estimation Value is crucial, the deep deterministic policy gradient algorithm introduces a target condition generator and a target evaluation network to evaluate the state The dynamics of the lower output And the corresponding value The evaluation network is updated according to the least square method, that is:

[0060] (7);

[0061] Where The reward at time t, Is the expert decision trajectory, , ,

[0062] In order to achieve the goal of minimizing formula (7), the gradient descent method is used to update the parameters of the evaluation network, that is:

[0063] (8);

[0064] Where The model parameters of the state-action evaluation network at time t, The target policy network and the target Value evaluation network updates the corresponding parameters according to the speed Mathematical expression is as follows:

[0065] (9);

[0066] (10).

[0067] Where The model parameters of the condition generator at time t, , The model parameters of the target condition generator at time t and , The model parameters of the target condition discriminator at time t and , , The model soft update rate. As a preferred technical solution of the present application: in step S4, the safety enhanced reinforcement learning decision constraint is as follows:

[0068] As a preferred technical solution of the present application: in step S4, the safety enhanced reinforcement learning decision constraint is as follows:

[0069] ​​Introducing prior knowledge about security during the training phase and enabling forward-looking assessment of long-term security performance avoids delayed outbreaks of security risks due to short-sighted decisions. The state-action value function in reinforcement learning is used to evaluate the long-term cumulative reward generated by making specific decisions in the current state.

[0070] Propose a network of security critics to evaluate the security performance of decision-making. The training data in the experience replay pool is then converted into quadruplets. Expanded into a quintuple ,in The safety performance indicators are designed based on on-site operating procedures and expert experience. The larger the value, the higher the safety performance of the decision-making process; condition generator. The goal is to generate actions that maximize expected safety. The loss function that maximizes expected safety is formally expressed as:

[0071] (11);

[0072] in , They are The state and decision vector at each moment, It is the policy network decision trajectory. For the network of security commentators,

[0073] To maximize the desired safety, the parameters of the condition generator are updated using gradient ascent, i.e.:

[0074] (12);

[0075] in It is a time step The condition generator parameters, It's the learning rate. It is the gradient operator.

[0076] Security Commentator Network It also needs to learn together with the policy network to ensure the safety of decisions. Similarly, security reinforcement learning also introduces corresponding target condition generators and target security evaluation networks for training. To accurately predict the long-term cumulative safety performance of actions taken under different conditions, the Security Commentator Network... The mathematical expression of the loss function is as follows:

[0077] (13);

[0078] in yes Momentary rewards It is the expert decision-making trajectory. To minimize equation (13), gradient descent is used to update the parameters of the evaluation network, i.e.:

[0079] (14);

[0080] Similarly, the corresponding target security review network is ranked by speed. The corresponding soft parameter update is performed, and its mathematical expression is as follows:

[0081] (15);

[0082] in Security commentators network Model parameters at time 10:00 , It is the target security commentator network in and Model parameters at time 10:00 It is the rate of soft updates to the model.

[0083] As a preferred technical solution of the present invention: In step S5, the expert constraint safety reinforcement learning decision optimization of state encoding is specifically as follows:

[0084] In expert-policy-constrained secure reinforcement learning, generative networks achieve knowledge inheritance through distribution alignment, drive performance improvement through reward optimization, and ensure risk control through explicit security constraints.

[0085] The final form of the objective function of the condition generator is as follows:

[0086] (16);

[0087] in It controls the weights for different optimization objectives. , , They are The state and decision vector at each moment, It is a policy network. For conditional discriminators, For condition generator, It is random noise. To evaluate the network, For the network of security commentators,

[0088] Minimizing the distribution difference between the policy network and the expert policy is equivalent to maximizing the negative distribution difference. To maximize the loss function of the generator network, the gradient ascent algorithm is used to update the parameters, i.e.:

[0089] (17);

[0090] where is the time step, is the conditional generator parameter, is the learning rate, is the gradient operator,

[0091] When the decision can only be made based on part of the state due to the limitations of the environment and sensors, In order to make up for the decision deviation caused by insufficient information, the model adopts LSTM memory unit to encode the observation sequence in the first layer to imitate the decision-making idea of the on-site expert and the decision sequence characterize the current state , that is:

[0092] (18);

[0093] After summarizing the historical information to solve the problem of incomplete information under partial observation state, the generation network parameter is updated as follows:

[0094] (19);

[0095] Similarly, the conditional discriminator, state-action evaluation network and safety evaluation network parameter update methods are as follows:

[0096] (20);

[0097] (21);

[0098] (22);

[0099] where is the time step, is the conditional discriminator parameter, , .

[0100] As a preferred technical solution of the present application: in step S6, the operation trajectory is randomly sampled to train the safety reinforcement learning framework under the constraint of the expert strategy, which is as follows:

[0101] A batch of data is sampled from the expert experience replay pool for training the proposed safety reinforcement learning model under the constraint of the expert strategy, and the trained strategy network , network structure and other parameters are saved, and when online testing, only the current state variable information needs to be provided to output the corresponding operation decision.

[0102] Compared with the prior art, the present application has the following beneficial effects:

[0103] The present application takes the blast furnace ironmaking process as the research object, and proposes a safety reinforcement learning decision optimization method with expert strategy constraints, which is used to realize safe and efficient decision optimization of blast furnace operation. In view of the contradiction that traditional reinforcement learning needs to explore learning through online trial and error, but the safety constraints of blast furnace operation prohibit online exploration, the present application innovatively relies only on offline expert trajectory for strategy learning. In order to ensure the safety and effectiveness of the strategy output, a conditional generative adversarial mechanism is introduced to realize the accurate alignment of the strategy distribution and the expert decision, while preserving the exploration ability within the verified safety domain while inheriting the expert experience. In addition, in view of the long-term cumulative risk problem caused by the lag characteristics of the blast furnace, an independent state-action safety evaluation network is designed and a discount factor is introduced, which effectively models and controls long-term safety risks. In order to solve the problem of incomplete information caused by partial observable states, a memory enhanced network is used to encode historical information to supplement the current state observation, which significantly reduces the information bias in decision making. The present application realizes the effective integration of the three guiding mechanisms of distribution alignment, reward optimization and safety constraints, and the decision network based on expert strategy constraints can provide fine regulation and control guidance for the furnace master in accordance with the actual working conditions, which improves the production efficiency and the quality of molten iron on the premise of ensuring the safe and stable operation of the blast furnace, and provides important technical support for the intelligent upgrading and sustainable development of the steel industry. BRIEF DESCRIPTION OF DRAWINGS

[0104] Figure 1 is expert strategy imitation learning based on generative adversarial network;

[0105] Figure 2 is a reinforcement learning decision model with expert strategy constraints;

[0106] Figure 3 is a safety enhanced reinforcement learning decision constraint model;

[0107] Figure 4 is a safety reinforcement learning decision optimization model with expert strategy constraints. DETAILED DESCRIPTION

[0108] The present application will be described in further detail below in combination with the drawings and specific embodiments:

[0109] The present application proposes a safety reinforcement learning decision optimization method with expert strategy constraints for blast furnace smelting, which specifically includes the following steps:

[0110] S1, acquire field data for preprocessing, including outlier rejection, missing value filling, mean value processing and standardization processing, and construct a reinforcement learning sample set based on the processed data.

[0111] Due to the influence of equipment degradation, human operation errors, high temperature and high pressure failures, and abnormal production conditions such as blast furnace wind, the collected data often has defects, including outliers and missing values. In order to improve the accuracy and reliability of the data, system preprocessing is needed.

[0112] The field data preprocessing and the construction of the reinforcement learning sample set are as follows:

[0113] S11, field data preprocessing:

[0114] S111, using the box plot method to identify and eliminate outliers caused by extreme working conditions or human input errors;

[0115] S112, for the problem of missing field data, the key information is filled in by using the average value before and after;

[0116] S113, considering the inconsistency of the sampling frequency of the process variables, the field data is averaged by hour and then time stamped;

[0117] S114, in order to eliminate the bias caused by the non-uniformity of the data dimension on the model training, the process data is standardized;

[0118] S12, construction of the reinforcement learning sample set:

[0119] For the processed data set, the corresponding sample set is prepared according to the training paradigm of reinforcement learning, which specifically includes:

[0120] Observation set: from the sensor data, the process variables used as observation variables during expert decision making are selected, including the four main categories of temperature, pressure, flow and quality indicators that the experts focus on during decision making;

[0121] Decision set: the lower thermal parameters of the blast furnace control, including oxygen enrichment rate, hot blast temperature, coal injection amount, etc.;

[0122] Reward set: hot metal temperature and chemical composition (silicon content, sulfur content and phosphorus content) are the core indicators of product quality. Based on the operation experience and process requirements of the field experts, a weighted comprehensive reward function is designed to calculate the decision return;

[0123] Safety set: silicon content as a sensitive indicator of the thermal state in the furnace, its trend can effectively predict the fluctuation of the furnace condition and potential safety risks. Based on the experience of the field experts, the safety score is calculated based on the fitted silicon content trend;

[0124] Expert experience replay pool: select the decision trajectory in a preset time to form a round, and use the sliding window to intercept the operation trajectory, store the prepared operation trajectory in the experience replay area for subsequent training. Since the on-site furnace master operates the blast furnace according to shifts, a shift (8 hours) of decision trajectory is selected to form a round. In order to improve the utilization efficiency of data, a sliding window with a length of 8 and a step of 1 is used to intercept the operation trajectory, that is: The prepared trajectory is stored in the experience replay area for subsequent training.

[0125] S2, learn the distribution pattern of expert data by conditional generative adversarial idea, realize distribution alignment by limiting the difference between strategy network and expert decision.

[0126] The complex system characteristics of the blast furnace smelting process make the on-site decision mainly based on the experience of the furnace master. Although the expert decision is not optimal, its safety, feasibility and stability have been verified in the field, so learning the expert strategy is the key means to ensure that the decision is reasonable and reliable. Generative adversarial learning forces the model to learn the overall distribution of expert data rather than simple point estimation through the game process of discriminator and generator, and thus better captures the subtle patterns and decision diversity in the expert strategy, so as to capture the deep logical relationship behind the expert decision. The expert strategy imitation learning framework based on conditional generative adversarial network is shown in Figure 1 .

[0127] Expert strategy imitation learning based on conditional generative adversarial network is as follows:

[0128] A conditional generator with a parameter of is used to represent the decision network that needs to be learned, which aims to generate actions related to the state according to the given condition state ;

[0129] A conditional discriminator with a parameter of is used to judge whether the state-action pair generated by the strategy network obeys the joint distribution probability of the state-action pair in the expert data set ;

[0130] The goal of the conditional generator is to generate as close as possible expert decisions based on the state condition to deceive the conditional discriminator. The loss function formula of the conditional generator is expressed as:

[0131] (1);

[0132] wherein , They are The state and decision vector at each moment, It is a policy network. It is random noise.

[0133] When fixing the condition discriminator, the parameters of the condition generator are updated using gradient descent, i.e.:

[0134] (2);

[0135] in It is a time step The condition generator parameters, It's the learning rate. It is the gradient operator.

[0136] Conditional discriminator The goal is to accurately determine the input smelting state – the action pair is derived from expert data. or policy network The formulaic expression of its loss function is as follows:

[0137] (3);

[0138] in For expert data, It is a policy network.

[0139] When the condition generator is fixed, the parameters of the condition discriminator are updated using gradient ascent, i.e.:

[0140] (4);

[0141] in It is a time step The condition discriminator parameters,

[0142] Conditional generative adversarial learning (GRE) introduces state information into both the generator and the discriminator to form conditional constraints. Combined with the dynamic feedback mechanism of the adversarial learning framework, it can effectively capture the multimodal characteristics of expert decision-making and alleviate the distribution drift problem. By aligning distribution differences, it can achieve effective inheritance of expert decision-making knowledge.

[0143] S3. Reinforcement learning is introduced to optimize the policy network's learning strategy. The policy network is guided by reward signals in a weakly supervised manner to improve the quality of molten iron.

[0144] If we only aim to align the distribution of the condition generator and the expert policy, then the upper limit of the learned policy quality is the level of the expert decision. Due to the complexity of blast furnace smelting decisions and the differences in the experts' skill levels, the label of the "best" policy is unclear on-site. More importantly, the ultimate goal of blast furnace smelting decision optimization is to improve the quality rate of molten iron under stable furnace conditions, rather than simply imitating a suboptimal expert decision. Therefore, additional reward-guided learning is needed for the condition generator to further help the policy network improve towards optimizing molten iron quality. Reinforcement learning guides the policy network's performance steadily through reward signals in a weakly supervised manner. Therefore, our work introduces an Actor-Critic framework to optimize the policy network based on generative adversarial networks, such as... Figure 2 As shown in the diagram. Specifically, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to optimize the policy network, further improving the high-return nature of decision-making based on expert distribution constraints.

[0145] The reinforcement learning decision optimization with expert policy constraints is as follows:

[0146] A deep deterministic policy gradient algorithm is used to optimize the policy network, whose main network consists of a parameterized condition generator. and parameterized state-action evaluation network Composition, condition generator The second objective is to generate high-return sequential decision trajectories through parameterized evaluation networks. To evaluate the long-term returns of its decisions, the loss function of the conditional generator, which maximizes the expected return, is formally expressed as:

[0147] (5);

[0148] in , They are The state and decision vector at each moment, It is random noise. It is the policy network decision trajectory. To evaluate the network,

[0149] To maximize the expected return, the parameters of the conditional generator are updated using gradient ascent, i.e.:

[0150] (6);

[0151] in It is a time step The condition generator parameters, It's the learning rate. It is the gradient operator.

[0152] evaluation network co-learns with the condition generator to estimate the action value associated with the policy network and guide the direction of policy learning, therefore, it is crucial to accurately estimate the value, the deep deterministic policy gradient algorithm introduces a target condition generator and a target evaluation network to evaluate the state dynamic of the output and the corresponding value , the evaluation network is updated according to the least square method, that is:

[0153] (7);

[0154] wherein is the reward at time t, is the expert decision trajectory, , in order to achieve the goal of minimizing formula (7), the gradient descent method is used to update the parameters of the evaluation network, that is:

[0155] (8);

[0156] wherein the model parameters of the state-action evaluation network at time t, the target policy network and the target value evaluation network update the corresponding parameters of the soft according to the speed , the mathematical expression is as follows:

[0157] (9);

[0158] (10);

[0159] wherein the model parameters of the condition generator at time t, , is the model parameters of the target condition generator at time t and , is the model parameters of the target condition discriminator at time t and , , is the model parameters of the target condition discriminator at time t and , is the model soft update rate. S4, design an independent state-action safety evaluation network to explicitly constrain the safety of decision-making, avoid the delayed outbreak of safety risks caused by short-sighted decision-making.

[0160]

[0161] ​​Although the expert-constrained reinforcement learning framework can improve the expected return of the policy based on imitating the expert policy, it cannot guarantee the safety of the decision. This is because the offline learning paradigm is prone to out-of-distribution generalization problems, i.e., giving overly optimistic estimates of the state-action value outside the expert data distribution, leading to significant safety risks in some unseen scenarios. Considering the strict requirements of the blast furnace smelting process for the safe and smooth operation of the furnace condition, we introduce prior knowledge about safety during the training phase and implement forward-looking evaluation of long-term safety performance to avoid the delayed outbreak of safety hazards caused by short-sighted decisions. The safety-enhanced reinforcement learning decision constraint is as shown in Figure 3 .

[0162] The safety-enhanced reinforcement learning decision constraint is as follows:

[0163] During the training phase, prior knowledge about safety is introduced, and forward-looking evaluation of long-term safety performance is implemented to avoid the delayed outbreak of safety hazards caused by short-sighted decisions. The state-action value function in reinforcement learning is used to evaluate the long-term cumulative return generated by making a specific decision in the current state.

[0164] A safety critic network is proposed to evaluate the safety performance of the decision , and the training data in the experience replay pool is expanded from the quadruple to the quintuple , where is a safety performance indicator designed based on the on-site operation system and expert experience, the larger the value, the higher the safety performance of the decision, and the goal of the conditional generator is to generate actions that maximize the expected safety. The formula of the loss function that maximizes the expected safety is as follows:

[0165] (11);

[0166] where , are the state and decision vectors at time , respectively, is the decision trajectory of the policy network, is the safety critic network,

[0167] To achieve the goal of maximizing the expected safety, the parameters of the conditional generator are updated using gradient ascent, i.e.:

[0168] (12);

[0169] where is the conditional generator parameter at time step , is the learning rate, is the gradient operator,

[0170] Security critic network It is also necessary to learn with the policy network to ensure the safety of the decision, and similarly, the corresponding goal condition generator and target safety evaluation network are introduced in security reinforcement learning for training To accurately estimate the long-term cumulative safety performance of taking action in different states, the loss function of the security critic network is mathematically expressed as follows:

[0171] (13);

[0172] Wherein is the reward at time t, is the expert decision trajectory, To achieve the goal of minimizing formula (13), the gradient descent method is used to update the parameters of the evaluation network, that is:

[0173] (14);

[0174] Similarly, the corresponding target safety critic network updates the parameters according to the speed The corresponding parameters are mathematically expressed as follows:

[0175] (15);

[0176] Wherein The model parameters of the security critic network at time t, , The model parameters of the target security critic network at time t and t+1, is the model soft update rate. S5, introduce a memory network to maintain historical information to enhance the state description ability, and combine the three mechanisms of distribution alignment, reward optimization and safety constraint to guide the policy network, realize the organic unity of knowledge inheritance, performance improvement and safety guarantee.

[0177] The generation network in the expert policy constraint security reinforcement learning realizes knowledge inheritance through distribution alignment, reward optimization drives performance improvement, and explicit safety constraint ensures risk prevention and control, and the final structure is shown in .

[0178] Figure 4 The state coding expert constraint security reinforcement learning decision optimization is as follows:

[0179] The state coding expert constraint security reinforcement learning decision optimization is as follows:

[0180] ​​​In expert-policy-constrained secure reinforcement learning, generative networks achieve knowledge inheritance through distribution alignment, drive performance improvement through reward optimization, and ensure risk control through explicit security constraints.

[0181] The final form of the objective function of the condition generator is as follows:

[0182] (16);

[0183] in It controls the weights for different optimization objectives. , , They are The state and decision vector at each moment, It is a policy network. For conditional discriminators, For condition generator, It is random noise. To evaluate the network, For the network of security commentators,

[0184] Minimizing the distribution difference between the policy network and the expert policy is equivalent to maximizing the negative distribution difference. To maximize the loss function of the generator network, the gradient ascent algorithm is used to update the parameters, i.e.:

[0185] (17);

[0186] in It is a time step The condition generator parameters, It's the learning rate. It is the gradient operator.

[0187] The preceding introduction assumed the state variables of the blast furnace. It is fully observable. However, due to environmental and sensor limitations, it can only be based on partial states. To compensate for decision-making biases caused by insufficient information during decision-making, and to mimic the decision-making process of on-site experts, the first layer of the model uses LSTM memory units to encode the observation sequence. and decision sequence Represent the current state ,Right now:

[0188] (18);

[0189] After summarizing historical information to address the issue of incomplete information under certain observation conditions, network parameters are generated. The update method is as follows:

[0190] (19);

[0191] Similarly, the conditional discriminator, state-action evaluation network and safety evaluation network parameter updating methods are as follows:

[0192] (20);

[0193] (21);

[0194] (22);

[0195] wherein is the conditional discriminator parameter of the time step , , .

[0196] S6, randomly sample the operation trajectory in the experience replay pool to train the safety reinforcement learning framework under the constraint of the expert policy, save the trained policy network structure and parameters, and use the trained model to provide real-time decision assistance for the furnace master.

[0197] Randomly sample the operation trajectory to train the safety reinforcement learning framework under the constraint of the expert policy, as follows:

[0198] Sample a batch of data from the expert experience replay pool to train the safety reinforcement learning model under the constraint of the expert policy, save the trained policy network related parameters and network structure, and when online testing, only the current state variable information needs to be provided, and the corresponding operation decision can be output.

[0199] Based on the above technical solutions:

[0200] (1) the safety reinforcement learning framework under the constraint of the expert policy is proposed, the policy network is learned and optimized offline through the expert trajectory, online trial and error exploration is not needed, and the fundamental conflict between the strict safety constraint of the blast furnace operation and the online exploration demand of reinforcement learning is solved;

[0201] (2) the conditional generative adversarial mechanism is introduced to realize the distribution alignment of the learned policy and the expert decision, the valuable experience of the expert is inherited, and the exploration ability in the verification safety domain is retained;

[0202] (3) the safety evaluation network of the state-action is designed independently, and the discount factor is introduced to evaluate the long-term cumulative risk caused by the hysteresis characteristics of the blast furnace, and the limitation that the traditional safety mechanism only focuses on the immediate risk is broken through;

[0203] (4) the long short-term memory network is used to encode the historical state and decision information to expand the incomplete state observation, the information missing problem in the partially observable environment is effectively solved, and the policy network can make decisions based on more complete state information;

[0204] (5) The present application realizes the effective integration of the triple guidance mechanism of distribution alignment, reward optimization and explicit safety constraint, ensures the coordination and unity of knowledge inheritance, performance improvement and risk prevention, and provides a systematic solution for intelligent operation of blast furnaces;

[0205] (6) The intelligent decision-making model constructed by the present application can provide real-time and accurate operation guidance suggestions for the furnace master, assist the on-site workers in fine-tuned control, improve the scientific nature of decision-making and the standardization of operation, and realize intelligent smelting management of man-machine cooperation.

[0206] The implementation case of the present application is verified in a 2650m 3 large blast furnace in a certain ironworks.

[0207] A blast furnace smelting safety reinforcement learning decision optimization method under expert strategy constraints, specifically comprising the following steps:

[0208] 1) Data preprocessing: the data collected on the blast furnace detection device is processed to improve the quality of the data, including outlier rejection, missing value filling, mean value processing and standardization processing.

[0209] 2) Reward function: in order to evaluate the quality of molten iron, the molten iron quality indicators (molten iron temperature, silicon content, sulfur content, phosphorus content) are graded according to the experience of on-site experts, and the detailed information is shown in Table 1:

[0210]

[0211] Table 1 Grade division rule table of molten iron quality indicators

[0212] According to different division grades, the corresponding returns are defined as follows:

[0213] (23);

[0214] Considering that the quality of molten iron needs to consider the influence between multiple indicators, the quantitative evaluation rule based on expert experience is shown in formula (24):

[0215] (24);

[0216] 3) Safety function: referring to the change trend of silicon content and the corresponding furnace condition, a silicon content change trend grade table is established according to the experience of on-site experts, and the trend category and corresponding safety score can be calculated by formula (23), and the detailed information is shown in Table 2:

[0217]

[0218] Table 2 Grade score table of silicon content change trend

[0219] 3) The safety reinforcement learning decision method for blast furnace smelting based on expert strategy constraint. The structure of the condition generator in this patent is input layer-LSTM layer-full connection layer-output layer, and the number of neurons and activation functions are: 33-256-128(R)-3(S). The structure of the condition discriminator and two evaluation networks is input layer-LSTM layer-full connection layer-output layer, and the number of neurons and activation functions are: 36-256-128(R)-1(S). The processed 5882 trajectories are used to train the offline reinforcement learning framework, and 100 trajectories are used to test the model effect. In order to quantitatively evaluate the reliability of the decision output by the trained strategy network, the decision vector output by the condition generator is input into the intelligent perception model of the multi-element molten iron quality and silicon content trend established in the previous work, and the corresponding return and safety score are counted according to the rules in Table 1 and Table 2. In order to evaluate the safety of the operation strategy given by the model, we take the expert experience as the benchmark, and take the return, safety and Jensen-Shannon divergence as the measurement standard, and calculate the difference between the decision provided by the condition generator and the expert decision, and the detailed results are shown in Table 3:

[0220]

[0221] Table 3 Performance indicators of different decision methods

[0222] From Table 3, it can be seen that the method proposed in this patent can obtain higher average return and safety performance than expert operation on the test set, and can exist difference with the expert decision distribution, which shows that introducing expert strategy constraint, safety evaluation mechanism and time sequence state modeling on the basis of reinforcement learning is an effective means to realize the optimization of blast furnace smelting decision, and the Jensen-Shannon divergence between the expert decision is 0.405, which shows that the method realizes the autonomous optimization and surpassing of strategy on the basis of understanding the logic of expert behavior. This further shows the feasibility and effectiveness of the safety reinforcement learning decision method based on expert strategy constraint in the optimization of blast furnace smelting operation.

[0223] The above is only a preferred embodiment of the present application, and is not intended to limit the present application in any other form, and any modification or equivalent change made according to the technical essence of the present application still falls within the scope of the present application.

Claims

1. A blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints, characterized in that, The method comprises the following steps: S1, obtaining field data for preprocessing, including outlier rejection, missing value filling, mean value processing and standardization processing, and constructing a reinforcement learning sample set based on the processed data; S2, learning the distribution mode of expert data using a conditional generative adversarial idea, and realizing distribution alignment by limiting the difference between the strategy network and the expert decision; S3, introducing reinforcement learning to optimize the strategy learned by the strategy network, and guiding the strategy network to realize molten iron quality improvement through a reward signal in a weakly supervised manner; S4, designing an independent state-action safety evaluation network to explicitly constrain the safety of the decision, and avoiding the delayed explosion of safety risks caused by short-sighted decisions; S5, introducing a memory network to maintain historical information to enhance state description capability, and comprehensively guiding the strategy network through the triple mechanism of distribution alignment, reward optimization and safety constraint to realize the organic unification of knowledge inheritance, performance improvement and safety guarantee; S6, training a safety reinforcement learning framework under the constraint of an expert strategy by randomly sampling operation trajectories in an experience replay pool, saving the trained strategy network structure and parameters, and providing real-time decision assistance for the furnace master using the trained model.

2. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S1, the field data preprocessing and the construction of the reinforcement learning sample set are as follows: S11, field data preprocessing: S111, using the box plot method to identify and eliminate outliers caused by extreme working conditions or human input errors; S112, for the problem of missing field data, using the mean value before and after to fill in the key information; S113, considering the inconsistency of the sampling frequency of the process variables, the field data is averaged by hour and then time stamped; S114, to eliminate the bias caused by the non-uniformity of the data dimension on the model training, the process data is standardized; S12, construction of the reinforcement learning sample set: For the processed data set, the corresponding sample set is prepared according to the training paradigm of reinforcement learning, which specifically includes: Observation set: process variables are selected from sensor data as observation variables when the expert makes a decision, including four main categories of temperature, pressure, flow and quality indicators that the expert focuses on when making a decision; Decision set: lower thermal parameters of blast furnace control; Reward set: molten iron temperature and chemical composition are the core indicators of product quality, based on the operation experience and process requirements of field experts, a weighted comprehensive reward function is designed to calculate the decision return; Safety set: based on the experience of field experts, the safety score is calculated based on the fitted silicon content trend; Expert experience replay pool: select the decision trajectory within a preset time to form a round, and use a sliding window to intercept the operation trajectory, and store the prepared operation trajectory in the experience replay area for subsequent training.

3. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S2, expert strategy imitation learning based on conditional generative adversarial network, specifically as follows: Using a parameterized condition generator represents a decision network to be learned, with the goal of generating actions related to a state given a condition state ;​​ A conditional discriminator with one parameter is used to judge state-action pairs produced by the policy network to conform to the joint distribution probability of state-action pairs in the expert data set ;​ Condition generator The goal of the condition generator is to produce the closest expert decision based on the state condition to fool the condition discriminator, and the loss function of the condition generator is formulated as: (1); where , are the state and decision vectors at time t, is the policy network, is random noise,​ When the condition discriminator is fixed, the parameters of the condition generator are updated using gradient descent, that is: (2); wherein is the time step the conditional generator parameter, is the learning rate, is the gradient operator, Conditional discriminator The goal is to accurately determine whether an input smelting state-action pair is from expert data or a policy network with the loss function formulated as: (3); wherein is expert data, is a policy network, When the condition generator is fixed, the parameters of the condition discriminator are updated using gradient ascent, that is: (4); wherein is the time step conditional discriminator parameters, The conditional generative adversarial learning forms a conditional constraint by introducing state information into the generator and discriminator simultaneously, combines the dynamic feedback mechanism of the adversarial learning framework, and can effectively capture the multi-modal characteristics of expert decision-making and alleviate the distribution drift problem. Through aligning the distribution difference, the expert decision-making knowledge is effectively inherited.

4. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S3, the reinforcement learning decision optimization of the expert policy constraint is as follows: A deep deterministic policy gradient algorithm is employed to optimize the policy network, which consists of a parameterized conditional generator and a parameterized state-action value network The conditional generator The second objective is to produce high-return sequence decision trajectories, which are evaluated by the parameterized value network for their long-term returns. The formulation of the loss function that the conditional generator maximizes for the expected return is: (5); where , are the state and decision vectors at time , is a random noise, is a policy network decision trajectory, is an evaluation network, To achieve the goal of maximizing the expected return, the parameters of the conditional generator are updated in the gradient ascent manner, that is: (6); wherein is a time step a conditional generator parameter, is a learning rate, is a gradient operator, evaluation network learns jointly with the condition generator to estimate the action values associated with the policy network and guide the direction of policy learning, therefore, accurate estimation of values is crucial, the deep deterministic policy gradient algorithm introduces a target condition generator and a target evaluation network to evaluate the state the dynamics of the next output and the corresponding value the evaluation network is updated in a least squares manner, i.e.: (7); wherein is a reward at time is an expert decision trajectory, , To achieve the goal of minimizing formula (7), the parameters of the evaluation network are updated in the gradient descent manner, that is: (8); wherein The state-action evaluation network is in The model parameters at the moment correspond to the target policy network and the target The value evaluation network performs a corresponding parameter soft update according to the speed The mathematical expression is as follows: (9); (10)。 5. wherein the condition generator is at the model parameters at the time instant, , the target condition generator is at and the model parameters at the time instant, , the target condition discriminator is at and the model parameters at the time instant, is the rate of model soft updates.

6. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S4, the safety-enhanced reinforcement learning decision constraint is as follows: In the training phase, the prior knowledge about safety is introduced and the forward-looking evaluation of long-term safety performance is realized, which avoids the delayed explosion of safety hazards caused by short-sighted decisions. The state-action value function in reinforcement learning is used to evaluate the long-term cumulative return of making a specific decision in the current state; Propose a network of security critics to evaluate the security performance of decision-making. The training data in the experience replay pool is then converted into quadruplets. Expanded into a quintuple ,in The safety performance indicators are designed based on on-site operating procedures and expert experience. The larger the value, the higher the safety performance of the decision-making process; condition generator. The goal is to generate actions that maximize expected safety. The loss function that maximizes expected safety is formally expressed as: (11); where , are respectively the state and decision vector at time is the policy network decision trajectory, is the safety critic network, To achieve the goal of maximizing the expected safety, the parameters of the conditional generator are updated in the gradient ascent manner, that is: (12); wherein is a time step a conditional generator parameter, is a learning rate, is a gradient operator, Security critic network It is also necessary to learn with the policy network to ensure the safety of the decision, similarly, the corresponding goal condition generator and target safety evaluation network are introduced in security reinforcement learning for training To accurately estimate the long-term cumulative safety performance of taking action in different states, the security critic network The loss function of the security critic network is mathematically expressed as follows: (13); wherein is the reward at time is the expert decision trajectory, To achieve the goal of minimizing formula (13), the parameters of the evaluation network are updated in the gradient descent manner, that is: (14); Similarly, the corresponding target security review network is updated according to the speed The corresponding parameter soft update is performed as follows: (15); wherein The network of security critics is in the model parameters at time , the target security critic network is in and the model parameters at time is the rate of model soft updates.

7. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S5, the state-encoding expert-constrained safety reinforcement learning decision optimization is as follows: The generation network in the expert policy-constrained safety reinforcement learning realizes knowledge inheritance through distribution alignment, reward optimization drives performance improvement, and explicit safety constraints ensure risk prevention and control, The final expression of the conditional generator objective function is as follows: (16); wherein is a weight controlling different optimization objectives, , , are state and decision vectors at time instant, is a policy network, is a conditional discriminator, is a conditional generator, is random noise, is a critic network, is a safety critic network, Minimizing the distribution difference between the policy network and the expert policy is equivalent to maximizing the negative distribution difference. In order to maximize the loss function of the generation network, the gradient ascent algorithm is used to update the parameters, that is: (17); wherein is a time step a conditional generator parameter, is a learning rate, is a gradient operator, When environmental and sensor limitations limit the ability to base decisions on only partial states... To compensate for decision-making biases caused by insufficient information during decision-making, and to mimic the decision-making process of on-site experts, the first layer of the model uses LSTM memory units to encode the observation sequence. and decision sequence Represent the current state ,Right now: (18); The history information is summarized to solve the problem of incomplete information in the partial observation state, and the network parameters are generated The update method is as follows: (19); Similarly, the parameter updating methods of the conditional discriminator, state-action evaluation network and safety evaluation network are as follows: (20); (21); (22); wherein is a time step conditional discriminator parameters, , .

8. The blast furnace smelting safety reinforcement learning decision optimization method with expert policy constraints according to claim 1, characterized in that, In step S6, the random sampling operation trajectory is used to train the safety reinforcement learning framework under the expert policy constraint, which is as follows: Sample batch data from the expert experience replay pool for training the proposed safety reinforcement learning model with expert policy constraints, save the trained policy network The relevant parameters and network structure, when online testing, only need to provide the current state variable information, and the corresponding operation decision can be output.

Citation Information

Patent Citations

  • Lazy learning-based self-adaptive robustness forecast control method of blast furnace molten iron quality

    CN109001979A

  • Indirect data driven blast furnace molten iron quality optimal tracking control method

    CN118112922A

  • Blast furnace smelting operation optimization method and system based on offline reinforcement learning

    CN116562127A

  • Active imitation learning in high dimensional continuous environments

    US20200082257A1

Cited By

  • Feedback enhancement and working condition guided wet leaching process reinforcement learning control method

    CN122239422A