Computer system and method for providing operational instructions for blast furnace heat control

By training a reinforcement learning model using a recurrent neural network, combined with domain-adaptive machine learning and transient models, optimized thermal control operation commands are generated, solving the problem of optimizing material and fuel consumption during blast furnace production, improving the efficiency and stability of the blast furnace, and extending its lifespan.

CN116261690BActive Publication Date: 2026-05-12PAUL WURTH SA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PAUL WURTH SA
Filing Date
2021-09-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The blast furnace production process is complex, making it difficult to optimize material and fuel consumption, which affects furnace efficiency and stability. Furthermore, existing technologies struggle to provide effective thermal control operation commands.

Method used

By training a reinforcement learning model of a recurrent neural network, combined with domain-adaptive machine learning and transient models, optimized thermal control operation commands are generated. The model is trained using multivariate time series data and manual operation data to achieve precise control of the blast furnace process.

Benefits of technology

It improved the operating efficiency and stability of the blast furnace, optimized fuel consumption, extended the blast furnace life, and achieved stable control over the quality of hot metal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116261690B_ABST
    Figure CN116261690B_ABST
Patent Text Reader

Abstract

A computer system (100), a computer-implemented method and a computer program product for training a reinforcement learning model (130) to provide operational instructions for blast furnace thermal control are provided. A domain-adaptive machine learning model (110) generates a first domain-invariant dataset (22) from historical operational data (21) obtained as multivariate time series and reflecting thermal states of respective blast furnaces (BF1 to BFn) of a plurality of domains. A transient model (121) of a generic blast furnace process is used to generate artificial operational data (24a) as multivariate time series reflecting thermal states of a generic blast furnace (BFg) for specific thermal control actions (26a). A generative deep learning network (122) generates a second domain-invariant dataset (23a) by transferring features learned from the historical operational data 21 to the artificial operational data (24a). A reinforcement learning model (130) determines (1400) a reward (131) for specific thermal control actions (26a) by processing the combined first and second domain-invariant datasets (22, 23a) in view of a given objective function. In accordance with the reward (131), the second domain-invariant dataset is regenerated based on modified parameters (123-2) and the determination of the reward is repeated to learn optimized operational instructions for optimized thermal control actions to be applied to respective operational states of one or more blast furnaces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to systems for controlling blast furnaces, and more specifically, to methods, computer program products, and systems for generating operating instructions for blast furnaces using machine learning methods. Background Technology

[0002] Blast furnaces are used to produce molten iron as a raw material for steelmaking. Blast furnaces involve highly complex processes that require modeling because they rely on multivariate process inputs and disturbances. The goal is to reduce material and fuel consumption to optimize the overall efficiency and stability of the furnace, the quality of the hot metal, and to extend the furnace's lifespan. Therefore, it is desirable to provide optimized operating instructions to define complex production objectives. Summary of the Invention

[0003] This technical problem is addressed by the features of the independent claims through training a reinforcement learning (RL) model implemented by a recurrent neural network to provide operating instructions for blast furnace thermal control. The operating instructions relate to corresponding thermal control actions. As used herein, a thermal control action refers to any action that affects an actuator in order to thermally control the blast furnace process. Depending on the level of control automation, the operating instructions may be directed to a human operator to provide guidance for corrective control of the blast furnace, or they may directly instruct the blast furnace's thermal controller to execute such instructions without human interaction.

[0004] Therefore, real-world (measured) operational data from multiple blast furnaces are used in conjunction with a simulation model (transient model) of the blast furnace process to train a recurrent neural network model via reinforcement learning. This can be understood as offline RL model training at both the data level and the simulation model level. Multiple additional features can be generated from historical data, providing better insights into the characterization of the blast furnace process. These features are either phenomena defined by rules implemented from the raw, recorded data, or predictions of process phenomena available in the form of predictions provided by machine learning models.

[0005] During training, the RL model provides recommendations for operating instructions to the blast furnace's main actuator, such as: tuyeres and blast setpoints, and parameters like pulverized coal injection (PCI) rate (kg / s) and blast flow rate (Nm³). 3 / s), oxygen enrichment (%), and / or load composition and loading setpoints, such as coke yield (kg / load), basicity, load distribution, etc. The recommendations provided ensure that the objective function will be optimized when the process is in thermal equilibrium, after the above recommendations are manually implemented by a virtual operator (autonomy level 5 to maximum autonomy level) or a human operator. The objective is defined by blast furnace experts and may consist of multiple objectives, such as (1) minimizing fuel consumption, (2) maximizing blast furnace life, (3) minimizing CO2 rejection, and (4) stabilizing the iron quality and quantity for blast furnace operation. Each objective is weighted (e.g., by experts) to define the global objective used to train the RL model. When the model is trained and deployed in production, it can continue to learn from the deviation between the global objective and the actual objective, which is achieved after the recommended operating instructions are executed for the thermal control of the corresponding blast furnace.

[0006] In one embodiment, a computer-implemented method is provided for training a reinforcement learning model to provide operating instructions for blast furnace thermal control. For example, the reinforcement learning model may be implemented by a recurrent neural network.

[0007] Domain-adaptive machine learning models trained through transfer learning process historical operational data as multivariate time series from multiple blast furnaces across multiple domains. This historical operational data reflects the thermal state of the corresponding blast furnaces in multiple domains. Typically, each blast furnace has thousands of sensors measuring operational parameters such as temperature, pressure, and chemical content. These parameters, measured at a specific point in time, define the corresponding thermal state of the blast furnace at that point. Due to the diverse characteristics of each blast furnace (e.g., operating mode, size, input materials (material composition), etc.), direct comparison of two blast furnaces (source and target blast furnaces) is not possible without specialized transformation of the multivariate time series data.

[0008] Domain-adaptive machine learning models generate a first domain-invariant dataset as output, representing the thermal state of any blast furnace, independent of its domain. Historical operating data is typically collected from multiple different blast furnaces (e.g., different sizes, operating under different conditions, etc.) in response to corresponding thermal control actions. Typically, each blast furnace corresponds to a specific domain, but the domain can also be a specific operation of the blast furnace. The domain-adaptive machine learning model is trained to perform a normalization operation on data obtained from different domains so that the data eventually become comparable.

[0009] Different transfer learning methods can be used. For example, a domain-adaptive machine learning model can be implemented by a deep learning neural network with convolutional and / or recurrent layers, trained to extract domain-invariant features from historical operational data as a first domain-invariant dataset. In this embodiment, transfer learning is implemented to extract domain-invariant features from historical operational data. The features in deep learning are abstract representations of specific blast furnace features extracted from multivariate time-series data generated by the operation of that particular blast furnace. By applying transfer learning, domain-invariant features can be extracted from multiple real-world blast furnaces independent of a specific furnace (i.e., independent of various domains).

[0010] In an alternative approach, a domain-adaptive machine learning model is trained to learn multiple mappings from multiple blast furnaces to a reference blast furnace corresponding to the original data. The reference blast furnace can be a virtual blast furnace representing a general blast furnace or an actual blast furnace. Each mapping is a representation of the transformation from the corresponding specific blast furnace to the reference blast furnace. In this approach, multiple mappings correspond to a first-domain-invariant dataset. For example, such a domain-adaptive machine learning model can be implemented using a generative deep learning architecture based on CycleGAN, which is popular in pseudo-image generation. CycleGAN is an extension of the GAN architecture involving the simultaneous training of two generator models and two discriminator models. One generator takes data from the first domain as input and outputs data for the second domain, while the other generator takes data from the second domain as input and generates data for the first domain. The discriminator models are then used to determine the credibility of the generated data and update the generator models accordingly. CycleGAN uses an additional extension to the architecture called cycle consistency. The idea behind this is that the data output by the first generator can be used as input to the second generator, and the output of the second generator should match the original data. The reverse is also true: the output of the second generator can be fed into the first generator as input, and the result should match the input of the second generator.

[0011] Cycle consistency is a concept in machine translation where a phrase translated from English to French should be translated back from French to English and be identical to the original phrase. The reverse process should also be correct. CycleGAN promotes cycle consistency by adding an additional loss to measure the difference between the generated output of the second generator and the original image, and vice versa. This acts as a regularization of the generator model, guiding the image generation process in new domains toward image translation. To adapt the original CycleGAN architecture from image processing to the processing of multivariate time series data to obtain first-domain invariant datasets, the following modifications can be implemented by combining recursive layers (e.g., LSTM) with convolutional layers to learn the temporal dependencies of multivariate time series data, as described by C. Schockaert, H. Hoyez, (2020) “MTS-CycleGAN: An Adversarial-based Deep Mapping Learning Network for MultivariateTime Series Domain Adaptation Applied to the Ironmaking Industry”, arXiv:2007.07518.

[0012] The obtained first domain-invariant dataset represents the thermal state of the blast furnace, which exists after the corresponding thermal control actions are applied to the respective blast furnace. After domain adaptation, this representation is no longer associated with a specific blast furnace (either in the form of a learned mapping of a reference blast furnace or in the form of extracted common features).

[0013] Simultaneously, transient models of the general-purpose blast furnace process are used to generate artificially manipulated data, serving as a multivariate time series reflecting the thermal state transition of the general-purpose blast furnace after the application of specific thermal control actions. The general-purpose blast furnace is a virtual device (similar to a reference blast furnace). The transient model is a transient numerical model with appropriate physical, chemical, thermal, and flow conditions, used to generate reasonable artificial data representing the thermal state of the general-purpose blast furnace. The transient model reflects the corresponding physical, chemical, thermal, and flow conditions of the general-purpose blast furnace and provides solutions for the upward gas flow and downward movement of the constructed solid layer within the general-purpose blast furnace during the exchange of heat, mass, and momentum transfer.

[0014] The model accepts the amount of loaded material and chemical analysis as input parameters, as well as hot blast conditions such as temperature, pressure, PCI rate, and oxygen enrichment. Transient models include: energy formulas for predicting hot metal temperature, formulas for calculating the chemical composition of hot metal, and gas-phase formulas for predicting top gas temperature, efficiency (Eta CO), and pressure. Due to the transient nature of the model, artificial dynamic time-series data can be generated by varying the input parameters over time, similar to real-world blast furnace operation. Advantageously, transient models can use a range of input parameters exceeding the range covered by historical operating data of real-world blast furnaces. In other words, the parameter range for a general blast furnace can be extended to an operating parameter space not covered by real-world blast furnace operating data.

[0015] A typical blast furnace is divided into a finite number of layers along its height. Each layer consists of a single load of raw materials (such as iron ore and coke). These layers represent computational units on which formulas are numerically solved. A raceway submodel is used to define boundary conditions for gas-phase properties (such as composition, velocity, and temperature), while the boundary conditions for the solid phase are defined as the composition of the loaded material at room temperature. This raceway model is described in “A reduced-order mathematical model of the blast furnace raceway with and without pulverized coal injection for real-time plant application”, International Journal of Modelling and Simulation, DOI:10.1080 / 02286203.2018.1435759.February 2018. Pulverized coal is injected into the blast furnace tuyeres to reduce coke consumption and lower the cost of hot metal production. Understanding the combustion behavior of pulverized coal and the accumulation of unburned coke in the blast furnace raceway zone is crucial. This paper describes a blast furnace reduced-order raceway model for real-time plant applications. The model is capable of predicting the radial temperature and gas composition distribution in the raceway region with and without pulverized coal injection (PCI). The effects of all key operating process parameters (such as PCI rate, blast temperature, blast volume, oxygen enrichment, and steam addition) on raceway combustion behavior, temperature and gas composition distribution, and raceway depth have been investigated and validated to the extent possible using literature and plant databases.

[0016] Fully resolving both the gas and solid phases is computationally very expensive. Therefore, according to this embodiment, to conserve computational resources (and thus energy), the gas phase can be treated as a steady state because the gas's drag time (approximately 3 seconds) is much smaller than the time step (approximately 2 minutes). However, the solid phase is considered a transient phase. The algorithm first solves the gas phase formula in an iterative sequence to satisfy the relative tolerances of the parameters at each time step. When the gas phase parameters converge to the defined tolerances, the solid phase formula is solved in the same time step sequence. The time loop continues until the end of the simulation. Gas and solid parameters, as well as transfer parameters such as heat and mass transfer, are updated at the beginning of each time step. In the sequential approach, once one parameter is solved, the others are considered known, meaning the old values ​​are used. This allows nonlinear terms and coupling parameters to be solved, avoiding complex and expensive block solvers.

[0017] In one implementation, the transient model has multiple computational units, each representing a corresponding layer of a general-purpose blast furnace consisting of a single load of raw materials. Each computational unit solves the gas-phase formula in an iterative sequence to satisfy the relative gas-phase parameter tolerances at each time step. When the gas-phase parameters converge to the predefined tolerance values, the solid-phase formula is solved in the same time-step sequence.

[0018] The iterative solution of the gas phase formula includes each iteration of the pressure-velocity correction loop: calculating the gas, solid, and liquid properties; calculating the reaction rate and heat transfer coefficient; and calculating the gas temperature, type, velocity, and pressure drop.

[0019] Once the gas phase parameters have converged to the predefined tolerance value, the calculation continues, continuously solving the solid phase formulas within the same time step, including: calculating the solid temperature and type; calculating the liquid temperature and type; and calculating the solid velocity.

[0020] The artificial operation data obtained from the transient model is then processed by a generative deep learning network trained on multivariate time series of historical operation data. This allows the artificial operation data to be augmented with features from real-world operation data to make them more realistic. A properly trained generative deep learning network can augment the artificial data in a way that makes the augmented synthetic operation data indistinguishable from real-world operation data for experts. This is beneficial for training reinforcement learning models with data that have features similar to real-world test inputs, which are anticipated when operating the reinforcement learning model during the prediction phase. In other words, the processing of the artificial operation data generates a second domain-invariant dataset augmented with features learned from the historical operation data. Although the second domain-invariant dataset is merely a synthetic dataset computed based on the transient model, it is still a domain-invariant dataset, displaying characteristic features present in the real-world historical operation data time series.

[0021] The reinforcement learning model is now trained using a combination of first- and second-domain invariant datasets. If training relies solely on the first dataset, the reinforcement learning model cannot learn control commands for optimizations not yet applied to multiple blast furnaces. By combining this real-world training dataset with artificially generated datasets, the transient model can be used to simulate the response of a general-purpose blast furnace to alternative control actions applied to a given thermal state of the general-purpose blast furnace under varying optimization objectives. When processing the combined first- and second-domain invariant datasets, the reinforcement learning model determines the reward for a specific thermal control action used by the transient model to compute the second-domain invariant dataset, given the objective function and the current state of the blast furnace. The reward functions describe how the reinforcement learning model (i.e., the agent) should behave. In other words, they have prescriptive content that dictates what the agent should accomplish. There are no absolute restrictions, but if the reward function “behaves better,” then the agent learns better. In practice, this means faster convergence and the agent is less likely to get stuck in local minima. For example, the reward function can measure “how far” from the Pareto front of the multi-objective function that a specific thermal control action is guiding the process. By definition, a Pareto front is a set of non-dominated solutions that is selected as optimal if no objective can be improved without sacrificing at least one other objective. For a given objective, the measure of the incremental improvement (delta) of the other objective can be, for example, by gradient analysis. The reward function can be a function of those measurements characterizing the Pareto front. Those skilled in the art can use other suitable reward functions.

[0022] If the determined reward is lower than the predetermined minimum reward, then the recommended thermal control action (control command) is not optimal in terms of its expected impact on the blast furnace thermal state. In this case, an alternative control action can be simulated using a transient model. For this purpose, based on the current environment of the reinforcement learning model and the output of the thermal control action at the current learning step (i.e., the control action that leads to an excessively low reward), a genetic search and / or Bayesian optimization algorithm guides the search for modified parameters (i.e., the input parameters of the transient model) for further (alternative) thermal control actions. The transient model then regenerates a second domain-invariant dataset (the updated second dataset) based on the modified parameters. The updated second dataset is then fed into the input layer of the reinforcement learning model, and a new reward is determined for the updated second dataset. This process is performed iteratively until the reinforcement learning model has learned to output optimized operating commands for optimized thermal control actions for any foreseeable situation.

[0023] Once the reinforcement learning model has been trained as described, it can be operated to predict optimized operating instructions for at least one actuator of a specific blast furnace in production, based on the current operating state data of that blast furnace. In other words, the trained reinforcement learning model receives test input data that includes operational data matching the input layer of the reinforcement learning model and specifies the current (thermal) state of the blast furnace. The model processes the test input data and provides predictions of optimized operating instructions corresponding to the thermal control actions applied to the blast furnace as output, to achieve the optimization result given a objective function.

[0024] Advantageously, each prediction dataset can be used to further improve the training of the reinforcement learning model. For this purpose, after applying a thermal control action to at least one actuator according to the optimized operating instructions (predicted output), the model determines the reward based on the new state of the specific blast furnace after executing the thermal control action. If the reward is below a predefined threshold, the transient model regenerates second-domain invariant data for one or more alternative operating instructions to retrain the reinforcement learning model. This retraining can be applied after applying any thermal control action according to the corresponding predicted optimized operating instructions.

[0025] Advantageously, the reinforcement learning model is trained to learn optimized operational instructions such that the associated target measurement lies within the predefined range of the Pareto front of the corresponding multidimensional objective function.

[0026] In one embodiment, the transient model has multiple computational units, each representing a corresponding layer of raw material in a general-purpose blast furnace, consisting of a single loading. Each computational unit solves the gas-phase formula in an iterative sequence to satisfy the relative gas-phase parameter tolerances at each time step. When the gas-phase parameters converge to a predefined tolerance value, the computational units solve the solid-phase formula in the same time-step sequence.

[0027] Table 1 - Levels of Thermal Control Automation for Blast Furnaces

[0028]

[0029] Table 1 describes the five levels of automation for blast furnace control. Recommendation reinforcement learning models and further associated machine learning models (e.g., process phenomenon prediction, hot metal temperature prediction, etc.) that generate high-level contextual information describing the process can be used to achieve Level 4 or 5 automation, while individual associated machine learning models can only contribute to Level 2 or 3 automation. Training the recommendation model without associated machine learning models may result in Level 3 automation. The method disclosed herein for training reinforcement learning models for recommending (predicting) optimal thermal control actions can be used to achieve Level 4 or 5 automation, provided that the process is accurately represented by high-level contextual data generated by the machine learning model and additional sensors for process characterization, as described in more detail in the specific implementation. These associated machine learning models can be used to add further data augmentation capabilities to improve the training dataset for reinforcement learning, as they provide predictions based on received operational data (raw sensor data), which is used as further input to train reinforcement learning models that go beyond domain-invariant process data. Through this additional “contextual” information, the reinforcement learning model gains knowledge about new dimensions that can be used to learn, more precisely, the optimal actions for thermal control.

[0030] When using such correlated machine learning models to predict information about the future thermal evolution of a specific blast furnace state based on historical operating data and / or further measured environmental data related to the blast furnace environment, the correlated machine learning models need to be trained accordingly to supplement the historical operating data (obtained from sensors) with future multivariate time series data related to future time points. The generated future multivariate time series can then be processed by a domain-adaptive machine learning model in the same manner as the historical operating data to augment the first domain-invariant dataset with data related to future time points.

[0031] The associated machine learning models can be trained as follows: In a first training step, multiple base models are trained using one or more machine learning algorithms, leveraging different selections of operational and / or environmental data, to provide base model-specific future multivariate time series data as training input to a particular model among the machine learning models. Thus, each base model focuses on a single, specific aspect of the blast furnace process (e.g., prediction of hot metal temperature trends over a given future time interval). In a second training step, the associated machine learning models are trained with base model-specific future multivariate time series data to learn which combination of base models is best suited for which state of the blast furnace.

[0032] Other aspects of the invention will be realized and obtained through the elements and combinations specifically described in the appended claims. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and not intended to limit the invention as described. Attached Figure Description

[0033] Figure 1 A simplified diagram of an embodiment of a computer system for training a reinforcement learning model to provide operating instructions for blast furnace thermal control is shown.

[0034] Figure 2 It is a simplified flowchart of a computer-implemented method that can be executed by an embodiment of a computer system;

[0035] Figure 3A A simplified flowchart reflecting the processing of a transient model according to an embodiment having multiple computing units is shown, where each unit represents a corresponding layer of a general blast furnace;

[0036] Figure 3B The computing unit is shown as a visual representation of a blast furnace;

[0037] Figure 4 An exemplary embodiment of reward calculation for a reinforcement learning model is shown;

[0038] Figure 5 The Pareto front of the point cloud in the target space, which serves as the reward function, is shown.

[0039] Figure 6A , Figure 6B An example of the Pareto front in the objective space of the corresponding two-dimensional objective function of a blast furnace is shown for visualization purposes;

[0040] Figure 7 The use of an additional machine learning model for training data augmentation is illustrated according to an embodiment;

[0041] Figure 8 The example illustrates using vent images for training data augmentation by employing an additional deep learning model.

[0042] Figure 9 This demonstrates the use of additional sensors to characterize the state of the blast furnace in order to train a reinforcement learning model; and

[0043] Figure 10 This is a diagram illustrating examples of general-purpose computer devices and general-purpose mobile computer devices that can be used with the technologies described herein. Detailed Implementation

[0044] Figure 1A simplified diagram of a basic embodiment of a computer system 100 for training a reinforcement learning model 130 to provide operating instructions for blast furnace thermal control is shown. Figure 1 Is Figure 2 In the context described, Figure 2 This is a simplified flowchart of a computer-implemented method 1000, which can be executed by an embodiment of computer system 100. Therefore, the following... Figure 2 In the context of Figure 1 The description involves the reference numerals of the two figures.

[0045] In one embodiment, computer system 100 is communicatively coupled to multiple blast furnaces BF1 to BFn. BF1 to BFn may belong to different domains and provide historical operating data 21, obtained as a multivariate time series, reflecting the thermal state of the respective blast furnaces. Examples of such historical operating data include, but are not limited to, loading material quantity and chemical analysis, temperature, pressure, PCI rate, and oxygen enrichment, with energy formulas for predicting hot metal temperature, one or more class formulas for calculating the hot metal chemical composition, and one or more gas-phase formulas for predicting top gas temperature, efficiency (Eta CO), and pressure.

[0046] In a real-world blast furnace, different domains can be associated with different combinations of parameter values ​​in historical operating data 21 describing the thermal state of the blast furnace in different domains, despite the similarities between these thermal states. Therefore, system 100 has a domain-adaptive machine learning model DAM 110 to generate a first domain-invariant dataset 22 representing the thermal state of any of blast furnaces BF1 to BFn, independent of the domain. DAM 110 has been trained using a transfer learning method TL 111. In one implementation, DAM 110 can be implemented by a generative deep learning neural network GDL1 113 with convolutional and / or recurrent layers, trained to extract domain-invariant features from historical operating data 21 as the first domain-invariant dataset 22.

[0047] In an alternative implementation, DAM 110 can be implemented using a generative deep learning architecture (e.g., based on the previously described CycleGAN architecture), which has been trained to learn multiple mappings 112 of the corresponding original data from multiple blast furnaces BF1 to BFn to a reference blast furnace BFr. Thus, each mapping is a representation of the transformation from the corresponding blast furnace (e.g., BF1) to the reference blast furnace BFr. In this implementation, the multiple mappings correspond to a first domain-invariant dataset 22.

[0048] System 100 further includes an Artificial Data Generator (ADG) module 120, which is configured to generate 1200 artificial operation data points 24a as a multivariate time series reflecting the thermal state of a general blast furnace (BFg) for a specific thermal control action 26a. For this purpose, ADG 120 uses a transient model 121 of the general blast furnace process. The transient model 121 is a simulation model that reflects the corresponding physical, chemical, thermal, and flow conditions of the general blast furnace and provides a solution for the upward gas flow and downward movement of the solid layer constructed in the general blast furnace during the exchange of heat, mass, and momentum transfer. Typically, the simulation model is based on simulation parameters corresponding to such real-world state parameters monitored in historical operation data.

[0049] Briefly turn Figure 3A The transient model 121 comprises multiple computational units, each representing a corresponding layer of a general blast furnace BFg consisting of a single-load feedstock. Each computational unit solves the gas-phase formula in an iterative sequence to satisfy the relative gas-phase parameter tolerances at each time step (iteration interval). Once the gas-phase parameters converge to the predefined tolerance values, the solid-phase formula is solved in the same time-step sequence. For each iteration of the pressure-rate correction cycle, the iterative steps for solving the gas-phase formula may include:

[0050] - Calculate the properties of 3300 gases, solids and liquids;

[0051] - Calculate the reaction rate and heat transfer coefficient at 3400; and

[0052] - Calculate the temperature, type, velocity, and pressure drop of a 3500 gas.

[0053] Sequential solutions to solid-state equations may include:

[0054] - Calculate the temperature and type of 3600 solids;

[0055] - Calculate the temperature and type of liquid 3700; and

[0056] - Calculate the velocity of the solid at 3800.

[0057] Now go to Figure 3BThe computational unit CC is shown in the visualization representation 300 of the blast furnace. The transient model 121 can receive one or more of the following input parameters 302: the amount and chemical analysis of the loading material 302-1, temperature, pressure, PCI rate 302-2, and oxygen enrichment. Furthermore, the furnace profile 302-3 describes the geometry of the blast furnace and therefore affects the loading material transfer time (e.g., the transfer time for a high-capacity blast furnace can be 8 hours, while for a short-capacity blast furnace it can be only 6 hours). The furnace profile 302-3 is a fixed parameter for each blast furnace used in the artificial data generation. It will be clear to those skilled in the art that the transient model takes into account the geometry of the blast furnace. The transient model produces outputs 303, such as an energy formula predicting the hot metal temperature, one or more class formulas calculating the hot metal chemical composition 303-2, and one or more gas-phase formulas predicting the top gas temperature, efficiency (Eta CO), and pressure (see top gas condition 303-1).

[0058] In other words, a transient (simulation) model is a numerical model with appropriate physical, chemical, thermal, and fluid conditions to generate plausible artificial data. Due to the transient nature of the model, artificial dynamic time-series data, similar to real-world furnace operations, can be generated by varying the input parameters over time. Therefore, the range of (parameter) data can be extended to a wider operating space that cannot be covered by actual blast furnace data obtained from real-world blast furnaces.

[0059] In the transient model, the furnace is divided into a finite number of layers along its height. Figure 3B In the blast furnace 300, each layer is separated by a solid horizontal line 301. Each layer consists of a single load of raw materials, in this case, iron ore and coke. These layers represent computational units CC310 as described above, on which formulas are numerically solved. In one embodiment, a roller sub-model 320 is used to define boundary conditions for gas phase characteristics (such as composition, velocity, and temperature), while the boundary conditions for the solid phase are defined as the composition of the loaded material at room temperature. The internal state 304 of the blast furnace 300 includes sub-states of gas, solid, and liquid phases. The sub-states of the gas phase can be characterized as: temperature (Tg, K), pressure (p, Pa), velocity (Vg, m / s), and species (CO, CO2, H2, H2O, N2). The sub-states of the solid phase can be characterized as: temperature (Ts, K), velocity (Vs, m / s), and species (Fe2O3, Fe3O4, FeO, Fe, slag, coke, coke ash). Furthermore, the sub-states of the liquid phase can be characterized by temperature (Tl, K) and species (Fe, slag, FeO).

[0060] Complete analysis of both the gas and solid phases is computationally very expensive (and time-consuming). Therefore, to save time and energy, the gas phase can be considered steady-state because the drag time of the gas (approximately 3 seconds) is much smaller than the time step constrained to the iteration interval (approximately 2 minutes). However, the solid phase is considered transient. Solution algorithm (see...) Figure 3A First, the gas-phase formulas are solved iteratively to satisfy the relative tolerances of the parameters at each time step. Once the gas-phase parameters converge to the predefined tolerance values, the solid-phase formulas are solved sequentially at the same time steps. The time loop continues until the simulation ends. Gas and solid parameters, as well as transfer parameters such as heat and mass transfer, are updated at the beginning of each time step. In the sequential approach, once a parameter is solved, the others are considered known, meaning the old values ​​are used. This allows for the solution of nonlinear terms and coupling parameters, avoiding complex and expensive block solvers.

[0061] As described above, the artificial operation data 24a generated by the transient model 121 is generated according to a mathematical formula that produces clean data because the artificial operation data 24a does not show any real-world characteristics, such as noise or offset caused by the corresponding measurement / sensor equipment. To train the reinforcement learning model RILM 130 for highly accurate predictions, it is desirable to provide model 130 with training data that reflects the features of the real-world operation data used as test inputs to RILM 130. Therefore, ADG 120 uses a generative deep learning network GDL2 122 with recursive layers to generate a second-domain invariant dataset 23a by passing features learned from historical operation data 21 to the artificial operation data 24a. GDL2 122 has been trained on the multivariate time series of historical operation data 21 to learn the real-world features from the historical operation data and apply the learned patterns to the simulated artificial operation data 24a. This produces a purely synthetic dataset 23a that reflects the thermal state of a general blast furnace BFg in response to thermal control actions 26a. It should be noted that learning the natural features of each signal in a context given other signals by utilizing generative deep learning networks with recursive layers is similar to well-known techniques applied to images to learn the style of a particular set of pictures and apply that style to any other image. Similar techniques can be applied to multivariate time series, and these methods can be solved, for example, when applicable to multivariate time series data, via the aforementioned CycleGAN-based algorithm.

[0062] The first domain-invariant dataset 22 and the synthesized second domain-invariant dataset 23 are provided as training data to the RILM 130. The RILM 130 determines the reward 131 for a specific thermal control action 26a by processing the combined first domain-invariant dataset 22 and second domain-invariant dataset 23a, given a objective function. Based on this training data, the RILM 130 learns thermal control actions that depend on the (general) blast furnace (environment) conditions. For example, this environment can be defined by the blast furnace operation, material composition, etc.

[0063] Depending on the reward 131, ADG 120 regenerates a second-domain invariant dataset of 1300 based on the modified parameters 123-2. The parameter generator PG 123 uses a genetic search and / or Bayesian optimization algorithm 123-1 to guide the search for modified parameters for further hot control actions based on the current environment 25a of RILM 130 and the output of the hot control action 26a of the current learning step. By modifying the parameters, the transient model simulates the hot state of the further control actions. The regenerated second-domain invariant dataset is then provided to RILM 130 as new training input, and the reward is determined again for the new training input. This process is repeated until the current reward exceeds a pre-defined reward threshold of 1500 to learn optimized operational instructions for the optimized hot control actions.

[0064] The following text describes examples of real-world scenarios used for reward calculation, such as... Figure 4 As shown. It should be noted that those skilled in the art can use other appropriate reward functions to implement reinforcement learning models. The following example scenario describes optimization to identify the optimal executor value for a simple biobjective function to be maximized using a genetic search algorithm.

[0065] Objectives: Maximize quality (with constant silicon content) and maximize yield.

[0066] Actuator: PCI speed (kg / s), airflow rate (nm) 3 / s), coke yield (kg / load)

[0067] award:

[0068] =1 / (Euclidian_dist_to_pareto_front)

[0069] As an example, an approximation is made through incremental analysis of improvements to each objective:

[0070] 1 / eucl_dist((quality_prev, prod_prev), (quality_new, prod_new))

[0071] In this example, the reward constraint is only effective if the genetic search algorithm is guaranteed to converge to the Pareto front. That is, in this example of maximizing a biobjective function, improvements in both quality and production are positive between two consecutive iterations.

[0072] Initial blast furnace thermal state (current environment): S_init

[0073] - Iteration 1:

[0074] Actuator value = [PCI_1, blast_flow_rate_1, coke_rate_1]

[0075] Target measurement = quality_1; prod_1

[0076] - Iteration 2:

[0077] Actuator value = [PCI_2, blast_flow_rate_2, coke_rate_2]

[0078] Target measurement = quality_2; prod_2

[0079] Reward=R_2=1 / eucl_dist((quality_1, prod_1), (quality_2, prod_2))

[0080] - Iteration 3:

[0081] Actuator value = [PCI_3, blast_flow_rate_3, coke_rate_3]

[0082] Target measurement = quality_3; prod_3

[0083] Reward=R_3=1 / eucl_dist((quality_2, prod_2),(quality_3,prod_3))

[0084] ………………………

[0085] - Iteration i:

[0086] Actuator value = [PCI_i, blast_flow_rate_i, coke_rate_i]

[0087] Target measurement = quality_i; prod_i

[0088] Reward=R_i=1 / eucl_dist((quality_i-1, prod_i-1),(quality_i,prod_i))

[0089] ………………………

[0090] - Iteration options: (Reaching the Pareto front)

[0091] Actuator value = [PCI_opt, blast_flow_rate_opt, coke_rate_opt]

[0092] Target measurement = quality_opt; prod_opt

[0093] Reward=R_opt=1 / eucl_dist((quality_opt-1, prod_opt-1), (quality_opt, prod_opt))

[0094] Without using a genetic search algorithm, a time-consuming random search can be performed. Figure 5 In this case, the Pareto front is represented by points (quality_i, prod_i) with a point-filled pattern (at the boundary of the point cloud in the target space). In this case, the reward for each point (quality_i, prod_i) can be computed as the reciprocal of the Euclidean distance, and is computed after the Pareto front has been identified (rather than during the random search process).

[0095] In summary, the reinforcement learning model 130 is trained to learn optimized operational instructions such that the associated target measurement lies within the predefined range of the Pareto front of the corresponding multidimensional objective function.

[0096] Once learning is complete, the RILM 130 has been trained to provide 1600 optimized operating instructions in response to test inputs and current operating data describing the current state of the blast furnace (see [link to relevant documentation]). Figure 2 ), used for thermal control of real-world blast furnaces. Optionally, training on RILM 130 can continue in online mode while the blast furnace is operating.

[0097] In online mode, reinforcement learning model 130 predicts 1700 based on the current operating state data of the blast furnace (see [link]). Figure 2Optimized operating instructions are used for at least one actuator of a specific blast furnace in production. It is assumed that the optimized operating instructions' thermal control actions are applied to the blast furnace (either by an operator or automatically via a corresponding control system). After applying the thermal control actions to at least one actuator according to the optimized operating instructions, a reward is now determined based on the new state of the blast furnace reached after the execution of the thermal control actions. Again, the determined reward is compared 1500 to a predefined reward threshold. If the reward is below this threshold, ADG 120 regenerates (using an instantaneous model 121) second-domain invariant data for one or more alternative operating instructions to retrain the reinforcement learning model 130.

[0098] Figure 6A and Figure 6B The diagram shows the Pareto front (dashed line) in the target space for the two-dimensional objective function (with two objectives O1, O2) for the corresponding blast furnace states BFS1, BFS2 (for visualization purposes). The RILM model needs to learn the optimal control commands for the blast furnace so that the relevant objective measurements lie on the Pareto front. In these diagrams, the objectives for each historical and artificial data sample have been computed. These diagrams illustrate the limitations of historical data, which is often limited to a few operating modes of the blast furnace, leading to clustering in the target space. Therefore, type 22-2 bullets are associated with domain-invariant datasets obtained from historical data. Type 23a squares are associated with domain-invariant datasets based on artificial (simulated) data. Type 22-1 bullets are associated with data generated by a deep generative model trained from historical data 21 (raw data) or 22 (domain-invariant raw data). This deep generative model acts as a high-level interpolation algorithm, providing new raw data generated from historical data. Therefore, the generated data can only be relatively close to the existing historical data. The generation of this data associated with type 23a is... Figure 1 and Figure 2 A more detailed description is provided below. Figure 6B In this study, the type 22-3 triangles were correlated with online data acquired during blast furnace operation and used for the online training model of RILM 130. Type 22-3 triangles are naturally closer to the Pareto front because they are generated by a training model that provides optimized operational instruction recommendations (see [link to relevant documentation]). Figure 2 (Prediction 1700). However, in order to further optimize the recommendations of operational instructions based on this data, an online retraining of RILM 130 was triggered.

[0099] In one embodiment, system 100 may include a data augmentation module DA 140 to augment raw operational data 21 measured by sensors on a blast furnace using one or more specially trained machine learning models ML1 to MLn to predict information about the future thermal evolution of the blast furnace state, or any other information related to the current thermal state (e.g., process phenomenon prediction, such as virtual sensors providing measurements at a higher frequency than actual sensors). This prediction serves the same purpose as the raw data used to train the RILM 130 model and is used in the same manner as the raw data 21 (historical operational data). An example of such a specially trained machine learning model is a model predicting the hot metal temperature over 3 hours. This prediction of the hot metal temperature can then be used to train the RILM 130. This data augmentation further improves the training dataset used for reinforcement training of the RILM 130 and results in improved prediction accuracy of the reinforcement learning model. Optionally, new sensors, such as… Figure 9 This allows for a more precise characterization of the blast furnace state used to train the RILM 130. For example, if some properties of the raw materials loaded in the furnace are missing (e.g., porosity, moisture), they can be measured (using additional sensors), or they can be potentially estimated using machine learning models ML1 to MLn.

[0100] Below is a list of machine learning (ML) models that facilitate data augmentation to allow for a more accurate representation of the blast furnace's state, and thus allow for more precise training of the RILM 130:

[0101] a) ML for advanced data validation: Any anomalies in the raw data provided by the blast furnace sensors can be detected before training a machine learning model, or can be used as input for the production of a deployed machine learning model.

[0102] b) ML used to predict blast furnace thermal state and hot metal production KPIs (Key Performance Indicators).

[0103] c) ML for loading matrix optimization

[0104] d) ML for vent camera-based process inspection

[0105] e) Recommended ML taps for optimal operation

[0106] f) ML for TMT SOMA phenomenon detection and KPI calculation / prediction

[0107] g) ML methods used to label phenomena through process rules defined by process engineers (potentially using outputs generated by machine learning models) or through supervised or unsupervised machine learning or pattern detection models.

[0108] h) ML used to predict phenomena based on the labels generated in g)

[0109] i) ML for process prediction

[0110] j) ML for predictive and prescriptive maintenance

[0111] k) ML for advanced contextual representation learning: Environmental sensors can be used to train unsupervised deep learning models for learning representations, thereby expanding the datasets required for the above use cases.

[0112] Figure 7 The method for implementing the DA140, which is used to train a machine learning model to predict the temperature of hot metal over 3 hours, is described in more detail. Figure 7 The process of training the machine learning model MLT (706) to predict the temperature of hot metal at a future time point (e.g., within 3 hours) based on predictions from multiple machine learning models (called the base model) is illustrated. The base model is trained (703) to generate predictions to augment the original measured data.

[0113] To this end, 703 multiple base models are trained using different variable selections (process variable 701 and / or contextual variable 702) and / or machine learning algorithms. Process variable 701 is raw data (operational data) measured directly on the blast furnace by corresponding sensors. Contextual variable 702 is measured by any other sensor measuring environmental variables such as noise, images, etc. Process and contextual variables are variables that can be used to train machine learning models.

[0114] Each base model provides outputs 704 and 705, which can be used to train the 706 MLT (typically used to train any machine learning model to predict parameters other than the hot metal temperature over 3 hours) to make better predictions for the parameters in question than any prediction from the base models. The purpose of the base models is to generate additional information for training more accurate predictive models (i.e., meta-models such as MLTs). The predictive model MLT also uses process variables 701 and contextual variables 702 as inputs to learn which combination of base models is best suited for which state of the blast furnace. In other words, the meta-model is learning how to combine the outputs of all the base models to make more accurate and precise predictions of specific blast furnace state parameters. Some base models may not predict the hot metal temperature over 3 hours, but they may predict trends in the hot metal temperature—for example, whether the temperature is rising, falling, or stabilizing—or they may predict the occurrence of specific process events in the near future, and so on. In other words, the base model generates additional information related to the process (process information PI 705) or additional information about the characteristics of the hot metal temperature, which is already used by the base model to predict the hot metal temperature (BMP 704), as output (e.g., trend prediction). Once the MLT has been trained based on the outputs of various base models 706, it provides more accurate predictions of MLTP than any base model (BMP 704) 705.

[0115] In this example, the process information PI 705 can provide input information for the MLT, such as feature predictions within the range [0, 6h], including but not limited to clustering, process phenomena, process / contextual variables, or features. These outputs are process-related and provide new inputs with potentially higher correlation to the hot metal temperature predicted by the MLT. The base model prediction for hot metal temperature, BMP 704, can provide information such as trends in hot metal temperature over 3 hours and 6 hours (e.g., high rise, medium rise, low rise, stable, low fall, medium fall, high fall) or predicted hot metal production quality within the time range. BMP 704 is the output of the base model directly related to the output of the MLT, or the same output or a feature of the MLT. An example of the same output is "hot metal temperature over 3 hours," and an example of this output feature could be "temperature trend" predicted by the base model.

[0116] The following text will describe some examples of the above list of machine learning models in more detail.

[0117] Advanced data validation:

[0118] The data validation pyramid can be defined by multiple data validation levels, as described below, starting from the lowest level and ending at the highest level.

[0119] - Sensor maintenance and calibration: Procedures for sensor maintenance and calibration can be implemented. Artificial intelligence (AI) can be involved to optimally schedule maintenance actions and specify the best actions to be performed to keep the sensor in operational mode for as long as possible.

[0120] - Minimum / maximum values ​​for individual sensor signals: Level 1 anomaly detection defines the minimum and maximum allowed values ​​for each sensor signal in the raw data. These minimum and maximum values ​​are constant and therefore independent of process operation. Conditional process minimum / maximum values ​​can be configured in rules defined by process experts to accommodate various scenarios.

[0121] - Outlier and anomaly detection in individual sensor signals: The typical methods listed below become increasingly complex:

[0122] i) Statistical outlier values ​​of amplitude:

[0123] A data analysis method for detecting point anomalies, wherein, by definition, point anomalies are specified by process experts and are based on the magnitude of the typical autocorrelation depth of a time series recorded by sensors, which deviates from the average value within a moving time window of length L.

[0124] ii) Supervised anomaly detection:

[0125] Supervised algorithms learn known patterns in sensor signals in order to detect anomalies.

[0126] iii) Unsupervised outlier detection:

[0127] These methods detect outliers by applying clustering algorithms after calculating features from sensor signals. Therefore, this approach is not limited to anomalous amplitude values ​​in a given scenario, but can also consider spectral information or any other features defined by the characteristics.

[0128] - Anomaly Detection in Multi-Sensor Signals: Due to the large number of sensors, manual cross-checking between redundant sensor signals is insufficient to detect complex contextual anomalies in the data. Rule-based methods are often limited because they only validate known relationships. The same limitation applies to supervised data-driven models that have been trained to detect known anomalies. Unsupervised data-driven methods serve as a complementary validation step to ensure that both known and unknown anomalies are detected. Contextual anomalies can be detected by data-driven models that have learned the correlations between sensor signals and are therefore able to detect whether sensor measurements deviate from normal operation in a given context defined by the process. Root cause analysis is achieved by combining machine learning in unsupervised data-driven anomaly detection to discover causal relationships.

[0129] - Cross-checking sensor and simulation model results: If a simulation model describing the process is available, cross-checking the model results with the raw sensor data provides expert-level autonomous data validation. However, this validation is limited to the operating conditions inherent in the simulation model assumptions.

[0130] The data validation pyramid aims to detect anomalies in received operational data (raw data). Anomalies may be related to faulty sensors, but also to the process. In cases of process anomalies, rare process events can be correctly labeled for developing specific machine learning models, such as "few-shot learning" (FSL), to ensure their correct detection or prediction. FSL is a known machine learning paradigm for learning from a limited number of examples with supervised information. One approach aimed at distinguishing between process anomalies and anomalies related to faulty sensors is root cause analysis. Analyzing the causal relationships leading to anomaly detection allows for the classification of anomalies as process- or sensor-related. To this end, process engineers are defining rules or machine learning models and training a semi-supervised classifier from the causal relationships and labels generated by these rules.

[0131] Blast furnace thermal status prediction:

[0132] This involves Figure 7 The machine learning model MLT is used as an example. MLT provides insights into the future thermal state of a blast furnace and the characteristics of hot metal production. Based on relevant process variables and other contextual variables that can be used to predict the thermal state or hot metal production characteristics of a blast furnace, MLT is trained to predict the following metrics over a given time range:

[0133] - Hot metal temperature trends over 3 and 6 hours: High increase, Medium increase, Low increase, Stable, Low decrease, Medium decrease, High decrease

[0134] - Prediction of thermal metal silicon content over multiple time ranges from 1 hour to 6 hours

[0135] - Hot metal mass over multiple time ranges from the next 1 hour to 6 hours

[0136] The model can be trained by manually measuring the hot metal temperature of each casting, or by continuously training autonomously using dedicated sensors. The meta-model MLT can be trained by combining predictions from multiple base models as new inputs, thus implementing an ensemble modeling approach that produces predictions with reduced prediction bias or variance.

[0137] Load matrix optimization:

[0138] Load allocation is one of the most important actuators that operators can use to optimize gas utilization (etaCO) to minimize coke yield and reduce CO2 emissions. Load allocation always needs to be adapted to blast furnace operation and is a trade-off between optimal gas utilization, smooth load reduction, and wall / diaphragm (skin flow) temperature.

[0139] Today, some plants use load distribution models to assess the impact of a given loading matrix on load distribution and to determine the C / (O+C) ratio of the entire blast furnace throat diameter. This information is valuable and provides fairly good hints about the temperature distribution in the viscous zone. However, constraining the loading matrix in the model is not straightforward, and the model offers only limited assistance in finding the optimal loading matrix for a given operation.

[0140] The operator defines a loading matrix to ensure optimal material distribution on the blast furnace. This loading matrix includes various parameters such as the inclination of the chute and the number of rotations for each material type. A machine learning model can be trained to predict the optimal loading matrix based on the furnace's current thermal state, its predicted evolution, and its production KPIs. If there is insufficient variation in the loading matrix elements for a single blast furnace, the loading matrix prediction model can be trained using raw data from multiple blast furnaces, thus training the machine learning model.

[0141] Process inspection based on vent cameras:

[0142] This example involves Figure 8 The image 801 provided by the vent camera is analyzed using a combination of a convolutional neural network (CNN) and computer vision 803. The CNN and computer vision 803 are designed to detect phenomena 804 by applying computer vision to regions detected by a CNN-based region classifier 802 (e.g., classifying circles, spray guns, and injection areas in image 801c). Together with the vent image 801, the detected phenomenon labels can then be used as input to a further deep learning model 805 trained to predict process phenomena.

[0143] Another application of machine learning in analyzing blast furnace image sequences is encoding spatial-temporal features to enrich representations of blast furnace states that define the environment of reinforcement learning models. For this purpose, multimodal learning 808 can be used as a method for learning representations 809 of the environment from heterogeneous data such as images 801, multivariate time series 806, and sound 807. This allows for more advanced methods compared to unimodal machine learning that assumes pattern independence.

[0144] Recommended taps for optimal operation:

[0145] Machine learning models can recommend tap scheduling and its parameterization (e.g., clay type, etc.).

[0146] Phenomenon detection and KPI calculation based on TMT SOMA:

[0147] SOMA is an instrument that provides two-dimensional information on the temperature distribution at the top of a blast furnace. Temperature mapping can be processed using machine vision algorithms, which may be combined with machine learning models for predictive purposes. Figure 8 The processing pipeline described in [the document] for camera-based vent inspection can also be applied to SOMA.

[0148] Phenomenon labeling and prediction:

[0149] Labeling process phenomena ensures the creation of rich information to improve the learning of the relationship between actions and environment in RILM 130. Labels can be generated by rules defined by the process engineer or by a pattern detection model trained on patterns selected by the process engineer from historical data. The emergence of patterns can be detected algorithmically, such as by dynamic time wrapping of univariate or multivariate time series data, or by training a corresponding machine learning model through feature constraints. In addition to providing high-level contextual information to the RILM model, these labels can be used to train machine learning models to detect the occurrence of combinations of phenomena or to predict the occurrence of individual phenomena or combinations of phenomena. Training a supervised machine learning model from the generated labels requires a sufficient number of labels with adequate variance.

[0150] Predictive and normative maintenance:

[0151] Machine learning models can be trained to predict maintenance and recommend actions to postpone it, thereby extending the lifespan of a blast furnace or any asset associated with it. Several methods are known for this, such as applying supervised learning to predict the asset's "remaining useful life" or "time of failure." Unsupervised learning models can be trained to detect rare events, and the training dataset can be temporally clustered to train supervised models that predict these rare events. Root cause analysis of the predictions allows an autonomous system trained with maintenance actions recorded from past maintenance to prescribe the most known actions to delay maintenance.

[0152] Advanced contextual representation learning:

[0153] Reinforcement learning models require a representation of the context in order to better model the environment and learn the optimal action to take in response to that environment. To this end, such as... Figure 9As shown, multiple sensors can be formed and placed around the furnace 90 to record images (camera sensor 91), sound waves (sound sensor 92), vibrations (vibration sensor 93), and analyze air at different locations (gas sensor 94). The corresponding multimodal time series can be analyzed using a deep learning network to extract meaningful representations of the scenario, which can potentially be combined with process data or material descriptive data from the blast furnace. The material descriptive data corresponds to the chemical analysis of the materials, as well as other properties that may affect the thermal regulation of the blast furnace.

[0154] Figure 10 This is a diagram illustrating examples of a general-purpose computer device 900 and a general-purpose mobile computer device 950 that can be used with the technologies described herein. The computing device 900 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The general-purpose computer device 900 can correspond to... Figure 1 Computer system 100. Computing device 950 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, driver assistance systems, or in-vehicle computers of vehicles (e.g., vehicles 401, 402, 403, see below). Figure 1 ) and other similar computing devices. For example, computing device 950 can be used as a front end by a user (e.g., a blast furnace operator) to interact with computing device 900. The components, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the implementation of the inventions described and / or claimed in this document.

[0155] Computing device 900 includes a processor 902, a memory 904, a storage device 906, a high-speed interface 908 connecting the memory 904 and a high-speed expansion port 910, and a low-speed interface 912 connecting a low-speed bus 914 and the storage device 906. Each of the components 902, 904, 906, 908, 910, and 912 is interconnected using various buses and can be mounted on a common motherboard or otherwise suitably mounted. Processor 902 can process instructions for execution within computing device 900, including instructions stored on memory 904 or storage device 906, to display graphical information of a GUI on an external input / output device, such as a display 916 coupled to high-speed interface 908. In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and memory types, can be suitably used. Furthermore, multiple computing devices 900 can be connected, with each device providing a portion of the necessary operation (e.g., as a server group, blade server group, or multiprocessor system).

[0156] The memory 904 stores information within the computing device 900. In one embodiment, the memory 904 is one or more volatile memory cells. In another embodiment, the memory 904 is one or more non-volatile memory cells. The memory 904 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0157] Storage device 906 provides massive storage for computing device 900. In one embodiment, storage device 906 may be or include computer-readable media such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices, or device arrays, including devices in storage area networks or other configurations. The computer program product may be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 904, storage device 906, or memory on processor 902.

[0158] High-speed controller 908 manages bandwidth-intensive operations of computing device 900, while low-speed controller 912 manages lower bandwidth-intensive operations. This functional allocation is merely exemplary. In one embodiment, high-speed controller 908 is coupled to memory 904, display 916 (e.g., via a graphics processor or accelerator), and high-speed expansion port 910, which can accept various expansion cards (not shown). In another embodiment, low-speed controller 912 is coupled to storage device 906 and low-speed expansion port 914. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), can be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices such as switches or routers, for example, via a network adapter.

[0159] As shown in the figure, computing device 900 can be implemented in a variety of different forms. For example, it can be implemented as a standard server 920, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 924. Furthermore, it can be implemented in a personal computer such as a laptop computer 922. Alternatively, components from computing device 900 can be combined with other components in a mobile device (not shown), such as device 950. Each of such devices can contain one or more of computing devices 900, 950, and the entire system can consist of multiple computing devices 900, 950 communicating with each other.

[0160] The computing device 950 includes a processor 952, a memory 964, input / output devices such as a display 954, a communication interface 966 and a transceiver 968, and other components. The device 950 may also be equipped with storage devices, such as microdrives or other devices, to provide additional storage. Each of the components 950, 952, 964, 954, 966, and 968 is interconnected using various buses, and some components may be mounted on a common motherboard or installed in other suitable ways.

[0161] Processor 952 can execute instructions within computing device 950, including instructions stored in memory 964. The processor can be implemented as a chipset comprising individual or multiple analog and digital processors. The processor can provide, for example, coordination of other components of device 950, such as control of the user interface, applications running by device 950, and wireless communication via device 950.

[0162] Processor 952 can communicate with the user via control interface 958 and display interface 956 coupled to display 954. Display 954 can be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technologies. Display interface 956 may include appropriate circuitry for driving display 954 to present graphics and other information to the user. Control interface 958 can receive commands from the user and translate them for submission to processor 952. Furthermore, an external interface 962 can be provided to communicate with processor 952, enabling device 950 to perform near-field communication with other devices. External interface 962 may provide wired communication in some embodiments, wireless communication in others, and multiple interfaces may be used.

[0163] Memory 964 stores information within computing device 950. Memory 964 can be implemented as one or more computer-readable media, one or more volatile memory cells, and one or more non-volatile memory cells. Extended memory 984 can also be provided and connected to device 950 via extended interface 982, which may include, for example, a SIMM (Single In-line Memory Module) card interface. This extended memory 984 can provide additional storage space for device 950, or it can store applications or other information for device 950. Specifically, extended memory 984 may include instructions for performing or supplementing the above processes, and may also include security information. Thus, for example, extended memory 984 can act as a security module of device 950 and can be programmed with instructions that allow secure use of device 950. Furthermore, secure applications and additional information, such as placing identification information on the SIMM card in a hackable manner, can be provided via a SIMM card.

[0164] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 964, extended memory 984, or memory on processor 952 that can be received, for example, by transceiver 968 or external interface 962.

[0165] Device 950 can communicate wirelessly via communication interface 966, which may include digital signal processing circuitry if necessary. Communication interface 966 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messages, CDMA, TDMA, PDC, WCDMA, CDMA 2000, or GPRS, etc. This communication can occur, for example, via radio frequency transceiver 968. Furthermore, short-range communication can occur, such as using Bluetooth, WiFi, or other transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 980 can provide device 950 with additional navigation and location-related wireless data, which can be appropriately used by an application running on device 950.

[0166] Device 950 can also use audio codec 960 for audio communication, which can receive voice information from the user and convert it into usable digital information. Audio codec 960 can also generate audible sounds for the user, such as through a speaker, for example, in the mobile phone of device 950. Such sounds can include sounds from voice call, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications operating on device 950.

[0167] As shown in the figure, the computing device 950 can be implemented in a variety of different forms. For example, it can be implemented as a cellular phone 980. It can also be implemented as part of a smartphone 982, a personal digital assistant, or other similar mobile device.

[0168] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These different implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive and transmit data and instructions from and to a storage system, at least one input device, and at least one output device.

[0169] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0170] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0171] The systems and technologies described herein can be implemented in computing devices that include back-end components (e.g., as a data server), middleware components (e.g., an application server), or front-end components (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with the implementation of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.

[0172] Computing devices can include clients and servers. Clients and servers are typically geographically separated and usually interact via communication networks. The client-server relationship arises from computer programs running on the respective computers, and they have a client-server relationship with each other.

[0173] Many embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention.

[0174] Furthermore, the logical flow shown in the figure does not require the specific order or sequence shown to achieve the desired result. Additionally, other steps can be provided from the described flow, or other steps can be eliminated from the described flow, and other components can be added to the described system or removed from the described system. Therefore, other embodiments are within the scope of the appended claims.

Claims

1. A computer-implemented method (1000) for training a reinforcement learning model (130) to provide operating instructions for blast furnace thermal control, the method being characterized by the following steps: A domain-adaptive machine learning model (110) trained by transfer learning processes historical operational data (21) obtained as a multivariate time series and reflecting the thermal state of the corresponding blast furnaces (BF1 to BFn) in multiple domains to generate a first domain-invariant dataset (22) representing the thermal state of any one of the blast furnaces (BF1 to BFn), independent of the domain. By using a transient model (121) of a general blast furnace process, artificial operation data (24a) is generated as a multivariate time series reflecting the thermal state of the general blast furnace (BFg) for a specific thermal control action (26a), wherein the transient model (121) reflects the corresponding physical, chemical, thermal and flow conditions of the general blast furnace and provides solutions for the upward gas flow and downward movement of the solid layer constructed in the general blast furnace during the exchange of heat, mass and momentum transfer; The artificial operation data (24a) is processed by a generative deep learning network (122) trained on the multivariate time series of the historical operation data (21) to generate a second domain-invariant dataset (23a) by passing features learned from the historical operation data (21) to the artificial operation data (24a). The reinforcement learning model (130) determines the reward (131) for the specific heat control action (26a) by processing a combination of a first domain-invariant dataset (22) and a second domain-invariant dataset (23a) given a target function; and Based on the reward (131), the second domain-invariant dataset is regenerated based on the modified parameters (123-2), wherein the genetic search and / or Bayesian optimization algorithm (123-1) guides the search for the modified parameters for further thermal control actions based on the current environment (25a) of the reinforcement learning model (130) and the output of the specific thermal control action (26a) of the current learning step, and the determination step is repeated to learn optimized operating instructions for optimized thermal control actions to be applied to the corresponding operating states of one or more blast furnaces.

2. The method according to claim 1, further comprising: The reinforcement learning model (130) predicts optimized operating instructions for at least one actuator of the specific blast furnace in production based on the current operating state data of the specific blast furnace. After applying the thermal control action according to the optimized operating instructions to the at least one actuator, the reward is determined based on the new state of the specific blast furnace after the thermal control action is executed; as well as If the reward is lower than a predetermined threshold, the transient model is used to regenerate second-domain invariant data for one or more alternative operation instructions, which is then used to retrain the reinforcement learning model.

3. The method according to claim 1, wherein, The domain-adaptive machine learning model (110) is implemented by a generative deep learning neural network with convolutional and / or recursive layers, which is trained to extract domain-invariant features from the historical operational data (21) as the first domain-invariant dataset.

4. The method according to claim 1, wherein, The domain-adaptive machine learning model (110) has been trained to learn multiple mappings of the corresponding original data from multiple blast furnaces (BF1 to BFn) to a reference blast furnace (BFr), wherein each mapping is a representation of the transformation from the corresponding blast furnace to the reference blast furnace, and the multiple mappings correspond to the first domain-invariant dataset.

5. The method according to claim 4, wherein, The domain-adaptive machine learning model (110) is implemented by a generative deep learning architecture based on CycleGAN.

6. The method according to claim 1, wherein, The reinforcement learning model is trained to learn the optimized operational instructions such that the associated target measurement lies within a predefined range of the Pareto front of the corresponding multidimensional objective function.

7. The method according to claim 1, wherein, The transient model (121) includes multiple computational units, each of which represents a corresponding layer of the general-purpose blast furnace consisting of raw materials loaded once. Each computational unit solves the gas phase formula in an iterative manner to satisfy the relative gas phase parameter tolerance in each iteration time interval, and when the gas phase parameters converge to a predetermined tolerance value, the solid phase formula is solved sequentially in the same iteration time interval.

8. The method according to claim 7, wherein, For each iteration of the pressure-velocity correction cycle, the iterative solution of the gas-phase formula includes: Calculate the properties of gases, solids, and liquids; Calculate the reaction rate and heat transfer coefficient; Calculate gas temperature, type, velocity, and pressure drop; and The sequential solution of the solid phase formula includes: Calculate the temperature and type of solid; Calculate the temperature and type of liquid; and Calculate the velocity of the solid.

9. The method according to claim 1, wherein, The transient model (121) receives one or more of the following input parameters: amount of loaded material and chemical analysis, temperature, pressure, PCI rate and oxygen enrichment, having an energy formula for predicting the temperature of the hot metal, one or more class formulas for calculating the chemical composition of the hot metal, and one or more gas-phase formulas for predicting the top gas temperature, efficiency (Eta CO) and pressure.

10. The method according to claim 1, wherein, The reinforcement learning model is implemented by a recurrent neural network.

11. The method of claim 1, further comprising: By using one or more separately trained associated machine learning models (ML1 to MLn), information about the future thermal evolution of a particular blast furnace state is predicted based on the historical operating data (21) and / or further measured environmental data related to the environment of the blast furnace, in order to supplement the historical operating data (21) with future multivariate time series data related to future time points. as well as The future multivariate time series are processed by the domain-adaptive machine learning model (110) to augment the first domain-invariant dataset (22) with data related to the future time points.

12. The method according to claim 11, wherein, Training the specific (MLT) model in the associated machine learning models (ML1 to MLn) includes: Using one or more machine learning algorithms, multiple base models are trained (703) using different selections of operational data (701) and / or environmental data (702) to provide future multivariate time series data specific to the base models as training inputs for specific models in the machine learning models; (706) Train a specific model in the associated machine learning model using future multivariate time series data specific to the base model to learn which combination of the base models is best suited for which state of the blast furnace.

13. The method according to claim 12, wherein, The specific model in the machine learning model (ML1 to MLn) is trained to predict one of the following parameters at the future time point: anomalies in the blast furnace process; the thermal state of the blast furnace and hot metal production KPIs; load matrix optimization; blast furnace phenomena based on process inspections using a tuyeres camera; and a recommended taper for optimal operation. Based on the phenomena and KPIs of TMT SOMA; Phenomena marked by process rules.

14. A computer program product, when loaded into the memory of a computer system and executed by at least one processor of said computer system, performs the steps of a computer-implemented method according to any one of the preceding claims.

15. A computer system (100) comprising a plurality of functional modules, wherein, when executed by the computer system, the functional modules perform the steps of the computer-implemented method according to any one of claims 1 to 13.