A method, system, device and medium for determining a lance position change state of an oxygen lance
By constructing a training set and utilizing a deep reinforcement learning neural network to automatically control the oxygen lance position, the problem of smelting quality relying on manual operation in converter steelmaking was solved, thereby improving the stability and safety of the steelmaking process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CISDI RES & DEV CO LTD
- Filing Date
- 2022-09-14
- Publication Date
- 2026-05-12
AI Technical Summary
In the converter steelmaking process, the control of the oxygen lance position depends on the experience of the operators, which leads to unstable smelting quality. In particular, the smelting quality of high-quality special steel grades is difficult to guarantee. Moreover, the differences between different operators lead to uneven steel composition, causing economic losses to steel plants.
By acquiring historical smelting data from the converter to construct a training set, and using a deep reinforcement learning neural network for training, real-time smelting data is acquired to determine the oxygen lance position change status and achieve automated control.
It improves the stability and safety of the steelmaking process, ensures the quality of the finished product, and reduces the variability caused by manual operation.
Smart Images

Figure CN115564020B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of steelmaking technology, specifically to a method, system, equipment, and medium for determining the changing state of an oxygen lance position. Background Technology
[0002] In the converter smelting process, the control of the oxygen lance position affects smelting quality indicators such as slag formation, decarburization, and temperature rise. It also has a direct impact on unstable production phenomena such as molten steel splashing and slag drying during the smelting process. Therefore, reasonable control of the oxygen lance position is very important in the converter steelmaking process.
[0003] Typically, during the entire converter smelting process, the oxygen lance position is manually controlled by workers. Therefore, smelting quality largely depends on the operator's skill and experience, which severely hinders the optimization and upgrading of steelmaking process parameters. This is especially true for some high-quality specialty steels, whose process parameters are relatively demanding and require strict quality control. Even experienced operators may make mistakes in operation or judgment, leading to substandard smelting quality and significant economic losses for the steel plant. Furthermore, the variability in manual operation often results in inconsistent steel composition between different shifts or by different operators, posing further challenges to product quality control. Summary of the Invention
[0004] In view of the shortcomings of the prior art described above, the present invention provides a method, system, equipment and medium for determining the oxygen lance position change state, so as to improve the stability and safety of production.
[0005] This invention provides a method for determining the changing state of an oxygen lance position, comprising:
[0006] A training set is constructed by acquiring historical smelting data of the converter. The historical smelting data includes historical data of smelting process parameters, historical information on changes in oxygen lance position, and historical furnace mouth combustion images.
[0007] The deep reinforcement learning neural network is trained using a well-constructed training set;
[0008] Obtain real-time smelting data from the converter;
[0009] Based on the real-time smelting data, the oxygen lance position is determined using a trained deep reinforcement learning neural network.
[0010] In one exemplary embodiment of this application, constructing a training set includes:
[0011] Based on the historical data of the smelting process parameters, a feature vector of the smelting process parameters is constructed;
[0012] Based on the historical information of the oxygen lance position change status, an oxygen lance position change status vector is constructed.
[0013] Based on the historical furnace combustion images, a reward variable is constructed;
[0014] A training set is constructed based on the feature vector of the smelting process parameters, the state vector of the oxygen lance position change, and the reward variable.
[0015] In an exemplary embodiment of this application, a reward variable is constructed based on the historical furnace combustion image, including:
[0016] Based on the historical furnace combustion images, the smelting status and the quality indicators of the smelted finished products are determined;
[0017] Based on the smelting state and the quality indicators of the smelted finished product, a reward variable is constructed.
[0018] Secondly, this application provides a method for controlling the position of an oxygen lance, comprising:
[0019] Obtain the current smelting process parameters;
[0020] The oxygen lance position change state is obtained based on the current smelting process parameters and the pre-trained deep reinforcement learning neural network; the deep reinforcement learning neural network is trained with the smelting process parameters as input and the oxygen lance position change state as output.
[0021] The oxygen lance position changes according to the oxygen lance position change status.
[0022] In an exemplary embodiment of this application, after changing the oxygen lance position according to the oxygen lance position change state, the method further includes:
[0023] The real-time quality indicators of the smelted finished product and the real-time combustion images of the furnace mouth during the smelting process are obtained. The real-time quality indicators of the smelted finished product and the real-time combustion images of the furnace mouth are obtained in real time based on the changed oxygen lance position and according to the current smelting process parameters.
[0024] The current real-time smelting data, the real-time smelting finished product quality indicators, and the real-time furnace mouth combustion image are used as training samples in the training set to update the trained deep reinforcement learning neural network.
[0025] Thirdly, this application provides a system for determining the state of oxygen lance position change, comprising:
[0026] The first acquisition module is used to acquire historical smelting data of the converter to construct a training set. The historical smelting data includes historical data of smelting process parameters, historical information on the change status of oxygen lance position, and historical furnace mouth combustion images.
[0027] The training module trains the deep reinforcement learning neural network using a pre-constructed training set.
[0028] The second acquisition module is used to acquire real-time smelting data from the converter;
[0029] The oxygen lance position determination module is used to determine the oxygen lance position based on the real-time smelting data using a trained deep reinforcement learning neural network.
[0030] Fourthly, this application provides an oxygen lance position control system, characterized in that it includes:
[0031] The first acquisition module is used to acquire historical smelting data of the converter to construct a training set. The historical smelting data includes historical data of smelting process parameters, historical information on the change status of oxygen lance position, and historical furnace mouth combustion images.
[0032] The training module trains the deep reinforcement learning neural network using a pre-constructed training set.
[0033] The second acquisition module is used to acquire real-time smelting data from the converter;
[0034] The oxygen lance position determination module is used to determine the oxygen lance position based on the real-time smelting data and through a trained deep reinforcement learning neural network.
[0035] The control module is used to change the oxygen lance position according to the oxygen lance position change status.
[0036] In another aspect, this application provides an electronic device comprising: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method described above.
[0037] In another aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer's processor, causes the computer to perform the method described above.
[0038] The beneficial effects of this invention are:
[0039] This invention obtains the current smelting process parameters and feeds them into a pre-trained deep reinforcement learning neural network to obtain the oxygen lance position change state. Then, it adjusts the oxygen lance position according to the oxygen lance position change state, thereby improving the stability and safety of production and ensuring the quality of the smelted finished product.
[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0042] Figure 1 A flowchart illustrating a method for determining the oxygen lance position change state, as shown in an exemplary embodiment of this application;
[0043] Figure 2 for Figure 1 The flowchart of step S120 in the illustrated embodiment, which involves constructing the training set, is shown in an exemplary embodiment.
[0044] Figure 3 for Figure 2 The flowchart of step S230 in the illustrated embodiment is shown in an exemplary embodiment.
[0045] Figure 4 A flowchart illustrating an exemplary embodiment of this application shows a method for controlling the position of an oxygen lance;
[0046] Figure 5 A flowchart illustrating a method for controlling the position of an oxygen lance, as shown in another exemplary embodiment of this application;
[0047] Figure 6 A flowchart illustrating a specific embodiment of an oxygen lance position control method;
[0048] Figure 7 for Figure 6 The illustrated embodiment shows a schematic diagram of the deep reinforcement learning neural network in the oxygen gun position control method.
[0049] Figure 8 A block diagram illustrating an oxygen lance position change state determination system as shown in an exemplary embodiment of this application;
[0050] Figure 9 A block diagram illustrating an oxygen lance position control system as an exemplary embodiment of this application;
[0051] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation
[0052] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0053] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0054] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0055] Please see Figure 1 , Figure 1 The flowchart illustrates a method for determining the oxygen lance position change state as an exemplary embodiment of this application.
[0056] like Figure 1 As shown in an exemplary embodiment of this application, the method for determining the oxygen lance position change state includes at least steps S110, S120, S130, and S140, which are described in detail below:
[0057] Step S110. Obtain historical smelting data from the converter to construct a training set;
[0058] Historical smelting data includes historical data on smelting process parameters, historical information on changes in oxygen lance position, and historical furnace mouth combustion images.
[0059] Historical data on smelting process parameters include initial molten iron weight, molten iron composition, molten iron temperature, charging time and weight, flue gas process data, sonar process curves, furnace mouth flame combustion image data, CO concentration curves, etc.
[0060] It should be noted that the oxygen lance position changes include the oxygen lance moving upward, moving downward, and remaining stationary.
[0061] Step S120. Train the deep reinforcement learning neural network using the constructed training set;
[0062] It should be noted that the deep reinforcement learning neural network is trained using historical data of smelting process parameters as input and historical data of oxygen lance position changes as output.
[0063] Step S130. Obtain real-time smelting data from the converter;
[0064] It should be noted that the real-time smelting data includes molten iron weight, molten iron composition, molten iron temperature, charging time and charging weight, flue gas process data, sonar process curve, furnace mouth flame combustion image data, CO concentration curve, etc.
[0065] Step S140. Determine the oxygen lance position using a trained deep reinforcement learning neural network based on real-time smelting data.
[0066] In related technologies, the oxygen lance position is typically controlled manually by workers, making smelting quality highly dependent on their skill level and experience. After researching these technologies, the inventors discovered that existing oxygen lance position control methods severely hinder the optimization and upgrading of steelmaking process parameters, especially for high-quality special steels. These processes have relatively high requirements for process parameters and necessitate strict quality control. Even experienced operators may make mistakes in operation or judgment, leading to substandard smelting quality and significant economic losses for steel mills. Furthermore, the variability in manual operation often results in inconsistent steel composition across different shifts or by different operators, posing further challenges to product quality control. Therefore, the inventors considered constructing a training set using historical converter smelting data. This training set was then used to train a deep reinforcement learning neural network, while real-time converter smelting data was acquired. Based on this real-time data, the trained deep reinforcement learning neural network determined the oxygen lance position and adjusted it according to changes in the lance's position, thus improving production stability.
[0067] Please see Figure 2 , Figure 2 for Figure 1 The flowchart of step S110 in the illustrated embodiment is shown in an exemplary embodiment.
[0068] like Figure 2 As shown in an exemplary embodiment of this application, the process of constructing a training set includes steps S210, S220, S230, and S240, which are described in detail below:
[0069] Step S210. Construct a feature vector of smelting process parameters based on historical data of smelting process parameters;
[0070] Specifically, the scalar characteristic data of molten iron information is standardized using the Z-score standardization method, and its data formula is expressed as follows:
[0071] (1);
[0072] in, For the i-th sample data, Indicates all On average, Indicates all The standard deviation of the composition.
[0073] Then, by combining the standardized scalar values of each molten iron information, the feature vector can be obtained.
[0074] For time-series data (including time point information) such as flue gas, sonar, and CO concentration curve features, a Long Short-Term Memory (LSTM) network is used for feature encoding, thereby converting the time-series data such as flue gas, sonar, and CO concentration curve features into feature vectors of a specified dimension.
[0075] Step S220. Construct an oxygen lance position change state vector based on the historical information of oxygen lance position change state;
[0076] For example, character conversion can be performed using one-hot encoding to transform the oxygen lance position change state into an oxygen lance position change state vector, the mathematical expression of which is:
[0077] Oxygen lance moved upwards (2);
[0078] oxygen lance moved down (3);
[0079] Oxygen lance stationary (4);
[0080] Step S230. Construct reward variables based on historical furnace combustion images;
[0081] Step S240. Construct a training set based on the feature vector of smelting process parameters, the state vector of oxygen lance position change, and the reward variable.
[0082] Please see Figure 3 , Figure 3 for Figure 2 The flowchart of step S230 in the illustrated embodiment is shown in an exemplary embodiment.
[0083] like Figure 3 As shown in an exemplary embodiment of this application, Figure 2In the illustrated embodiment, step S230, which involves constructing a reward variable based on historical furnace combustion images, includes steps S310 and S320, which are detailed below:
[0084] Step S310. Determine the smelting status and the quality indicators of the smelted finished product based on historical furnace combustion images.
[0085] It should be noted that the smelting state includes splashing, slag jumping, and normal. Normal refers to the smelting state other than splashing and slag jumping. During the smelting process, the smelting state can be identified based on the changes in state in historical furnace combustion images.
[0086] Specifically, at the end of the smelting process, the quality indicators of the finished smelting product can be determined based on historical furnace combustion images.
[0087] Specifically, the quality indicators of the smelted finished product can be determined by using automatic identification equipment based on historical furnace combustion images.
[0088] The quality indicators of smelted finished products include the composition of the smelted finished products (such as iron content, calcium content, silicon content, sulfur content), temperature, etc.
[0089] Step S320. Construct reward variables based on the smelting state and the quality indicators of the smelted finished product.
[0090] For example, if the smelting state is normal, the reward variable is 1; if the smelting state is splashing, the reward variable is -100; if the smelting state is slag jumping, the reward variable is determined according to the preset slag jumping severity level-reward variable correspondence. The preset slag jumping severity level-reward variable correspondence can be implemented in a gradient manner. For example, within the first preset slag jumping severity level range, the reward variable is at the first preset reward variable value; within the second preset slag jumping severity level range, the reward variable is at the second preset reward variable value, and so on. The value of the reward variable decreases sequentially from 0 to -5 as the slag jumping severity level increases. The preset slag jumping severity level-reward variable correspondence can be set by the user and will not be elaborated here.
[0091] For example, based on historical data of smelted finished product quality indicators, the reward variable can be constructed as follows:
[0092] If the quality indicators of the smelted finished product meet the standard (preset standard), the reward variable value is 10; if the quality indicators of the smelted finished product do not meet the standard, the reward variable value is -100.
[0093] Please see Figure 4 , Figure 4 The flowchart illustrates an oxygen lance position control method as an exemplary embodiment of this application.
[0094] like Figure 4As shown in an exemplary embodiment of this application, the oxygen lance position control method includes at least steps S410, S420, S430, S440, and S450, which are described in detail below:
[0095] Step S410. Obtain historical smelting data from the converter to construct a training set;
[0096] Step S420. Train the deep reinforcement learning neural network using the constructed training set;
[0097] Step S430. Obtain real-time smelting data from the converter;
[0098] Step S440. Determine the oxygen lance position using a trained deep reinforcement learning neural network based on real-time smelting data;
[0099] Step S450. Change the oxygen lance position according to the oxygen lance position change status.
[0100] like Figure 5 As shown in another exemplary embodiment of this application, Figure 4 In the embodiment shown, after changing the oxygen lance position according to the oxygen lance position change state, steps S510 and S570 are further included, which are described in detail below:
[0101] Step S560. Obtain real-time quality indicators of smelted finished products and real-time furnace mouth combustion images during the smelting process;
[0102] The real-time smelting finished product quality indicators and real-time furnace mouth combustion images are obtained in real time based on the changed oxygen lance position and according to the current smelting process parameters.
[0103] Step S570. Use the current real-time smelting data, real-time smelting finished product quality indicators, and real-time furnace mouth combustion images as training samples in the training set to update the trained deep reinforcement learning neural network.
[0104] Specifically, the trained deep reinforcement learning neural network can be iteratively updated using stochastic gradient descent algorithms such as SGD, Adam, Adagrad, and RMSProp until it converges.
[0105] By updating the trained deep reinforcement learning neural network, the network can be continuously optimized to further improve the accuracy of oxygen lance position control, thereby enhancing production stability.
[0106] Please see Figure 6 , Figure 6 The oxygen lance position control method shown in a specific embodiment includes the following steps:
[0107] Historical data on converter smelting process parameters were collected, including initial molten iron weight, molten iron composition, molten iron temperature, charging time and weight, flue gas process data, sonar process curves, furnace mouth flame combustion image data, CO concentration curves, etc.
[0108] The scalar characteristic data of molten iron information is standardized using the Z-score standardization method, and its data formula is expressed as follows:
[0109] (1);
[0110] in, For the i-th sample data, For all On average, Indicates all The standard deviation of the composition.
[0111] For time-series data (including time point information) such as flue gas, sonar, and CO concentration curve features, a Long Short-Term Memory (LSTM) network is used to encode the features and convert them into feature vectors of a specified dimension.
[0112] Then, a model-free reinforcement learning framework is constructed, namely a Markov process with unknown state transition distribution, whose triple is defined as (S, A, R), where S represents the state variable, A represents the decision variable, and R represents the reward variable. Considering the oxygen lance position control problem within the reinforcement learning framework means transforming oxygen lance position control into a temporal decision-making process. At each time step, based on the current smelting state characteristics, a decision is made regarding the oxygen lance operation at that time step, i.e., the oxygen lance position change state. The oxygen lance operation decision includes three possibilities: moving the oxygen lance upwards, moving it downwards, or keeping it stationary. The process then proceeds to the next time step, while simultaneously observing the reward variable value resulting from the decision at that time step.
[0113] The initial state of this time-series decision problem includes initial information about molten iron, including composition, weight, and temperature. The smelting state s at the current time step is composed of features such as charging signal records (including charging time, charging weight, and charging type), sonar process curves, CO concentration curve characteristics (including peak values, trough values, and average rate of rise), and oxygen lance operation curves since the start of smelting. Its mathematical representation is as follows:
[0114] (5);
[0115] in, This is the initial information of the molten iron; , which is the characteristic vector of material added since the start of smelting; , which is the characteristic vector of the sonar process curve; , which is the characteristic vector of the CO concentration curve since the start of smelting; It is the characteristic vector of the oxygen lance operation curve since the start of smelting.
[0116] For the decision variable, the three oxygen lance operation decisions 'a' are converted into characters using One-hot encoding, and their mathematical representation is as follows:
[0117] Oxygen lance moved upwards (2);
[0118] oxygen lance moved down (3);
[0119] Oxygen lance stationary (4);
[0120] Within the deep reinforcement learning framework, the value of the reward variable R depends on the smelting state identified from the current furnace mouth image data. Existing converter splash recognition equipment is used to obtain the recognition type of the current furnace mouth combustion image mapping. At the current smelting moment, if the smelting state identified from the furnace mouth image is splashing, the reward variable is -100; if the smelting state identified from the furnace mouth image is slag jumping, the reward variable decreases from 0 to -5 according to the severity of slag jumping, decreasing by 1 each time; if the smelting state identified from the furnace mouth image is normal, the reward variable is 1. At the smelting endpoint, the reward variable is determined based on the smelting product quality indicators confirmed by the furnace mouth image. If the smelting product quality indicators confirmed by the furnace mouth image meet the standards, the reward variable is 10; if the smelting product quality indicators confirmed by the furnace mouth image do not meet the standards, the reward variable is -100. Based on the reward variables defined above, the oxygen lance position decision learned under the reinforcement learning framework maximizes the cumulative reward variable (i.e.,) of the smelting process, meaning that the smelting process is stable, does not splash, and meets the smelting standards.
[0121] Construct a training set for the reinforcement learning algorithm. Take 3 seconds as a time step, preprocess the data for each time step according to the above requirements, and store them in chronological order using the following tuple format:
[0122] (6);
[0123] Specifically, data from multiple smelting processes are collected. For each smelting process, a time step is defined as every 3 seconds from the start of smelting, and this time step is extracted. The feed signal recording, sonar process curve, CO concentration curve features, and oxygen lance operation curve are input into the corresponding LSTM encoding networks to form feature vectors. They are pieced together to form state information. Based on the comparison between the oxygen lance position in the next time step and the current oxygen lance position, a lance position decision is given. Based on the current time step Based on the furnace mouth image data at that time, determine the molten steel splashing situation, and give the reward value according to the reward variable value selection criteria. Due to the state information of the next time step and The formation process is the same, so it will not be elaborated here;
[0124] Then, a neural network is constructed, and the neural network is trained based on training samples to obtain a deep reinforcement learning neural network. The specific steps are as follows:
[0125] structure Figure 8 The deep reinforcement learning neural network shown outputs a state. The expected cumulative reward for decision-making is denoted by the mathematical symbol . It is a vector with the same dimension as the number of decisions.
[0126] The applied deep reinforcement learning optimization algorithm is the Deep Q-Network (DQN) algorithm, and its optimization function is mathematically expressed as follows:
[0127] (7);
[0128] in, The output of the neural network is a three-dimensional vector. These are neural network parameters. For periodically updated target network parameters, yes One of the constants, typically with a value of 0.995.
[0129] Applying stochastic gradient descent algorithms, such as SGD, Adam, Adagrad, RMSProp, etc., to the neural network parameters The process involves iterative updates until convergence. In this specific example, the SGD optimizer is used to optimize the neural network parameters. The process is iteratively updated until convergence. The SGD optimizer randomly selects a set of training samples at each update step. The neural network parameters are updated using the following expression:
[0130] (8);
[0131] in, It is the learning rate parameter, which is updated using an exponential decay strategy.
[0132] After training the aforementioned neural network, its output is used to guide the oxygen lance position control during the smelting process, and new training samples are acquired. Specifically, starting from the smelting control, still using 3-second time steps, the current state information is preprocessed at each smelting time step and input into the trained deep reinforcement learning neural network. The probability of each decision is calculated based on the network model mapping results, and samples are taken accordingly. The mathematical expression is:
[0133] (9);
[0134] in, The output of a neural network is a three-dimensional vector; the softmax function is defined as follows:
[0135] (10);
[0136] At the next time step, repeat the above process to make gun position decisions until the smelting is completed.
[0137] The process data for this batch was compiled into The new training data is used as a random replacement of some old training data with new experience samples. The parameters of the deep reinforcement learning neural network are then iteratively updated using the updated training dataset, allowing the deep reinforcement learning neural network to iterate through online learning. This process of generating new experience samples and updating the neural network parameters online is repeated until the neural network converges.
[0138] The resulting deep reinforcement learning neural network can provide the optimal oxygen lance position operation decision at each time step, ensuring a stable and safe smelting process and the quality of the smelted product.
[0139] Please see Figure 8 This application also provides an oxygen lance position change determination system 800, including:
[0140] The first acquisition module 810 is used to acquire historical smelting data of the converter to construct a training set. The historical smelting data includes historical data of smelting process parameters, historical information on the change status of oxygen lance position, and historical furnace mouth combustion images.
[0141] Training module 820 is used to train a deep reinforcement learning neural network using a pre-constructed training set;
[0142] The second acquisition module 830 is used to acquire real-time smelting data from the converter;
[0143] The oxygen lance position determination module 840 is used to determine the oxygen lance position based on real-time smelting data using a trained deep reinforcement learning neural network.
[0144] Please see Figure 9 This application also provides an oxygen lance position control system 900, comprising:
[0145] The first acquisition module 910 is used to acquire historical smelting data of the converter to construct a training set. The historical smelting data includes historical data of smelting process parameters, historical information on the change status of oxygen lance position, and historical furnace mouth combustion images.
[0146] Training module 920 trains the deep reinforcement learning neural network using the constructed training set;
[0147] The second acquisition module 930 is used to acquire real-time smelting data from the converter;
[0148] The oxygen lance position determination module 940 is used to determine the oxygen lance position based on real-time smelting data and a trained deep reinforcement learning neural network.
[0149] The control module 950 is used to change the oxygen lance position according to the oxygen lance position change status.
[0150] It should be noted that the oxygen lance position change state determination system and the oxygen lance position change state determination method provided in the above embodiments belong to the same concept, and the oxygen lance position control system and the oxygen lance position control method provided in the above embodiments belong to the same concept. The specific methods by which each module and unit performs its operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the oxygen lance position change state determination system and oxygen lance position control system provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.
[0151] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device enables the oxygen lance position change state determination method or oxygen lance position control method provided in the above embodiments.
[0152] Figure 10 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 10 The computer system 1000 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0153] like Figure 10As shown, the computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from storage portion 1008 into Random Access Memory (RAM) 1003, such as performing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1003. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An Input / Output (I / O) interface 1005 is also connected to bus 1004.
[0154] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.
[0155] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.
[0156] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0158] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0159] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer's processor, causes the computer to perform a method for determining the change state of the oxygen lance position or a method for controlling the oxygen lance position.
[0160] The computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently without being assembled into the electronic device.
[0161] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the oxygen lance position change determination method or oxygen lance position control method provided in the above embodiments.
[0162] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for determining the changing state of an oxygen lance position, characterized in that, include: A training set is constructed by acquiring historical smelting data of the converter. The historical smelting data includes historical data of smelting process parameters, historical information on changes in oxygen lance position, and historical furnace mouth combustion images. The deep reinforcement learning neural network is trained using a well-constructed training set; Obtain real-time smelting data from the converter; Based on the real-time smelting data, the oxygen lance position is determined by a trained deep reinforcement learning neural network. Constructing the training set includes: Based on the historical data of the smelting process parameters, a feature vector of the smelting process parameters is constructed; Based on the historical information of the oxygen lance position change status, an oxygen lance position change status vector is constructed. Based on the historical furnace combustion images, a reward variable is constructed; A training set is constructed based on the feature vector of the smelting process parameters, the state vector of the oxygen lance position change, and the reward variable.
2. The method for determining the oxygen lance position change state according to claim 1, characterized in that, Based on the historical furnace combustion images, a reward variable is constructed, including: Based on the historical furnace combustion images, the smelting status and the quality indicators of the smelted finished products are determined; Based on the smelting state and the quality indicators of the smelted finished product, a reward variable is constructed.
3. A system for determining the position change state of an oxygen lance, characterized in that, include: The first acquisition module is used to acquire historical smelting data of the converter to construct a training set. The historical smelting data includes historical data of smelting process parameters, historical information on oxygen lance position changes, and historical furnace combustion images. Constructing the training set includes: constructing a feature vector of smelting process parameters based on the historical data of smelting process parameters; constructing a vector of oxygen lance position changes based on the historical information on oxygen lance position changes; constructing a reward variable based on the historical furnace combustion images; and constructing the training set based on the feature vector of smelting process parameters, the vector of oxygen lance position changes, and the reward variable. The training module trains the deep reinforcement learning neural network using a pre-constructed training set. The second acquisition module is used to acquire real-time smelting data from the converter; The oxygen lance position determination module is used to determine the oxygen lance position based on the real-time smelting data using a trained deep reinforcement learning neural network.
4. An electronic device, characterized in that, The electronic device includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method as described in claim 1 or 2.
5. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the method described in claim 1 or 2.