Iron ore magnetic separation roasting temperature control method based on double-ring reinforcement learning
The temperature control method using dual-loop reinforcement learning solves the problem of inaccurate temperature control in traditional iron ore magnetic separation roasting, achieves precise control of the temperature of the main roasting furnace, improves response speed and energy efficiency, and has strong adaptability and is suitable for different production scenarios.
Patent Information
- Application Number
- CN202510900742.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-30
AI Technical Summary
The traditional temperature control method of iron ore magnetic separation roasting has problems such as inaccurate temperature control, slow response speed and high energy consumption, which makes it difficult to meet the temperature control requirements of suspended magnetization roasting.
A temperature control method based on dual-loop reinforcement learning is adopted. By constructing an inner loop for gas flow control and an outer loop for main furnace temperature control, the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature are synchronously generated using the reinforcement learning on-policy algorithm to achieve precise control of the main furnace temperature.
It improves the quality and production efficiency of iron ore concentrate products, reduces energy consumption, broadens the adaptability and robustness of the method, and can achieve effective control under complex working conditions.
Smart Images

Figure CN120720879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a temperature control method and system, and in particular to an iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning. Background Art
[0002] Suspended magnetization roasting technology is one of the most advanced iron ore processing techniques currently available. It involves suspending iron ore particles in an airflow, where a chemical reaction occurs under suitable temperature and atmospheric conditions, altering the iron ore's physical properties and imparting strong magnetic properties. Specifically, the iron ore concentrate is filtered and then conveyed to a feed silo by a belt conveyor. The dried ore then enters a roaster, where it undergoes a thermal decomposition reaction at approximately 580°C. This thermal decomposition reaction is a key step in suspended magnetization roasting, transforming weakly magnetic iron ore into a strongly magnetic mineral for subsequent reactions and operations.
[0003] During the thermal decomposition reaction process, the temperature control of the main roasting furnace is crucial, and the temperature of the main furnace needs to be maintained at around 580°C. Traditional temperature control usually relies on fixed operating parameters or manual adjustments based on experience. These methods have many shortcomings, including inaccurate temperature control in the heating decomposition stage, slow response speed, high energy consumption and other problems, which directly affect the quality and production efficiency of iron ore concentrate products and cannot meet the current temperature control requirements of suspended magnetization roasting. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method and system for controlling the temperature of iron ore magnetic separation and roasting based on double-loop reinforcement learning, which can achieve precise control of the temperature of the roasting main furnace, improve the response speed, reduce energy consumption, and thus improve the quality and production efficiency of iron ore concentrate products.
[0005] According to the technical solution provided by the present invention, a temperature control method for iron ore magnetic separation and roasting based on double-loop reinforcement learning, the temperature control method includes:
[0006] A target roasting furnace for magnetic separation roasting of iron ore is provided, and a roasting temperature control loop adapted to the target roasting furnace is constructed, wherein the roasting temperature control loop includes a gas flow control inner loop for controlling the gas flow rate and a main furnace temperature control outer loop for controlling the temperature of the roasting main furnace.
[0007] The inner loop of gas flow control is configured with an optimal gas flow control strategy, and the outer loop of main furnace temperature control is configured with an optimal main furnace temperature control strategy. These optimal gas flow control strategies and main furnace temperature control strategies are synchronously constructed and generated based on a reinforcement learning on-policy algorithm.
[0008] During roasting temperature control, outer loop operating status information is obtained, and a main furnace temperature optimal control strategy is used to generate a gas reference flow rate corresponding to the current outer loop operating status information, wherein the outer loop operating status is generated based on at least the roasting reference temperature and the current roasting operating temperature of the target roasting furnace;
[0009] The inner loop operation status information is obtained, and the gas flow optimal control strategy is used to generate a gas flow control signal corresponding to the current inner loop operation status information, so as to match the gas delivery flow delivered to the burner with the gas reference flow based on the gas flow control signal, thereby stabilizing the roasting operating temperature of the target roasting furnace at the roasting reference temperature, wherein,
[0010] The inner loop operating state information is generated based on at least a gas reference flow rate and a gas flow rate currently fed into the burner.
[0011] For the optimal control strategy of gas flow, we have:
[0012]
[0013] Among them, u q is the gas flow control signal, K q is the optimal control gain matrix of gas flow, ξ q is the inner loop running status information, K q1 , K q2 is the gain coefficient of the gas flow optimal control gain matrix, x q is the gas flow rate currently delivered to the burner, e q The gas flow tracking error is generated based on the gas reference flow and the gas delivery flow delivered to the burner;
[0014] For the optimal control strategy of the main furnace temperature control, we have:
[0015]
[0016] Among them, u w is the gas reference flow rate, K w is the optimal control gain matrix for the main furnace temperature control, ξ w is the outer loop operation status information, K w1 , K w2 is the gain coefficient of the optimal control gain matrix for the main furnace temperature control, x w is the current roasting temperature of the target roasting furnace, e w The temperature tracking error is generated based on the roasting reference temperature and the current roasting operating temperature of the target roasting furnace;
[0017] When the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control are synchronously constructed based on the reinforcement learning on-policy algorithm, at least the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control are synchronously generated.
[0018] When synchronously determining and generating the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control based on the reinforcement learning on-policy algorithm, the following steps are included:
[0019] For the inner loop of gas flow control, an inner loop optimal control strategy corresponding to the gas flow optimal control strategy is constructed based on the gas flow optimal control, wherein the inner loop optimal control strategy includes an inner loop optimal control gain matrix to be solved;
[0020] For the main furnace temperature control outer loop, an outer loop optimal control strategy corresponding to the main furnace temperature optimal control strategy is constructed based on the main furnace temperature optimal control, wherein the outer loop optimal control strategy includes an outer loop optimal control gain matrix to be solved;
[0021] Based on the outer loop optimal control strategy and inner loop optimal control strategy constructed above, the roasting temperature of the target roasting furnace is pre-controlled, and the gain matrix solution based on reinforcement learning is performed during the roasting temperature pre-control process, so that the corresponding inner loop optimal control gain matrix and outer loop optimal control gain matrix are obtained synchronously after the gain matrix solution processing.
[0022] The inner loop optimal control gain matrix obtained by the solution is configured as the gas flow optimal control gain matrix, and the gas flow optimal control strategy is generated based on the gas flow optimal control gain matrix;
[0023] The outer loop optimal control gain matrix obtained by solving the problem is configured as the main furnace temperature control optimal control gain matrix, and the main furnace temperature control optimal control strategy is generated based on the main furnace temperature control optimal control gain matrix.
[0024] When performing the gain matrix solution process based on reinforcement learning, at least one pre-control process is performed, and the pre-control process includes data acquisition processing, equation solving processing and gain matrix updating processing performed in sequence, wherein,
[0025] When performing data acquisition processing, inner loop pre-operation state information satisfying the inner loop full rank state and outer loop pre-operation state information satisfying the outer loop full rank state are acquired;
[0026] When performing equation solving, the inner loop pre-running state information is used to solve the inner loop algebraic Riccati equation to obtain the inner loop equation solution matrix, and the outer loop pre-running state information is used to solve the outer loop algebraic Riccati equation to obtain the outer loop equation solution matrix, where,
[0027] The inner loop algebraic Riccati equation is generated based on at least the inner loop optimal control strategy.
[0028] The outer loop algebraic Riccati equation is generated based on at least the outer loop optimal control strategy;
[0029] When performing the gain matrix update process, the inner loop optimal control gain matrix is updated based on the inner loop equation solution matrix and the inner loop gain matrix update strategy, and the outer loop optimal control gain matrix is updated based on the outer loop equation solution matrix and the outer loop gain matrix update strategy;
[0030] Repeat the above pre-control processing until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions. Thereafter, the updated inner-loop optimal control gain matrix is used as the solved inner-loop optimal control gain matrix, and the updated outer-loop optimal control gain matrix is used as the solved outer-loop optimal control gain matrix.
[0031] When pre-controlling the roasting temperature of the target roasting furnace, it includes:
[0032] Constructing an initial inner loop optimal control gain matrix that satisfies the inner loop Hurwitz matrix state, and an initial outer loop optimal control gain matrix that satisfies the outer loop Hurwitz matrix state;
[0033] The initial inner loop optimal control gain matrix is configured within the inner loop optimal control strategy, and the initial outer loop optimal control gain matrix is configured within the outer loop optimal control strategy, so as to pre-control the temperature of the target roasting furnace using the corresponding inner loop optimal control strategy and outer loop optimal control strategy, and then perform a pre-control process;
[0034] After executing the pre-control processing, when the updated inner loop optimal control gain matrix and the outer loop optimal control gain matrix meet the pre-control convergence conditions, the roasting temperature pre-control of the target roasting furnace is exited. Otherwise,
[0035] The updated inner-loop optimal control gain matrix is configured in the inner-loop optimal control strategy, and the outer-loop optimal control gain matrix is configured in the outer-loop optimal control strategy. The roasting temperature of the target roasting furnace is continued to be pre-controlled, and then the pre-control processing is performed until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions.
[0036] For the pre-control convergence condition, we have:
[0037]
[0038] in, is the inner loop optimal control gain matrix when executing the i-th pre-control process, is the updated inner loop optimal control gain matrix after executing the i-th pre-control process, is the outer loop optimal control gain matrix when executing the i-th pre-control process, is the updated optimal control gain matrix of the inner loop after executing the i-th pre-control process, and σ is the pre-control convergence threshold.
[0039] When the inner loop optimal control strategy and the outer loop optimal control strategy are used to pre-control the roasting temperature of the target roasting furnace, the following equations are obtained:
[0040]
[0041] in, is the gas flow control signal under the jth inner loop pre-operation status sub-information when executing the i-th pre-control processing, is the jth inner loop pre-operation state sub-information when executing the i-th pre-control processing, ρ1 is the measurable noise of the inner loop, is the gas reference flow rate under the jth outer loop pre-operation status sub-information when executing the i-th pre-control processing, is the jth outer loop pre-operation state sub-information when executing the i-th pre-control processing, and ρ2 is the measurable noise of the outer loop.
[0042] When constructing the inner loop algebraic Riccati equation, include:
[0043] Based on the gas flow control state of the gas flow control inner loop, a gas flow control state equation corresponding to the gas flow control inner loop is constructed, and a gas flow control performance index of the gas flow control inner loop is determined, wherein:
[0044] The gas flow control inner loop includes at least a gas flow regulating valve for regulating the gas flow delivered to the burner and a gas flow transmitter for obtaining the current gas flow delivered to the burner;
[0045] For the determined gas flow control performance index, the tracking state of the gas delivery flow tracking the gas reference flow is converted into a quadratic linear regulation state of the gas delivery flow;
[0046] Generating a quadratic linear regulation state of the gas delivery flow rate based on the conversion, generating an inner loop optimal control gain matrix associated with the quadratic linear regulation state based on the gas flow rate optimal control target, constructing an inner loop optimal control strategy based on the generated inner loop optimal control gain matrix, and constructing an inner loop gain matrix update strategy based on a calculation form of the inner loop optimal control gain matrix;
[0047] Based on the quadratic linear regulation state of gas delivery flow and the inner loop optimal control gain matrix, the inner loop algebraic Riccati equation is constructed.
[0048] When constructing the outer loop algebraic Riccati equation, include:
[0049] Based on the main furnace temperature control state of the main furnace temperature control outer loop, a main furnace temperature control state equation corresponding to the main furnace temperature control outer loop is constructed, and the main furnace temperature control performance index of the main furnace temperature control outer loop is determined, wherein,
[0050] The main furnace temperature control outer loop includes at least a burner, a roasting main furnace, and a temperature transmitter for the current roasting working temperature of the roasting main furnace. When the main furnace temperature control state is established, the burner is used as a delay link;
[0051] For the determined main furnace temperature control performance index, the tracking state of the roasting working temperature tracking the roasting reference temperature is converted into a quadratic linear regulation state of the roasting working temperature;
[0052] A quadratic linear regulation state of the roasting operating temperature is generated based on the conversion, an outer loop optimal control gain matrix associated with the quadratic linear regulation state is generated based on the optimal control target of the main furnace temperature, and an outer loop optimal control strategy is constructed based on the generated outer loop optimal control gain matrix, wherein an inner loop gain matrix update strategy is constructed based on the calculation form of the outer loop optimal control gain matrix;
[0053] Based on the quadratic linear regulation state of the roasting working temperature and the outer loop optimal control gain matrix, the outer loop algebraic Riccati equation is constructed.
[0054] When the roasting temperature of the target roasting furnace is controlled by using the inner loop of gas flow control and the outer loop of main furnace temperature control, the inner loop of gas flow control and the outer loop of main furnace temperature control are interacted with each other, and based on the mutual interaction state, the optimal control strategy of gas flow in the inner loop of gas flow control is fine-tuned, and / or the optimal control strategy of main furnace temperature control in the outer loop of main furnace temperature control is fine-tuned.
[0055] The advantages of the present invention are as follows: the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control are synchronously constructed and generated based on the reinforcement learning on-policy algorithm, thereby eliminating the need to obtain the working parameters of the target roasting furnace and not relying on prior information of the target roasting furnace. This is of great significance for the temperature control of the target roasting furnace for which it is difficult to obtain accurate model parameters, reduces the requirements for understanding the working parameters of the target roasting furnace, broadens the application scope of the method in different production scenarios, has strong adaptability and robustness, and can achieve effective control under complex working conditions.
[0056] By interacting the inner loop of gas flow control with the outer loop of main furnace temperature control, it is possible to further achieve precise control of the temperature of the roasting main furnace, improve response speed, reduce energy consumption, and thereby improve the quality and production efficiency of iron ore concentrate products, and solve the problems of inaccurate control existing in traditional control. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 The figure is a flow chart of an embodiment of the method for controlling the temperature of iron ore magnetic separation and roasting according to the present invention.
[0058] Figure 2 This is a control block diagram of an embodiment of the present invention for controlling the temperature of a target roasting furnace.
[0059] Figure 3 The present invention is a flowchart diagram of an embodiment of the present invention for synchronously constructing and generating the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control based on the reinforcement learning on-policy algorithm. DETAILED DESCRIPTION
[0060] The present invention will be further described below with reference to specific drawings and embodiments.
[0061] In order to achieve precise control of the temperature of the roasting main furnace, improve response speed, and reduce energy consumption, the present invention provides an iron ore magnetic separation roasting temperature control method based on double-loop reinforcement learning. Specifically, the temperature control method includes:
[0062] A target roasting furnace for magnetic separation roasting of iron ore is provided, and a roasting temperature control loop adapted to the target roasting furnace is constructed, wherein the roasting temperature control loop includes a gas flow control inner loop for controlling the gas flow rate and a main furnace temperature control outer loop for controlling the temperature of the roasting main furnace.
[0063] The inner loop of gas flow control is configured with an optimal gas flow control strategy, and the outer loop of main furnace temperature control is configured with an optimal main furnace temperature control strategy. These optimal gas flow control strategies and main furnace temperature control strategies are synchronously constructed and generated based on a reinforcement learning on-policy algorithm.
[0064] During roasting temperature control, outer loop operating status information is obtained, and a main furnace temperature optimal control strategy is used to generate a gas reference flow rate corresponding to the current outer loop operating status information, wherein the outer loop operating status is generated based on at least the roasting reference temperature and the current roasting operating temperature of the target roasting furnace;
[0065] The inner loop operation status information is obtained, and the gas flow optimal control strategy is used to generate a gas flow control signal corresponding to the current inner loop operation status information, so as to match the gas delivery flow delivered to the burner with the gas reference flow based on the gas flow control signal, thereby stabilizing the roasting operating temperature of the target roasting furnace at the roasting reference temperature, wherein,
[0066] The inner loop operating state information is generated based on at least a gas reference flow rate and a gas flow rate currently fed into the burner.
[0067] Figure 1FIG1 shows a flow chart of an embodiment of the roasting temperature control of the present invention. As can be seen from the figure, the object of temperature control is the target roasting furnace, wherein the target roasting furnace is a roasting furnace commonly used for magnetic separation roasting of iron ore. Figure 2 A control block diagram of an embodiment of controlling the roasting temperature of a target roasting furnace is shown in FIG. Figure 2 It can be seen that the target roasting furnace generally includes at least a roasting main furnace and a burner. When the iron ore is magnetically roasted, coal gas is transported into the burner, and the coal gas burns in the burner, thereby forming a temperature environment corresponding to the magnetic roasting of the iron ore in the roasting main furnace. The target roasting furnace and the method of using the target roasting furnace for magnetic roasting of iron ore are consistent with the existing technology and will not be repeated here.
[0068] As can be seen from the above description, when the target roasting furnace is temperature controlled, the main purpose is to make the target roasting furnace provide the temperature required for stabilizing the magnetic separation roasting of iron ore. Therefore, when temperature control is performed, a roasting reference temperature should be set, and the roasting operating temperature in the roasting main furnace should be stabilized at the roasting reference temperature. Specifically, the roasting reference temperature can generally be the 580°C mentioned above. It can be understood that when the roasting operating temperature of the roasting main furnace is stabilized at the roasting reference temperature, precise control of the temperature of the roasting main furnace is achieved, that is, roasting temperature control is achieved. The roasting operating temperature is stabilized at the roasting reference temperature, which specifically refers to that the roasting operating temperature is consistent with the roasting reference temperature, or the temperature error between the roasting operating temperature and the roasting reference temperature is within an allowable numerical range. The allowable numerical range can be selected as needed to meet the needs of iron ore magnetic separation roasting.
[0069] Depend on Figure 1 It can be seen that when the target roasting furnace is temperature controlled, a roasting temperature control loop should be constructed, and the roasting temperature control loop includes an inner loop of gas flow control and an outer loop of main furnace temperature control, wherein the main furnace temperature optimal control strategy is set in the outer loop of main furnace temperature control, and the gas flow optimal control strategy is set in the inner loop of gas flow control. In one embodiment of the present invention, the gas flow optimal control strategy and the main furnace temperature control optimal control strategy are synchronously constructed and generated based on the reinforcement learning on-policy algorithm. Based on the characteristics of the reinforcement learning on-policy algorithm, it can be seen that when the target roasting furnace is temperature controlled by the present invention, it is not necessary to accurately obtain the operating parameters of the target roasting furnace, nor to rely on the prior information of the target roasting furnace, thereby reducing the difficulty of temperature control of the target roasting furnace. Specifically, when the gas flow optimal control strategy and the main furnace temperature control optimal control strategy are synchronously constructed and generated by the reinforcement learning on-policy algorithm, the corresponding construction method and process can refer to the following corresponding description.
[0070] When controlling the roasting temperature, it is necessary to obtain the outer loop operating status information. After that, the main furnace temperature optimal control strategy is used to generate a gas reference flow corresponding to the outer loop operating status information. The generated gas reference flow should be used as the control target of the inner loop of the gas flow control. The outer loop operating status information should be generated based on at least the roasting reference temperature and the current roasting operating temperature of the target roasting furnace. The roasting operating temperature is the current temperature in the roasting main furnace. Figure 2 An embodiment of obtaining the roasting working temperature is shown in the figure. In the figure, the roasting working temperature of the roasting main furnace can be obtained through a temperature transmitter. The temperature transmitter can adopt an existing commonly used form. The method of obtaining the roasting working temperature using the temperature transmitter can be consistent with the existing technology. The situation of the outer loop operation status information will be specifically described below.
[0071] After the outer loop of the main furnace temperature control generates the gas reference flow, it should also obtain the inner loop operating status information. Thereafter, the gas flow optimal control strategy can be used to generate a gas flow control signal corresponding to the current inner loop operating status information. The gas flow control signal can be used to adjust the gas flow delivered to the burner, and the gas delivery flow delivered to the burner can be matched with the gas reference flow. The gas delivery flow is the gas flow delivered to the burner. The gas delivery flow matches the gas reference flow, specifically, the gas delivery flow is consistent with the gas reference flow, or the difference between the gas delivery flow and the gas reference flow is within an allowable numerical range. It should be understood that the allowable numerical range should be based on the demand for controlling the roasting temperature.
[0072] It is understood that the gas flow control signal should be related to the control of the gas flow delivered to the burner. Figure 2 When the gas flow rate to the burner is regulated by a flow control valve, the gas flow control signal should be a control signal that controls the valve opening of the flow control valve. For other situations, please refer to the description here. In addition, for the inner loop operating status information, please refer to the corresponding description below.
[0073] Based on the method of generating the gas reference flow rate of the present invention, it can be seen that the present invention controls the gas delivery flow rate to adapt to the gas reference flow rate, thereby stabilizing the roasting operating temperature at the roasting reference temperature, and thus achieving precise temperature control of the target roasting furnace, thereby improving the quality and production efficiency of iron ore concentrate products.
[0074] In one embodiment of the present invention, the optimal control strategy for gas flow is as follows:
[0075]
[0076] Among them, u q is the gas flow control signal, K qis the optimal control gain matrix of gas flow, ξ q is the inner loop running status information, K q1 , K q2 is the gain coefficient of the gas flow optimal control gain matrix, x q is the gas flow rate currently delivered to the burner, e q The gas flow tracking error is generated based on the gas reference flow and the gas delivery flow delivered to the burner;
[0077] From the above description, it can be seen that the gas flow optimal control strategy is a control method that can generate a gas flow control signal based on the inner loop operating status information. In order to implement the gas flow optimal control strategy, an inner loop controller should be set in the gas flow control inner loop, and the gas flow optimal control strategy should be configured in the inner loop controller, such as Figure 2 As shown, the inner loop controller can adopt an existing commonly used form, which is based on whether it can meet the requirements of executing the optimal control strategy for gas flow.
[0078] When implementing it specifically, z q is the cumulative error of gas flow tracking. It can be seen that the inner loop operation status information can include the current gas flow delivered to the burner and the cumulative error of gas flow tracking. The gas flow tracking error is generally the difference between the gas reference flow and the gas delivery flow, so: e q =q * -q, where q * is the gas reference flow rate, and q is the gas delivery flow rate. When calculating the cumulative gas flow tracking error, t is the time the target roaster temperature is controlled. Therefore, the cumulative gas flow tracking error can be calculated based on the roasting temperature control time. Therefore, after obtaining the current gas delivery flow rate, the corresponding inner loop operating status information can be obtained.
[0079] It can be understood that when the gas flow optimal control strategy is set in the inner loop controller, the gas flow optimal control gain matrix is set in the inner loop controller. Thereafter, after obtaining the inner loop operating status information, the inner loop operating status information is matrix multiplied by the gas flow optimal control gain matrix, and the result of the matrix multiplication is used as the output after control processing by the gas flow optimal control strategy, and the gas flow control signal can be obtained.
[0080] In one embodiment of the present invention, the optimal control strategy for the main furnace temperature control is:
[0081]
[0082] Among them, u w is the gas reference flow rate, K wis the optimal control gain matrix for the main furnace temperature control, ξ w is the outer loop operation status information, K w1 , K w2 is the gain coefficient of the optimal control gain matrix for the main furnace temperature control, x w is the current roasting temperature of the target roasting furnace, e w The temperature tracking error is generated based on the roasting reference temperature and the current roasting operating temperature of the target roasting furnace;
[0083] It should be noted that the method of executing the optimal control strategy for the main furnace temperature can refer to the above description of executing the optimal control strategy for the gas flow rate, such as the outer loop controller should be built into the outer loop of the main furnace temperature control, such as Figure 2 As shown, thereafter, the optimal control strategy for the main furnace temperature is set in the outer loop controller.
[0084] Specifically, z w The cumulative temperature tracking error is calculated using the same method as described above for the cumulative gas flow tracking error. Once the cumulative temperature tracking error is obtained, the outer loop operating status information can be obtained. This information may include the roasting operating temperature and the cumulative temperature tracking error. As can be seen from the above description, after obtaining the outer loop operating status information, it is matrix-multiplied with the main furnace temperature control optimal control gain matrix. The result of this matrix multiplication is used as the output of the main furnace temperature control optimal control strategy, resulting in the gas reference flow rate.
[0085] It should be noted that the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control are synchronously constructed based on the reinforcement learning on-policy algorithm. Specifically, based on the reinforcement learning on-policy algorithm, at least the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control can be synchronously generated. Therefore, constructing the optimal control strategy for gas flow at least includes constructing the optimal control gain matrix for gas flow; similarly, constructing the optimal control strategy for main furnace temperature control at least includes constructing the optimal control gain matrix for main furnace temperature control.
[0086] In one embodiment of the present invention, when synchronously determining and generating the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control based on a reinforcement learning on-policy algorithm, the method includes:
[0087] For the inner loop of gas flow control, an inner loop optimal control strategy corresponding to the gas flow optimal control strategy is constructed based on the gas flow optimal control, wherein the inner loop optimal control strategy includes an inner loop optimal control gain matrix to be solved;
[0088] For the main furnace temperature control outer loop, an outer loop optimal control strategy corresponding to the main furnace temperature optimal control strategy is constructed based on the main furnace temperature optimal control, wherein the outer loop optimal control strategy includes an outer loop optimal control gain matrix to be solved;
[0089] Based on the outer loop optimal control strategy and inner loop optimal control strategy constructed above, the roasting temperature of the target roasting furnace is pre-controlled, and the gain matrix solution based on reinforcement learning is performed during the roasting temperature pre-control process, so that the corresponding inner loop optimal control gain matrix and outer loop optimal control gain matrix are obtained synchronously after the gain matrix solution processing.
[0090] The inner loop optimal control gain matrix obtained by the solution is configured as the gas flow optimal control gain matrix, and the gas flow optimal control strategy is generated based on the gas flow optimal control gain matrix;
[0091] The outer loop optimal control gain matrix obtained by solving the problem is configured as the main furnace temperature control optimal control gain matrix, and the main furnace temperature control optimal control strategy is generated based on the main furnace temperature control optimal control gain matrix.
[0092] When constructing a gas flow optimal control strategy, an inner-loop optimal control strategy should be constructed first. The inner-loop optimal control strategy should correspond to the aforementioned gas flow optimal control strategy. Specifically, the inner-loop optimal control strategy corresponds to the gas flow optimal control strategy. Specifically, the form of the inner-loop optimal control strategy should be similar to that of the gas flow optimal control strategy. The difference is that the inner-loop optimal control strategy includes an inner-loop optimal control gain matrix to be solved. After solving the inner-loop optimal control gain matrix, the solved inner-loop optimal control gain matrix can be configured as the gas flow optimal control gain matrix, thereby obtaining the aforementioned gas flow optimal control strategy. It should be noted that when constructing the inner-loop optimal control strategy, it should be constructed at least based on the control objective of the gas flow optimal control. The specific construction method and process will be described in detail below.
[0093] Similarly, when constructing the optimal control strategy for the main furnace temperature control, the outer loop optimal control strategy should be constructed first. The situation between the constructed outer loop optimal control strategy and the main furnace temperature control optimal control strategy can refer to the corresponding description of the above-mentioned gas flow optimal control strategy, which will not be repeated here.
[0094] It can be seen from the above description that the roasting temperature control of the present invention does not require accurate acquisition of the operating parameters of the target roasting furnace, nor does it rely on the prior information of the target roasting furnace. During the specific implementation, in order to construct and solve the above-mentioned inner loop optimal control gain matrix and outer loop optimal control gain matrix, and then realize the construction of the gas flow optimal control strategy and the main furnace temperature control optimal control strategy, in one embodiment of the present invention, the target roasting furnace should be pre-controlled for roasting temperature based on the outer loop optimal control strategy and the inner loop optimal control strategy constructed above. Thereafter, in the process of roasting temperature pre-control, a gain matrix solution based on reinforcement learning is performed, so that the corresponding inner loop optimal control gain matrix and outer loop optimal control gain matrix can be obtained synchronously after the gain matrix solution is processed. Specifically, the method and process of performing roasting temperature pre-control on the target roasting furnace and the gain matrix solution based on reinforcement learning will be described in detail below.
[0095] In specific implementation, the inner-loop optimal control gain matrix obtained by solution is configured as the gas flow optimal control gain matrix, and the gas flow optimal control strategy can be generated based on the gas flow optimal control gain matrix; similarly, the outer-loop optimal control gain matrix obtained by solution is configured as the main furnace temperature control optimal control gain matrix, and the main furnace temperature control optimal control strategy is generated based on the main furnace temperature control optimal control gain matrix. At this time, the gas flow optimal control strategy and the main furnace temperature control optimal control strategy are simultaneously constructed and generated based on the reinforcement learning on-policy algorithm.
[0096] In one embodiment of the present invention, when performing the gain matrix solution process based on reinforcement learning, at least one pre-control process is performed, and the pre-control process includes data acquisition processing, equation solving processing, and gain matrix updating processing performed in sequence, wherein:
[0097] When performing data acquisition processing, inner loop pre-operation state information satisfying the inner loop full rank state and outer loop pre-operation state information satisfying the outer loop full rank state are acquired;
[0098] When performing equation solving, the inner loop pre-running state information is used to solve the inner loop algebraic Riccati equation to obtain the inner loop equation solution matrix, and the outer loop pre-running state information is used to solve the outer loop algebraic Riccati equation to obtain the outer loop equation solution matrix, where,
[0099] The inner loop algebraic Riccati equation is generated based on at least the inner loop optimal control strategy.
[0100] The outer loop algebraic Riccati equation is generated based on at least the outer loop optimal control strategy;
[0101] When performing the gain matrix update process, the inner loop optimal control gain matrix is updated based on the inner loop equation solution matrix and the inner loop gain matrix update strategy, and the outer loop optimal control gain matrix is updated based on the outer loop equation solution matrix and the outer loop gain matrix update strategy;
[0102] Repeat the above pre-control processing until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions. Thereafter, the updated inner-loop optimal control gain matrix is used as the solved inner-loop optimal control gain matrix, and the updated outer-loop optimal control gain matrix is used as the solved outer-loop optimal control gain matrix.
[0103] Figure 3 A flow chart of an embodiment of the gain matrix solution processing. Specifically, the gain matrix solution processing should perform pre-control processing at least once. It can be seen from the above description that when performing the pre-control processing, the roasting temperature of the target roasting furnace should be pre-controlled. The following is a specific description of the situation of pre-controlling the roasting temperature of the target roasting furnace.
[0104] In one embodiment of the present invention, when performing pre-control of the roasting temperature of a target roasting furnace, the method includes:
[0105] Constructing an initial inner loop optimal control gain matrix that satisfies the inner loop Hurwitz matrix state, and an initial outer loop optimal control gain matrix that satisfies the outer loop Hurwitz matrix state;
[0106] The initial inner loop optimal control gain matrix is configured within the inner loop optimal control strategy, and the initial outer loop optimal control gain matrix is configured within the outer loop optimal control strategy, so as to pre-control the temperature of the target roasting furnace using the corresponding inner loop optimal control strategy and outer loop optimal control strategy, and then perform a pre-control process;
[0107] After executing the pre-control processing, when the updated inner loop optimal control gain matrix and the outer loop optimal control gain matrix meet the pre-control convergence conditions, the roasting temperature pre-control of the target roasting furnace is exited. Otherwise,
[0108] The updated inner-loop optimal control gain matrix is configured in the inner-loop optimal control strategy, and the outer-loop optimal control gain matrix is configured in the outer-loop optimal control strategy. The roasting temperature of the target roasting furnace is continued to be pre-controlled, and then the pre-control processing is performed until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions.
[0109] It should be noted that, compared with the above-mentioned roasting temperature control, roasting temperature pre-control is an informal control of the target roasting furnace. However, during roasting temperature pre-control, the target roasting furnace still needs to be in operation. The difference is that during roasting temperature pre-control, the inner loop optimal control strategy and the outer loop optimal control strategy are mainly used for temperature control, while during the above-mentioned roasting temperature control, the gas flow optimal control strategy and the main furnace temperature control optimal control strategy are used for roasting temperature control.
[0110] From the above description, it can be seen that in order to perform roasting temperature pre-control, at least the inner-loop optimal control gain matrix within the inner-loop optimal control strategy and the outer-loop optimal control gain matrix within the outer-loop optimal control strategy should be given. Therefore, in the initial case, the initial inner-loop optimal control gain matrix and the initial outer-loop optimal control gain matrix should be given, wherein the initial inner-loop optimal control gain matrix can be determined by satisfying the state of the inner-loop Hurwitz matrix. Similarly, the initial outer-loop optimal control gain matrix can be determined by satisfying the state of the outer-loop Hurwitz matrix. The method of satisfying the state of the inner-loop Hurwitz matrix and the state of the outer-loop Hurwitz matrix will be specifically described below.
[0111] After the initial inner-loop optimal control gain matrix and the initial outer-loop optimal control gain matrix are given, the initial inner-loop optimal control gain matrix should be configured in the inner-loop optimal control strategy, and the initial outer-loop optimal control gain matrix should be configured in the outer-loop optimal control strategy. Thereafter, the corresponding inner-loop optimal control strategy and outer-loop optimal control strategy can be used to perform temperature pre-control on the target roasting furnace. The method of temperature pre-control can refer to the above description of roasting temperature control on the target roasting furnace, which is specifically: obtaining the outer-loop pre-operation status information, and using the outer-loop optimal control strategy to generate a pre-control gas reference flow corresponding to the current outer-loop pre-operation status information. After generating the pre-control gas reference flow, obtain the corresponding inner-loop pre-operation status information. Thereafter, the inner-loop optimal control strategy can be used to generate pre-control gas control information corresponding to the current inner-loop pre-operation status information.
[0112] It should be noted that when the initial inner loop optimal control gain matrix is configured in the inner loop optimal control strategy, and the initial outer loop optimal control gain matrix is configured in the outer loop optimal control strategy, then the moment when the target roasting furnace is pre-controlled is the moment t=0 of the present invention, that is, the zero point of the present invention.
[0113] During specific implementation, a pre-control process should be performed during the process of pre-controlling the temperature of the target roasting furnace using the corresponding inner-loop optimal control strategy and outer-loop optimal control strategy. For example, the above-mentioned initial inner-loop optimal control gain matrix and the initial outer-loop optimal control gain matrix are respectively configured in the inner-loop optimal control strategy and the outer-loop optimal control strategy, and during the process of pre-controlling the roasting temperature, a pre-control process should be performed. After performing the pre-control process, the inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix are updated to obtain a new inner-loop optimal control gain matrix and an outer-loop optimal control gain matrix.
[0114] In order to update the inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix after one pre-control processing, in one embodiment of the present invention, the pre-control processing should at least include data acquisition processing, equation solving processing and gain matrix update processing, wherein, in each pre-control processing, data acquisition processing, equation solving processing and gain matrix update processing will be executed in sequence. The data acquisition processing, equation solving processing and gain matrix are explained in detail below.
[0115] When executing data collection and processing, the inner loop pre-operation state information that satisfies the inner loop full rank state and the outer loop pre-operation state information that satisfies the outer loop full rank state are collected. Specifically, the inner loop pre-operation state information and the outer loop pre-operation state information should correspond to the above-mentioned inner loop operation state information and the outer loop operation state information. For details, please refer to the above description. In specific implementation, the inner loop pre-operation state information should satisfy the inner loop full rank state, and at the same time, the outer loop pre-operation state information should satisfy the outer loop full rank state. It can be seen that Figure 3 The inner and outer ring full rank conditions in are simultaneously met, specifically referring to the inner ring full rank state and the outer ring full rank state being met at the same time. The situations of the inner ring full rank state and the outer ring full rank state will be explained in detail below.
[0116] When performing equation solving, the inner-loop pre-operational state information is used to solve the inner-loop algebraic Riccati equation to obtain an inner-loop equation solution matrix, wherein the inner-loop algebraic Riccati equation is generated based on at least the inner-loop optimal control strategy. Therefore, the inner-loop pre-operational state information collected to satisfy the inner-loop full-rank state is primarily used to solve the inner-loop equation solution matrix of the inner-loop algebraic Riccati equation, wherein the inner-loop equation solution matrix specifically refers to the solution of the inner-loop algebraic Riccati equation, and the solution of the inner-loop algebraic Riccati equation is in matrix form. Similarly, the outer-loop pre-operational state information is used to solve the outer-loop equation solution matrix of the outer-loop algebraic Riccati equation.
[0117] In one embodiment of the present invention, constructing an inner loop algebraic Riccati equation includes:
[0118] Based on the gas flow control state of the gas flow control inner loop, a gas flow control state equation corresponding to the gas flow control inner loop is constructed, and a gas flow control performance index of the gas flow control inner loop is determined, wherein:
[0119] The gas flow control inner loop includes at least a gas flow regulating valve for regulating the gas flow delivered to the burner and a gas flow transmitter for obtaining the current gas flow delivered to the burner;
[0120] For the determined gas flow control performance index, the tracking state of the gas delivery flow tracking the gas reference flow is converted into a quadratic linear regulation state of the gas delivery flow;
[0121] Generating a quadratic linear regulation state of the gas delivery flow rate based on the conversion, generating an inner loop optimal control gain matrix associated with the quadratic linear regulation state based on the gas flow rate optimal control target, constructing an inner loop optimal control strategy based on the generated inner loop optimal control gain matrix, and constructing an inner loop gain matrix update strategy based on a calculation form of the inner loop optimal control gain matrix;
[0122] Based on the quadratic linear regulation state of gas delivery flow and the inner loop optimal control gain matrix, the inner loop algebraic Riccati equation is constructed.
[0123] It should be understood that the constructed inner loop algebraic Riccati equation should be related to the gas flow control state adopted by the inner loop of gas flow control. Figure 2 As can be seen from the description of the gas flow control state given in the above description, the present invention adjusts the gas flow delivered to the burner by the valve opening of the gas flow regulating valve, and the current gas flow can be obtained by the gas flow transmitter; specifically, when using Figure 2 When the gas flow control state is shown, the constructed gas flow control state equation can be:
[0124]
[0125] Among them, x q (t)=q(t),u q (t) = η(t), A q ,B q are all unknown parameter matrices of the inner loop of gas flow control, C q =[-1],q * Indicates the gas reference flow rate.
[0126] When constructing the gas flow control state equation, the gas flow regulating valve can be regarded as a first-order inertia link, and the gas flow variable can be regarded as a proportional link with a proportional coefficient of 1. In the constructed gas flow control state equation, is xq The differential of (t), q(t) represents the gas delivery flow at time t, u q (t) represents the gas flow control signal at time t, that is, the valve opening at time t, e q (t) represents the gas flow tracking error between the gas delivery flow q(t) at time t and the gas reference flow.
[0127] From the above description, it can be seen that the inner loop of gas flow control mainly makes the gas delivery flow match the gas reference flow. Therefore, when constructing the gas flow control performance index, we have:
[0128]
[0129] Among them, Q q With R q are all positive definite matrices, for u q The differential of st(1) is the above formula (1), and the positive definite matrix Q q , positive definite matrix R q It can be determined by using common technical means in this technical field.
[0130] It is understood that the above formula (2) represents the tracking state of the gas delivery flow rate tracking the gas reference flow rate. In order to convert the tracking state of the gas delivery flow rate tracking the gas reference flow rate into a quadratic linear regulation state of the gas delivery flow rate, in one embodiment of the present invention, the following is obtained:
[0131] The gas flow control auxiliary system is constructed as follows:
[0132]
[0133] Where,
[0134] According to the constructed gas flow control auxiliary system, the gas flow control performance index of the above formula (2) can be written as:
[0135]
[0136] in,
[0137] It can be understood that the above formula (4) is the converted quadratic linear regulation state of the gas delivery flow, that is, the tracking of the gas delivery flow to the gas reference flow can be converted into the optimal regulation of formula (4). Specifically, according to the optimal control theory, for the above gas flow control auxiliary system (3), the corresponding control state under the quadratic linear regulation of minimizing the gas delivery flow can be obtained as follows:
[0138]
[0139] Specifically, in the above formula (5) That is the inner loop optimal control gain matrix, where the inner loop optimal control gain matrix It can be expressed as:
[0140]
[0141] Among them, P q is the matrix solution of the inner loop algebraic Riccati equation, R q -1 is a positive definite matrix R q The inverse matrix of for In the description of the present invention, "T" represents the transposition operation.
[0142] In specific implementation, the inner loop algebraic Riccati equation can be constructed based on the quadratic linear regulation state of the gas delivery flow and the inner loop optimal control gain matrix. Specifically, the constructed inner loop algebraic Riccati equation is:
[0143]
[0144] From the above description, it can be seen that the inner-loop algebraic Riccati equation is related to the above-constructed gas flow control auxiliary system and the inner-loop optimal control gain matrix. That is, the inner-loop algebraic Riccati equation of formula (7) can be constructed based on the quadratic linear regulation state of the gas delivery flow and the inner-loop optimal control gain matrix.
[0145] Furthermore, as can be seen from the above description, after obtaining the inner loop optimal control gain matrix, an inner loop optimal control strategy can be constructed. The inner loop optimal control strategy can refer to the aforementioned gas flow optimal control strategy. In addition, based on the calculation form of the inner loop optimal control gain matrix, an inner loop gain matrix update strategy is constructed. Specifically, the inner loop optimal control gain matrix update strategy can be:
[0146]
[0147] Among them, K q1 i+1 is the inner loop optimal control gain matrix obtained after the i-th pre-control update, P q i is the solution matrix of the inner loop equation obtained after the i-th pre-control, For details, please refer to the following instructions.
[0148] It can be understood that after collecting the outer loop pre-operation state information that meets the outer loop full rank condition, the outer loop pre-operation state information can be used to solve the outer loop equation solution matrix of the outer loop algebraic Riccati equation, wherein the outer loop algebraic Riccati equation is at least constructed and generated based on the outer loop optimal control strategy. The outer loop algebraic Riccati equation and the outer loop equation solution matrix can refer to the above-mentioned description of the inner loop algebraic Riccati equation and the inner loop equation solution matrix. The method and process of constructing the outer loop algebraic Riccati equation are specifically described below.
[0149] In one embodiment of the present invention, constructing an outer loop algebraic Riccati equation includes:
[0150] Based on the main furnace temperature control state of the main furnace temperature control outer loop, a main furnace temperature control state equation corresponding to the main furnace temperature control outer loop is constructed, and the main furnace temperature control performance index of the main furnace temperature control outer loop is determined, wherein,
[0151] The main furnace temperature control outer loop includes at least a burner, a roasting main furnace, and a temperature transmitter for the current roasting working temperature of the roasting main furnace. When the main furnace temperature control state is established, the burner is used as a delay link;
[0152] For the determined main furnace temperature control performance index, the tracking state of the roasting working temperature tracking the roasting reference temperature is converted into a quadratic linear regulation state of the roasting working temperature;
[0153] A quadratic linear regulation state of the roasting operating temperature is generated based on the conversion, an outer loop optimal control gain matrix associated with the quadratic linear regulation state is generated based on the optimal control target of the main furnace temperature, and an outer loop optimal control strategy is constructed based on the generated outer loop optimal control gain matrix, wherein an inner loop gain matrix update strategy is constructed based on the calculation form of the outer loop optimal control gain matrix;
[0154] Based on the quadratic linear regulation state of the roasting working temperature and the outer loop optimal control gain matrix, the outer loop algebraic Riccati equation is constructed.
[0155] During specific implementation, the main furnace temperature control state can refer to the relevant description of the gas flow control state of the gas flow control inner loop. Specifically, Figure 2 The above description also provides an embodiment of the main furnace temperature control state. The main furnace temperature control outer loop includes at least a burner, a roasting main furnace, and a temperature transmitter for the current roasting working temperature of the roasting main furnace. That is, the main furnace temperature control state can be obtained through the burner, the roasting main furnace, and the temperature transmitter.
[0156] During the actual roasting process, the temperature inside the main roasting furnace doesn't change immediately when the gas flow rate changes, meaning there's a lag in the roasting temperature control process. Therefore, it's essential to consider the delay element when developing the main furnace temperature control equation. In the first embodiment of the present invention, the burner is considered the delay element, while the main roasting furnace is considered the first-order inertia element.
[0157] Based on the above description, the state equation of the main furnace temperature control is constructed as follows:
[0158]
[0159] Among them, x w (t)=r(t),u w (t) = q * , A w ,A d ,B w The parameter matrix of the outer loop of the main furnace temperature control is unknown, C w =[-1].
[0160] Specifically, r(t) is the calcination temperature at time t, r * is the roasting reference temperature. As can be seen from the above description, the roasting reference temperature r * Generally 580℃, q * is the reference gas flow rate, τ is the delay time, which is generally related to the working parameters of the burner. w It is the temperature tracking error between the coal roasting reference temperature and the roasting working temperature.
[0161] Since there is a time lag in the temperature generation in the outer loop of the main furnace temperature control, that is, there is a delay link in the main furnace temperature control state equation, in the specific implementation, the state variables with the delay link are processed as follows. Specifically,
[0162] First, define the lag operator D τ , and have:
[0163] x w (t-τ)=D τ x w (t) (10)
[0164] Among them, the lag operator D τ Should meet Therefore, the above formula (9) can be rewritten as:
[0165]
[0166] in,
[0167] Based on the main furnace temperature control state of formula (11), the main furnace temperature control performance index is:
[0168]
[0169] Among them, Q w With R w Are all positive definite matrices, positive definite matrix Q w , positive definite matrix R w For the case of , please refer to the corresponding description above. In addition, "(11)" in st(11) specifically refers to the above formula (11).
[0170] In order to convert the tracking state of the roasting working temperature tracking the roasting reference temperature into the quadratic linear regulation state of the roasting working temperature, it can be seen from the above description that the main furnace temperature control auxiliary system should be constructed. Specifically,
[0171]
[0172] Where, for u w The differential of .
[0173] With reference to the above description, the main furnace temperature control performance index of formula (12) can be converted to obtain the quadratic linear regulation state of the roasting working temperature, specifically:
[0174]
[0175] Where,
[0176] It can be understood that the above formula (14) is the quadratic linear regulation state of the roasting working temperature generated by the conversion, that is, the tracking of the roasting working temperature to the roasting reference temperature can be converted into the optimal regulation of formula (14). Specifically, according to the optimal control theory, for the above main furnace temperature control auxiliary system (13), the corresponding control state under the quadratic linear regulation of minimizing the roasting working temperature should be:
[0177]
[0178] Specifically, in the above formula (15) That is the outer loop optimal control gain matrix, where the outer loop optimal control gain matrix It can be expressed as:
[0179]
[0180] Among them, P w is the matrix solution of the inner loop algebraic Riccati equation.
[0181] In a specific implementation, an outer-loop algebraic Riccati equation can be constructed based on the quadratic linear regulation state of the roasting working temperature and the outer-loop optimal control gain matrix. Specifically, the inner-loop algebraic Riccati equation is constructed as follows:
[0182]
[0183] From the above description, it can be seen that the outer-loop algebraic Riccati equation is related to the main furnace temperature control auxiliary system and the outer-loop optimal control gain matrix constructed above, that is, the outer-loop algebraic Riccati equation of formula (17) can be constructed based on the quadratic linear adjustment state of the roasting working temperature and the outer-loop optimal control gain matrix.
[0184] Furthermore, as can be seen from the above description, after obtaining the outer loop optimal control gain matrix, an outer loop optimal control strategy can be constructed. The outer loop optimal control strategy can refer to the above-mentioned main furnace temperature control optimal control strategy. In addition, based on the calculation form of the inner loop optimal control gain matrix, an outer loop gain matrix update strategy is constructed. Specifically, the outer loop optimal control gain matrix strategy can be:
[0185]
[0186] in, is the outer loop optimal control gain matrix obtained after the i-th pre-control update, is the solution matrix of the outer loop equation obtained after the i-th pre-control. In the initial state, i is 0, that is, i should be 0 during the first control processing.
[0187] The above provides an embodiment of constructing the inner loop algebraic Riccati equation and the outer loop algebraic Riccati equation. The following example illustrates a method of using the inner loop pre-operation state information to solve the inner loop algebraic Riccati equation to obtain the inner loop equation solution matrix. Specifically,
[0188] For the inner loop algebraic Riccati equation constructed by formula (7), the inner loop augmented state is constructed, and based on the constructed inner loop augmented state system, we can have:
[0189]
[0190] in, It is the inner loop pre-operation status information, Please refer to the above description of the inner loop operating status information.
[0191] Based on the above inner loop augmented state system and Kleiman algorithm, the above inner loop algebraic Riccati equation is iteratively solved. The specific equation to be iteratively solved is:
[0192]
[0193] Depend on Figure 3 As can be seen from the above iterative solution description, the value of the inner loop optimal control gain matrix should be given during the iterative solution, that is, the initial inner loop optimal control gain matrix should be given. In specific implementation, in order to meet the subsequent gas flow optimal control strategy so that the gas delivery flow can accurately track the gas reference flow, the initial inner loop optimal control gain matrix should satisfy the inner loop Hurwitz matrix state. When the inner loop Hurwitz matrix state is satisfied, then: is the inner ring Hurwitz matrix, where is the initial inner loop optimal control gain matrix. Generally, the initial inner loop optimal control gain matrix can be determined by trial and error.
[0194] It should be noted that the above-mentioned iterative solution equation can be solved by the commonly used method in this technical field to obtain the iterative equation. Specifically, the method and process of iteratively solving equation (20) by using the inner loop pre-operation state information can be consistent with the existing technology, such as using an online model-free RL algorithm. The online model-free RL algorithm can refer to the corresponding description in "Proceedings of the Chinese Society of Electrical Engineering, Vol. 44, No. 9, published on May 5, 2024: Model-free Optimal Coordinated Control of a Rigid-Connected Dual-Motor System Based on Reinforcement Learning". Thereafter, according to the update strategy of the inner loop optimal control gain matrix, it can be obtained After the first inner loop optimal control gain matrix is updated, the inner loop optimal control gain matrix can be obtained: The same applies to other situations, and no examples will be given here one by one.
[0195] From the above description, it can be seen that the gas flow optimal control gain matrix corresponds to the inner loop optimal control gain matrix, and the inner loop operation state information corresponds to the inner loop pre-operation state. Therefore, according to the iterative solution method of the above iterative solution equation and the method of updating the inner loop optimal control gain matrix, it can be seen that the z of the inner loop operation state information q It can be regarded as a dynamic compensator containing a constant signal. Therefore, when the gas flow control inner loop is used to utilize the above-mentioned gas flow optimal control strategy, the flow tracking error e q It will gradually converge to 0, which means that the gas delivery flow rate matches the gas reference flow rate.
[0196] In specific implementations, based on the gas flow control state of the inner gas flow control loop described above, the inner gas flow control loop can be considered a first-order system. In this case, the rank corresponding to the full-rank state of the inner gas flow control loop should be 7. Therefore, when collecting the inner loop pre-operational state information, the rank corresponding to the inner loop pre-operational state should be 7. This allows the time corresponding to the collection of the inner loop pre-operational state information to be determined, and thus the inner loop pre-operational state information under each pre-control process to be determined.
[0197] It should be noted that, when collecting the outer loop pre-operation state information to satisfy the outer loop full rank state, the above description of the inner loop pre-operation state information satisfying the inner loop full rank state can be referred to, and no further details will be given here. After collecting the outer loop pre-operation state information that satisfies the outer loop full rank state, the outer loop pre-operation state information can be used to solve the outer loop algebraic Riccati equation to obtain the outer loop equation solution matrix. The specific method of solving the outer loop equation solution matrix can refer to the description of solving the inner loop equation solution matrix. For example, the above method can be used to construct an inner loop augmented system, and then the Clayman algorithm can be used to construct an iterative solution equation. The specific method of constructing the outer loop iterative solution equation can refer to the above description, and no further details will be given here.
[0198] Depend on Figure 3 As can be seen from the above description, in the iterative solution process, the initial outer-loop optimal control gain matrix should also be given, and the initial outer-loop optimal control gain matrix should also satisfy the outer-loop Hurwitz matrix state. The outer-loop Hurwitz matrix state can refer to the description of the inner-loop Hurwitz matrix above.
[0199] Furthermore, after obtaining the optimal control gain matrix and optimal control strategy of the main furnace temperature control of the present invention by the above method, the temperature tracking error e between the roasting working temperature and the roasting reference temperature can be reduced to w It will gradually converge to 0, which means that the roasting working temperature will be stabilized at the roasting reference temperature.
[0200] In order to improve the anti-interference ability, when the inner loop optimal control strategy and the outer loop optimal control strategy are used to pre-control the roasting temperature of the target roasting furnace, the following equations are obtained:
[0201]
[0202] in, is the gas flow control signal under the jth inner loop pre-operation status sub-information when executing the i-th pre-control processing, is the jth inner loop pre-operation state sub-information when executing the i-th pre-control processing, ρ1 is the measurable noise of the inner loop, is the gas reference flow rate under the hth outer loop pre-operation status sub-information when executing the i-th pre-control processing, is the hth outer loop pre-operation state sub-information when executing the i-th pre-control processing, and ρ2 is the measurable noise of the outer loop.
[0203] Specifically, the inner loop measurable noise ρ1 and the outer loop measurable noise ρ2 can be determined according to the working conditions of the target roasting furnace, that is, the inner loop measurable noise ρ1 and the outer loop measurable noise ρ2 should be based on the interference / noise that can characterize the inner loop of the gas flow control and the outer loop of the main furnace temperature control of the target roasting furnace. It is understandable that when the inner loop measurable noise ρ1 and the outer loop measurable noise ρ2 are added during the pre-control of the roasting temperature, it can also ensure that the collected inner loop pre-operation state sub-information and the outer loop pre-operation state sub-information have sufficient incentives. It is understandable that after adding the inner loop measurable noise ρ1 and the outer loop measurable noise ρ2, the collected inner loop pre-operation state information and the outer loop pre-operation state information will be affected.
[0204] From the above description, it can be seen that during a pre-control process, the inner loop pre-operation state information collected may include multiple ring pre-operation state sub-information. Therefore, the j-th inner loop pre-operation state sub-information is the j-th inner loop pre-operation state sub-information obtained according to the sampling time in the inner loop pre-operation state information; similarly, the h-th outer loop pre-operation state sub-information can be obtained. The corresponding meanings will not be elaborated here.
[0205] It should be understood that when pre-controlling the roasting temperature, the inner loop optimal control strategy and the outer loop optimal control strategy are respectively added with the inner loop measurable noise ρ1 and the outer loop measurable noise ρ2. After the gas flow optimal control strategy and the main furnace temperature control optimal control strategy are obtained by the above method, the anti-interference ability of the roasting temperature control can be effectively improved.
[0206] It should be noted that when performing the gain matrix update process, the inner-loop optimal control gain matrix is updated based on the inner-loop equation solution matrix and the inner-loop gain matrix update strategy, and the outer-loop optimal control gain matrix is updated based on the outer-loop equation solution matrix and the outer-loop gain matrix update strategy. Specifically, after updating the inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix, it is necessary to determine whether the pre-control convergence condition is met. If the pre-control convergence condition is met, the pre-control process is stopped; otherwise, the pre-control process should be continued until the pre-control convergence condition is met.
[0207] When the pre-control convergence conditions are met, the updated inner-loop optimal control gain matrix is used as the solved inner-loop optimal control gain matrix, and the updated outer-loop optimal control gain matrix is used as the solved outer-loop optimal control gain matrix. From the above description, it can be seen that at this time, the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control can be constructed.
[0208] In one embodiment of the present invention, the pre-control convergence condition is:
[0209]
[0210] in, is the inner loop optimal control gain matrix when executing the i-th pre-control process, is the updated inner loop optimal control gain matrix after executing the i-th pre-control process, is the outer loop optimal control gain matrix when executing the i-th pre-control process, is the updated optimal control gain matrix of the inner loop after executing the i-th pre-control process, and σ is the pre-control convergence threshold.
[0211] Specifically, the pre-control convergence condition is Figure 3 The double loop in converges. Figure 3 It can be seen that when performing pre-control processing, the pre-control convergence threshold σ should be set. Generally, the pre-control convergence threshold σ can be a smaller value, which is based on the accuracy of roasting temperature control. Figure 3 It can be seen that the sampling time δt should also be set. The inner loop pre-operation status information and the outer loop pre-operation status information can be obtained through the sampling time δt. If the inner loop full rank state is met, the collected inner loop pre-operation status information should include multiple inner loop pre-operation status sub-information. The time interval between two adjacent inner loop pre-operation status sub-information should be the sampling time δt. The situation of each inner loop pre-operation status sub-information should be consistent with the above-mentioned inner loop operation status information. Similarly, the situation of the outer loop pre-operation status information can refer to the corresponding description here, and will not be repeated here.
[0212] Specifically, when and Of course, in specific implementation, the roasting temperature control requirements can also be set, and the pre-control convergence condition can be set to other situations, which will not be listed here one by one.
[0213] It can be understood that after executing a pre-control process, when the above-mentioned pre-control convergence conditions are met, the subsequent pre-control process will be stopped. At this time, the inner loop optimal control gain matrix after the first update should be configured as the gas flow optimal control gain matrix. At the same time, the outer loop optimal control gain matrix after the first update should be configured as the main furnace temperature control optimal control gain matrix. For other situations, please refer to the above and this explanation, and no examples will be given one by one.
[0214] In one embodiment of the present invention, when the roasting temperature of the target roasting furnace is controlled by using the inner loop of gas flow control and the outer loop of main furnace temperature control, the inner loop of gas flow control and the outer loop of main furnace temperature control are interacted with each other, and based on the mutual interaction state, the optimal control strategy of gas flow in the inner loop of gas flow control is fine-tuned, and / or the optimal control strategy of main furnace temperature control in the outer loop of main furnace temperature control is fine-tuned.
[0215] From the above explanation, it can be seen that after obtaining the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control, the optimal gas flow control strategy should be deployed in the inner loop controller, and the optimal main furnace temperature control strategy should be deployed in the outer loop controller. In the actual complex industry, the key link of iron ore magnetic separation roasting temperature control is the ability to quickly adapt to the complex and changing operating conditions of the suspended magnetization roasting process. Whether it is fluctuations in raw material characteristics, interference from the external environment, or changes in equipment performance, it must rely on its efficient learning ability to adjust the control in a timely manner to ensure that the roasting operating temperature of the roasting furnace is always stable at the roasting reference temperature, such as stabilizing the roasting operating temperature of the roasting furnace at 580°C.
[0216] In order to further improve the accuracy of roasting temperature control, in one embodiment of the present invention, the inner loop of gas flow control is configured to interact with the outer loop of main furnace temperature control, wherein the interaction specifically refers to at least allowing the inner loop controller and the outer loop controller to interact with each other, such as Figure 2 shown.
[0217] Figure 2 In the process, the outer loop controller should receive the temperature tracking error e w , roasting operating temperature, gas delivery flow rate and gas flow optimal control strategy, and the outer loop controller receives the flow tracking error e q , roasting operating temperature, gas delivery flow rate and optimal control strategy for main furnace temperature control.
[0218] In specific implementation, the temperature tracking error e is received w After determining the optimal control strategy for the roasting operating temperature, gas delivery flow, and gas flow, combined with the gas reference flow output by the outer loop controller, the performance of the current main furnace temperature control optimal control strategy in minimizing the preset temperature control outer loop cost function can be constructed and evaluated. The constructed temperature control outer loop cost function should include factors such as temperature error and control consumption. The temperature control outer loop cost function can be selected as needed to characterize the current control consumption and temperature control accuracy. In specific implementation, the constructed temperature control outer loop cost function should be:
[0219] It is understandable that when the gas flow control inner loop fails to effectively achieve the gas delivery flow accurately tracking the gas reference flow, it affects the tracking of the roasting working temperature to the roasting reference temperature, which will lead to a higher value of the temperature control outer loop cost function. At this time, the evaluation part of the outer loop controller will observe the flow tracking error e. q and the resulting temperature tracking error e w , it is determined that the current optimal control strategy for the main furnace temperature control does not perform well in minimizing the outer loop cost function of the temperature control. Afterwards, in the execution part of the outer loop controller, the gain coefficient K in the optimal control gain matrix of the main furnace temperature control isw1 , gain coefficient K w2 Perform fine-tuning and observe the temperature control status of the target roasting furnace and the performance of the above-mentioned temperature control test function after fine-tuning. During fine-tuning, the gain coefficient K can be adjusted. w1 and / or gain factor K w2 The corresponding value, of course, adjusts the gain factor K w1 and / or gain factor K w2 When selecting the corresponding value, the adjustment direction should be based on minimizing the cost function of the temperature control outer loop.
[0220] In specific implementation, the inner loop controller can adopt the same interactive fine-tuning method as above, such as constructing a corresponding temperature control inner loop cost function, and determining the adjustment gain coefficient K in the roasting temperature control process according to the constructed inner loop cost function. q1 and / or gain factor K q2 For the adjustment of the corresponding values, please refer to the description of the interaction of the outer loop controller. Specifically, the constructed temperature control inner loop cost function can be:
[0221] It should be understood that the outer loop of the main furnace temperature control does not blindly control the roasting working temperature, but rather indirectly guides and coordinates the inner loop of the gas flow control and the outer loop of the main furnace temperature control by tracking the gas delivery flow in the inner loop of the gas flow control to the gas reference flow, and evaluating the performance of the current main furnace temperature control optimal control strategy in minimizing the cost function of the outer loop of the temperature control.
[0222] Furthermore, the inner loop of the gas flow control is not a simple flow tracking control, but controls the gas delivery flow based on the information transmitted by the outer loop of the main furnace temperature control and the optimal gas flow control strategy. Therefore, on the inner loop of the gas flow control, the gas delivery flow can match the gas reference flow, and on the outer loop of the main furnace temperature control, the roasting operating temperature can be stabilized at the roasting reference temperature, and the cost function of the outer loop of the temperature control can be minimized. Therefore, while being able to accurately control the temperature of the roasting main furnace, energy consumption can be further reduced and the quality and production efficiency of iron ore products can be improved.
Claims
1. A method for controlling the temperature of iron ore magnetic separation and roasting based on double-loop reinforcement learning, characterized in that: The temperature control method comprises: A target roasting furnace for magnetic separation roasting of iron ore is provided, and a roasting temperature control loop adapted to the target roasting furnace is constructed, wherein the roasting temperature control loop includes a gas flow control inner loop for controlling the gas flow rate and a main furnace temperature control outer loop for controlling the temperature of the roasting main furnace. The inner loop of gas flow control is configured with an optimal gas flow control strategy, and the outer loop of main furnace temperature control is configured with an optimal main furnace temperature control strategy. These optimal gas flow control strategies and main furnace temperature control strategies are synchronously constructed and generated based on a reinforcement learning on-policy algorithm. During roasting temperature control, outer loop operating status information is obtained, and a main furnace temperature optimal control strategy is used to generate a gas reference flow rate corresponding to the current outer loop operating status information, wherein the outer loop operating status is generated based on at least the roasting reference temperature and the current roasting operating temperature of the target roasting furnace; The inner loop operation status information is obtained, and the gas flow optimal control strategy is used to generate a gas flow control signal corresponding to the current inner loop operation status information, so as to match the gas delivery flow delivered to the burner with the gas reference flow based on the gas flow control signal, thereby stabilizing the roasting operating temperature of the target roasting furnace at the roasting reference temperature, wherein, The inner loop operating state information is generated based on at least a gas reference flow rate and a gas flow rate currently fed into the burner.
2. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 1 is characterized in that: For the optimal control strategy of gas flow, we have: Among them, u q is the gas flow control signal, K q is the optimal control gain matrix of gas flow, ξ q is the inner loop running status information, K q1 , K q2 is the gain coefficient of the gas flow optimal control gain matrix, x q is the gas flow rate currently delivered to the burner, e q The gas flow tracking error is generated based on the gas reference flow and the gas delivery flow delivered to the burner; For the optimal control strategy of the main furnace temperature control, we have: Among them, u w is the gas reference flow rate, K w is the optimal control gain matrix for the main furnace temperature control, ξ w is the outer loop operation status information, K w1 , K w2 is the gain coefficient of the optimal control gain matrix for the main furnace temperature control, x w is the current roasting temperature of the target roasting furnace, e w The temperature tracking error is generated based on the roasting reference temperature and the current roasting operating temperature of the target roasting furnace; When the optimal control strategy for gas flow and the optimal control strategy for main furnace temperature control are synchronously constructed based on the reinforcement learning on-policy algorithm, at least the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control are synchronously generated.
3. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 2 is characterized in that: When synchronously determining and generating the optimal control gain matrix for gas flow and the optimal control gain matrix for main furnace temperature control based on the reinforcement learning on-policy algorithm, the following steps are included: For the inner loop of gas flow control, an inner loop optimal control strategy corresponding to the gas flow optimal control strategy is constructed based on the gas flow optimal control, wherein the inner loop optimal control strategy includes an inner loop optimal control gain matrix to be solved; For the main furnace temperature control outer loop, an outer loop optimal control strategy corresponding to the main furnace temperature optimal control strategy is constructed based on the main furnace temperature optimal control, wherein the outer loop optimal control strategy includes an outer loop optimal control gain matrix to be solved; Based on the outer loop optimal control strategy and inner loop optimal control strategy constructed above, the roasting temperature of the target roasting furnace is pre-controlled, and the gain matrix solution based on reinforcement learning is performed during the roasting temperature pre-control process, so that the corresponding inner loop optimal control gain matrix and outer loop optimal control gain matrix are obtained synchronously after the gain matrix solution processing. The inner loop optimal control gain matrix obtained by the solution is configured as the gas flow optimal control gain matrix, and the gas flow optimal control strategy is generated based on the gas flow optimal control gain matrix; The outer loop optimal control gain matrix obtained by solving the problem is configured as the main furnace temperature control optimal control gain matrix, and the main furnace temperature control optimal control strategy is generated based on the main furnace temperature control optimal control gain matrix.
4. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 3 is characterized in that: When performing the gain matrix solution process based on reinforcement learning, at least one pre-control process is performed, and the pre-control process includes data acquisition processing, equation solving processing and gain matrix updating processing performed in sequence, wherein, When performing data acquisition processing, inner loop pre-operation state information satisfying the inner loop full rank state and outer loop pre-operation state information satisfying the outer loop full rank state are acquired; When performing equation solving, the inner loop pre-running state information is used to solve the inner loop algebraic Riccati equation to obtain the inner loop equation solution matrix, and the outer loop pre-running state information is used to solve the outer loop algebraic Riccati equation to obtain the outer loop equation solution matrix, where, The inner loop algebraic Riccati equation is generated based on at least the inner loop optimal control strategy. The outer loop algebraic Riccati equation is generated based on at least the outer loop optimal control strategy; When performing the gain matrix update process, the inner loop optimal control gain matrix is updated based on the inner loop equation solution matrix and the inner loop gain matrix update strategy, and the outer loop optimal control gain matrix is updated based on the outer loop equation solution matrix and the outer loop gain matrix update strategy; Repeat the above pre-control processing until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions. Thereafter, the updated inner-loop optimal control gain matrix is used as the solved inner-loop optimal control gain matrix, and the updated outer-loop optimal control gain matrix is used as the solved outer-loop optimal control gain matrix.
5. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 4 is characterized in that: When pre-controlling the roasting temperature of the target roasting furnace, it includes: Constructing an initial inner loop optimal control gain matrix that satisfies the inner loop Hurwitz matrix state, and an initial outer loop optimal control gain matrix that satisfies the outer loop Hurwitz matrix state; The initial inner loop optimal control gain matrix is configured within the inner loop optimal control strategy, and the initial outer loop optimal control gain matrix is configured within the outer loop optimal control strategy, so as to pre-control the temperature of the target roasting furnace using the corresponding inner loop optimal control strategy and outer loop optimal control strategy, and then perform a pre-control process; After executing the pre-control processing, when the updated inner loop optimal control gain matrix and the outer loop optimal control gain matrix meet the pre-control convergence conditions, the roasting temperature pre-control of the target roasting furnace is exited. Otherwise, The updated inner-loop optimal control gain matrix is configured in the inner-loop optimal control strategy, and the outer-loop optimal control gain matrix is configured in the outer-loop optimal control strategy. The roasting temperature of the target roasting furnace is continued to be pre-controlled, and then the pre-control processing is performed until the updated inner-loop optimal control gain matrix and the outer-loop optimal control gain matrix meet the pre-control convergence conditions.
6. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 4 is characterized in that: For the pre-control convergence condition, we have: in, is the inner loop optimal control gain matrix when executing the i-th pre-control process, is the updated inner loop optimal control gain matrix after executing the i-th pre-control process, is the outer loop optimal control gain matrix when executing the i-th pre-control process, is the updated optimal control gain matrix of the inner loop after executing the i-th pre-control process, and σ is the pre-control convergence threshold.
7. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 6 is characterized in that: When the inner loop optimal control strategy and the outer loop optimal control strategy are used to pre-control the roasting temperature of the target roasting furnace, the following equations are obtained: in, is the gas flow control signal under the jth inner loop pre-operation status sub-information when executing the i-th pre-control processing, is the jth inner loop pre-operation state sub-information when executing the i-th pre-control processing, ρ1 is the measurable noise of the inner loop, is the gas reference flow rate under the jth outer loop pre-operation status sub-information when executing the i-th pre-control processing, is the jth outer loop pre-operation state sub-information when executing the i-th pre-control processing, and ρ2 is the measurable noise of the outer loop.
8. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 6 is characterized in that: When constructing the inner loop algebraic Riccati equation, include: Based on the gas flow control state of the gas flow control inner loop, a gas flow control state equation corresponding to the gas flow control inner loop is constructed, and a gas flow control performance index of the gas flow control inner loop is determined, wherein: The gas flow control inner loop includes at least a gas flow regulating valve for adjusting the gas flow delivered to the burner and a gas flow transmitter for obtaining the current gas flow delivered to the burner. For the determined gas flow control performance index, the tracking state of the gas delivery flow tracking the gas reference flow is converted into a quadratic linear regulation state of the gas delivery flow; Generating a quadratic linear regulation state of the gas delivery flow rate based on the conversion, generating an inner loop optimal control gain matrix associated with the quadratic linear regulation state based on the gas flow rate optimal control target, constructing an inner loop optimal control strategy based on the generated inner loop optimal control gain matrix, and constructing an inner loop gain matrix update strategy based on a calculation form of the inner loop optimal control gain matrix; Based on the quadratic linear regulation state of gas delivery flow and the inner loop optimal control gain matrix, the inner loop algebraic Riccati equation is constructed.
9. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to claim 4 is characterized in that: When constructing the outer loop algebraic Riccati equation, include: Based on the main furnace temperature control state of the main furnace temperature control outer loop, a main furnace temperature control state equation corresponding to the main furnace temperature control outer loop is constructed, and the main furnace temperature control performance index of the main furnace temperature control outer loop is determined, wherein, The main furnace temperature control outer loop includes at least a burner, a roasting main furnace, and a temperature transmitter for the current roasting working temperature of the roasting main furnace. When the main furnace temperature control state is established, the burner is used as a delay link; For the determined main furnace temperature control performance index, the tracking state of the roasting working temperature tracking the roasting reference temperature is converted into a quadratic linear regulation state of the roasting working temperature; A quadratic linear regulation state of the roasting operating temperature is generated based on the conversion, an outer loop optimal control gain matrix associated with the quadratic linear regulation state is generated based on the optimal control target of the main furnace temperature, and an outer loop optimal control strategy is constructed based on the generated outer loop optimal control gain matrix, wherein an inner loop gain matrix update strategy is constructed based on the calculation form of the outer loop optimal control gain matrix; Based on the quadratic linear regulation state of the roasting working temperature and the outer loop optimal control gain matrix, the outer loop algebraic Riccati equation is constructed.
10. The iron ore magnetic separation and roasting temperature control method based on double-loop reinforcement learning according to any one of claims 1 to 9, characterized in that: When the roasting temperature of the target roasting furnace is controlled by using the inner loop of gas flow control and the outer loop of main furnace temperature control, the inner loop of gas flow control and the outer loop of main furnace temperature control are interacted with each other, and based on the mutual interaction state, the optimal control strategy of gas flow in the inner loop of gas flow control is fine-tuned, and / or the optimal control strategy of main furnace temperature control in the outer loop of main furnace temperature control is fine-tuned.