Tin smelting chemical reaction optimization control method based on reinforcement learning
By applying reinforcement learning-based control methods during tin smelting and optimizing process parameters in combination with physical models, the problem that traditional control methods cannot effectively optimize the tin generation amount and reduce carbon monoxide emissions is solved, and a more efficient and environmentally friendly tin smelting process is achieved.
Patent Information
- Application Number
- CN202510668665.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The traditional tin smelting control method cannot effectively consider the nonlinear coupling relationship between multivariate parameters, and lacks adaptability to complex reaction mechanisms and dynamic working conditions, resulting in problems such as difficulty in stably increasing the tin generation, difficulty in reducing carbon monoxide emissions, and excessive reaction time.
The optimization control method of tin smelting chemical reaction based on reinforcement learning is adopted, and by combining reinforcement learning algorithms with physical models in the metallurgy field, the coupling relationship between process parameters and chemical reactions is dynamically learned, so as to accurately control the parameters such as oxygen flow, coal flow, and spray gun position during tin smelting.
It increases the tin generation amount, reduces carbon monoxide emissions, shortens reaction time, and improves energy utilization efficiency, adapts to the multi-target optimization needs in complex industrial environments, providing a new solution for the intelligent and green development of tin smelting processes.
Smart Images

Figure CN120178693A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an optimized control method for the chemical reaction of tin smelting based on reinforcement learning, belonging to the cross - technical field of metallurgical engineering and artificial intelligence. Background Art
[0002] Tin smelting is an extremely important technological process in the metallurgical industry. Its core is to reduce stannic oxide in tin ore to metallic tin through chemical reactions. This process involves complex multi - variable technological parameters, such as oxygen flow rate, coal - burning flow rate, lance position, furnace pressure, smelting temperature, etc. The synergistic effect of these parameters directly affects the amount of tin produced, energy utilization efficiency, and the emission of reaction by - products (such as carbon monoxide). Therefore, how to achieve precise control of chemical reactions under dynamically changing technological conditions is the key issue for the optimization of the tin smelting process.
[0003] Traditional tin smelting control methods mainly rely on empirical rules or simple feedback control systems. These methods often cannot comprehensively consider the non - linear coupling relationship between multi - variable parameters and lack the adaptive ability to complex reaction mechanisms and dynamic working conditions. In addition, the lack of data and noise in actual production further exacerbate the control difficulty, resulting in problems such as the unstable increase in tin production, the difficulty in reducing carbon monoxide emissions, and the excessive reaction time, which not only affect production efficiency but also bring environmental burdens and energy waste.
[0004] In recent years, with the development of artificial intelligence technology, reinforcement learning has gradually become an important means for industrial process optimization due to its excellent performance in dynamic decision - making problems. Through continuous interaction with the environment, reinforcement learning can automatically adjust control strategies according to real - time states to maximize the objective function (such as tin production, energy consumption, environmental indicators, etc.). However, the application of traditional reinforcement learning algorithms in complex industrial scenarios still faces many challenges, including: how to use physical knowledge in the metallurgical field to guide model learning, how to efficiently explore the optimal strategy in a high - dimensional parameter space, and how to ensure the rapid adaptation of the model to real - time working condition changes, etc. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an optimized control method for the chemical reaction of tin smelting based on reinforcement learning, which can optimize the operation parameters in the smelting process in real - time, thus solving the above problems.
[0006] The technical solution of the present invention is: an optimized control method for the chemical reaction of tin smelting based on reinforcement learning. First, the reinforcement learning algorithm is combined with the physical model in the metallurgical field. By dynamically learning the coupling relationship between process parameters and chemical reactions, precise control of parameters such as oxygen flow rate, coal flow rate, and lance position during the tin smelting process is achieved. This method can increase the tin production, reduce carbon monoxide emissions, shorten the reaction time, improve the energy utilization efficiency at the same time, meet the multi-objective optimization requirements in complex industrial environments, and provide a new solution for the intelligent and green development of the tin smelting process.
[0007] An optimized control method for the chemical reaction of tin smelting based on reinforcement learning, the specific steps are as follows:
[0008] Step1: Collect parameter data during the smelting process;
[0009] Step2: Preprocess the collected original parameter data, clean outliers and fill in missing values, and use an improved Fourier synthesis method to simulate and generate the preprocessed parameter data to obtain simulated data;
[0010] Step3: Build a reinforcement learning model, define the state space, action space and design the reward function;
[0011] Step4: Design a dual-network architecture for the Actor-Critic network in the reinforcement learning model, and introduce the main network and the auxiliary network;
[0012] Step5: Based on the tin smelting process parameters, establish a physical and chemical model, and calculate the carbon monoxide generation amount in combination with the thermodynamic equilibrium formula and the chemical reaction formula in the furnace;
[0013] Step6: Input the calculated carbon monoxide generation amount as the reward function of the optimized reinforcement learning model to guide the agent in the reinforcement learning to select and execute actions;
[0014] Step7: Use the generated simulated data and the collected parameter data to train the optimized reinforcement learning model, and save the parameters of the final reinforcement learning model.
[0015] The specific content of Step2 is as follows:
[0016] Step2.1: According to the sensor measurement range and combined with the actual operation rules of the metallurgical process, set the parameter measurement range and the actual operation rule threshold, and eliminate outliers that exceed the parameter measurement range threshold and the parameter change speed exceeds the actual operation rule threshold;
[0017] Step2.2: Add an adaptive bottom feature detection algorithm to the original Fourier synthesis method to identify the steady-state interval of the parameter data;
[0018] Step 2.3: Use the improved Fourier synthesis algorithm to simulate and generate tin smelting parameter data, and expand the scale of the reinforcement learning training set.
[0019] The specific content of Step 3 is as follows:
[0020] Step 3.1: Design the parameter values affecting the chemical reactions in the tin smelting furnace as the state space, and define the controllable optimization variables during the smelting process as the action space;
[0021] Step 3.2: The reward function is used as the optimization objective of reinforcement learning, and is designed based on the three objectives of maximizing tin production, minimizing carbon monoxide emissions, and minimizing reaction time;
[0022] Step 3.3: Use the data and metallurgy dual-drive method as the state update model of reinforcement learning.
[0023] The specific content of Step 4 is as follows:
[0024] Step 4.1: Introduce two neural networks, the main network and the auxiliary network, into the Actor network. The main network is responsible for generating policies and outputting the optimized control actions in the current state; the auxiliary network interacts with the Critic network to provide action evaluation signals for the main network and guide the main network to generate actions;
[0025] Step 4.2: Introduce two neural networks, the main network and the auxiliary network, into the Critic network. The main network directly provides the value evaluation of the current action, which is used to guide the Actor network to optimize the policy; the auxiliary network provides a definite target value through the soft update mechanism.
[0026] The specific content of Step 5 is as follows:
[0027] Step 5.1: According to the actual operating conditions of the tin smelting process, determine the key process parameters related to the carbon monoxide generation amount and the tin generation amount. The key process parameters include smelting temperature, furnace pressure, oxygen concentration, stannic oxide addition amount, and carbon addition amount;
[0028] Step 5.2: Establish a thermodynamic equilibrium model according to the chemical reaction equation of the tin smelting process, which is used to calculate the carbon monoxide generation amount and the tin generation amount.
[0029] The specific content of Step 6 is as follows:
[0030] Step 6.1: Use the calculated tin generation amount and carbon monoxide generation amount as the input of the reward function. The reward function is:
[0031]
[0032] In the formula, is the reward function, , , are weight parameters used to balance the importance of different optimization objectives, and Sn t represents the amount of tin generated in the furnace at the current moment, and CO t represents the amount of carbon monoxide generated in the furnace at the current moment, and Time is the reaction time;
[0033] Step 6.2: Collect real-time status data of the furnace, and the optimal strategy model recommends the optimal action based on the real-time status data of the furnace.
[0034] Specifically, Step 2.2 is as follows:
[0035]
[0036] In the formula, is the window variance ratio, is the variance of data within the sliding window centered at time t,
[0037] is the global data variance.
[0038] Specifically, Step 3.3 is as follows:
[0039] Organize the industrial data of the tin smelting process into a triple form , in the formula, represents the state at the current moment, represents the operation action executed at the current moment, represents the state at the next moment;
[0040] Design an LSTM network, use the current state and action as input features, and the state at the next moment as the target variable for training;
[0041] Integrate metallurgical knowledge, use the parameter values calculated by chemical reactions to limit the parameter values predicted by the LSTM network, and realize the dynamic optimization of state update.
[0042] The beneficial effects of the present invention are as follows:
[0043] (1) Improvement in automation and intelligence: Replace the traditional manual regulation and optimization methods with automated intelligent control, greatly improving the control accuracy and decision-making efficiency, reducing the errors and dependencies of manual operations, and enhancing the intelligent level of the production process;
[0044] (2) Reduce data acquisition costs: Through high-fidelity spectrum reconstruction and dynamic noise injection for analog data, while preserving the physical laws of the real process, it significantly expands the diversity of training samples and enhances the model's adaptability to extreme working conditions and data missing scenarios. Description of the Drawings
[0045] Figure 1 is the overall framework diagram of the present invention;
[0046] Figure 2 is the execution process diagram of the reinforcement learning model;
[0047] Figure 3 is the experimental result diagram of the original Fourier synthesis method;
[0048] Figure 4 is the experimental result diagram of the improved Fourier synthesis method. Detailed Embodiments
[0049] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0050] Example 1: As Figure 1 shown, a method for optimizing the control of the chemical reaction in tin smelting based on reinforcement learning, the specific steps are as follows:
[0051] Step1: Collect parameter data during the smelting process.
[0052] Specifically, collect key smelting process data from the tin smelting plant database, including oxygen flow rate, lance position, coal flow rate, furnace pressure, smelting temperature, etc., to ensure the comprehensiveness and timeliness of the data and provide accurate basic data for subsequent model training.
[0053] Furthermore, to verify the real-time dynamic optimization effect of the present invention on simulating tin smelting process data and the chemical reaction in the furnace during tin smelting, this example is based on the actual operation data of the Ausmel furnace in a certain tin smelting plant.
[0054] Step2: Preprocess the collected original parameter data, clean outliers and fill in missing values to ensure data integrity and avoid data quality problems affecting subsequent model training and prediction. Then, use the improved Fourier synthesis method to simulate and generate the preprocessed parameter data to obtain simulated data.
[0055] Step2.1: Set the parameter measurement range and the threshold of the actual operation law according to the sensor measurement range and in combination with the actual operation law of the metallurgical process, and eliminate outliers that exceed the parameter measurement range threshold and the parameter change speed exceeds the actual operation law threshold;
[0056] Step2.2: Add an adaptive bottom feature detection algorithm to the original Fourier synthesis method to identify the steady-state interval of parameter data. The formula is as follows:
[0057]
[0058] In the formula, is the window variance ratio, is the data variance within the sliding window centered at time t,
[0059] is the global data variance. In this way, the improved Fourier synthesis method can better simulate the extreme working conditions in the original data.
[0060] Step2.3: Use the improved Fourier synthesis algorithm to simulate and generate tin smelting parameter data to expand the scale of the reinforcement learning training set.
[0061] Specifically, perform improved Fourier synthesis on the preprocessed tin smelting parameters. First, use the sliding window Fourier transform to decompose time series data such as temperature and pressure, screen the main frequency components through the dynamic threshold formula, and retain the original phase information; then, based on the window variance ratio, perform outlier replacement and missing value filling. For the steady-state interval, use the mean-preserving method to generate data, and for the transition interval, compensate through phase extrapolation to obtain simulation data that retains the true metallurgical process law. As Figure 3 and Figure 4 The experimental results show that the improved Fourier synthesis method can simulate and generate more realistic tin smelting parameter data.
[0062] Step3: Build a reinforcement learning model, define the state space, action space, and design the reward function.
[0063] Step3.1: Design the parameter values that affect the chemical reactions in the tin smelting furnace (such as oxygen flow rate, furnace pressure, bottom furnace temperature, oxygen concentration, etc.) as the state space, and define the controllable optimization variables during the smelting process (such as oxygen flow rate adjustment, coal flow rate adjustment, lance position adjustment, etc.) as the action space;
[0064] Step3.2: The reward function, as the optimization goal of reinforcement learning, is designed according to the three aspects of maximizing tin production, minimizing carbon monoxide emissions, and minimizing reaction time;
[0065] Step3.3: Use the data and metallurgy dual-drive method as the state update model of reinforcement learning.
[0066] Specifically, organize the industrial data of the tin smelting process into a triple form , In the formula, represents the state at the current moment, Represents the operation action executed at the current moment, represents the state at the next moment;
[0067] Design an LSTM network, use the current state and action as input features, and the state at the next moment as the target variable for training;
[0068] Integrate metallurgical knowledge, and use the parameter values calculated by chemical reactions to limit the parameter values predicted by the LSTM network to achieve dynamic optimization of state updates.
[0069] Furthermore, during the tin smelting process, the main chemical reaction is:
[0070] SnO + C → Sn + CO
[0071] In the formula, SnO represents tin monoxide. Subsequently, by calculating the thermodynamic equilibrium constant K, the direction of the reaction and the ratio of reactants and products can be predicted. The calculation formula for the thermodynamic equilibrium constant K is:
[0072]
[0073] In the formula, is the standard free energy change (related to the reaction temperature), K represents the thermodynamic equilibrium constant, R = 8.314 J / (mol⋅K) is the gas constant, and T is the temperature in the furnace. The calculated K is a quantity that describes the relationship between reactants and products when a chemical reaction reaches equilibrium at a specific temperature and pressure. In the following formulas, K always represents the thermodynamic equilibrium constant. Since the reaction occurs in the same furnace, the calculated K is the same value.
[0074] By combining data-driven (LSTM model prediction) and metallurgical knowledge (thermodynamic equilibrium and chemical reaction principles), dynamic optimization of state updates is achieved, providing more reliable environmental feedback for reinforcement learning.
[0075] Step4: Design a dual-network architecture for the Actor-Critic network in the reinforcement learning model, introducing a main network and an auxiliary network.
[0076] Step4.1: Introduce two neural networks, the main network and the auxiliary network, into the Actor network. The main network is responsible for generating policies and outputting optimized control actions in the current state; the auxiliary network provides action evaluation signals to the main network by interacting with the Critic network to guide the main network to generate actions;
[0077] Step 4.2: Dual-network design of the Critic. Introduce two neural networks in the Critic: the main network and the auxiliary network. The main network directly provides the value evaluation of the current action, which is used to guide the Actor network to optimize the strategy. The auxiliary network provides a stable target value through a soft update mechanism to avoid overestimation or oscillation of the Q value. This dual-network architecture separates action generation from action evaluation by introducing an auxiliary network in both the Actor and the Critic, strengthening the task targeting of the network and improving the learning efficiency and stability of the model.
[0078] Step 5: Based on the tin smelting process parameters, establish a physical and chemical model, and calculate the carbon monoxide generation amount by combining the thermodynamic equilibrium formula and the chemical reaction formula in the furnace.
[0079] Step 5.1: According to the actual operating conditions of the tin smelting process, determine the key process parameters related to the carbon monoxide generation amount and the tin generation amount. The key process parameters include the smelting temperature, the furnace chamber pressure, the oxygen concentration, the addition amount of tin oxide, and the addition amount of carbon.
[0080] Step 5.2: Establish a thermodynamic equilibrium model according to the chemical reaction equation of the tin smelting process, which is used to calculate the carbon monoxide generation amount and the tin generation amount.
[0081] Specifically, the tin generation amount is directly proportional to the addition amount of the reactant SnO. The tin generation amount generated at equilibrium can be calculated in combination with the thermodynamic equilibrium constant K. The thermodynamic equilibrium constant K can be defined as:
[0082]
[0083] The equilibrium concentration of carbon monoxide can be calculated by the following formula:
[0084]
[0085]
[0086] Because the stoichiometric ratio is 1:1, the tin generation amount is equal to the carbon monoxide generation amount. The ratio of carbon monoxide to O2 is 2:1, so the tin generation amount = .
[0087] Step 6: Input the calculated carbon monoxide generation amount as the reward function of the optimized reinforcement learning model to guide the agent in the reinforcement learning to select and execute actions.
[0088] Step 6.1: Input the calculated tin generation amount and carbon monoxide generation amount as the input of the reward function. The reward function is:
[0089]
[0090] In the formula, is the reward function, , , are weight parameters used to balance the importance of different optimization objectives. Sn t represents the amount of tin generated in the furnace at the current moment, and CO t represents the amount of carbon monoxide generated in the furnace at the current moment. Time is the reaction time;
[0091] Step6.2: Collect the real-time status data of the furnace. The optimal strategy model recommends the optimal actions based on the real-time status data of the furnace to ensure the optimality and stability of the operation.
[0092] Specifically, after constructing the reinforcement learning model, start training the model through interaction with the Ausmel furnace simulation environment. The reinforcement learning agent continuously explores and exploits in this environment, and continuously adjusts the control strategy according to the feedback of the reward function to optimize objectives such as the amount of tin generated, carbon monoxide emissions, and reaction time. Through multiple rounds of training, the model gradually converges and can automatically adjust the operation parameters under different working conditions to achieve the optimization of the full-automatic and real-time tin smelting process. Among them, the operation process of the reinforcement learning is as Figure 2 shown.
[0093] Step7: Use the generated simulation data and the collected parameter data to train the optimized reinforcement learning model and save the parameters of the final reinforcement learning model.
[0094] Specifically, based on the trained reinforcement learning model, verify it in an actual tin smelting factory. By comparing the process data before and after optimization, evaluate the improvement degree of the amount of tin generated, the reduction amplitude of carbon monoxide emissions, and the shortening effect of the reaction time. The comparative test based on the actual production data shows that after applying the method of the present invention, the tin content in the discharged waste residue is reduced by 0.5 percentage points, the average value of carbon monoxide emissions in the discharged waste gas is reduced to 598 ppm (original 1600 ppm), and the reaction equilibrium time is shortened to 108 ± 15 seconds (traditional control 142 ± 18 seconds).
[0095] The specific implementation manners of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above implementation manners, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. An optimized control method for the chemical reaction of tin smelting based on reinforcement learning, characterized in that, The method specifically comprises: Step 1: Collect parameter data during the smelting process; Step 2: Preprocess the collected original parameter data, clean outliers and fill missing values, and use the improved Fourier synthesis method to simulate and generate the preprocessed parameter data to obtain simulated data; Step 3: Build a reinforcement learning model, define the state space, action space, and design the reward function; Step 4: Design a dual network architecture for the Actor-Critic network in the reinforcement learning model, introducing the main network and the auxiliary network; Step 5: Based on the tin smelting process parameters, a physical and chemical model is established, and the amount of carbon monoxide generated is calculated by combining the thermodynamic equilibrium formula and the chemical reaction formula in the furnace; Step 6: Use the calculated carbon monoxide generation as the reward function input of the optimized reinforcement learning model to guide the agent in the reinforcement learning to select the execution action; Step 7: Use the generated simulation data and the collected parameter data to train the optimized reinforcement learning model and save the final reinforcement learning model parameters.
2. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 1, characterized in that, The Step 2 is specifically as follows: Step 2.1: According to the sensor measurement range and the actual operation law of the metallurgical process, the parameter measurement range and the actual operation law threshold are set, and the abnormal values exceeding the parameter measurement range threshold and the parameter change speed exceeding the actual operation law threshold are eliminated; Step 2.2: Add an adaptive bottom feature detection algorithm based on the original Fourier synthesis method to identify the steady-state interval of parameter data; Step 2.3: Use the improved Fourier synthesis algorithm to simulate and generate tin smelting parameter data to expand the scale of reinforcement learning training set.
3. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 1, characterized in that, The Step 3 is specifically as follows: Step 3.1: Design the parameter values that affect the chemical reaction in the tin smelting furnace as the state space, and define the controllable optimization variables in the smelting process as the action space; Step 3.2: The reward function is used as the optimization goal of reinforcement learning, and is designed based on the three goals of maximizing tin production, minimizing carbon monoxide emissions, and minimizing reaction time; Step 3.3: Use the data and metallurgy dual-driven method as the state update model of reinforcement learning.
4. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 1, characterized in that, The Step 4 is specifically as follows: Step 4.1: Two neural networks, the main network and the auxiliary network, are introduced into the Actor network. The main network is responsible for generating strategies and outputting the optimal control actions under the current state; the auxiliary network interacts with the Critic network to provide action evaluation signals to the main network and guide the main network to generate actions; Step 4.2: Introduce two neural networks, the main network and the auxiliary network, into the Critic network. The main network directly provides a value assessment of the current action to guide the Actor network optimization strategy. The auxiliary network provides a certain target value through a soft update mechanism.
5. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 1, characterized in that, The Step 5 is specifically as follows: Step 5.1: Determine the key process parameters related to carbon monoxide generation and tin generation according to the actual operating conditions of the tin smelting process; Step 5.2: Based on the chemical reaction equation of the tin smelting process, establish a thermodynamic equilibrium model to calculate the amount of carbon monoxide generated and the amount of tin generated.
6. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 1, characterized in that, The specific content of Step 6 is as follows: Step 6.1: Use the calculated amount of tin generated and the amount of carbon monoxide generated as the input of the reward function, and the reward function is: ; In the formula, is the reward function, , , are weight parameters used to balance the importance of different optimization objectives, and Sn t represents the amount of tin generated in the furnace at the current moment, and CO t represents the amount of carbon monoxide generated in the furnace at the current moment, and Time is the reaction time; Step 6.2: Collect the real-time status data in the furnace, and the optimal strategy model recommends the optimal action based on the real-time status data in the furnace.
7. The optimized control method for the chemical reaction of tin smelting based on reinforcement learning according to claim 2, characterized in that, The specific content of Step 2.2 is as follows: ; In the formula, is the window variance ratio, is the data variance within the sliding window centered at time t, is the global data variance.
8. The optimized control method for the tin smelting chemical reaction based on reinforcement learning according to claim 3, wherein, The specific content of Step 3.3 is as follows: Organize the industrial data of the tin smelting process into triple form , where represents the state at the current moment, represents the operation action executed at the current moment, represents the state at the next moment; Design an LSTM network, use the current state and action as input features, and the state at the next moment as the target variable for training; Integrate metallurgical knowledge, use the parameter values calculated by chemical reactions to limit the parameter values predicted by the LSTM network, and achieve dynamic optimization of state updates.
Citation Information
Patent Citations
Layout design system, and layout design method
CN110998585A
Furnace control system, furnace control method, and furnace provided with same
CN112304106A
Data center energy consumption joint optimization method, system, medium and equipment
CN112966431A
Optimization parameter generation method and device
CN116485168A
SMT production line key process parameter optimization method based on man-machine fusion and storage medium
CN117312802A