Dynamic optimization control method for tin smelting process
Through deep reinforcement learning combined with dynamic optimization control methods of data-driven models, the problem that traditional tin smelting control methods are difficult to optimize dynamic changes in the furnace in real time is solved, and the effects of improving tin purity, reducing CO emissions and optimizing energy consumption are achieved.
Patent Information
- Application Number
- CN202510677737.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Traditional tin smelting control methods are difficult to optimize the dynamic changes and nonlinear complex relationships in the furnace in real time, resulting in problems such as high energy consumption, large emissions, and fluctuations in product quality.
Using a dynamic optimization control method with deep reinforcement learning combined with data-driven model, by constructing a reinforcement learning environment, the agent can independently learn and dynamically optimize the gun operating parameters, perceive process conditions changes in real time, and design reward functions to improve tin purity and reduce CO emissions.
Dynamic optimization control of the tin smelting process is realized, tin purity is improved, CO emissions is reduced, energy consumption is optimized, and process efficiency and stability are improved.
Smart Images

Figure CN120215281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dynamic optimization control method for the tin smelting process, belonging to the cross - technical field of metallurgical engineering and artificial intelligence. Background Art
[0002] Tin smelting, as an important process in non - ferrous metal smelting, has complex physical and chemical reactions and multi - variable coupling characteristics. Traditional tin smelting control methods usually rely on empirical rules and static optimization strategies, making it difficult to cope with the dynamic changes in the furnace and non - linear complex relationships. This method faces many challenges in practical applications, such as the difficulty in accurately predicting the dynamic changes of key indicators in the smelting process (such as tin purity, CO emission concentration, etc.), and the difficulty in real - time optimizing control strategies, resulting in problems such as high energy consumption, large emissions, and fluctuating product quality.
[0003] In recent years, with the rapid development of intelligent manufacturing technology, deep learning and reinforcement learning have provided new solutions for the intelligent optimization of the tin smelting process. Among them, the excellent performance of recurrent neural networks and their improved models (such as long short - term memory networks) in time - series modeling makes it possible to accurately predict key parameters in the tin smelting process; while deep reinforcement learning technology provides a theoretical basis for the construction of dynamic optimization control strategies. However, relying solely on data - driven machine learning models is likely to ignore physical constraints such as stoichiometry and energy balance in the metallurgical process, which may lead to insufficient reliability of prediction and optimization results.
[0004] Therefore, researching a dynamic optimization control method that integrates data - driven and metallurgical knowledge can fully exploit the data potential in the tin smelting process, while embedding domain knowledge to enhance the interpretability and accuracy of the model. By constructing a reinforcement learning environment, taking the lance operation parameters as the action space, the key states of the tin smelting process as the state space, and designing a reward function that can sense changes in process conditions in real - time, the intelligent agent can autonomously learn and dynamically optimize control strategies, thereby achieving the improvement of tin purity, the reduction of CO emissions, and the minimization of energy consumption. This technical background lays the theoretical and practical foundation for the intelligent optimization of the tin smelting process. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a dynamic optimization control method for the tin smelting process, which can optimize the operation parameters in the smelting process in real - time, thereby solving the above problems.
[0006] The technical solution of the present invention is: a dynamic optimization control method for tin smelting process, first, collect relevant data from the tin smelting plant database, including state variables and spray gun operating parameters at each moment in the smelting process. Secondly, preprocess the data, remove the abnormal values recorded by the sensor, and use the conditional generative adversarial network (Conditional GAN, cGAN) to generate key variables for low-frequency measurement to make up for data gaps. Subsequently, the long short-term memory network (Long Short-Term Memory, LSTM) model is used to model the key state variables (such as tin purity, CO concentration, etc.) in the tin smelting process to capture the time series change law, and at the same time embed metallurgical knowledge (such as oxygen-coal stoichiometric ratio, oxygen flow rate change, etc.) as constraints to improve the accuracy and reliability of model prediction. In the construction of the reinforcement learning environment, a data-driven state update model is adopted, and the key state variables and spray gun operating parameters are used as inputs. The state at the next moment is predicted by the deep learning model, and the credibility of the state transfer is improved by combining metallurgical knowledge. Finally, a deep reinforcement learning algorithm is introduced, with the spray gun operation parameters as the action space and the tin smelting state as the state space. The reward function is designed to reflect the optimization goals of improving tin purity and reducing CO emissions. The intelligent agent learns the optimal strategy, which can guide the intelligent agent to generate the best spray gun operation suggestions under different environmental conditions.
[0007] The specific steps are:
[0008] Step 1: Collect relevant parameter data during the tin smelting process, including the state variable values and spray gun operation parameter values at each moment in the smelting furnace, to provide a comprehensive data basis for subsequent analysis and modeling;
[0009] Step 2: Preprocess the relevant parameter data, delete the abnormal values recorded by the sensor, and use cGAN to generate variables with a measurement frequency lower than a preset threshold, so as to fill the data gaps and improve data integrity;
[0010] Step 3: Use the improved LSTM model to model the tin purity and CO concentration in the tin smelting process, capture the time series change rules of state variable values and spray gun operation parameter values, and improve the accuracy and reliability of prediction;
[0011] Step 4: Build a reinforcement learning environment, adopt a data-driven state update model, take the state variable values and spray gun operation parameter values in the tin smelting process as input, use the deep learning model to predict the state value at the next moment, and enhance the reliability of state transfer and the actual performance of the environment model by integrating metallurgical knowledge constraints;
[0012] Step 5: Design the reward function. Based on the captured time series variation patterns, guide the agents of reinforcement learning to learn in the direction of improving tin purity and reducing CO emissions. The reward function provides real-time feedback on the decision-making effects of the agents, offering effective guidance for the optimization process.
[0013] Step 6: Introduce the Deep Reinforcement Learning (DRL) algorithm. Design the spray gun operation parameters as the action space and the states in the tin smelting process as the state space. The agents interact and learn with the state update model, continuously optimizing the control strategy to achieve the dynamic optimization goal. This strategy can guide the agents to generate the best spray gun operation suggestions under different environmental states.
[0014] The specific content of Step 2 is as follows:
[0015] Step 2.1: According to the measurement range of the sensors and the experience of process experts, set the threshold values of each variable, and directly eliminate the outliers beyond the thresholds to ensure the accuracy and reliability of the subsequent modeling data.
[0016] Step 2.2: Conduct data partitioning. Use the complete high-frequency measurement variables as the conditional inputs of the model and the low-frequency measurement variables as the target generation variables. Generate missing values using the cGAN model. Among them, the high-frequency measurement variables are used as conditional inputs into the generator and discriminator to provide additional constraint information, thereby improving data integrity and consistency.
[0017] The specific content of Step 3 is as follows:
[0018] Step 3.1: Based on the existing industrial data in the tin smelting process, extract the core variables reflecting the dynamic changes in the furnace, and generate new feature columns related to the smelting process through mathematical transformation. These features not only enrich the data dimension but also provide more meaningful input information for subsequent modeling.
[0019] Step 3.2: Embed the physical constraints in the metallurgical field into the loss function of the original LSTM model, and guide the LSTM model to learn prediction results that conform to the actual process laws through the constraint penalty term, improving the scientificity and interpretability of the prediction results.
[0020] Step 3.3: Use the improved LSTM model to predict and train the index variables in the tin smelting process, capture complex time-dependent relationships through the gating mechanism, and accurately model the dynamic changes between variables.
[0021] Step 3.4: After the training is completed, use an independent test data set to verify the prediction performance of the model, evaluate the prediction accuracy of tin purity and CO concentration, and save the LSTM model parameters to ensure its applicability in the industrial scenario.
[0022] Step 4 is specifically as follows:
[0023] Step 4.1: Obtain the state changes of consecutive time steps from time series data, including state variables and action variables, and organize the data into a triple form:
[0024]
[0025] In the formula, represents the state variable at the current moment, represents the action variable at the current moment, represents the state variable at the next moment, and this data format lays the foundation for the training of the subsequent state update model;
[0026] Step 4.2: Construct a state update model of a fully connected neural network, take the current state and the current action as inputs, predict the state at the next moment, and use the constructed triple data to train the state update model;
[0027] Step 4.3: Evaluate the performance of the state update model on the validation set, save the state update model parameters that meet the preset index requirements, and provide a reliable state update mechanism for subsequent embedding in the reinforcement learning environment;
[0028] Step 4.4: In the reinforcement learning environment, use the trained state update model parameters, predict the state of the reinforcement learning environment at the next moment through the current environment and the currently executed action, guide the agent to make action selections. This embedding process enhances the ability of the environment to describe complex dynamic changes, provides more real - time state feedback for the agent, guides it to optimize action selections, and realizes the efficient control of the tin smelting process.
[0029] Step 5 is specifically as follows:
[0030] Step 5.1: Define the optimization objectives: improve the tin purity and reduce CO emissions. At the same time, set auxiliary objectives, including reducing energy consumption and maintaining the stability of key process parameters such as furnace pressure, so as to achieve multi - objective collaborative optimization and ensure the efficient operation of the tin smelting process;
[0031] Step 5.2: Use the saved LSTM model parameters to predict the impact of the adjustment of current process parameters on the tin purity and CO emissions in a preset future time period, and use the prediction results as the basis for calculating the reward function to accurately reflect the contribution of the optimization strategy to the long - term objectives;
[0032] Step 5.3: Design the reward function, and the formula is:
[0033] ( )
[0034] In the formula, R is the reward function, 、 、 are weight coefficients used to balance the importance of multiple objectives, is the tin purity content at the current moment, is the carbon monoxide content at the current moment, is the tin purity content at the next moment, is the carbon monoxide content at the next moment, is the energy consumption value at the next moment;
[0035] To enhance real-time performance and robustness, the reward function combines the prediction results of key parameters based on the LSTM model during design, decomposes the long-term optimization goal of the tin smelting process into several short-term optimization sub-goals, more precisely reflects the impact of current process adjustments on future states, and thus effectively guides the agent to learn and continuously optimize the control strategy.
[0036] Specifically, Step6 is as follows:
[0037] Step6.1: Regulate the lance operation parameters based on the DDPG algorithm. Define the operations performed by the lance in the furnace as the action space, and define ten core variables in the furnace: total concentration, bath accumulation, oxygen content percentage, temperature in the middle of the furnace bottom, external temperature of the furnace bottom, internal temperature of the furnace bottom, temperature increase of the furnace, furnace pressure, waste gas CO analysis, and total energy consumption as the state space;
[0038] Step6.2: Use the Actor network in the DDPG algorithm to predict the optimal action based on the current state for adjusting the lance operation parameters, and use the Critic network to evaluate the Q value of the current policy for guiding the optimization of the Actor;
[0039] Step6.3: Store the interaction data between the agent in reinforcement learning and the furnace environment in the experience pool, randomly sample several furnace states and lance operation parameter data from it for training, break the time correlation, enable the reinforcement learning agent to learn operation experiences not available in the previous continuous furnace period data, and improve the training stability and generalization ability of the model;
[0040] Step6.4: Enhance the exploration ability of the agent for the action space by adding noise, avoid getting stuck in local optima, and at the same time simulate the situation where the factory encounters failures to complete the model training;
[0041] Step 6.5: Utilize the trained model to achieve intelligent dynamic optimization control during the tin smelting process. By real-time monitoring of state variables, combining with the optimization results of the model, provide optimization suggestions to the lance operator, dynamically adjust operating parameters, thereby improving tin purity, reducing CO emissions, and maximizing process energy efficiency.
[0042] Specifically, the said Step 3.2 is as follows:
[0043] Add metallurgical knowledge constraints to the loss function of the LSTM model, and the constraint formula is:
[0044]
[0045] Wherein, is the total loss value, is the predicted loss value of the LSTM model, is a constant, The function needs to combine metallurgical knowledge, and the specific formula is as follows:
[0046]
[0047] Wherein, is a constant, is the difference between the predicted value and the theoretical value, and the theoretical value of CO is calculated as follows:
[0048]
[0049] In the formula, represents the combustion coal flow rate, is a constant related to the calorific value and combustion efficiency of the fuel coal, is the oxygen flow rate, is the theoretical stoichiometric ratio of the combustion coal.
[0050] The beneficial effects of the present invention are:
[0051] (1) Dynamic optimization control ability: Compared with the existing static optimization methods, the present invention realizes the dynamic optimization control of the tin smelting process by combining the deep reinforcement learning algorithm and the data-driven model. The intelligent agent can real-time monitor the state and autonomously optimize the lance operation strategy to cope with complex process changes;
[0052] (2) Precise modeling integrating metallurgical knowledge: Different from the traditional data-driven model, the present invention introduces metallurgical physical constraints such as the oxygen-coal stoichiometric ratio, and combines with the improved recurrent neural network (RNN) to accurately capture the dynamic change rules of key parameters, improving the prediction accuracy and reliability;
[0053] (3)Efficient environment construction: The present invention adopts a data-driven state update model to accurately reflect the dynamic characteristics of tin smelting, providing a reliable state transition mechanism for reinforcement learning, avoiding the invalidation of the optimization strategy caused by the accumulation of environmental errors, and thus improving the learning efficiency and optimization performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is the overall framework diagram of the present invention;
[0055] Figure 2 is the comparison diagram of the predicted value and the true value of the total concentration in the embodiment of the present invention; Figure 3 is the comparison diagram of the predicted value and the true value of the molten pool accumulation in the embodiment of the present invention; Figure 4 is the comparison diagram of the predicted value and the true value of the oxygen content percentage in the embodiment of the present invention; Figure 5 is the comparison diagram of the predicted value and the true value of the temperature in the middle of the furnace bottom in the embodiment of the present invention; Figure 6 is the comparison diagram of the predicted value and the true value of the external temperature of the furnace bottom in the embodiment of the present invention; Figure 7 is the comparison diagram of the predicted value and the true value of the internal temperature of the furnace bottom in the embodiment of the present invention; Figure 8 is the comparison diagram of the predicted value and the true value of the temperature increase of the furnace in the embodiment of the present invention; Figure 9 is the comparison diagram of the predicted value and the true value of the furnace pressure in the embodiment of the present invention; Figure 10 is the comparison diagram of the predicted value and the true value of the CO content in the waste gas in the embodiment of the present invention; Figure 11 is the comparison diagram of the predicted value and the true value of the total energy consumption in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0056] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0057] Example 1: As Figure 1 shown, a dynamic optimization control method for the tin smelting process, the specific steps are as follows:
[0058] Step1: Collect relevant parameter data in the tin smelting process. The relevant parameter data includes the state variable values and lance operation parameter values at each moment in the smelting furnace, providing a comprehensive data basis for subsequent analysis and modeling.
[0059] Specifically, to verify the real-time dynamic optimization effect of the model of the present invention in the tin smelting process, in this embodiment, based on the actual operation data of a tin smelting factory, the interaction between the simulation reinforcement learning model and the industrial production environment is carried out to evaluate the effects of improving tin purity, reducing CO emissions, and optimizing energy consumption. The smelting data for 30 consecutive days is extracted from the data system of the Ausmel furnace in a tin smelting factory, and the time resolution is 1 minute. Data scale: a total of 43,200 records, including 10 state variables and 8 operating parameters. Specifically, as shown in Table 1 and Table 2.
[0060] Table 1 Action variables executed by the lance in the tin factory data
[0061] Action variable name Unit Total material flow rate t / h Fuel coal flow rate kg / h Coal-carrying air stream Nm³ / h Coal-carrying air pressure kPa Back pressure of the lance kPa Lance position mm Oxygen flow rate Nm³ / h Air flow rate Nm³ / h
[0062] Table 2 State variables of the Ausmel furnace in the tin factory data
[0063] State variable name Unit Total concentration t Accumulation in the bath t Oxygen concentration % Temperature at the middle of the furnace bottom Temperature Temperature outside the furnace bottom Temperature Temperature inside the furnace bottom Temperature Furnace heating-up temperature Temperature Furnace internal pressure Pa CO content in the waste gas Content per million Energy consumption kW
[0064] Step2: Preprocess the relevant parameter data, delete the outliers recorded by the sensors, and use cGAN to generate variables with a measurement frequency lower than the preset threshold, so as to fill the data gaps and improve data integrity.
[0065] Step2.1: According to the measurement range of the sensors and the experience of process experts, set the threshold values of each variable, and directly eliminate the outliers exceeding the threshold to ensure the accuracy and reliability of the subsequent modeling data;
[0066] Step2.2: Perform data partitioning. Take the complete high-frequency measurement variables (such as oxygen flow rate, furnace pressure, etc.) as the conditional inputs of the model, and the low-frequency measurement variables (such as tin purity, etc.) as the target generation variables, and use the cGAN model to generate missing values. Among them, the high-frequency measurement variables are used as conditional inputs to the generator and discriminator to provide additional constraint information, thereby improving data integrity and consistency.
[0067] Step3: Use an improved LSTM model to model the tin purity and CO concentration in the tin smelting process, capture the time series change rules of the state variable values and the lance operation parameter values, and improve the accuracy and reliability of the prediction.
[0068] Step3.1: Based on the existing industrial data in the tin smelting process, extract the core variables (such as oxygen flow rate, fuel coal flow rate, furnace pressure, etc.) that reflect the dynamic changes in the furnace, and generate new feature columns related to the smelting process through mathematical transformation. These features not only enrich the data dimension but also provide more meaningful input information for subsequent modeling;
[0069] Specifically, the new feature columns are specifically: the change in oxygen flow rate, the change in the flow rate of combusted coal, and the stoichiometric ratio of oxygen to coal. Among them, the change in oxygen flow rate and the change in the flow rate of combusted coal help the model capture the temporal dynamics and non-linear relationships of the combustion reaction in the furnace, and the stoichiometric ratio of oxygen to coal reflects the supply-demand relationship between oxygen and fuel. The specific calculation formula is:
[0070]
[0071] In the formula, represents the change in oxygen flow rate, represents the oxygen flow rate at the current time step, represents the oxygen flow rate at the previous time step;
[0072]
[0073] In the formula, represents the change in the flow rate of combusted coal, represents the flow rate of combusted coal at the current time step, represents the flow rate of combusted coal at the previous time step;
[0074]
[0075] In the formula, represents the stoichiometric ratio of oxygen to coal, represents the oxygen flow rate, represents the oxygen purity, represents the density of oxygen, represents the flow rate of combusted coal, represents the purity of combusted coal.
[0076] Step3.2: Embed the physical constraints in the metallurgical field into the loss function of the original LSTM model, and guide the LSTM model to learn the prediction results that conform to the actual process laws through constraint penalty terms (such as the deviation of the stoichiometric ratio of oxygen to coal), so as to improve the scientificity and interpretability of the prediction results;
[0077] Specifically, add metallurgical knowledge constraints to the loss function of the LSTM model, and the constraint formula is:
[0078]
[0079] Among them, is the total loss value, is the prediction loss value of the LSTM model, is a constant, The function needs to combine metallurgical knowledge, and the specific formula is as follows:
[0080]
[0081] Among them, is a constant, is the difference between the predicted value and the theoretical value. The theoretical value of CO is calculated as follows:
[0082]
[0083] In the formula, represents the flow rate of the combusted coal, is a constant related to the calorific value of the fuel coal and the combustion efficiency, is the oxygen flow rate, is the theoretical stoichiometric ratio of the combusted coal.
[0084] Step3.3: Use the improved LSTM model to predict and train the index variables (such as tin purity, CO concentration, etc.) in the tin smelting process, capture complex time-dependent relationships through the gating mechanism, and accurately model the dynamic changes between variables;
[0085] Step3.4: After the training is completed, use an independent test data set to verify the prediction performance of the model, evaluate the prediction accuracy of tin purity and CO concentration, and save the LSTM model parameters to ensure its applicability in industrial scenarios.
[0086] Step4: Build a reinforcement learning environment, adopt a data-driven state update model, use the state variable values and lance operation parameter values in the tin smelting process as inputs, use a deep learning model to predict the state value at the next moment, and enhance the reliability of state transition and the actual performance of the environment model by integrating metallurgical knowledge constraints.
[0087] Step4.1: Obtain the state changes at consecutive time steps through time series data, including state variables and action variables, and organize the data into a triple form:
[0088]
[0089] In the formula, represents the state variable at the current moment, represents the action variable at the current moment, represents the state variable at the next moment. This data format lays the foundation for the training of the subsequent state update model;
[0090] Step4.2: Build a state update model of a fully connected neural network, use the current state and the current action as inputs to predict the state at the next moment, and use the constructed triple data to train the state update model;
[0091] Step4.3: Evaluate the performance of the state update model on the validation set, save the state update model parameters that meet the preset metric requirements, and provide a reliable state update mechanism for the subsequent embedded reinforcement learning environment. The state update effect is as follows Figure 2 - Figure 11 shown, which shows the comparison between the predicted values and the true values of the reinforcement learning state update model for ten different states. The ten states are specifically total concentration, accumulated molten pool, oxygen content percentage, temperature in the middle of the furnace bottom, external temperature of the furnace bottom, internal temperature of the furnace bottom, temperature increase of the furnace, furnace pressure, CO content in the waste gas, and total energy consumption. Among them, the predicted value represents the state at the next moment predicted using the state and action at the current moment, and the true value represents the actual state at the next moment;
[0092] Step4.4: In the reinforcement learning environment, use the trained state update model parameters to predict the state of the reinforcement learning environment at the next moment through the current environment and the currently executed action, and guide the agent to make action selections. This embedding process enhances the environment's ability to describe complex dynamic changes, provides more real state feedback for the agent, guides it to optimize action selections, and realizes the efficient control of the tin smelting process.
[0093] Step5: Design a reward function to guide the reinforcement learning agent to learn in the direction of improving tin purity and reducing CO emissions based on the captured time series change rules. The reward function provides real-time feedback on the decision-making effect of the agent and provides effective guidance for the optimization process.
[0094] Step5.1: Set optimization goals: improve tin purity and reduce CO emissions. At the same time, set auxiliary goals, including reducing energy consumption and maintaining the stability of key process parameters such as furnace pressure, so as to achieve multi-objective collaborative optimization and ensure the efficient operation of the tin smelting process;
[0095] Step5.2: Use the saved LSTM model parameters to predict the impact of the adjustment of current process parameters on tin purity and CO emissions in the next 3 minutes, and use the prediction results as the basis for calculating the reward function to accurately reflect the contribution of the optimization strategy to the long-term goal;
[0096] Step5.3: Design a reward function, and the formula is:
[0097] ( )
[0098] In the formula, R is the reward function, 、 、 are weight coefficients used to balance the importance of multiple goals, is the tin purity content at the current moment, is the carbon monoxide content at the current moment, is the tin purity content at the next moment, is the carbon monoxide content at the next moment, is the energy consumption value at the next moment.
[0099] Furthermore, when the prediction result of the LSTM model shows that the adjustment of the current operating parameters can effectively improve the tin purity and reduce the CO concentration, the reward function will give a positive incentive; on the contrary, if the operation causes the CO concentration to increase or the tin purity to decrease, the reward function will give a negative feedback.
[0100] Step6: Introduce the Deep Reinforcement Learning (DRL) algorithm. Design the spray gun operation parameters as the action space and the states in the tin smelting process as the state space. The agent interacts and learns with the state update model, continuously optimizing the control strategy to achieve the dynamic optimization goal. This strategy can guide the agent to generate the best spray gun operation suggestions in different environmental states.
[0101] Step6.1: Based on the DDPG algorithm, regulate the spray gun operation parameters. Define the operations performed by the spray gun in the furnace (such as oxygen flow rate, coal-carrying wind pressure, etc.) as the action space, and define ten core variables in the furnace: total concentration, molten pool accumulation, oxygen content percentage, temperature in the middle of the furnace bottom, temperature outside the furnace bottom, internal temperature of the furnace bottom, increased furnace temperature, furnace pressure, waste gas CO analysis, and total energy consumption as the state space;
[0102] Step6.2: Use the Actor network in the DDPG algorithm to predict the optimal action based on the current state for adjusting the spray gun operation parameters, and use the Critic network to evaluate the Q value of the current policy for guiding the optimization of the Actor;
[0103] Step6.3: Store the interaction data between the agent in the reinforcement learning and the furnace environment in the experience pool, randomly sample several furnace states and spray gun operation parameter data from it for training, break the time correlation, enable the reinforcement learning agent to learn operation experiences not available in the previous continuous furnace period data, and improve the training stability and generalization ability of the model;
[0104] Step6.4: By adding noise, enhance the exploration ability of the agent for the action space, avoid getting stuck in local optima, and at the same time simulate the situation where the factory encounters failures to complete the model training;
[0105] Step6.5: Use the trained model to achieve intelligent dynamic optimization control in the tin smelting process. By real-time monitoring the state variables, combined with the model optimization results, provide optimization suggestions to the spray gun operator, dynamically adjust the operation parameters, thereby improving the tin purity, reducing the CO emissions, and achieving the maximum process energy efficiency.
[0106] Step 7: Verify the performance of the reinforcement learning model on the test data set and evaluate the effects of improved tin purity, reduced CO emissions, and optimized energy consumption. The experimental results are as follows. Current furnace state: Total concentration: 5.777605 tons, molten pool accumulation: 5.4335 tons, oxygen concentration: 1.2%, temperature in the middle of the furnace bottom: 594.0 temperature units, external temperature of the furnace bottom: 518.0 temperature units, internal temperature of the furnace bottom: 518.0 temperature units, furnace heating temperature: 740.0 temperature units, furnace pressure: 3.490 Pa, CO content in the waste gas: 2.000 ppm (content per million), energy consumption: 1.780 kW. Actions recommended by the algorithm: Total material flow rate: 86.5 tons / hour, fuel coal flow rate: 9960 kg / hour, coal-carrying air flow: 0.106 standard cubic meters / hour, coal-carrying air pressure: 32.5 kPa, back pressure of the lance: 410 kPa, lance position: 7031 mm, oxygen flow rate: 17832 standard cubic meters / hour, air flow rate: 17832 standard cubic meters / hour. When the agent executes the lance operation parameters according to the optimal strategy, the average tin purity is increased by 1.5% and the CO emission concentration is reduced by 5%.
[0107] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A dynamic optimization control method for the tin smelting process, characterized in that The method specifically includes: Step1: Collect relevant parameter data during the tin smelting process. The relevant parameter data includes the state variable values and lance operation parameter values at each moment in the smelting furnace. Step2: Preprocess the relevant parameter data, delete the outliers recorded by the sensors, and use cGAN to generate variables with a measurement frequency lower than the preset threshold. Step3: Use an improved LSTM model to model the tin purity and CO concentration during the tin smelting process, and capture the time series change rules of the state variable values and lance operation parameter values. Step4: Construct a reinforcement learning environment, adopt a data-driven state update model, use the state variable values and lance operation parameter values during the tin smelting process as inputs, and use a deep learning model to predict the state value at the next moment. Step5: Design a reward function to guide the agent of reinforcement learning based on the captured time series change rules. Step6: Introduce a deep reinforcement learning algorithm, design the lance operation parameters as the action space, design the state during the tin smelting process as the state space, the agent interacts and learns with the state update model, and continuously optimizes the control strategy to achieve the dynamic optimization goal.
2. The dynamic optimization control method for the tin smelting process according to claim 1, wherein, The specific content of Step2 is as follows: Step2.1: According to the sensor measurement range and the experience of process experts, set the threshold values of each variable, and directly eliminate the outliers that exceed the threshold. Step2.2: Conduct data partitioning. Use the complete high-frequency measurement variables as the conditional inputs of the model, and the low-frequency measurement variables as the target generation variables. Use the cGAN model to generate missing values. Among them, the high-frequency measurement variables are used as conditions and input into the generator and discriminator to provide additional constraint information.
3. A dynamic optimization control method for the tin smelting process according to claim 1, characterized in that The specific content of Step3 is as follows: Step3.1: Based on the existing industrial data during the tin smelting process, extract the core variables reflecting the dynamic changes in the furnace, and generate new feature columns related to the smelting process through mathematical transformation. Step3.2: Embed the physical constraints in the metallurgical field into the loss function of the original LSTM model, and guide the LSTM model to learn the prediction results that conform to the actual process rules through the constraint penalty term. Step3.3: Use an improved LSTM model to predict and train the index variables during the tin smelting process. Step3.4: After the training is completed, use an independent test data set to verify the prediction performance of the model, evaluate the prediction accuracy of tin purity and CO concentration, and save the LSTM model parameters.
4. A dynamic optimization control method for the tin smelting process according to claim 1, characterized in that, The specific content of Step4 is as follows: Step4.1: Obtain the state changes of consecutive time steps through time series data, including state variables and action variables, and organize the data into a triple form: ; In the formula, represents the state variable at the current moment, represents the action variable at the current moment, represents the state variable at the next moment; Step 4.2: Construct a state update model of a fully connected neural network, and use the current state and the current action as inputs to predict the state at the next moment, and use the constructed triple data to train the state update model; Step4.3: Evaluate the performance of the state update model on the validation set, and save the state update model parameters that meet the preset index requirements. Step4.4: In the reinforcement learning environment, use the trained state update model parameters, predict the state of the reinforcement learning environment at the next moment through the current environment and the currently executed action, and guide the agent to select actions.
5. A dynamic optimization control method for the tin smelting process according to claim 1, characterized in that, The specific content of Step5 is as follows: Step5.1: Formulate an optimization goal. Step5.2: Use the saved LSTM model parameters to predict the impact of the adjustment of the current process parameters on the tin purity and CO emissions in a preset future time period, and use the prediction results as the basis for calculating the reward function; Step5.3: Design the reward function, and the formula is: ( ) ; where R is the reward function, , , are weight coefficients used to balance the importance of multiple objectives, is the tin purity content at the current moment, is the carbon monoxide content at the current moment, is the tin purity content at the next moment, is the carbon monoxide content at the next moment, is the energy consumption value at the next moment.
6. The dynamic optimization control method for the tin smelting process according to claim 1, wherein The specific content of Step6 is as follows: Step6.1: Based on the DDPG algorithm, regulate the spray gun operation parameters. Define the operations performed by the spray gun in the furnace as the action space, and define the following ten core variables in the furnace: total concentration, molten pool accumulation, oxygen content percentage, temperature in the middle of the furnace bottom, external temperature of the furnace bottom, internal temperature of the furnace bottom, temperature increase of the furnace, furnace pressure, waste gas CO analysis, and total energy consumption as the state space; Step6.2: Use the Actor network in the DDPG algorithm to predict the optimal action based on the current state for adjusting the spray gun operation parameters, and use the Critic network to evaluate the Q value of the current policy for guiding the optimization of the Actor; Step6.3: Store the interaction data between the agent in reinforcement learning and the furnace environment in the experience pool, randomly sample several furnace states and spray gun operation parameter data from it for training, so that the reinforcement learning agent can learn operation experiences that are not available in the previous continuous furnace period data; Step6.4: Complete the model training by adding noise to simulate the situation where the factory encounters failures; Step6.5: Use the trained model to achieve intelligent dynamic optimization control during the tin smelting process.
7. A dynamic optimization control method for the tin smelting process according to claim 3, characterized in that The specific content of Step3.2 is as follows: Add metallurgical knowledge constraints to the loss function of the LSTM model, and the constraint formula is: ; Among them, is the total loss value, is the predicted loss value of the LSTM model, is a constant, The function needs to combine metallurgical knowledge, and the specific formula is as follows: ; Among them, is a constant, is the difference between the predicted value and the theoretical value. The theoretical value of CO is calculated as follows: ; In the formula, represents the flow rate of the combusted coal, is a constant related to the calorific value and combustion efficiency of the fuel coal, is the oxygen flow rate, is the theoretical stoichiometric ratio of the combusted coal.
Citation Information
Patent Citations
Systems and methods for operations a robotic system and executing robotic interactions
CN112088070A
Furnace control system, furnace control method, and furnace provided with same
CN112304106A
Blast furnace smelting operation optimization method and system based on offline reinforcement learning
CN116562127A
Tin smelting production scheduling optimization method based on graph convolutional network and reinforcement learning
CN118917571A
Automatic driving decision planning method based on deep reinforcement learning and deep learning
CN118966311A
Cited By
Converter steelmaking end point intelligent control method based on artificial intelligence
CN120719081A
Feedback enhancement and working condition guided wet leaching process reinforcement learning control method
CN122239422A
Electrolytic copper key parameter control method and system based on reinforcement learning
CN122466523A