Nonferrous metallurgy intelligent optimization method and system

By fusing multi-source heterogeneous data and hybrid modeling, combined with multi-timescale hierarchical collaborative optimization and cross-process collaborative optimization, the dynamic adaptability and collaborative optimization problems in non-ferrous metal metallurgical production were solved, achieving stability and efficiency improvement in the production process, reducing energy consumption and environmental pollution, and achieving a balance between short-term and long-term benefits.

CN121902017AInactive Publication Date: 2026-04-21GUANGXI MODERN VOCATIONAL & TECH COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGXI MODERN VOCATIONAL & TECH COLLEGE
Filing Date
2025-12-08
Publication Date
2026-04-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing non-ferrous metal metallurgical production suffers from problems such as insufficient dynamic adaptability of raw materials, weak cross-process collaborative optimization capabilities, lack of multi-timescale optimization and collaboration mechanisms, and insufficient uncertainty handling capabilities. These problems lead to unstable production processes, large fluctuations in product quality, excessive use of equipment, and increased maintenance costs.

Method used

By employing methods such as multi-source heterogeneous data fusion, mechanism-driven and data-driven hybrid modeling, multi-timescale hierarchical collaborative optimization, and cross-process collaborative optimization, combined with reinforcement learning and robust optimization, an intelligent optimization system for non-ferrous metals metallurgy is established. This system enables quality assessment and fusion of multi-source heterogeneous data, identification of key variables, causal inference and hybrid prediction, and multi-objective comprehensive benefit assessment and dynamic balance.

Benefits of technology

It has improved the stability and efficiency of non-ferrous metal metallurgical production, reduced energy consumption and environmental pollution, achieved a dynamic balance between short-term and long-term benefits, and enhanced the robustness and automation level of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121902017A_ABST
    Figure CN121902017A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of nonferrous metallurgy, and discloses an intelligent optimization method and system for nonferrous metallurgy, and the method comprises the steps: obtaining multi-source heterogeneous data, and carrying out the quality evaluation and fusion of the multi-source heterogeneous data; performing correlation analysis and causal inference to obtain a key variable set; mechanism and data drive hybrid modeling is carried out to obtain a lightweight hybrid drive prediction model; constructing a Gaussian process regression agent model, and introducing an uncertainty penalty term to obtain an optimal scheme; performing multi-time scale hierarchical collaborative optimization to obtain a hierarchical control scheme; performing cross-process material energy flow collaborative optimization to obtain a cross-process collaborative optimization scheme; performing offline and online hybrid reinforcement learning optimization to obtain an optimization strategy; performing multi-target comprehensive benefit evaluation and dynamic balance to obtain a final intelligent optimization strategy; the method can effectively cope with the dynamic change of raw materials, improve the product quality stability and reduce the energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of non-ferrous metal metallurgy technology, and more specifically, to a smart optimization method and system for non-ferrous metal metallurgy. Background Technology

[0002] Non-ferrous metals metallurgy is a crucial foundational industry of the national economy, with copper, aluminum, lead, zinc, and other non-ferrous metals widely used in power, electronics, construction, and transportation. Traditional non-ferrous metals metallurgical production relies heavily on manual experience and simple feedback control, facing problems such as high energy consumption, large fluctuations in product quality, low production efficiency, and severe environmental pollution. With the introduction of "dual-carbon" targets and the rapid development of intelligent manufacturing, how to utilize next-generation information technologies such as artificial intelligence, big data, and the Internet of Things to achieve intelligent optimization of non-ferrous metals metallurgical processes has become a key issue that the industry needs to address.

[0003] Existing intelligent optimization technologies for metallurgical processes suffer from the following technical problems in practical applications: insufficient dynamic adaptability to raw materials. Non-ferrous metallurgical raw materials are widely available, and the composition of different batches varies significantly. Existing technologies struggle to adapt quickly to frequent changes in raw material composition, leading to production instability and large fluctuations in product quality. Weak cross-process collaborative optimization capabilities. Non-ferrous metallurgical production involves multiple continuous processes with close material and energy flow coupling. Existing technologies primarily focus on local optimization of single processes, lacking a comprehensive systemic optimization method for the entire process, making it difficult to achieve optimal overall efficiency. A lack of multi-timescale optimization collaboration mechanisms. Metallurgical production processes involve decision-making problems at different time scales. This includes long-term production planning, medium-term scheduling optimization, and short-term process control. Existing technologies struggle to effectively coordinate optimization objectives across different time scales, leading to conflicting optimization decisions and poor execution results. Furthermore, there is insufficient uncertainty handling capability. Metallurgical production processes involve numerous uncertainties, including data measurement errors, model prediction biases, and external disturbances. Existing technologies are inadequate in quantifying and handling these uncertainties, resulting in poor robustness of optimization schemes and their tendency to fail in actual implementation. Finally, the long-term and short-term benefit balancing mechanism is imperfect. Existing technologies often excessively pursue the optimization of short-term production indicators, neglecting long-term benefits such as equipment lifespan and maintenance costs, leading to overuse of equipment, increased maintenance costs, and a decline in overall economic efficiency.

[0004] Therefore, there is a need to provide an intelligent optimization method and system for non-ferrous metal metallurgy that can effectively cope with dynamic changes in raw materials, realize cross-process collaborative optimization and multi-timescale collaborative control, improve the robustness of optimization schemes, and achieve a dynamic balance between short-term and long-term benefits. Summary of the Invention

[0005] This invention provides a smart optimization method and system for non-ferrous metal metallurgy, which solves the technical problems of insufficient dynamic adaptability of raw materials, weak cross-process collaborative optimization capability, lack of multi-timescale optimization collaborative mechanism, and insufficient uncertainty handling capability in related technologies.

[0006] This invention provides a smart optimization method for non-ferrous metal metallurgy, comprising the following steps:

[0007] Acquire multi-source heterogeneous data, perform quality assessment and fusion of the multi-source heterogeneous data, establish material conservation constraints and energy conservation constraints, and obtain a unified dataset with credibility weights;

[0008] Based on a unified dataset with confidence weights, correlation analysis and causal inference are performed to obtain a set of key variables;

[0009] Based on the key variable set, a cascade fusion method is used to perform hybrid modeling of mechanism and data-driven approaches, resulting in a lightweight hybrid-driven prediction model.

[0010] A Gaussian process regression surrogate model was trained using a lightweight hybrid-driven prediction model as the evaluator. An uncertainty penalty term was introduced, and a reference-point guided multi-objective optimization algorithm was used for iteration to obtain the optimal solution.

[0011] Based on the optimal solution, a long-term planning layer, a medium-term rolling layer, and a short-term feedback layer are established, and multi-timescale hierarchical collaborative optimization is carried out to obtain a hierarchical control scheme.

[0012] Based on the hierarchical control scheme, the alternating direction multiplier method is used to optimize the material energy flow across processes, resulting in a cross-process collaborative optimization scheme.

[0013] Based on the cross-process collaborative optimization scheme, offline and online hybrid reinforcement learning optimization is performed to obtain the optimization strategy;

[0014] Based on the optimization strategy, a multi-objective comprehensive benefit assessment and dynamic balance are performed to obtain the final intelligent optimization strategy.

[0015] In a preferred embodiment, the quality assessment and fusion of multi-source heterogeneous data includes:

[0016] Collect heterogeneous data from multiple sources, including data from production management systems, distributed control systems, laboratory information management systems, online analyzers, and environmental monitoring systems;

[0017] The data quality is assessed from four dimensions: completeness, accuracy, consistency, and timeliness. A weighted comprehensive evaluation method is used to calculate the comprehensive data quality score, and the comprehensive quality score is mapped to the credibility weight.

[0018] A uniform sampling period is set, and a weighted average method is used for time alignment of high-frequency data, while a case-based reasoning imputation method is used for time alignment of low-frequency data.

[0019] Establish material conservation constraints and energy conservation constraints, calculate the constraint violation degree, and trigger the data correction process when the violation degree is greater than the threshold. Use constraint optimization methods to establish a data correction optimization model to minimize the data correction amount while satisfying physical constraints.

[0020] Obtain a unified dataset with credibility weights that satisfies the physical consistency constraint.

[0021] In a preferred embodiment, obtaining the key variable set includes:

[0022] Based on time series data of candidate variables and target variables, Pearson correlation coefficient and Spearman rank correlation coefficient are used to assess the correlation between variables. A correlation coefficient threshold is set for screening to obtain a set of candidate key variables.

[0023] The Granger causality test method for multivariate time series was used to identify causal relationships. Restricted autoregressive models and unrestricted autoregressive models were established. The model parameters were estimated based on the least squares method. The F-statistic was constructed for significance testing to obtain the set of key causal variables.

[0024] Based on metallurgical mechanism knowledge, the set of key causal variables was verified and supplemented, including oxidation reaction mechanism, slag formation mechanism and heat balance mechanism, to obtain the final set of key variables.

[0025] In a preferred embodiment, obtaining the lightweight hybrid-driven prediction model includes:

[0026] A simplified mechanism model was established based on the material balance mechanism and the energy balance mechanism, including a copper distribution model, an energy consumption calculation model and an SO2 concentration model. The nonlinear least squares method was used to calibrate the mechanism model parameters.

[0027] A data-driven prediction model is constructed using deep learning methods, employing a multilayer perceptron architecture and reducing model complexity through knowledge distillation techniques.

[0028] Based on the cascade fusion method, the mechanism model and the data-driven prediction model are integrated, and the prediction output of the mechanism model is used as an extended feature. A lightweight MLP model is used to learn the residual correction.

[0029] By employing an incremental learning method, the model can be updated online using newly generated production data, resulting in a lightweight hybrid-driven prediction model.

[0030] In a preferred embodiment, obtaining the optimal solution includes:

[0031] The Latin hypercube sampling method is used to generate initial sample points within the feasible region of the decision variables, and the mixed prediction model is called to calculate the objective function value to obtain the initial training sample set;

[0032] Based on the Gaussian process regression method, a surrogate model is established for each objective function, and the hyperparameters are optimized by the maximum likelihood estimation method. During the optimization iteration process, the expected improvement criterion is used to select points for accurate evaluation.

[0033] Based on data credibility information, an uncertainty penalty term is introduced into the optimization objective function to construct a robust optimization objective.

[0034] Establish hard constraints for the process and use constraint violation degree function and hierarchical selection strategy to handle the constraints;

[0035] A reference-point-guided multi-objective optimization algorithm is adopted, which sets the reference point position based on the decision-maker's preferences and performs evolutionary iteration. The weighted Tchebycheff method is used to select the best compromise solution from the Pareto solution set, thus obtaining the optimal solution that takes into account data credibility, computational efficiency and decision-maker preferences.

[0036] In a preferred embodiment, the resulting hierarchical control scheme includes:

[0037] Establish a long-term planning layer, based on raw material composition prediction and production plan, and call a robust multi-objective optimization method to generate the optimal batching scheme and process parameter setpoints for each time period, taking into account raw material inventory constraints;

[0038] A medium-term rolling layer is established, which generates control sequences based on long-term planning goals and the current state using model predictive control methods. The first control action is executed in each control cycle and the sequence is updated on a rolling basis.

[0039] A short-term feedback layer is established, and a PID controller is used to quickly adjust key variables to compensate for model errors and external disturbances.

[0040] A multi-layered negotiation mechanism is established, whereby the optimization results of the upper layer are used as the target settings for the lower layer, and the execution deviations of the lower layer are fed back to the upper layer for correction; thus, a multi-time-scale hierarchical collaborative control scheme is obtained.

[0041] In a preferred embodiment, obtaining the cross-process collaborative optimization scheme includes:

[0042] The cross-process collaborative optimization problem is decomposed into two sub-problems using the alternating direction multiplier method. Information exchange is achieved through coordination variables. The decision variables and optimization objectives of each process sub-problem are defined. During iteration, the coordination variables are fixed to optimize each sub-problem, the coordination variables and Lagrange multipliers are updated, and convergence is checked.

[0043] The material properties between processes are calculated based on a lightweight hybrid prediction model and an energy balance model to achieve collaborative optimization of material flow.

[0044] An energy flow network topology model and an optimization model for waste heat recovery and steam distribution are established and solved using linear programming methods to achieve coordinated optimization of energy flow.

[0045] Weakly coupled iterations of material flow and energy flow are performed to update process parameters and objective function values;

[0046] Establish an inter-process buffer mechanism to reduce the real-time coupling intensity between processes by utilizing the buffering capacity of intermediate tanks; and obtain a full-process optimization scheme for the joint optimization of material and energy flow.

[0047] In a preferred embodiment, the optimized strategy includes:

[0048] A digital twin simulation environment is constructed based on a lightweight hybrid prediction model, defining the state space, action space, state transition function, and reward function.

[0049] Offline training is performed using reinforcement learning algorithms. The Q-function is trained using historical production data, and a greedy policy is extracted as the initial policy.

[0050] The initial strategy is deployed to actual production for supervised operation, a confidence assessment mechanism is established, multi-level decision-making modes are set based on the confidence level, and human-machine collaborative decision-making data is collected.

[0051] The strategy is updated by using online fine-tuning methods based on online data, and by using a mix of offline and online data, giving higher weight to online data and gradually increasing the proportion of automatic execution.

[0052] When a significant change in operating conditions is detected, a meta-learning method is used for rapid adaptation, and a small number of samples are collected for gradient updates in a few steps.

[0053] An adaptive decision-making strategy based on offline training, human-machine collaboration, progressive automation, and meta-learning was obtained.

[0054] In a preferred embodiment, the final intelligent optimization strategy includes:

[0055] Based on furnace lining condition monitoring data and equipment maintenance records, furnace lining erosion characteristics are extracted, and a furnace lining failure risk model is established using survival analysis to predict the remaining furnace service time.

[0056] Establish a unit product lifecycle cost model, including raw material costs, energy costs, labor costs, equipment depreciation and maintenance costs, and calculate the comprehensive cost under different optimization strategies;

[0057] Establish a decision-making model that balances short-term and long-term benefits, define short-term and long-term benefit indicators, and use dynamic programming to solve for the optimal decision sequence to maximize the cumulative benefits within the remaining time of the current furnace operation.

[0058] An interpretability analysis module is established, and feature importance analysis and counterfactual reasoning methods are used to generate explanatory information for optimization decisions; thus, a short-term and long-term balance scheme that takes into account both short-term production indicators and long-term equipment life is obtained.

[0059] This invention provides an intelligent optimization system for non-ferrous metal metallurgy, used to execute the aforementioned intelligent optimization method for non-ferrous metal metallurgy, comprising:

[0060] The data fusion and quality assessment module is used to acquire multi-source heterogeneous data, perform quality assessment and fusion on the multi-source heterogeneous data, and obtain a unified dataset with credibility weights.

[0061] The key variable identification module performs correlation analysis and causal inference based on a unified dataset with confidence weights to obtain a set of key variables;

[0062] The hybrid predictive modeling module performs mechanism-driven and data-driven hybrid modeling based on the key variable set, resulting in a lightweight hybrid-driven predictive model.

[0063] The robust multi-objective optimization module uses a lightweight hybrid-driven prediction model as an evaluator to train a Gaussian process regression surrogate model, introduces an uncertainty penalty term, and obtains the optimal solution.

[0064] The multi-timescale collaborative optimization module performs multi-timescale hierarchical collaborative optimization based on the optimal solution to obtain a hierarchical control scheme.

[0065] The cross-process collaborative optimization module, based on a hierarchical control scheme, performs cross-process material and energy flow collaborative optimization to obtain a cross-process collaborative optimization scheme.

[0066] The reinforcement learning optimization module, based on the cross-process collaborative optimization scheme, performs offline and online hybrid reinforcement learning optimization to obtain the optimization strategy;

[0067] The comprehensive benefit assessment module, based on the optimization strategy, performs multi-objective comprehensive benefit assessment and dynamic balancing to obtain the final intelligent optimization strategy.

[0068] The beneficial effects of this invention are as follows:

[0069] By introducing a multi-dimensional data quality assessment mechanism, the system calculates the credibility weight of each data item and employs a weighted fusion method to reduce the impact of low-quality data. Consistency checks and corrections are performed on the fused data using material and energy conservation constraints to ensure data conformity to physical laws. Adaptive data quality management is achieved through dynamically updating the credibility weights. Even in the event of partial sensor failure or data loss, the system can still provide reliable fused data, providing a high-quality data foundation for subsequent analysis and optimization.

[0070] Unlike traditional correlation analysis, which can only identify statistical associations, this invention uses Granger causality test to identify process parameters that have a real causal effect on the target variable, eliminating the interference of spurious correlations and confounding variables; it also combines metallurgical mechanism knowledge to verify and supplement the causal analysis results, ensuring that the identified key variables conform to the physical mechanism. Attached Figure Description

[0071] Figure 1 This is a flowchart of an intelligent optimization method for non-ferrous metal metallurgy according to the present invention;

[0072] Figure 2 This is a block diagram of an intelligent optimization system for non-ferrous metal metallurgy according to the present invention. Detailed Implementation

[0073] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0074] At least one embodiment of the present invention discloses a smart optimization method for non-ferrous metal metallurgy, such as... Figure 1 As shown, it includes:

[0075] Step 1: Acquire multi-source heterogeneous data, perform quality assessment and fusion on the multi-source heterogeneous data, establish material conservation constraints and energy conservation constraints, and obtain a unified dataset with confidence weights.

[0076] Based on the industrial internet architecture of this copper smelting production line, the following five types of multi-source heterogeneous data are collected: Production Management System (MES) data, including raw material batch information (source mine, batch number, theoretical composition), batching plan data (concentrate ratio, flux addition), and production shift information, with a sampling frequency of once per batch; Distributed Control System (DCS) data, including flash furnace reaction zone temperature (12 measuring points distributed at different heights of the furnace body), settling zone temperature (6 measuring points), concentrate feed flow rate, oxygen flow rate, oxygen enrichment concentration, fuel oil flow rate, furnace pressure, flue gas temperature, and flue gas SO2 concentration, with a preferred sampling frequency of 5 seconds; Laboratory information management system data; and other data. The system includes LIMS data, which includes the analysis of the composition of the feed concentrate mixture (elemental content such as Cu, Fe, S, SiO2, As, Pb, Zn, etc.), the analysis of the composition of the produced matte (Cu, Fe, S content), and the analysis of the composition of the slag (Cu, Fe, SiO2 content), with a preferred analysis frequency of once every 2 hours; online analyzer data, including laser-induced breakdown spectroscopy (LIBS) analysis of the concentrate on the conveyor belt (main elements such as Cu, Fe, and S), with a preferred sampling frequency of 1 minute; and environmental monitoring system data, including SO2 concentration, particulate matter concentration, and nitrogen oxide concentration at the chimney emission outlet, with a preferred sampling frequency of 1 minute.

[0077] At the same time, historical reliability statistics of each data source are collected, including: measurement deviation statistics of LIBS analyzer in the past 30 days (relative to laboratory standard values), fault records and data missing rate of DCS sensor, and sample representativeness assessment of LIMS test data.

[0078] Data quality is assessed across four dimensions: completeness, accuracy, consistency, and timeliness. The specific calculation methods are as follows: Completeness is calculated as the complement of the missing data rate, i.e., completeness equals 1 minus the number of missing data points divided by the total number of data points; Accuracy is based on historical reliability statistics, and for online analyzers, accuracy equals 1 minus the average relative deviation; Consistency is calculated by comparing measurements of the same variable from different data sources, using the correlation coefficient as the consistency indicator; Timeliness is based on the difference between the data timestamp and the current time, and is equal to the ratio of the negative time difference of the natural exponential function to the decay constant. A weighted comprehensive evaluation method is used to calculate the overall data quality score, which equals 0.3 multiplied by completeness plus 0.3 multiplied by accuracy plus 0.2 multiplied by consistency plus 0.2 multiplied by timeliness. The comprehensive score is mapped to a credibility weight using a logistic function, where the credibility weight equals 1 divided by 1 plus the negative adjustment coefficient of the natural exponential function multiplied by the comprehensive score minus 0.5, where the adjustment coefficient is set to 10 in one embodiment. This results in a quality-labeled dataset, where each data item contains three attributes: data value, credibility weight, and timestamp.

[0079] A standardized sampling period of 1 minute was set to generate a standard time series. For high-frequency data, a weighted average method was used for time alignment. Within each standard time grid interval, all data points within that interval were weighted according to their confidence level, with higher-confidence data points having a larger weight in the weighted average. For low-frequency data, a case-based reasoning imputation method was used for time alignment. The specific steps were as follows: Historical cases similar to the current operating conditions were retrieved from the historical database, with similarity calculated based on the Euclidean distance of high-frequency variables such as temperature and flow rate; the five historical cases with the highest similarity were selected; the low-frequency variable values, such as laboratory components, were extracted from these five cases at the corresponding time points; the low-frequency variable values ​​at the current time point were estimated using a weighted average method, with the weight proportional to the similarity; the estimated values ​​were then imputed to the current time point to complete the time alignment, resulting in a time-aligned multi-source dataset.

[0080] Material conservation constraints and energy conservation constraints are established, and the constraint violation degree is calculated. When the violation degree exceeds a threshold, a data correction process is triggered. A data correction optimization model is established using constraint optimization methods, with the goal of minimizing the data correction amount while satisfying physical constraints. A sequential quadratic programming algorithm is used to solve the model, resulting in a fused dataset that satisfies the physical consistency constraints.

[0081] A unified dataset was obtained after quality assessment, time alignment, and physical constraint verification. Each standard time point corresponds to a vector containing the fused values ​​of all variables (including 58 variables such as DCS process parameters, LIMS test components, online analysis components, and environmental monitoring data), as well as a corresponding confidence weight vector. This dataset serves as the unified data foundation for subsequent causal analysis, model building, and optimization decision-making.

[0082] Step 2: Based on the unified dataset with confidence weights, perform correlation analysis and causal inference to obtain the set of key variables;

[0083] Based on the high-quality unified dataset output in step 1, this dataset contains time series data for 58 candidate variables (covering 19 DCS process parameters, 20 LIMS analysis components, 10 online analysis components, and 9 environmental monitoring data) and 4 target variables (copper matte grade, copper content in slag, energy consumption, and SO2 concentration). The data spans 6 months and contains a total of 259,200 time points. The Pearson correlation coefficient method is used to evaluate the linear correlation between variables. The linear correlation coefficient is calculated for each pair of variables. The sample mean is calculated for each combination of candidate and target variables. The deviation of each data point from the corresponding mean is calculated. The deviations of the two variables are multiplied to obtain the covariance numerator. All covariance numerators are summed to obtain the covariance numerator. The sum of squares of the deviations of the two variables is calculated and the square root is taken to obtain the standard deviation. The covariance numerator is divided by the product of the two standard deviations to obtain the Pearson correlation coefficient. The coefficient ranges from -1 to 1, and the larger the absolute value, the stronger the linear correlation.

[0084] For nonlinear relationships, Spearman's rank correlation coefficient is used to supplement the analysis. The original data is converted into ranks, and the Pearson correlation coefficient between the ranks is calculated to obtain the rank correlation coefficient.

[0085] Based on correlation threshold screening, in one embodiment, the correlation coefficient threshold is set to 0.25. For candidate variables, if there is at least one target variable such that the absolute value of the correlation coefficient or the absolute value of the rank correlation coefficient is greater than the threshold, then the variable is included in the candidate key variable set. After screening, 32 candidate key variables are obtained, including key process parameters such as concentrate feed flow rate, oxygen flow rate, oxygen enrichment concentration, fuel oil flow rate, average temperature of the reaction zone, average temperature of the settling zone, furnace pressure, flue gas temperature, copper grade of the concentrate fed into the furnace, sulfur content of the concentrate fed into the furnace, iron content of the concentrate fed into the furnace, quartz flux addition, matte production flow rate, slag production flow rate, flue gas flow rate, and flue gas SO2 concentration.

[0086] The Granger causality test for multivariate time series variables was used to identify causal relationships. The specific steps were as follows: For candidate and target variables, a restricted autoregressive model was established, where the current value of the target variable equals a linear combination of its past values ​​plus an error term; an unrestricted autoregressive model was established, where the current value of the target variable equals a linear combination of its past values ​​plus a linear combination of the candidate variable's past values ​​plus an error term; the lag order was determined using the Akaike Information Criterion, and in one embodiment, the lag order was set to 5; the parameters of the two models were estimated using the least squares method; an F-statistic was constructed, which equals the difference between the sum of squared residuals of the restricted model and the sum of squared residuals of the unrestricted model, divided by the lag order, then divided by the sum of squared residuals of the unrestricted model divided by the sample size minus twice the lag order minus 1; an F-test was performed at a significance level of 0.05. If the F-statistic was greater than the critical value, a Granger causal relationship was considered to exist between the candidate variable and the target variable; each of the 32 candidate key variables was tested, and variables with significant causal relationships to at least one target variable were selected, resulting in a set of 23 key causal variables.

[0087] Based on metallurgical mechanism knowledge, the set of key causal variables was verified and supplemented, including oxidation reaction mechanism, slag formation mechanism, and heat balance mechanism. According to the oxidation reaction mechanism, oxygen flow rate and oxygen enrichment concentration are key factors affecting the oxidation reaction rate and furnace temperature. According to the slag formation mechanism, the amount of quartz flux added affects slag fluidity and copper distribution in the slag. According to the heat balance mechanism, fuel oil flow rate affects the furnace heat balance and temperature control. After verification based on mechanism knowledge, 4 variables out of 23 key causal variables were removed because the mechanism analysis determined they were not direct influencing factors. Meanwhile, variables that are important in mechanism but have no significant statistical correlation, such as furnace lining temperature, were added, resulting in a final set of 19 key variables, specifically: concentrate feed flow rate, oxygen flow rate, oxygen enrichment concentration, fuel oil flow rate, quartz flux flow rate, average temperature of the reaction zone, average temperature of the settling zone, furnace pressure, flue gas temperature, copper grade of the concentrate fed into the furnace, sulfur content of the concentrate fed into the furnace, iron content of the concentrate fed into the furnace, silica content of the concentrate fed into the furnace, matte production flow rate, slag production flow rate, flue gas flow rate, sulfur dioxide concentration in flue gas, oxygen concentration in flue gas, and furnace lining temperature.

[0088] Step 3: Based on the key variable set, a cascade fusion method is used to perform mechanism-driven and data-driven hybrid modeling to obtain a lightweight hybrid-driven prediction model;

[0089] Based on the 19 key variables and their historical data identified in step 2, and the historical data of the 4 target variables, combined with the metallurgical mechanism model of flash melting, hybrid modeling is carried out under the constraints of computational resources (the time for a single prediction inference must be less than 0.5 seconds, and the number of model parameters must not exceed 500,000).

[0090] Simplified mechanistic models were established based on material and energy balance mechanisms. The copper distribution model was established based on the relationship between copper recovery rate and oxygen supply intensity, furnace temperature, and slag type index. A correction function was used to adjust the baseline recovery rate, and the copper grade of matte and the copper content in the slag were calculated. The energy consumption calculation model was established based on fuel oil consumption, oxygen consumption, and the exothermic reaction of sulfide ore oxidation, calculating the energy consumption per unit product. The SO2 concentration model was established based on the sulfur content of the concentrate and the degree of flue gas dilution. Uncertain parameters in the mechanistic models were used as parameters to be estimated. A parameter estimation optimization problem was established using the nonlinear least squares method, and solved using the Levenberg-Marquardt algorithm to obtain the optimized mechanistic model parameters.

[0091] A data-driven prediction model is constructed using deep learning methods. A multilayer perceptron architecture is employed, with the network structure including an input layer, hidden layers, and an output layer, using ReLU as the activation function. To reduce model complexity, knowledge distillation is used to train a large-scale teacher network and a small-scale student network to simulate the teacher network's output. The student network's loss function includes the mean squared error with the true label and the KL divergence with the teacher network's output distribution. The Adam optimizer is used to train the model, and an early stopping strategy is employed to prevent overfitting, resulting in a lightweight data-driven prediction model.

[0092] The cascade fusion method integrates the mechanistic model and the data-driven prediction model. The specific steps are as follows: 19 key variables are input as basic features into the mechanistic model, resulting in four predicted outputs: predicted value for matte grade, predicted value for copper content in slag, predicted value for energy consumption, and predicted value for sulfur dioxide concentration. The 19 key variables and the four predicted outputs are concatenated to form a 23-dimensional extended feature vector. A lightweight multilayer perceptron model is constructed to learn residual correction. The multilayer perceptron structure consists of 23 neurons in the input layer, 64 neurons in one hidden layer, and 4 neurons in the output layer. The activation function is a modified linear unit function. The training objective of the residual correction model is to minimize the mean square error of the true value minus the predicted value of the mechanistic model minus the residual correction value. The final predicted value of the hybrid model equals the predicted value of the mechanistic model plus the residual correction value.

[0093] An incremental learning approach is adopted to enable the model to be updated online using newly generated production data. An incremental update is triggered every time a certain number of new samples are accumulated. Mini-batch gradient descent is used to fine-tune the model parameters with a small learning rate. The number of update iterations is limited to a certain range to ensure that the update time is controllable. This results in a lightweight hybrid-driven prediction model that is computationally efficient, real-time, and capable of online updates. Given the input of key variables, this model can quickly output high-precision predicted values ​​of the target variable, meeting the requirements of prediction accuracy and computational speed for real-time optimization.

[0094] Step 4: Train the Gaussian process regression surrogate model using the lightweight hybrid-driven prediction model as the evaluator, introduce an uncertainty penalty term, and use a reference point-guided multi-objective optimization algorithm for iteration to obtain the optimal solution;

[0095] Based on the lightweight hybrid-driven prediction model constructed in step 3, the Latin hypercube sampling method is used to generate initial sample points in the feasible domain of the decision variables, and the hybrid prediction model is called to calculate the objective function value to obtain the initial training sample set.

[0096] The Gaussian process regression method establishes a surrogate model for each objective function. The specific steps are as follows: First, a kernel function is selected, using the squared exponential kernel function, also known as the radial basis function kernel. The kernel function value equals the signal variance multiplied by the square of the L2 norm of the difference between the negative input vectors of the natural exponential function, divided by twice the square of the length scale. Here, the signal variance and length scale are hyperparameters. Second, based on the initial training sample set, a log-likelihood function is constructed. The log-likelihood function equals -0.5 multiplied by the transpose of the objective function value vector multiplied by the inverse of the kernel matrix multiplied by the objective function value vector, minus 0.5 multiplied by the logarithm of the kernel matrix determinant, and then minus 0.5 multiplied by the number of samples multiplied by the logarithm of pi. Third, the maximum likelihood estimation method is used to optimize the hyperparameters, and the gradient ascent method is used to solve for the hyperparameter values ​​that maximize the log-likelihood function. Fourth, for a new input point, the predicted mean equals the kernel vector transpose of that point and the training point multiplied by the inverse of the kernel matrix multiplied by the objective function value vector. The predicted variance equals the kernel function value at that point minus the kernel vector transpose multiplied by the inverse of the kernel matrix multiplied by the kernel vector. During the optimization iteration process, the expected improvement criterion is used to select key points for precise evaluation. Points with larger expected improvement values ​​are more likely to be the optimal solution. These points are prioritized for precise evaluation using the original hybrid model, and the surrogate model is updated accordingly. Based on data reliability information, an uncertainty penalty term is introduced into the optimization objective function. The uncertainty penalty term equals the prediction variance multiplied by the penalty coefficient. A robust optimization objective is constructed, which equals the prediction mean plus the uncertainty penalty term. Hard process constraints are established, including concentrate feed flow rate between 80 and 110 tons per hour, oxygen flow rate between 15,000 and 22,000 standard cubic meters per hour, oxygen enrichment concentration between 45% and 60%, fuel oil flow rate between 1 and 3 tons per hour, and quartz flux addition between 4 and 10 tons per hour. Constraints are handled using a constraint violation function and a hierarchical selection strategy. Feasible solutions that satisfy all constraints are screened. If the number of feasible solutions is insufficient, they are selected by sorting them from smallest to largest constraint violation degree.

[0097] A reference-point-guided multi-objective optimization algorithm is adopted, which sets the reference point position based on the decision-maker's preference and performs evolutionary iteration. Simulated binary crossover and polynomial mutation operators are used to generate the offspring population. Feasible solutions are screened according to a hierarchical selection strategy and non-dominated sorting is performed on the feasible solutions. A reference point association method is used to select solutions to enter the next generation population. The perpendicular distance from each solution to each reference direction is calculated, and solutions corresponding to reference directions with fewer associated solutions are selected first. Key solutions are periodically selected for accurate evaluation using the original mixture model, and the Gaussian process regression surrogate model is updated.

[0098] The termination condition is set as 150 iterations or the objective function improvement is less than 0.0001 for 20 consecutive iterations, resulting in a Pareto optimal solution set containing 100 solutions. The weighted Tchebycheff method is used to select the best compromise solution from the Pareto solution set. For each Pareto optimal solution, the difference between each objective function value and the ideal value is calculated, divided by the difference between the ideal value and the worst value, and then multiplied by the corresponding preference weight. The maximum value among all objectives is taken as the Tchebycheff distance of the solution, and the solution with the smallest Tchebycheff distance is selected as the best compromise solution.

[0099] The following example of a specific optimization illustrates the implementation process of this method. Based on the current operating conditions, the copper grade of the concentrate fed into the furnace is 25.3%, the sulfur content is 28.5%, and the iron content is 26.2%. Using the method of this invention for optimization, the operating parameters corresponding to the optimal compromise solution are: concentrate feed flow rate 95.2 tons per hour, oxygen flow rate 18,500 standard cubic meters per hour, oxygen enrichment concentration 52.3%, fuel oil flow rate 1.8 tons per hour, and quartz flux addition 6.5 tons per hour.

[0100] The optimal solution, considering data reliability, computational efficiency, and consideration of decision-maker preferences, was obtained. The predicted target values ​​after implementing this solution are: 66.8% matte grade, 0.65% copper content in slag, 0.42 tons of standard coal equivalent per ton of copper, and 4.2% sulfur dioxide concentration in flue gas. This optimized solution was implemented in actual production. After two hours of stable operation, the measured results were: 66.7% matte grade, 0.66% copper content in slag, 0.43 tons of standard coal equivalent per ton of copper, and 4.3% sulfur dioxide concentration in flue gas. The deviations from the predicted values ​​were all within 5%, verifying the effectiveness of the optimized solution. This optimal solution serves as the long-term planning target for multi-timescale collaborative control.

[0101] Step 5: Based on the optimal solution, establish a long-term planning layer, a medium-term rolling layer, and a short-term feedback layer, and perform multi-time-scale hierarchical collaborative optimization to obtain a hierarchical control scheme;

[0102] Based on the optimal solution obtained in step 4 and the real-time production data stream, a long-term planning layer is established under the constraint of equipment adjustment capacity parameters. The planning time domain is the next 120 minutes, and the time domain is divided into 4 planning periods. Based on the future raw material batch information and composition prediction provided by the raw material management system, the predicted sequence of concentrate composition for each period is obtained.

[0103] For each time period, the raw material composition prediction and production plan for that time period are used as inputs. The robust multi-objective optimization method in step 4 is called to generate the optimal batching scheme and process parameter settings for that time period. Considering the raw material inventory constraints, the batching structure is adjusted using a linear programming method. The long-term optimization plan is obtained, in which the target parameters of the first time period are used as reference settings for medium-term rolling optimization.

[0104] A mid-term rolling layer is established, with a rolling time domain of the next 20 minutes, a sampling period of 2 minutes, and a prediction time domain step count of 10. A control sequence is generated based on the model predictive control method. Specifically, an explicit model predictive control method is adopted, establishing a state-space model. The value of the state vector at the next time step equals the state transition matrix multiplied by the current state vector plus the control input matrix multiplied by the current control input. The output vector equals the output matrix multiplied by the state vector. The model predictive control optimization problem is defined, with the objective function being the weighted sum of squares of the output deviations within the prediction time domain plus the weighted sum of squares of the control increments. The constraints are the upper and lower bounds of the control input, control increment, and output. In the offline stage, the model predictive control problem is transformed into a multi-parameter quadratic programming problem. Based on the different values ​​of the current state, the state space is divided into multiple polyhedral regions, each corresponding to an affine control law. The control input equals the state feedback matrix multiplied by the state vector plus the constant vector. The region boundaries and control law parameters are stored as lookup tables. During the online phase, based on the current state variables including reaction zone temperature, settling zone temperature, furnace pressure, and matte production flow rate, the pre-calculated control law table is queried to determine which region the current state belongs to, and the optimal control increment corresponding to the current state region is obtained. The calculation time is less than 0.1 seconds. Based on the control increment, the control input is updated. At the same time, considering the equipment adjustment rate limit, the control increment is limited to between the minimum adjustment rate multiplied by the sampling period and the maximum adjustment rate multiplied by the sampling period. The intermediate rolling control output that satisfies the adjustment rate constraint is obtained. The first control action is executed in each control cycle and the output is updated on a rolling basis.

[0105] A short-term feedback layer is established, and a PID controller is used to quickly adjust key variables to compensate for model errors and external disturbances. An incremental PID control algorithm is used to calculate the adjustment amount. A feedforward control term is introduced to predict the impact on furnace temperature based on changes in concentrate feed flow rate and concentrate sulfur content, and adjust the fuel oil flow rate in advance. The control output is limited and smoothed to obtain a smooth and executable real-time control output, which is then sent to the DCS system for execution.

[0106] An inter-layer negotiation mechanism is established, using the optimization results of the upper layer as the target setpoints of the lower layer, and feeding back the execution deviations of the lower layer to the upper layer for correction. When the short-term control layer detects that it cannot effectively track the intermediate setpoints, the inter-layer negotiation process is triggered. The short-term control layer feeds back constraint conflict information to the intermediate optimization layer. After receiving the feedback information, the intermediate optimization layer relaxes the constraints or adjusts other variables and regenerates the corrected control sequence. If the intermediate optimization layer still cannot meet the constraints after correction, it reports to the long-term planning layer, and the long-term planning layer adjusts the production plan.

[0107] A multi-timescale hierarchical collaborative control scheme was obtained, which includes: the batching plan and process parameter target trajectory output by the long-term planning layer, the control sequence output by the medium-term rolling layer, and the setpoint tracking control signal output by the short-term control layer. This hierarchical control scheme, as the output of single-process optimization, will be passed to the cross-process collaborative optimization step.

[0108] Step 6: Based on the hierarchical control scheme, the alternating direction multiplier method is used to perform cross-process material energy flow collaborative optimization to obtain a cross-process collaborative optimization scheme;

[0109] Based on the multi-timescale collaborative control scheme for the flash smelting process obtained in step 5 (focusing on the control sequence for the next 20 minutes output from the intermediate rolling layer), combined with the current status information of the converter blowing process (progress of the current blowing batch, crude copper temperature in the converter of 1180℃, oxygen flow rate of 12000 standard cubic meters / hour) and the converter's demand information for upstream matte (expected matte grade of 66%–68%, expected matte temperature of 1180–1220℃, expected matte supply flow rate of 35–40 tons / hour), as well as the topology and parameters of the energy flow network (including the flash furnace flue gas waste heat recovery system, the converter flue gas waste heat recovery system, and the concentrate preheating system), The connection relationship and heat exchange efficiency parameters of the waste heat boiler are considered. The Alternating Directional Multiplier Method (ADMM) is used to decompose the cross-process collaborative optimization problem of flash furnace-converter into two sub-problems. Information exchange is achieved through coordination variables (copper matte properties). The decision variables of the flash furnace sub-problem are defined as flash furnace process parameters (concentrate ratio, oxygen flow rate, etc.). The optimization objective is the multi-objective performance index of the flash furnace (copper content in slag, energy consumption, etc.) plus a penalty term for deviation from the coordination variables. The penalty term includes the square of the L2 norm of the difference between the copper matte property vector determined by the decision variables and the coordination consistency constraint value, multiplied by the penalty parameter (0.5), and the inner product of the Lagrange multiplier vector and the difference.

[0110] The decision variables of the converter subproblem are defined as converter process parameters (oxygen flow rate, blowing time, etc.), and the optimization objective is the converter's multi-objective performance index (crude copper quality, blowing time, energy consumption, etc.) plus a penalty term for deviation from the coordination variables. The penalty term includes the square of the L2 norm of the difference between the converter's demand for matte properties and the coordination constraint value, multiplied by the penalty parameter, and the negative inner product of the Lagrange multiplier vector and the difference.

[0111] In the k-th iteration, the coordination variables and Lagrange multipliers are fixed, and the flash furnace decision variables are optimized to obtain the flash furnace process parameters and corresponding matte properties. The coordination variables and Lagrange multipliers are also fixed, and the converter decision variables are optimized to obtain the converter process parameters and corresponding matte demand. Based on the flash furnace output and converter demand, the coordination variables are updated to a weighted average of the two. The Lagrange multipliers are updated by adding the penalty parameter to the Lagrange multipliers from the previous iteration and multiplying them by the difference between the flash furnace matte properties and the coordination variables. A convergence test is performed. If the original residual (the norm of the difference between the flash furnace matte properties and the coordination variables) is less than the threshold 0.01 and the dual residual (the norm of the penalty parameter multiplied by the change in the coordination variables) is less than the threshold 0.01, the algorithm converges; otherwise, it returns to continue iterating. The maximum number of iterations is limited to 5 to ensure computational efficiency.

[0112] Based on the process parameters obtained from solving the flash furnace problem, the lightweight hybrid prediction model constructed in step 3 is called to quickly calculate the matte grade; based on material balance, the matte output flow rate is determined by the copper recovery rate and the amount of copper fed into the furnace, which has been calculated in the prediction model in step 3.

[0113] The temperature of matte is calculated based on an energy balance model. The principle of energy balance is that the heat released in the reaction zone is equal to the heat carried away by the matte and slag, plus the heat carried away by the flue gas, plus the heat loss from the furnace body. The temperature of the matte is mainly determined by the temperature of the reaction zone in the furnace and the residence time in the settling zone. During the settling process, the matte cools down through heat dissipation from the furnace wall and heat exchange with the slag. The matte temperature is equal to the reaction zone temperature minus the temperature drop during the settling process. The temperature drop during the settling process is calculated based on the residence time and heat loss, using a simplified Newton's law of cooling. The temperature drop is equal to the cooling coefficient multiplied by the residence time multiplied by the difference between the reaction zone temperature and the ambient temperature. The cooling coefficient is determined based on historical data statistics and thermodynamic calculations. In one embodiment, it is preferably taken as 0.002 degrees Celsius to the power of negative 1 multiplied by minutes to the power of negative 1. The residence time in the settling zone is determined based on the liquid level in the furnace and the copper discharge rate. Under the current operating conditions, the residence time is approximately 15 minutes, the reaction zone temperature is 1250 degrees Celsius, and the ambient temperature is 30 degrees Celsius. The calculated temperature drop during the settling process is approximately 55 degrees Celsius, and the matte temperature is approximately 1195 degrees Celsius. The predicted matte properties include matte grade, temperature, and flow rate, with a calculation time of approximately 0.5 seconds.

[0114] Based on the converter process requirements, the desired properties of the matte in the converter are 67% grade, 1200℃ temperature, and 38 tons / hour flow rate, which serve as the coordination target. In the ADMM iteration, the output of the flash furnace and the requirements of the converter gradually become consistent, eventually converging to the coordination variables of 66.9% grade, 1195℃ temperature, and 37.5 tons / hour flow rate, thus achieving collaborative optimization of material flow.

[0115] Based on the energy flow network topology, a simplified model of the main waste heat recovery path is established. The flue gas from the flash furnace (temperature about 1300℃, flow rate about 150,000 Nm³ / h) enters the waste heat boiler to recover high-temperature waste heat and generate steam (pressure 3.5 MPa, temperature 450℃). The steam generated by the waste heat boiler has three uses: for drying and preheating concentrate, for process steam in the converter area, and for power generation by the waste heat generator set.

[0116] A steam distribution optimization model was established with the goal of maximizing the overall energy utilization efficiency. The model was calculated by multiplying the energy utilization efficiency of steam used for concentrate preheating (0.75) by the distribution amount, adding the energy utilization efficiency of steam used for converter process (0.65) by the distribution amount, and adding the energy utilization efficiency of steam used for waste heat power generation (0.25) by the distribution amount. The constraints were that the total steam volume was balanced, the sum of the distribution amounts for the three uses was equal to the total steam production of the waste heat boiler (calculated as 120 tons / hour based on the flue gas parameters of the flash furnace), and the demand limits for each use: steam for concentrate preheating was between 30 and 50 tons / hour, steam for converter process was between 15 and 25 tons / hour, and steam for waste heat power generation was non-negative.

[0117] The optimization model is a linear programming problem, which is solved quickly using the simplex method to obtain the optimal steam distribution scheme: 45 tons / hour of steam for concentrate preheating, 20 tons / hour of steam for converter process, and 55 tons / hour of steam for waste heat power generation, with a calculation time of less than 0.1 seconds. Based on the steam distribution scheme, the comprehensive energy utilization rate is calculated by multiplying the energy utilization efficiency of each purpose by the sum of the distribution amounts and dividing by the total steam production, resulting in a comprehensive energy utilization rate of 56.7%, which is 4.7 percentage points higher than the baseline value of 52% without optimization. The energy optimization scheme is obtained, and the optimized configuration of waste heat recovery increases the concentrate preheating temperature to 180℃ (originally 120℃), reducing the fuel oil consumption of the flash furnace.

[0118] Based on the material flow optimization results (copper matte properties) and energy flow optimization results (steam distribution scheme and concentrate preheating temperature of 180℃), a weakly coupled iteration is performed. After increasing the concentrate preheating temperature from the original 120℃ to 180℃, the heat content of the concentrate entering the furnace increases. According to the energy balance, the fuel oil consumption can be reduced. The fuel oil saving is equal to the concentrate feed flow rate multiplied by the specific heat capacity of the concentrate (0.8 kcal / (kg·℃)) multiplied by the preheating temperature difference (60℃), and finally divided by the lower heating value of fuel oil (10500 kcal / kg). The fuel oil saving is calculated to be 0.44 tons / hour.

[0119] After reducing fuel oil consumption, the flash furnace problem was resolved to obtain updated process parameters. The predictive model was then used to calculate the new matte properties and objective function value. It was found that the matte grade slightly increased to 67.1% (due to a slight decrease in furnace temperature caused by reduced fuel oil, which helps reduce copper volatilization loss), the copper content in the slag decreased to 0.63%, and the energy consumption decreased to 0.38 tons of standard coal / ton of copper. Since the objective function improved to 8.5% after one iteration, and the calculation time had accumulated to about 2 minutes, the weakly coupled iteration was terminated to ensure real-time performance. The current solution was adopted as the final cross-process collaborative optimization scheme. The material-energy flow joint optimization scheme was obtained, the process parameters of the flash furnace and converter were coordinated, the waste heat recovery configuration of the energy system was optimized, and the overall energy consumption of the entire process was reduced by about 12%.

[0120] Based on the buffering capacity of the matte tank (capacity 150 tons, current level 80 tons), it can provide a buffer time of about 2 hours. A buffering strategy is established: when the properties (grade, temperature) of the matte produced by the flash furnace deviate significantly from the converter's requirements (grade deviation greater than 2% or temperature deviation greater than 40°C), it is not immediately sent to the converter. Instead, it is temporarily stored in the matte tank until the properties return to the normal range before being supplied to the converter, thus avoiding upstream fluctuations being directly transmitted downstream.

[0121] When the level in the copper matte tank approaches the upper limit (greater than 120 tons) or the lower limit (less than 30 tons), an emergency coordination mechanism is triggered: if the level is too high, the converter's requirements for copper matte quality are relaxed, accelerating copper matte consumption; if the level is too low, the flash furnace is notified to increase the output flow rate or appropriately reduce the quality requirements, increasing the supply of copper matte. Through the buffer mechanism, the real-time coupling strength between the flash furnace and the converter is reduced, and the two processes can be optimized independently to a certain extent, ensuring material balance and quality matching only on a longer time scale (hourly), thus improving the robustness and flexibility of the system.

[0122] A computationally efficient (total optimization time approximately 3 minutes), distributed implementation, and buffered cross-process collaborative optimization scheme was obtained, including: flash furnace process parameters (concentrate feed flow rate 95.2 tons / hour, oxygen flow rate 18200 standard cubic meters / hour, oxygen enrichment concentration 52.5%, fuel oil flow rate 1.36 tons / hour, quartz flux 6.5 tons / hour), converter process parameters (blowing oxygen flow rate 12500 standard cubic meters / hour, blowing time adjusted to 43 minutes), and a waste heat distribution scheme for the energy system (45 tons / hour of steam for concentrate preheating, 20 tons / hour of steam for converter process, 55 tons / hour of steam for waste heat power generation, and concentrate preheating temperature 180℃). This scheme achieves material coordination between the flash furnace and converter (copper matte grade 67.1%, temperature 1195℃, and flow rate 37.5 tons / hour meeting the needs of both) and energy optimization (overall energy consumption reduced by 12%, and comprehensive energy utilization rate increased to 56.7%). This collaborative optimization scheme, as the output of static optimization, performs online adaptive optimization on the input reinforcement learning module.

[0123] Step 7: Based on the cross-process collaborative optimization scheme, perform offline and online hybrid reinforcement learning optimization to obtain the optimization strategy;

[0124] Based on the cross-process collaborative optimization scheme and historical production data obtained in step 6, a digital twin simulation environment is constructed using the lightweight hybrid prediction model in step 3, defining the state space, action space, state transition function, and reward function.

[0125] Define the state space, including the current process parameters of the flash furnace, furnace state variables, product quality indicators, equipment status, and upstream and downstream process information.

[0126] The motion space is defined as the adjustment amount of process parameters, and continuous motion space is used to achieve fine control.

[0127] Define a state transition function based on the lightweight hybrid prediction model in step 3. Input the current state and action, and predict the state at the next moment. The state transition function contains a deterministic part and a stochastic part. The stochastic part adds noise based on historical data statistics to simulate the uncertainty in actual production.

[0128] Define a reward function that comprehensively considers multiple production objectives. The reward function equals the output reward item minus the quality penalty item minus the energy consumption penalty item minus the equipment wear penalty item. By adjusting the weight coefficients of each item, a trade-off can be made between different production objectives to adapt to different production strategy requirements.

[0129] Based on the constructed digital twin simulation environment, reinforcement learning algorithms were used for offline training, specifically the deep Q-network algorithm. The Q-function was represented by a deep neural network with an input layer of 42 neurons corresponding to the state dimension, three hidden layers of 256, 128, and 64 neurons respectively, and an output layer representing the action value function. Historical production data was used to train the Q-function, which contained approximately 500,000 state transition samples from the past six months of production records. The dataset was divided into a training set (80%) and a validation set (20%), and an experience replay mechanism was used for training. The replay buffer size was 100,000 samples, and 256 samples were randomly sampled from the training set each time for training.

[0130] A target network technique was employed to stabilize the training process, with the target network updated every 500 iterations. The loss function was the expected value of the immediate reward plus a discount factor multiplied by the maximum Q-value of the next state minus the square of the Q-value of the current state action pair, where the discount factor was set to 0.99. An adaptive moment estimation optimizer was used for parameter updates, with a learning rate of 0.0003, a batch size of 256, and 10,000 training iterations. Performance was evaluated on the validation set every 100 iterations. After offline training, a greedy policy was extracted from the trained Q-function as the initial policy, i.e., for a given state, the action that maximizes the Q-value was selected as the policy output. The initial policy was tested on the validation set, and the average cumulative reward was improved by approximately 15% compared to the manual operation policy in historical data, indicating that the offline training obtained an initial policy superior to human experience.

[0131] The initial strategy is deployed to actual production for supervised operation. During this stage, the optimization suggestions output by the strategy are not directly executed, but are provided to the operators as a reference, and the operators decide whether to adopt them. A confidence evaluation mechanism is established to calculate a confidence score for each optimization suggestion. The confidence score is based on a comprehensive evaluation of multiple factors: the similarity between the current state and historical training data, the prediction uncertainty of the Q function, and the degree of deviation between the optimization suggestion and the current operation.

[0132] A multi-level decision-making model is set based on confidence levels: high-confidence recommendations, medium-confidence recommendations, and low-confidence recommendations correspond to different decision-making models; through the multi-level decision-making model, operators' trust in the optimization system is gradually built up while ensuring production safety.

[0133] During the supervised operation phase, human-machine collaborative decision-making data is continuously collected, and each optimization suggestion, the operator's actual decision, and the actual effect after implementation are recorded.

[0134] The strategy is updated using an online fine-tuning method based on online data. The online fine-tuning adopts a hybrid data training approach, which uses both offline and online data, and gives higher weight to online data. A hybrid experience replay buffer is constructed, in which offline and online data are stored separately. When calculating the loss function, the loss of online data samples is multiplied by the weight factor, so that online data has a greater impact on the strategy update.

[0135] As online data accumulates and strategy performance improves, the proportion of automatic execution is gradually increased. During the gradual automation process, system performance and safety indicators are continuously monitored. When anomalies are detected, the proportion of automatic execution is automatically reduced, and the system reverts to a more conservative decision-making mode to ensure production safety. A manual intervention mechanism is established, allowing operators to pause automatic execution at any time and take over control.

[0136] When a significant change in operating conditions is detected, a meta-learning method is used for rapid adaptation. The detection of changes in operating conditions is based on the statistical characteristics of the state distribution. The difference between the state distribution of the current time window and the state distribution of the historical normal operating conditions is calculated. When the difference exceeds a threshold, it is determined to be a significant change in operating conditions.

[0137] The core idea of ​​meta-learning methods is to learn how to learn quickly, that is, to train a meta-policy so that it can quickly adapt to new working conditions based on a small amount of new data; after detecting changes in working conditions in actual production, a small number of samples are collected to perform gradient updates in a few steps, and a policy that adapts to the new working conditions is quickly obtained.

[0138] An adaptive decision-making strategy based on offline training, human-machine collaboration, progressive automation, and meta-learning was obtained.

[0139] Step 8: Based on the optimization strategy, conduct multi-objective comprehensive benefit evaluation and dynamic balancing to obtain the final intelligent optimization strategy;

[0140] Based on the optimization strategy obtained in step 7, and combined with furnace lining condition monitoring data, equipment maintenance records, production cost data, and market demand information, a multi-objective comprehensive benefit assessment and dynamic balance are conducted. Furnace lining erosion characteristics are extracted. The furnace lining condition monitoring data includes furnace lining thickness measurements, furnace lining temperature distribution, and furnace lining material composition analysis; characteristic variables are extracted from the monitoring data.

[0141] A survival analysis method was used to establish a furnace lining failure risk model to predict the remaining furnace service time. Specifically, the survival analysis method employed a Cox proportional hazards model, whose risk function is expressed as the product of a baseline risk function and a linear combination of covariate indices. Covariates included the current furnace lining thickness, cumulative operating time, average operating temperature, temperature fluctuation range, raw material erosion index, and operating intensity index. Based on historical furnace service data, a partial likelihood function was constructed, and the regression coefficients of the covariates were estimated using the maximum likelihood estimation method. An iterative algorithm was then used to solve for the parameter values ​​that maximized the partial likelihood function. The current furnace lining state and operating conditions were substituted into the fitted Cox model to calculate the survival function, which represents the probability that the furnace lining has not failed at a given time point. The median of the remaining furnace service time was extracted from the survival function as the predicted value. Based on the predicted remaining furnace service time, maintenance plans and optimization strategies could be rationally arranged.

[0142] Establish a unit product lifecycle cost model, including raw material costs, energy costs, labor costs, equipment depreciation and maintenance costs, and calculate the comprehensive cost under different optimization strategies. Different optimization strategies will affect various costs, and it is necessary to weigh multiple objectives to find the optimization strategy with the lowest comprehensive cost.

[0143] Establish a decision-making model that balances short-term and long-term benefits, and define short-term benefit indicators and long-term benefit indicators. Short-term benefit indicators include the output of the current batch, the quality pass rate, energy consumption, and direct production costs. Long-term benefit indicators include the remaining life of the equipment, cumulative maintenance costs, and the comprehensive cost of the entire life cycle.

[0144] The optimal decision sequence is solved using dynamic programming to maximize the cumulative benefit within the remaining time of the furnace operation. The specific steps are as follows: Define the state as the equipment state on day t, including furnace lining thickness and cumulative operating time; define the action as the optimal strategy selected on day t, corresponding to a set of reward function weight coefficients; define the immediate benefit as the benefit gained from executing the strategy on day t; define the value function as the maximum cumulative benefit from day t to the end of the furnace operation; establish the Bellman equation, where the value function equals the maximum value of the immediate benefit among all possible actions plus the value function of the next state, where the next state is determined by the state transition function; the boundary condition is that the value function is zero at the end of the furnace operation; discretize the remaining furnace operation time into multiple decision stages, each stage being one day, and select an optimal decision sequence in each stage. The optimal strategy is implemented to obtain the daily benefit while affecting equipment status, such as reducing furnace lining thickness, and entering the next stage. A reverse recursive method is used to solve the problem, starting from the end of the furnace service and progressively calculating the optimal strategy for each stage. This strategy maximizes the immediate benefit plus the value function of the next state, along with the optimal value function. For example, if the remaining furnace service time is 95 days, dynamic programming provides the following strategies: a balanced strategy with a short-term weight of 0.5 and a long-term weight of 0.5 for the first 30 days; a slightly conservative strategy with a short-term weight of 0.4 and a long-term weight of 0.6 for the middle 40 days; and a conservative strategy with a short-term weight of 0.3 and a long-term weight of 0.7 for the last 25 days. This yields the optimal strategy sequence from the current moment to the end of the furnace service, which is expected to increase the cumulative benefit by approximately 8% compared to the fixed strategy.

[0145] The weight coefficients in the reinforcement learning reward function are dynamically adjusted based on the current production context, which includes factors such as furnace operation stage, market demand, raw material supply, and equipment status. By dynamically adjusting the weights of the reward function, the reinforcement learning strategy can adapt to different production stages and external conditions. The weight adjustment adopts a smooth transition method to avoid abrupt changes that could lead to strategy instability. This results in a dynamically optimized strategy that adapts to different production stages and external conditions.

[0146] An interpretability analysis module is established, employing feature importance analysis and counterfactual reasoning methods to generate explanatory information for optimization decisions. Feature importance analysis is used to identify which state variables have the greatest impact on optimization decisions, assigning an importance score to each feature. The results of feature importance analysis are presented to operators in the form of visual charts to help them understand the basis for optimization decisions.

[0147] Counterfactual reasoning is used to answer questions such as "What would happen if we didn't adopt the optimization suggestion?" or "What would happen if we adopted other options?". Through comparative analysis, it helps operators evaluate the value of optimization suggestions. Based on a digital twin simulation environment, it simulates the execution results of different decision-making options. By comparing the simulation results of different options, it calculates the benefit improvement of the optimization suggestion compared to maintaining the status quo, as well as its advantages and disadvantages compared to other alternative options.

[0148] A balanced solution that takes into account both short-term production targets and long-term equipment lifespan was obtained.

[0149] The intelligent optimization system for non-ferrous metallurgy of this invention adopts a layered distributed architecture, including a data layer, a model layer, an optimization layer, a decision layer, and an execution layer. The data layer is responsible for the acquisition, fusion, and quality assessment of multi-source heterogeneous data, and is deployed on a field data acquisition server. The model layer deploys core artificial intelligence models such as lightweight hybrid-driven prediction models and digital twin simulation environments, and is deployed on an edge computing server. The optimization layer runs multi-timescale optimization algorithms, cross-process collaborative optimization algorithms, and robust multi-objective optimization algorithms, and is deployed on an optimization computing server. The decision layer integrates reinforcement learning strategies, comprehensive benefit evaluation modules, and interpretability analysis modules, and is deployed on a decision server. The execution layer interfaces with the distributed control system, converting optimization decisions into control commands and issuing them to field equipment for execution, and is deployed on the distributed control system interface server. The layers communicate through standardized interfaces and employ a message queue mechanism to achieve asynchronous decoupling, improving the system's scalability and fault tolerance.

[0150] The resulting intelligent optimization strategy comprehensively considers multiple dimensions of benefits, including output, quality, cost, energy consumption, and equipment lifespan, and possesses context-adaptive capabilities and interpretability. This strategy dynamically adjusts the weights of optimization objectives to achieve the optimal balance between short-term and long-term benefits under different production stages and external conditions. It realizes a closed-loop intelligent optimization process, encompassing data collection, model building, multi-objective optimization, multi-scale collaboration, cross-process collaboration, online learning, and comprehensive balancing.

[0151] A non-ferrous metal metallurgical intelligent optimization system is used to execute the aforementioned non-ferrous metal metallurgical intelligent optimization method, such as... Figure 2 As shown, it includes:

[0152] The data fusion and quality assessment module is used to acquire multi-source heterogeneous data, perform quality assessment and fusion on the multi-source heterogeneous data, establish material conservation constraints and energy conservation constraints, and obtain a unified dataset with credibility weights.

[0153] The key variable identification module performs correlation analysis and causal inference based on a unified dataset with confidence weights to obtain a set of key variables;

[0154] The hybrid predictive modeling module, based on the key variable set, uses a cascade fusion method to perform hybrid modeling of mechanism and data-driven approaches, resulting in a lightweight hybrid-driven predictive model.

[0155] The robust multi-objective optimization module uses a lightweight hybrid-driven prediction model as the evaluator to train a Gaussian process regression surrogate model, introduces an uncertainty penalty term, and uses a reference-point guided multi-objective optimization algorithm for iteration to obtain the optimal solution.

[0156] The multi-timescale collaborative optimization module establishes a long-term planning layer, a medium-term rolling layer, and a short-term feedback layer based on the optimal solution, and performs multi-timescale hierarchical collaborative optimization to obtain a hierarchical control scheme.

[0157] The cross-process collaborative optimization module, based on the hierarchical control scheme, uses the alternating direction multiplier method to perform cross-process material energy flow collaborative optimization, and obtains the cross-process collaborative optimization scheme.

[0158] The reinforcement learning optimization module, based on the cross-process collaborative optimization scheme, performs offline and online hybrid reinforcement learning optimization to obtain the optimization strategy;

[0159] The comprehensive benefit assessment module, based on the optimization strategy, performs multi-objective comprehensive benefit assessment and dynamic balancing to obtain the final intelligent optimization strategy.

[0160] In one embodiment of the present invention, a specific example is provided:

[0161] The intelligent optimization method of this invention underwent a 90-day field test on a flash smelting-blowing production line in a large copper smelting enterprise. During the test, a complete intelligent optimization system was deployed, including a multi-source data acquisition system, edge computing units, a hybrid-driven predictive model, a robust multi-objective optimization engine, and a multi-scale collaborative control system. The production line is designed to produce 300,000 tons of electrolytic copper annually, with a flash furnace daily concentrate processing capacity of 2,400 tons, and is equipped with a comprehensive DCS control system and online monitoring instruments. The test collected multi-dimensional information such as raw material composition, process parameters, production indicators, and energy consumption data. Typical data from a week in July 2025 is used as an example for illustration.

[0162] Due to changes in raw material inventory and supply plans, the daily ingredient mixing plan needs to be dynamically adjusted. Table 1 shows the daily raw material ratios and mixture composition data for the test week:

[0163] Table 1: Raw material composition and ingredient data (a week in July 2025);

[0164]

[0165] As shown in Table 1, the composition of the mixture fluctuated significantly within a week: copper grade fluctuated between 21.9% and 23.6% (fluctuation range 7.8%), and arsenic content fluctuated between 1590 and 2180 ppm (fluctuation range 37.1%). This fluctuation reflects the uncertainty of raw material supply and the complexity of batching adjustments in actual production. Traditional methods struggle to adapt quickly to such dynamic changes, while the intelligent optimization system of this invention, through multi-source data fusion in step 1 and key variable identification in step 2, can accurately grasp the changing patterns of raw material composition and adjust process parameters in advance.

[0166] The intelligent optimization system dynamically adjusts process parameters based on changes in raw material composition to ensure stable production indicators. Table 2 shows the daily setpoints for key process parameters and actual production indicators during the test week.

[0167] Table 2: Process parameters and production indicators (a week in July 2025);

[0168]

[0169] As shown in Table 2, despite significant fluctuations in raw material composition, the intelligent optimization system dynamically adjusted process parameters to stabilize the matte grade within the range of 66.3%–67.2% (standard deviation approximately 0.32%), control the copper content in the slag within the range of 0.62%–0.72%, and maintain the overall energy consumption at a low level of 0.408–0.436 tons of standard coal equivalent per ton of copper. Data shows that when the raw material copper grade is high (e.g., Tuesdays and Fridays), the system automatically reduces fuel oil flow and oxygen concentration to achieve energy conservation and consumption reduction; when the raw material grade is low (e.g., Wednesdays and Saturdays), the system appropriately increases the reaction intensity to ensure matte quality. This demonstrates the effectiveness of robust multi-objective optimization in step 4 and multi-timescale collaborative control in step 5.

[0170] The intelligent optimization method and system for non-ferrous metal metallurgy proposed in this invention organically combines key technologies such as multi-source data fusion, causal inference, hybrid modeling, robust optimization, multi-scale collaboration, cross-process collaboration, reinforcement learning, and comprehensive balancing to construct a full-process intelligent optimization system from data to decision-making. This method overcomes the shortcomings of traditional optimization methods in dealing with complex problems such as dynamic changes in raw materials, multi-objective conflicts, uncertainties, and long-term and short-term balances, achieving efficient, stable, economical, and environmentally friendly operation of non-ferrous metal metallurgical processes.

[0171] Practical application verification shows that the method of this invention can improve product quality stability, reduce material loss, save energy consumption, and extend equipment life. This method has good scalability and can be applied to other non-ferrous metal metallurgical processes such as aluminum electrolysis, lead-zinc smelting, and nickel smelting, providing an effective technical solution for the intelligent transformation and upgrading of the industry.

[0172] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments based on the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A smart optimization method for non-ferrous metal metallurgy, characterized in that, Includes the following steps: Acquire multi-source heterogeneous data, perform quality assessment and fusion of the multi-source heterogeneous data, establish material conservation constraints and energy conservation constraints, and obtain a unified dataset with credibility weights; Based on a unified dataset with confidence weights, correlation analysis and causal inference are performed to obtain a set of key variables; Based on the key variable set, a cascade fusion method is used to perform hybrid modeling of mechanism and data-driven approaches, resulting in a lightweight hybrid-driven prediction model. A Gaussian process regression surrogate model was trained using a lightweight hybrid-driven prediction model as the evaluator. An uncertainty penalty term was introduced, and a reference-point guided multi-objective optimization algorithm was used for iteration to obtain the optimal solution. Based on the optimal solution, a long-term planning layer, a medium-term rolling layer, and a short-term feedback layer are established, and multi-timescale hierarchical collaborative optimization is carried out to obtain a hierarchical control scheme. Based on the hierarchical control scheme, the alternating direction multiplier method is used to optimize the material energy flow across processes, resulting in a cross-process collaborative optimization scheme. Based on the cross-process collaborative optimization scheme, offline and online hybrid reinforcement learning optimization is performed to obtain the optimization strategy; Based on the optimization strategy, a multi-objective comprehensive benefit assessment and dynamic balance are performed to obtain the final intelligent optimization strategy.

2. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The quality assessment and fusion of multi-source heterogeneous data includes: Collect heterogeneous data from multiple sources, including data from production management systems, distributed control systems, laboratory information management systems, online analyzers, and environmental monitoring systems; The data quality is assessed from four dimensions: completeness, accuracy, consistency, and timeliness. A weighted comprehensive evaluation method is used to calculate the comprehensive data quality score, and the comprehensive quality score is mapped to the credibility weight. A uniform sampling period is set, and a weighted average method is used for time alignment of high-frequency data, while a case-based reasoning imputation method is used for time alignment of low-frequency data. Establish material conservation constraints and energy conservation constraints, calculate the constraint violation degree, and trigger the data correction process when the violation degree is greater than the threshold. Use constraint optimization methods to establish a data correction optimization model to minimize the data correction amount while satisfying physical constraints. Obtain a unified dataset with credibility weights that satisfies the physical consistency constraint.

3. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The obtained set of key variables includes: Based on time series data of candidate variables and target variables, Pearson correlation coefficient and Spearman rank correlation coefficient are used to assess the correlation between variables. A correlation coefficient threshold is set for screening to obtain a set of candidate key variables. The Granger causality test method for multivariate time series was used to identify causal relationships. Restricted autoregressive models and unrestricted autoregressive models were established. The model parameters were estimated based on the least squares method. The F-statistic was constructed for significance testing to obtain the set of key causal variables. Based on metallurgical mechanism knowledge, the set of key causal variables was verified and supplemented, including oxidation reaction mechanism, slag formation mechanism and heat balance mechanism, to obtain the final set of key variables.

4. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The obtained lightweight hybrid-driven prediction model includes: Simplified mechanism models were established based on material balance and energy balance mechanisms, including a copper distribution model, an energy consumption calculation model, and an SO2 concentration model. The nonlinear least squares method was used to calibrate the mechanism model parameters. A data-driven prediction model is constructed using deep learning methods, employing a multilayer perceptron architecture and reducing model complexity through knowledge distillation techniques. Based on the cascade fusion method, the mechanism model and the data-driven prediction model are integrated, and the prediction output of the mechanism model is used as an extended feature. A lightweight MLP model is used to learn the residual correction. By employing an incremental learning method, the model can be updated online using newly generated production data, resulting in a lightweight hybrid-driven prediction model.

5. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The optimal solution includes: The Latin hypercube sampling method is used to generate initial sample points within the feasible region of the decision variables, and the mixed prediction model is called to calculate the objective function value to obtain the initial training sample set; Based on the Gaussian process regression method, a surrogate model is established for each objective function, and the maximum likelihood estimation method is used to optimize the hyperparameters. During the optimization iteration process, the expected improvement criterion is used to select points for accurate evaluation. Based on data credibility information, an uncertainty penalty term is introduced into the optimization objective function to construct a robust optimization objective. Establish hard constraints for the process and use constraint violation function and hierarchical selection strategy to handle the constraints; A reference-point-guided multi-objective optimization algorithm is adopted, which sets the reference point position based on the decision-maker's preferences and performs evolutionary iteration. The weighted Tchebycheff method is used to select the best compromise solution from the Pareto solution set, thus obtaining the optimal solution that takes into account data credibility, computational efficiency and decision-maker preferences.

6. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The resulting hierarchical control scheme includes: Establish a long-term planning layer, based on raw material composition prediction and production plan, and call a robust multi-objective optimization method to generate the optimal batching scheme and process parameter set values ​​for each time period, taking into account raw material inventory constraints; A medium-term rolling layer is established, which generates control sequences based on long-term planning goals and the current state using model predictive control methods. The first control action is executed in each control cycle and the sequence is updated on a rolling basis. A short-term feedback layer is established, and a PID controller is used to quickly adjust key variables to compensate for model errors and external disturbances. A multi-layered negotiation mechanism is established, whereby the optimization results of the upper layer are used as the target settings for the lower layer, and the execution deviations of the lower layer are fed back to the upper layer for correction; thus, a multi-time-scale hierarchical collaborative control scheme is obtained.

7. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The obtained cross-process collaborative optimization scheme includes: The cross-process collaborative optimization problem is decomposed into two sub-problems using the alternating direction multiplier method. Information exchange is achieved through coordination variables. The decision variables and optimization objectives of each process sub-problem are defined. During iteration, the coordination variables are fixed to optimize each sub-problem, the coordination variables and Lagrange multipliers are updated, and convergence is checked. The material properties between processes are calculated based on a lightweight hybrid prediction model and an energy balance model to achieve collaborative optimization of material flow. An energy flow network topology model and an optimization model for waste heat recovery and steam distribution are established and solved using linear programming methods to achieve coordinated optimization of energy flow. Weakly coupled iterations of material flow and energy flow are performed to update process parameters and objective function values; Establish an inter-process buffer mechanism to reduce the real-time coupling intensity between processes by utilizing the buffering capacity of intermediate tanks; and obtain a full-process optimization scheme for the joint optimization of material and energy flow.

8. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The optimized strategies include: A digital twin simulation environment is constructed based on a lightweight hybrid prediction model, defining the state space, action space, state transition function, and reward function. Offline training is performed using reinforcement learning algorithms. The Q-function is trained using historical production data, and a greedy policy is extracted as the initial policy. The initial strategy is deployed to actual production for supervised operation, a confidence assessment mechanism is established, multi-level decision-making modes are set based on the confidence level, and human-machine collaborative decision-making data is collected. The strategy is updated by using online fine-tuning methods based on online data, and by using a mix of offline and online data, giving higher weight to online data and gradually increasing the proportion of automatic execution. When a significant change in operating conditions is detected, a meta-learning method is used for rapid adaptation, and a small number of samples are collected for gradient updates in a few steps. An adaptive decision-making strategy based on offline training, human-machine collaboration, progressive automation, and meta-learning was obtained.

9. The intelligent optimization method for non-ferrous metal metallurgy according to claim 1, characterized in that, The final intelligent optimization strategy obtained includes: Based on furnace lining condition monitoring data and equipment maintenance records, furnace lining erosion characteristics are extracted, and a furnace lining failure risk model is established using survival analysis to predict the remaining furnace service time. Establish a unit product lifecycle cost model, including raw material costs, energy costs, labor costs, equipment depreciation and maintenance costs, and calculate the comprehensive cost under different optimization strategies; Establish a decision-making model that balances short-term and long-term benefits, define short-term and long-term benefit indicators, and use dynamic programming to solve for the optimal decision sequence to maximize the cumulative benefits within the remaining time of the current furnace operation. An interpretability analysis module is established, and feature importance analysis and counterfactual reasoning methods are used to generate explanatory information for optimization decisions; thus, a short-term and long-term balance scheme that takes into account both short-term production indicators and long-term equipment life is obtained.

10. A smart optimization system for non-ferrous metal metallurgy, characterized in that, A method for implementing a non-ferrous metal metallurgical intelligent optimization method according to any one of claims 1-9 includes: The data fusion and quality assessment module is used to acquire multi-source heterogeneous data, perform quality assessment and fusion on the multi-source heterogeneous data, establish material conservation constraints and energy conservation constraints, and obtain a unified dataset with credibility weights. The key variable identification module performs correlation analysis and causal inference based on a unified dataset with confidence weights to obtain a set of key variables; The hybrid predictive modeling module, based on the key variable set, uses a cascade fusion method to perform hybrid modeling of mechanism and data-driven approaches, resulting in a lightweight hybrid-driven predictive model. The robust multi-objective optimization module uses a lightweight hybrid-driven prediction model as the evaluator to train a Gaussian process regression surrogate model, introduces an uncertainty penalty term, and uses a reference-point guided multi-objective optimization algorithm for iteration to obtain the optimal solution. The multi-timescale collaborative optimization module establishes a long-term planning layer, a medium-term rolling layer, and a short-term feedback layer based on the optimal solution, and performs multi-timescale hierarchical collaborative optimization to obtain a hierarchical control scheme. The cross-process collaborative optimization module, based on the hierarchical control scheme, uses the alternating direction multiplier method to perform cross-process material energy flow collaborative optimization, and obtains the cross-process collaborative optimization scheme. The reinforcement learning optimization module, based on the cross-process collaborative optimization scheme, performs offline and online hybrid reinforcement learning optimization to obtain the optimization strategy; The comprehensive benefit assessment module, based on the optimization strategy, performs multi-objective comprehensive benefit assessment and dynamic balancing to obtain the final intelligent optimization strategy.