Prediction modeling method for copper electrodeposition effluent concentration based on double-layer residual correction
By constructing a mechanism model based on mass conservation and combining linear regression residual correction with machine learning algorithms, the accuracy and robustness issues of predicting the concentration of copper electrolyte were solved, and the optimization and precise control of the copper electrowinning process were achieved.
Patent Information
- Application Number
- CN202511476913.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-30
AI Technical Summary
Existing technologies struggle to accurately predict the concentration of Me in copper electrolytes, making it difficult to optimize and control the production process. Traditional mechanism modeling and data-driven methods suffer from large prediction errors and poor generalization performance in complex and data-scarce scenarios.
By constructing a mechanism model based on mass conservation and combining it with parameters obtained by least squares inversion, linear regression residual correction is performed. Then, machine learning algorithms are used to train the residual model and perform double-layer residual correction to finally form the target model to predict the effluent concentration.
It improves the accuracy and robustness of copper electrowinning solution concentration prediction, adapts to scarce data scenarios, enhances the model's generalization ability and prediction accuracy, and supports energy consumption optimization and yield improvement in the copper electrowinning process.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of copper metallurgy clean production, and particularly relates to a prediction modeling method for copper electrowinning effluent concentration based on double-layer residual correction. BACKGROUND
[0002] The copper electrowinning technology is the most mature copper electrolyte purification process at present and is widely used in the copper smelting industry. Due to the strong acidity, high concentration and many impurity interferences of the copper electrolyte, it is difficult to control and regulate the electrowinning reaction process in actual production. The concentration of metals / metalloids Me such as copper, arsenic, antimony, bismuth and nickel in the electrowinning effluent directly affects the treatment load and purification removal effect of the subsequent process. How to accurately predict the Me concentration in the effluent is of great significance to realize the optimization of energy consumption, the improvement of copper electrowinning yield and the regulation of the subsequent reaction process.
[0003] At present, the detection of the Me concentration in the effluent mainly relies on the method of inductively coupled plasma emission spectrometry. This method needs a series of processes such as on-site sampling, sample sending and dilution detection. The detection process is complicated, time-consuming and high in detection cost. Mechanism modeling based on the operating condition parameters of the electrowinning can roughly predict the effluent concentration. However, due to the complexity, time-varying nature and strong coupling between parameter characteristics of the electrowinning process, it is difficult to accurately predict by traditional mechanism modeling. Especially in the scene of data scarcity, because of the existence of detection errors, the traditional data-driven machine learning modeling method often has large prediction errors and poor model generalization performance. SUMMARY
[0004] The purpose of the present application is to provide a prediction modeling method for copper electrowinning effluent concentration based on double-layer residual correction. By setting a combination mechanism model, the prediction accuracy, generalization ability and robustness of the copper electrowinning effluent concentration can be effectively improved. The application scenario of scarce data can be effectively expanded, and data support can be provided for the energy consumption optimization of the copper electrowinning process and the provision of copper electrowinning yield and efficient operation.
[0005] It should be noted that Me in the present embodiment is a general abbreviation of different metal ions, which can be one or more of copper, arsenic, antimony, bismuth and nickel. Copper is only an example, and the present method is applicable to different metal ions.
[0006] To achieve the above object, the present application provides the following technical scheme: A prediction modeling method for copper electrowinning effluent concentration based on double-layer residual correction, comprising: S1: Based on the historical data of the control system and the detection system of the copper electrowinning, relevant feature variables of continuous time series are extracted to form an input variable data set, and the data set is cleaned and processed; S2: Me concentration data of the liquid component detected at equal time intervals is used to form an output variable data set, and the data set is cleaned and processed; S3: The input variable data set is filtered based on the correlation coefficient between the input variable data and the output variable data to obtain a first input data set; S4: The output variable data set and the first input data set are subjected to data conversion processing to obtain a second input-output data set, and the second input-output data set is divided into a training set and a test set; S5: A mechanism model is constructed based on the conservation of mass, the training set is input into the mechanism model, and the model parameters of the mechanism model are inverted based on the least squares method until the sum of squared errors between the theoretical prediction value and the actual observation value is minimized; S6: Linear regression residual correction is performed on the theoretical prediction value to obtain a first corrected prediction value; S7: The difference between the first corrected prediction value and the actual observation value is used as the output variable data set to construct a prediction model between the above input variable data set and the output data set; S8: Based on the residual prediction model established in S7, different machine learning algorithms are trained until the sum of squared errors between the predicted value of the residual and the observed value of the residual is minimized, and the trained model is applied to the test set until the requirements are met; S9: The residual model is used to correct the first corrected prediction value to obtain a second corrected prediction value, and a target model is obtained based on the second corrected prediction value, and the target model is used to obtain the liquid Me concentration prediction value.
[0007] As a further scheme of the present application: the correlation characteristic variables include time characteristics, inlet liquid copper ion concentration, inlet liquid arsenic ion concentration, inlet liquid nickel ion concentration, inlet liquid flow, electrolyte sulfuric acid concentration, electrolyte volume and current density; As a further scheme of the present application: in S1 and S2, the data set is cleaned and processed, including: Abnormal values of the input variable data and the output variable data that do not meet the preset value range are removed to obtain first cleaned data; The missing values of the first cleaned data are filled by using the upper and lower average value method to obtain second cleaned values, and the second cleaned values are used as the input-output variable data set.
[0008] As a further scheme of the present application: in S3, the cleaned and processed input variable data set is filtered based on the correlation coefficient between the input variable data and the output variable data, including: The calculation formula is: The correlation coefficient between the input variable data and the Me concentration data is calculated, wherein, For input variable data, For Me concentration data, express and The Pearson product-moment correlation coefficient; express and covariance; express The variance; express The variance; Input variables with a correlation coefficient greater than 0.9 were removed as redundant variables, and the input variables after removing redundant variables were used as the first input dataset.
[0009] As a further aspect of the present invention: in step S4, the output variable dataset and the first dataset undergo data transformation processing, including: The MinMaxScaler method is used for data normalization, converting both the output variable dataset and the first dataset into supervised learning data in the [0,1] interval. The calculation formula for the MinMaxScaler method is as follows: In the formula, To output the variable dataset or the first dataset, and They are respectively The maximum and minimum values, This is the new value after normalization.
[0010] As a further aspect of the present invention: In step S5, the mechanism model based on the Me mass conservation includes the use of the following calculation formula: C theory The formula =(x1*x2-a*x1**b*x3**c+d) / (x2+0.00001) calculates the copper ion concentration at the inlet and outlet of the copper electrodeposition solution, where C... theory X1 represents the Me concentration (g / L) at time t for copper electrowinning before and after the electrolyte, X2 represents the Me concentration at the inlet and outlet of the electrolyte, and X3 represents the liquid volume (m³). 3 X3 represents the electrodeposition current density (A / m³). 2 ), a, b, c, and d are all parameters.
[0011] As a further aspect of the present invention: the inversion of model parameters of the mechanism model based on the least squares method includes employing an error function: To minimize the sum of squared errors between theoretical predictions and actual observations, where, , These represent the actual observed value and the theoretical predicted value of Me ion concentration, respectively. For the sample size, It is the value that minimizes the sum of squared errors between the theoretical prediction and the actual observation.
[0012] As a further aspect of the present invention, the machine learning algorithm involved in S8 can be GBR algorithm, ExtraTrees algorithm, CatBoost algorithm, XGBoost algorithm or SVR algorithm, etc.
[0013] As a further aspect of the present invention: the formula for calculating the determination coefficient of the target model is: , The formula for calculating the mean absolute percentage error loss of the target model is as follows: The formula for calculating the root mean square error loss of the target model is as follows: In the formula, , , These represent the observed value, the average value of the observed values, and the predicted value of the Me ion concentration, respectively. For the number of data points.
[0014] As a further aspect of the present invention: the method is used in the copper electrolyte electrowinning purification process of hydrometallurgical copper electrolytic refining process, and the metal ions Me in the effluent can be copper, arsenic, antimony, bismuth, and nickel.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs a mechanistic model based on the mass conservation law of Me, and uses the least squares method to invert the model parameters of the mechanistic model until the sum of squared errors between the theoretical prediction and the actual observation is minimized. This ensures that the target component in the liquid inlet and outlet system of the dynamic electrowinning unit can follow the constraints of the mass conservation law, so that the overall trend of the result data meets expectations. By combining a secondary correction process, the prediction error is reduced and the accuracy of the target model prediction is improved.
[0016] 2. This invention calculates the Me ion concentration at the inlet and outlet of copper electrowinning by developing an empirical expression for the electrowinning reaction. This can reflect the dynamic characteristics of Me ion migration and deposition during the electrowinning process. At the same time, based on the principle of mass conservation, the concentration value is predicted by inputting real-time variables, providing an initial prediction benchmark for subsequent double-layer residual correction, thereby reducing systematic errors and improving the robustness of the overall model.
[0017] 3. This invention combines residual distribution characteristics for a first linear fitting correction and uses machine learning algorithms to capture the residual characteristics of the predicted values after linear correction for training and prediction. Combined with a second correction process for the theoretical predicted values, it can effectively improve the accuracy of prediction and the precision of fitting. At the same time, it is beneficial to improve the interpretability, generalization ability, sensitivity to abnormal data and robustness of data. By adapting to application scenarios with scarce data, it can help expand the scenario adaptability of the prediction model. Attached Figure Description
[0018] Figure 1 This is a comparison chart showing the prediction effect of copper concentration in an electrowinning copper stripping unit of the present invention and existing data-driven modeling methods. Figure 2 This is a comparison chart showing the prediction effect of the copper concentration of the two-stage electrowinning copper removal unit of the present invention and the existing data-driven modeling method. Figure 3 This is a diagram illustrating the method steps of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example: Please see Figure 3 In this embodiment of the invention, a predictive modeling method for copper electrodeposition solution concentration based on double-layer residual correction includes the following steps: S1: Based on the historical data of the control and detection systems of copper electrowinning, relevant feature variables of continuous time series are extracted to form an input variable dataset, and the dataset is cleaned and processed. S2: The Me concentration data in the effluent components detected at equal time intervals constitute the output variable dataset, and the dataset is cleaned and processed. S3: Based on the correlation coefficient between the input variable data and the output variable data, the input variable dataset is filtered to obtain the first input dataset; S4: Perform data transformation on the output variable dataset and the first input dataset to obtain the second input-output dataset, and divide the second input-output dataset into a training set and a test set; S5: Construct a mechanism model based on mass conservation, input the training set into the mechanism model, and invert the model parameters of the mechanism model based on the least squares method until the sum of squared errors between the theoretical predictions and the actual observations is minimized; S6: Correct the theoretical prediction value using linear regression residuals to obtain the first corrected prediction value; S7: Using the difference between the first corrected predicted value and the actual observed value as the output variable dataset, construct a prediction model between the above input variable dataset and the output dataset; S8: Based on the residual prediction model established in S7, different machine learning algorithms are used for training until the sum of squared errors between the predicted residual value and the observed residual value is minimized. The trained model is then used in the test set until the requirements are met. S9: The first corrected prediction value is corrected using the residual model to obtain the second corrected prediction value. The target model is obtained based on the second corrected prediction value, and the predicted value of the effluent Me concentration is obtained using the target model.
[0021] Preferably, the relevant characteristic variables include time characteristics, influent copper ion concentration, influent arsenic ion concentration, influent nickel ion concentration, influent flow rate, electrolyte sulfuric acid concentration, electrolyte volume, and current density; Historical data from the control and detection systems based on copper electrowinning includes: The time features of the continuous time series, the concentrations of copper ions, arsenic ions, nickel ions, flow rate, sulfuric acid concentration, electrolyte volume, and current density are extracted to form the target feature variables.
[0022] In this embodiment, copper electrowinning is a technique used for purifying copper electrolyte in the hydrometallurgical copper electrolytic refining process.
[0023] Preferably, in steps S1 and S2, cleaning and processing the dataset involves cleaning the Me concentration data, including: The first cleaned data is obtained by removing outliers from the input and output variable data that do not conform to the preset value range; The missing values in the first cleaned data are filled in using the upper and lower averaging method to obtain the second cleaned value, which is then used as the input and output variable dataset.
[0024] Preferably, in step S3, the correlation coefficient between the input variable data and the output variable data is based on the correlation coefficient between the input variable data and the Me concentration data. The cleaned and processed input variable dataset is then filtered, including: Using the calculation formula: Calculate the correlation coefficient between the input variable data and the Me concentration data, where, For input variable data, For Me concentration data, express and The Pearson product-moment correlation coefficient; express and covariance; express The variance; express The variance; Input variables with a correlation coefficient greater than 0.9 were removed as redundant variables, and the input variables after removing redundant variables were used as the first input dataset.
[0025] Preferably, in step S4, the output variable dataset and the first dataset undergo data transformation processing, including: The MinMaxScaler method is used for data normalization, transforming both the output variable dataset and the first dataset into supervised learning data within the [0,1] interval. The MinMaxScaler method is calculated as follows: In the formula, To output the variable dataset or the first dataset, and They are respectively The maximum and minimum values, This is the new value after normalization.
[0026] Preferably, in step S5, a mechanism model is constructed based on the mass conservation of Me, including the use of the following calculation formula: C theory The formula =(x1*x2-a*x1**b*x3**c+d) / (x2+0.00001) calculates the copper ion concentration at the inlet and outlet of the copper electrodeposition solution, where C... theory X1 represents the Me ion concentration at time t during copper electrowinning, X2 represents the liquid volume, X3 represents the electrowinning current density, and a, b, c, and d are all parameters.
[0027] Preferably, the model parameters of the mechanistic model are inverted based on the least squares method, including the use of an error function: To minimize the sum of squared errors between theoretical predictions and actual observations, where, , These represent the actual observed value and the theoretical predicted value of Me ion concentration, respectively. For the sample size, It is the value that minimizes the sum of squared errors between the theoretical prediction and the actual observation.
[0028] Preferably, the machine learning algorithm involved in S8 can be the GBR algorithm, ExtraTrees algorithm, CatBoost algorithm, XGBoost algorithm or SVR algorithm.
[0029] In this embodiment, the least squares method provides the basic prediction under physical constraints, and the GBR model learns its residual characteristics to enhance it. The two are trained together to achieve two corrections of the residuals, forming a prediction system of "mechanism-based, data-enhanced, and dual-optimized".
[0030] Preferably, the formula for calculating the coefficient of determination of the target model is: , The formula for calculating the mean absolute percentage error loss of the target model is: The formula for calculating the root mean square error loss of the target model is: In the formula, , , These represent the observed value, the average value of the observed values, and the predicted value of the Me ion concentration, respectively. For the number of data points.
[0031] Preferably, the method is used in the copper electrolyte electrowinning purification process of hydrometallurgical copper electrolytic refining, where the metal ions Me in the effluent can be used to predict the concentration of copper, arsenic, antimony, bismuth, and nickel components.
[0032] Example 1: In this embodiment, by performing linear regression error correction on the prediction results, the prediction bias can be further reduced and the prediction accuracy of the third dataset can be improved by fitting the linear relationship between the residuals and the observed values.
[0033] In this embodiment, taking the first-stage electrowinning copper removal process of a copper electrolytic refining and purification system as an example, the existing technology used is a data-driven machine learning (GBR algorithm) modeling method. Compared with the method provided by this invention, the prediction results are shown in Table 1 below: Table 1 In Table 1, R 2 R0 is the coefficient of determination, used to measure the model's ability to explain variations in the data. 2 The higher the value, the better the model fits the data. As can be seen from Table 1 above, the determination coefficient of the training set of the present invention is 94.2%, which is higher than the 87.4% of the existing GBR algorithm. Moreover, the determination coefficient of the present invention is 7.78% higher than that of the existing technology.
[0034] The coefficient of determination for the test set of this invention is 93.5%, while the coefficient of determination for the GBR algorithm on the test set is 81.8%. In the test set, the coefficient of determination for this invention is 14.3% higher than that of the prior art, indicating that the prediction model provided by this invention has a better fit. Figure 1 As shown; In Table 1, MAPE (Mean Absolute Percentage Error) is used to evaluate the deviation between the predicted and actual values. A smaller MAPE value indicates higher prediction accuracy. Table 1 shows that the MAPE of this invention is smaller than that of existing technologies. For the training set, the MAPE of this invention is 1.25%, while that of the GBR algorithm is 2.86%. Therefore, the MAPE of this invention in the training set is reduced by 56.29% compared to the GBR algorithm. For the test set, the MAPE of this invention is 1.52%, while that of the GBR algorithm is 3.05%. Therefore, the MAPE of this invention in the training set is reduced by 50.16% compared to the GBR algorithm. This demonstrates that this invention effectively improves prediction accuracy compared to existing technologies.
[0035] Example 2: In this embodiment, taking the two-stage electrowinning copper removal process of the copper electrolytic refining and purification system as an example, the existing technology used is a data-driven machine learning (XGBoost algorithm) modeling method. Compared with the method provided by this invention, the prediction results are shown in Table 2 below: Table 2 In Table 2, R 2 R0 is the coefficient of determination, used to measure the model's ability to explain variations in the data. 2 The higher the value, the better the model fits the data. As can be seen from Table 2 above, the determination coefficient of the training set of the present invention is 100%, which is higher than the 89.3% of the existing XGBoost algorithm. Moreover, the determination coefficient of the present invention is 10.7% higher than that of the existing technology.
[0036] The coefficient of determination (COD) on the test set of this invention is 99.5%, while that of the XGBoost algorithm is 87%. In the test set, the COD of this invention is 14.37% higher than that of the prior art, indicating that the prediction model provided by this invention has a better fit. Figure 2 As shown; In Table 2, MAPE represents the Mean Absolute Percentage Error, used to evaluate the deviation between the predicted and actual values. A smaller MAPE value indicates higher prediction accuracy. Table 2 shows that the present invention has a smaller mean absolute percentage error than existing technologies. For the training set, the MAPE of the present invention is 1.2e-16, while that of the GBR algorithm is 20.6%. Therefore, the present invention reduces the MAPE by 100.0% compared to the XGBoost algorithm in the training set. For the test set, the MAPE of the present invention is 6.8%, while that of the XGBoost algorithm is 22.47%. Therefore, the present invention reduces the MAPE by 69.74% compared to the XGBoost algorithm in the training set. This demonstrates that the present invention effectively improves prediction accuracy compared to existing technologies. Figure 2 As shown.
[0037] As can be seen from the two embodiments above, the present invention significantly improves the prediction accuracy and generalization ability of the model. Compared with the scarce scenario of 4136 electrotrophic data points in Embodiment 1 and only 1812 electrotrophic data points in Embodiment 2, the improvement in prediction performance is more significant, fully demonstrating the superiority of the method of the present invention.
[0038] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting modeling of copper electrowinning effluent concentration based on double-layer residual correction, characterized in that, The method comprises the following steps: S1: Based on the historical data of the control system and the detection system of copper electrodeposition, the relevant characteristic variables of continuous time series are extracted to form an input variable data set, and the data set is cleaned and processed; S2: The Me concentration data in the liquid outlet component detected at equal time intervals form an output variable data set, which is cleaned and processed; S3: Based on the correlation coefficient between the input variable data and the output variable data, the input variable data set is screened to obtain a first input data set; S4: The output variable data set and the first input data set are subjected to data conversion processing to obtain a second input-output data set, which is divided into a training set and a test set; S5: Based on the mass conservation, a mechanism model is constructed, the training set is input into the mechanism model, and the model parameters of the mechanism model are inverted based on the least squares method until the sum of squared errors between the theoretical prediction value and the actual observation value is minimized; S6: The linear regression residual correction is performed on the theoretical prediction value to obtain a first corrected prediction value; S7: The difference between the first corrected prediction value and the actual observation value is used as the output variable data set to construct a prediction model between the above input variable data set and the output data set; S8: Based on the residual prediction model established in S7, different machine learning algorithms are trained until the sum of squared errors between the predicted value of the residual and the observed value of the residual is minimized, and the trained model is used in the test set; S9: The residual model is used to correct the first corrected prediction value to obtain a second corrected prediction value, and a target model is obtained based on the second corrected prediction value, and the target model is used to obtain the prediction value of the liquid outlet Me concentration.
2. The method of claim 1, wherein the method is characterized by: The relevant characteristic variables include time characteristics, inlet liquid copper ion concentration, inlet liquid arsenic ion concentration, inlet liquid nickel ion concentration, inlet liquid flow, electrolyte sulfuric acid concentration, electrolyte volume, and current density.
3. The method of claim 2, wherein the method is characterized by: In S1 and S2, the data set is cleaned and processed, including: Removing abnormal values of the input variable data and the output variable data that do not meet the preset value range to obtain first cleaning data; Using the upper and lower average value method, the missing values of the first cleaning data are filled to obtain second cleaning values, and the second cleaning values are used as input-output variable data sets.
4. The method of claim 3, wherein the method is characterized by: In S3, based on the correlation coefficient between the input variable data and the output variable data, the cleaned and processed input variable data set is screened, including: Using the calculation formula: a correlation coefficient between the input variable data and the Me concentration data, wherein, for the input variable data, for the Me concentration data, denotes the Pearson product-moment correlation coefficient; denotes the covariance; denotes the variance; denotes the variance; denotes the variance; The input variables with a correlation coefficient greater than 0.9 are removed as redundant variables, and the input variables after removing the redundant variables are used as the first input data set.
5. The method of claim 4, wherein the method is characterized by: In S4, the output variable data set and the first data set are subjected to data conversion processing, including: Using the MinMaxScaler method for data normalization, the output variable data set and the first data set are converted into supervised learning data in the [0, 1] interval, wherein the calculation formula of the MinMaxScaler method is: wherein is the output variable dataset or first dataset, and are the maximum and minimum values, respectively, of the input variable dataset or second dataset, is the normalized new value.
6. The method of claim 5, wherein the method is characterized by: In S5, the mechanism model is constructed based on the mass conservation of Me, including using the calculation formula: C theory = (x1*x2-a*x1**b*x3**c+d) / (x2+0.00001) calculates the copper ion concentration of the copper electrodeposition inlet and outlet liquid, wherein C theory represents the Me ion concentration (g / L) of the copper electrodeposition inlet and outlet liquid at t time, X1 represents the inlet Me concentration (g / L), X2 represents the liquid volume (m 3 ), X3 represents the electrodeposition current density (A / m 2 ), and a, b, c, and d are all parameter items.
7. The method of claim 6, wherein the method is characterized by: The model parameter inversion of the mechanism model based on the least square method comprises adopting an error function: minimizes the sum of squares of errors between the theoretical prediction value and the actual observation value, , are the actual observation value and the theoretical prediction value of the Me ion concentration, respectively, is the number of samples, is the value that minimizes the sum of squares of errors between the theoretical prediction value and the actual observation value.
8. The method of claim 1, wherein the method is characterized by: The machine learning algorithm involved in S8 can be a GBR algorithm, an ExtraTrees algorithm, a CatBoost algorithm, an XGBoost algorithm or an SVR algorithm.
9. The method for predicting modeling of copper electrowinning effluent concentration based on double-layer residual correction according to claim 1, characterized in that: The calculation formula of the determination coefficient of the target model is: , The calculation formula of the mean absolute percentage error loss of the target model is, The calculation formula of the root mean square error loss of the target model is, wherein , , are the observed values, the average of the observed values, the predicted values, respectively, of the Me ion concentration, is the number of data.
10. The method for predicting modeling of copper electrowinning effluent concentration based on double-layer residual correction according to claim 1, characterized in that: The method is used in the copper electrolyte electrodeposition purification process in the hydrometallurgical copper electrolytic refining process, and the metal ion Me of the effluent can be copper, arsenic, antimony, bismuth and nickel.