Copper electrolysis cross-border energy consumption optimization method based on agent model and optimizer cooperation
Patent Information
- Application Number
- CN202611072269.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-20
AI Technical Summary
[0005]本发明的目的在于解决铜电解跨域数据难共享导致能耗模型泛化不足,以及现有优化难以实现优化经验共享的问题,并提出一种基于代理模型与优化器协同的铜电解跨域能耗优化方法
[0034]本发明提出一种代理模型与优化器协同训练的跨域优化架构,在服务器端对各域上传的能耗代理模型区中与优化器的网络权重进行聚合,使各客户端域既能够吸收其他域中的工艺参数到能耗映射知识,又能够继承其他客户端域在寻优过程中形成的有效策略方向与经验信息,从而在不共享原始生产数据的前提下实现跨域知识与优化经验共享,提升整体寻优效率。
Smart Images

Figure CN122595861B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of energy consumption optimization, specifically relating to a cross-domain energy consumption optimization method for copper electrolysis based on the collaboration of a surrogate model and an optimizer. Background Technology
[0002] Non-ferrous metal electrolysis, especially copper electrolysis, is a typical complex electrochemical preparation process involving ion migration, electron transfer, and interfacial reactions under the coupling of multiple physical fields. This process not only needs to balance current efficiency and energy consumption but also must meet key performance indicators such as the tensile strength and elongation of copper foil. Meanwhile, the persistent inefficiency caused by side reactions such as hydrogen evolution makes high energy consumption a core bottleneck restricting the industry's green and sustainable development.
[0003] Existing methods for optimizing copper electrolysis consumption mainly fall into two categories: one is electrolyte parameter control, such as adjusting additive ratios, copper ion concentrations, acid concentrations, or temperatures to improve deposition behavior and suppress side reactions; the other is non-electrolyte process optimization, such as adjusting current intensity, electrode spacing, electrode materials, and power supply strategies to improve current efficiency and maintain stable production. However, these methods still have significant limitations: on the one hand, process parameters exhibit strong nonlinearity and complex coupling, and improper parameter settings can easily lead to increased polarization, increased ohmic losses, and even equipment corrosion, making it difficult to achieve precise control based on manual experience; on the other hand, traditional optimization often relies on the experience of on-site personnel and repeated experiments, lacking a reusable and transferable systematic adjustment mechanism.
[0004] With the development of artificial intelligence technology, data-driven modeling and intelligent optimization have gradually become important directions in this field. For example, ensemble learning is used to build energy consumption prediction models, or intelligent optimization algorithms are combined to search for better process parameters. Although related research has made some progress in prediction accuracy and optimization capabilities, it still faces problems such as limited data scale and distribution differences of data across enterprises / production lines in real industrial scenarios. A single data source is difficult to support a global model with both accuracy and generalization. Federated learning can achieve collaborative modeling under the premise of privacy protection, but traditional federated methods mostly focus on the average aggregation of surrogate model parameters, ignoring the scale shift and optimization instability caused by the differences in local boundaries of each domain, and also lacking a mechanism for sharing and fusion of optimizer strategies. At the same time, the surrogate model continuously evolves with the communication rounds, introducing time-varying uncertainty. Without quantitative evaluation of stability and cross-domain generalization error and adaptive adjustment of hyperparameters, it is easy to cause convergence oscillations or efficiency decline. Therefore, it is necessary to propose a collaborative optimization method for cross-domain scenarios, so as to obtain more robust low-energy-consumption process parameter solutions while ensuring quality and production stability. Summary of the Invention
[0005] The purpose of this invention is to address the problems of insufficient generalization of energy consumption models due to the difficulty in sharing cross-domain data in copper electrolysis, and the difficulty in sharing optimization experience in existing optimization methods. This invention proposes a cross-domain energy consumption optimization method for copper electrolysis based on the collaboration of a surrogate model and an optimizer. This method centers on an energy consumption surrogate model and a reinforcement learning optimizer: the surrogate model learns the nonlinear mapping from process parameters to energy consumption, and the optimizer treats the surrogate model as an interactive environment to perform continuous action searches within the feasible domains of each domain; the server-side performs cross-domain aggregation of the reinforcement learning optimizer network weights, realizing the sharing and iterative evolution of cross-domain optimization strategies.
[0006] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0007] A method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration includes:
[0008] Each domain collects multi-dimensional process parameter vectors and corresponding actual energy consumption during the copper electrolysis production process to form a local dataset;
[0009] Each domain calculates the local boundary of each dimension of the process parameters in the local dataset. The server aggregates the local boundaries of each domain. Each domain normalizes the process parameter vector based on the global feasible domain boundary after boundary aggregation and maps the local boundary to the normalized space to obtain the local sub-feasible domain boundary.
[0010] In each round of communication, each domain trains and updates the energy consumption proxy model on the local dataset to obtain the weights of the candidate proxy models;
[0011] Each domain uses the updated energy consumption proxy model as the interaction environment, constructs a reinforcement learning optimizer, and optimizes process parameters within the boundary of the local sub-feasible domain to obtain the candidate optimal process parameter vector under this round of communication, and updates the network weights of the reinforcement learning optimizer simultaneously.
[0012] The server performs weighted aggregation of the candidate agent model weights and reinforcement learning optimizer network weights for each domain, generates aggregated agent model weights and aggregated optimizer weights for each domain, and sends them to the corresponding domain as initialization parameters for the next communication round.
[0013] The server performs cross-validation and historical best-in-class selection on candidate proxy models for each domain, and then distributes the determined historical best-in-class proxy models to the corresponding domains.
[0014] Each domain substitutes the candidate optimal process parameter vectors from previous communications into the historical optimal proxy model for backtracking evaluation, and selects the process parameter vector with the lowest predicted energy consumption as the optimal process parameter vector for final application in production control.
[0015] Furthermore, the energy consumption proxy model is constructed using a deep neural network, and the mean squared error loss function is minimized in each domain using the Adam optimization algorithm.
[0016] Furthermore, each domain trains and updates its energy consumption proxy model on a local dataset to obtain candidate proxy model weights, including:
[0017] Each domain uses the aggregated proxy model weights issued by the previous round of communication server as the initial values for local training, and executes training on the training set. The local iterative update aims to minimize the mean square error between predicted and actual energy consumption, thereby updating the surrogate model weights and obtaining the candidate surrogate model weights.
[0018] Furthermore, the reinforcement learning optimizer models the process parameter optimization process as a Markov decision process, where the state is a process parameter vector and the action is the adjustment amount of each dimension of the process parameter, and the action satisfies... ;
[0019] After an action is performed, the state transition is performed according to the following rules: ;
[0020] in, For domain No. The state of the interactive step, For domain No. The state of the interactive step, For projection operators, This represents the normalized boundary of the local subfeasible region. This is the scaling factor for the action step size. For domain No. The action of interactive step, For the dimensions of process parameters;
[0021] The reward function consists of a main reward, an improvement reward, and a diversity reward weighted according to preset reward weights. The main reward is calculated based on the relative difference between the current predicted energy consumption and the lowest predicted energy consumption of the current communication round, using a hyperbolic tangent function. The improvement reward is determined based on whether the predicted energy consumption of the current interaction step is lower than the predicted energy consumption of the previous interaction step. The diversity reward is calculated based on the deviation between each action component in a preset historical window and the mean of the corresponding window.
[0022] Furthermore, the server performs weighted aggregation on the candidate agent model weights and reinforcement learning optimizer network weights for each domain, generating aggregated agent model weights and aggregated optimizer weights for each domain, expressed by the formula:
[0023]
[0024] in, Indicates the first Round communication domain Aggregate proxy model weights and aggregate optimizer weights The aggregated weight parameter vector is formed. Representation domain In the The proxy model weights obtained after round-robin communication Network weights of reinforcement learning optimizer The resulting weight parameter vector , Representation domain In the The proxy model weights obtained after round-robin communication Network weights of reinforcement learning optimizer The resulting weight parameter vector ; For the first Hybrid coefficients of round-robin communication For domain The number of samples, For domain The number of samples.
[0025] Furthermore, the server performs cross-validation and historical best-in-class selection on candidate proxy models for each domain, and distributes the determined historical best-in-class proxy models to the corresponding domains, including:
[0026] The server will have a domain Candidate proxy model weights broadcast to the domain Other domains ;
[0027] By domain Use local validation sets for the domain The candidate agent model weights are validated, and the domain is calculated. Candidate agent models in the domain Cross-validation error under different data distributions;
[0028] Based on all other domain pairs The cross-validation error of the candidate surrogate model weights is calculated in the computational domain. Cross-domain weighted generalization error;
[0029] Domain based on all communication rounds The cross-domain weighted generalization error is calculated, and the communication round with the smallest cross-domain weighted generalization error is selected as the domain. The historical best round is used to determine the weights of the candidate proxy models corresponding to the historical best round as the domain. The historical best proxy model weights.
[0030] Furthermore, it also includes:
[0031] The server obtains the mean squared error of validation for each domain in the current communication round and the previous communication round, calculates the rate of change of validation error between adjacent rounds based on the magnitude of change between the two, and performs an exponential moving average on the rate of change between adjacent rounds to obtain the smoothed rate of change of validation error. This is then weighted with the mixing coefficient to determine the degree of fluctuation of the surrogate model as a reinforcement learning interaction environment, and calculates the instability score of the surrogate model.
[0032] The server performs hyperparameter adaptive tuning of the reinforcement learning optimizer based on the agent model instability score. The hyperparameter adaptive tuning includes at least adjusting the number of environmental interaction steps, learning rate, and exploration intensity of the reinforcement learning optimizer.
[0033] Compared with the prior art, the significant advantages of this invention are:
[0034] This invention proposes a cross-domain optimization architecture that coordinates the training of a proxy model and an optimizer. On the server side, the network weights of the energy consumption proxy model and the optimizer uploaded by each domain are aggregated. This allows each client domain to absorb the knowledge of energy consumption mapping from process parameters in other domains, and to inherit the effective strategy directions and experience information formed by other client domains in the optimization process. Thus, cross-domain knowledge and optimization experience sharing can be achieved without sharing the original production data, thereby improving the overall optimization efficiency.
[0035] To mitigate scale shifts and optimization instability caused by differences in cross-domain data distribution, this invention obtains a unified feature scale through boundary aggregation and maps local sub-feasible domains of each domain within a unified space. Simultaneously, the search process is constrained within "practically operable" local sub-feasible domains, achieving safer and more stable parameter search and accelerating convergence.
[0036] To address the time-varying uncertainty of the reinforcement learning interaction environment introduced by the dynamic evolution of the surrogate model with each communication round, this invention introduces an instability score and hyperparameter adaptive adjustment mechanism based on the surrogate model. This mechanism dynamically adjusts the number of interaction steps, learning rate, and exploration intensity, ensuring sufficient exploration in the early stages of training and focusing on fine convergence in the later stages. Furthermore, by combining cross-domain weighted generalization error and cumulative optimal decision-making mechanisms, the invention dynamically locks the historically optimal surrogate model with the strongest generalization ability and re-evaluates the sequence of optimization solutions from each iteration. This approach achieves more robust, reusable, and low-energy-consumption process parameter solutions while protecting privacy, reducing the optimization risks caused by insufficient data or distribution bias.
[0037] This invention achieves collaborative training of the energy consumption surrogate model and optimizer by merging multi-domain data without uploading the original data. It employs boundary aggregation and unified normalization to mitigate scale shifts caused by differences in value ranges across domains, while confining the search to local sub-feasible domains within each domain to improve optimization stability and efficiency. Furthermore, it overcomes the limitations of traditional methods that only aggregate surrogate model parameters by aggregating the network weights of the reinforcement learning optimizer on the server side, enabling cross-domain optimization experience sharing and policy transfer. It also introduces a hyperparameter adaptive adjustment mechanism based on model instability scores to dynamically adjust the learning rate, exploration intensity, and interaction steps to address time-varying uncertainties arising from the evolution of the surrogate model with each communication round. Simultaneously, it uses cross-domain weighted generalization error for cross-validation and historical best surrogate model selection, and performs backtracking evaluation and cumulative optimal decision-making on each optimization solution. This allows for more robust and reusable low-energy-consumption process parameter solutions while protecting privacy, achieving continuous, stable, and reliable reduction of energy consumption in complex electrodeposition processes. Attached Figure Description
[0038] Figure 1 The flowchart shows a cross-domain energy consumption optimization method for copper electrolysis based on the collaboration of a proxy model and an optimizer, according to the present invention.
[0039] Figure 2 R is the energy consumption proxy model of this invention. 2 Graph showing score improvement;
[0040] Figure 3 The graph shows the improvement effect of the algorithms in each domain of this invention compared to the baseline model. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0042] like Figure 1 As shown, this embodiment proposes a cross-domain energy consumption optimization method for copper electrolysis based on the collaboration of a surrogate model and an optimizer, including the following steps:
[0043] Step 1, Each Domain By deploying sensors, multi-dimensional process parameters and corresponding energy consumption indicators in the copper electrolysis production process are collected, and a local dataset is formed. ,in, For domain Local dataset, For domain The number of samples, For sample index, For domain No. A vector of process parameters for each sample. For domain No. Energy consumption indicators for each sample. Process parameters may include at least key variables such as current intensity, acid concentration, copper ion concentration, temperature, and liquid flow rate.
[0044] Step 2: To eliminate scale shifts caused by differences in cross-domain value ranges, it is necessary to normalize the vectors of each process parameter.
[0045] Step 2.1, each domain corresponds to the first 3D process parameters are calculated locally and uploaded to the server:
[0046]
[0047] in, Represents a vector of process parameters The 3D process parameters, For domain No. Minimum values of process parameters, For domain No. The maximum value of the process parameters.
[0048] Server aggregation yields the global feasible domain boundary;
[0049]
[0050] in, This represents the boundary of the globally feasible region. These are the minimum and maximum values of the global feasible region, respectively.
[0051] domain Based on this, a unified linear normalization is performed on the process parameter vector: for the domain any sample Its normalized first 3D process parameters Represented as:
[0052]
[0053] in, For normalized post-domain The The first sample Dimensional process parameters.
[0054] Step 2.2, transfer the domain The original local boundary is mapped to the normalized space to form the local sub-feasible region boundary (used to constrain the optimization search range):
[0055]
[0056] in, For domain No. The maximum value of the normalized process parameters. For domain No. The minimum value of the normalized process parameters.
[0057] Step 3: Construct energy consumption proxy models for each domain. In communication rounds Using training set The energy consumption proxy model is updated. Specifically, The training set portion of the local dataset is provided by The samples are randomly selected, and the proportion can be set by the user. The energy consumption proxy model uses a deep neural network to learn the nonlinear mapping between normalized process parameters and energy consumption. The deep neural network can be one or more combinations of multilayer perceptron networks, convolutional neural networks, recurrent neural networks, long short-term memory networks, gated recurrent unit networks, or Transformer networks. The training objective of the energy consumption proxy model uses the mean squared error loss function. :
[0058]
[0059] in, Representation domain An energy consumption proxy model is used to characterize the nonlinear mapping relationship between the normalized process parameter vector and the actual energy consumption. Represents a deep neural network. For domain The number of training set samples; This is the normalized vector of process parameters; This represents the actual energy consumption. For domain The network weight parameters to be optimized.
[0060] In the When round communication begins, the domain Based on the aggregated proxy model weights issued by the server in the previous round As initial values for local training, in the training set Execution This is the second local iteration update. When t=1, the initial local training values are the initial proxy model weights randomly initialized and distributed by the server. Specifically, the domain... Using learning rate Minimize the local mean squared error loss function and set the initial weights Updated to the weights of the candidate proxy models in this round. After training, the domain In the validation set The mean square error was calculated and verified above. :
[0061]
[0062] in, For domain In the Verification mean square error in round-robin communication; For domain The validation set is randomly sampled from the local dataset, and its samples are not the same as those in the training set. repeat; For domain The number of validation set samples; This is the vector of process parameters after unified boundary normalization; This corresponds to the actual energy consumption; Indicates the first Round communication start time domain Initial weights of the proxy model; Representation domain The weights of the candidate surrogate model obtained after this round of local training. Validation of mean squared error. Used for subsequent model stability evaluation and adaptive hyperparameter adjustment.
[0063] Step 4, Domain Updated energy consumption proxy model Considering the energy consumption evaluation function of the reinforcement learning environment, the process parameters that minimize energy consumption are searched within a local sub-feasible region of the normalized space. Its expression is as follows:
[0064]
[0065] in, Indicates in the current domain Under the constraints of the local subfeasible region, the optimal normalized process parameter vector that minimizes the predicted value of the energy consumption surrogate model is the vector of the reinforcement learning optimizer in the th... Candidate optimal process parameters obtained through round-robin communication. Represents normalized input Energy consumption proxy model, For the normalized first Dimensional process parameters.
[0066] Step 4.1: Model the optimization process as a Markov decision process: the state is the current process parameter vector, and the action is the continuous adjustment of each dimension of the process parameter.
[0067] (1) Definition of state and action:
[0068]
[0069] in, For domain In the The state under each interaction step; Representation domain In the Normalized process parameter vectors under each interaction step For domain In the Actions under each interactive step; This refers to the dimension of process parameters.
[0070] (2) State transition:
[0071] In the interactive step In the middle, the reinforcement learning optimizer adjusts according to the current state. Output continuous action This action indicates the direction and magnitude of adjustment to the current normalized process parameter vector. To avoid excessively large adjustments in a single step, an action step scaling factor is introduced. This converts the action into actual process parameter increments. Because the copper electrolysis process parameters need to meet the domain The actual operational range of production, and the updated process parameter vectors need to be further projected onto the domain. Within the local feasible region, the state for the next interaction step is obtained:
[0072]
[0073] in, For domain Interactive Steps state, This is the motion step scaling factor, used to control the adjustment range of process parameters for each motion. For domain Interactive Steps The state; This is a projection operator used to prune the updated process parameter vector dimensionally to the domain. Within the normalized local sub-feasible region interval, This represents the normalized boundary of the local subfeasible region.
[0074] (3) Energy consumption prediction:
[0075] In the interactive step In the middle, domain Current state Input the updated energy consumption proxy model, and the model will output the predicted energy consumption value corresponding to the current combination of process parameters. Due to the state... This represents a vector of process parameters in the normalized space, and therefore can be directly used as input to the energy consumption proxy model. The energy consumption prediction process is expressed as:
[0076]
[0077] in, For domain In the Round communication, the first Predicted energy consumption values under each interaction step; For domain After the first The parameters of the candidate agent model obtained after one round of local training; This represents the current normalized process parameter state vector. This predicted value is used to evaluate the energy consumption level of the current process parameter combination and further participates in the reward function calculation to guide the reinforcement learning optimizer to search for low-energy process parameter regions.
[0078] (4) Reward function: To drive energy consumption minimization while taking into account search efficiency, the reward is composed of a weighted average of the main reward, progress reward, and diversity reward:
[0079]
[0080] in, This is the reward weighting coefficient; Representation domain The total reward obtained in the current interaction step; This indicates the main reward, reflecting the current energy consumption level; This indicates a progress reward, reflecting the energy consumption improvement in adjacent steps; This represents a diversity reward, used to evaluate the variation range of recent actions across various process parameters, to avoid premature convergence caused by actions being concentrated in a small range for an extended period. (The specific function form can be set according to production needs; for example, the tanh function can be used to smoothly truncate the main reward to enhance numerical stability). In this embodiment, the reward settings are as follows:
[0081]
[0082] in, Representation domain The lowest energy consumption for the current communication round. Representation domain No. Energy consumption prediction for the interaction step. This represents the tanh(⋅) function. The main reward smoothing amplification factor is used to adjust the sensitivity of the relative improvement in energy consumption in the tanh(⋅) function, so as to avoid drastic fluctuations in the reward value due to excessive differences in energy consumption scale. d represents the number of action dimensions, typically 5 in copper electrolysis, and N is the history window, recording N consecutive actions. Indicates time The Dimensional motion components, Indicates time The Action components.
[0083] Step 4.2, Reinforcement Learning Optimizer and Environment The policy is learned through continuous interaction in several learning rounds. This learning process typically involves different networks. In this embodiment, the Proximal Policy Optimization (PPO) algorithm is used as the optimizer to optimize the process parameters. It can be replaced by any reinforcement learning algorithm, including Actor-Critic (AC), Deep Deterministic Policy Gradient (DDPG), Group Relative Policy Optimization (GRPO), Soft Actor-Critic (SAC), TwinDelayed Deep Deterministic Policy Gradient (TD3), etc.
[0084] PPO consists of two networks: an actor network and a critic network. The critic network is used to output the state value. These are the weight parameters of the actor network. Here are the weight parameters of the critic network. To constrain the policy update magnitude and prevent drastic performance degradation, a pruning alternative objective is designed:
[0085]
[0086] in Let be the policy loss function of the Actor network; This represents the expectation of all sampled interaction steps. Indicates the pruning factor in the PPO strategy update. For the first The ratio of the probability of the new policy versus the old policy for the same state-action sample at each interaction step is expressed as:
[0087]
[0088] Used to measure the effectiveness of new strategies Compared to the old strategy For the same state With action The rate of change of probability density. This represents the probability ratio. Cut to This limits the extent of strategy changes in a single update. The dominant function is designed as follows:
[0089]
[0090] in, Representation domain In the The advantage function estimate for each interaction step is used to evaluate the action. Advantages and disadvantages relative to the current average level of the strategy; For domain In the The time difference residuals of each interaction step are used to represent the deviation of the Critic network from the state value estimate in the current interaction step; For domain In the Instant rewards earned per interactive step; and These represent the Critic network's value estimates for the current state and the next state, respectively. This is a discount factor used to adjust the importance of future rewards in the current advantage estimate; The smoothing coefficient for dominance estimation is used to adjust the trade-off between single-step time difference estimation and multi-step Monte Carlo estimation; l is the cumulative index, and n is the maximum multi-step length used in dominance estimation.
[0091]
[0092] in, The value loss function of the Critic network; This represents the expectation of all sampled interaction step data; For Critic network domain In the Each interaction step state Value estimation; The value supervision labels, constructed based on the sampling rewards, are used as target values for training the Critic network. According to the definition of the advantage function, the value supervision labels can be expressed as:
[0093]
[0094] in, Representation domain In the The objective value of each interaction step. This objective value is used to supervise the Critic network so that its output state value estimate gradually approaches the actual reward level reflected by the sampled trajectory.
[0095] Furthermore, to increase the exploratory nature of the strategy and prevent it from getting trapped in local optima too early, policy entropy is introduced:
[0096]
[0097] in, For the Actor network in state The policy entropy corresponding to the action distribution; Indicates the state of the Actor network. Select action The probability density; Indicates action In state obtained by downsampling; This represents the expectation of the action distribution generated by the current strategy; This is the logarithm of the probability density of the action under the current policy; This is the policy entropy weight coefficient, used to adjust the influence of the exploration term in policy optimization. Since optimization during training is typically performed by minimizing the loss function, a negative sign is added before the policy entropy to minimize... This is equivalent to encouraging higher policy entropy, thereby enhancing action exploration capabilities. The total loss is:
[0098]
[0099] in, The total loss function for the reinforcement learning optimizer; This is the policy loss function of the Actor network, used to update the policy network weights. ; The value loss function for the Critic network is used to update the value network weights. ; This is the policy entropy loss term, used to enhance policy exploration capabilities; The policy entropy loss weights are used to jointly optimize the Actor and Critic networks, and are applied during gradient updates. Sample a batch of trajectories and keep them fixed during updates. For reference, and It is continuously updated using gradient descent; the process is repeated after the update is complete. Moving on to the next iteration update, the domain In the Execution within round communication This process involves several interactions and updates to the network weights of the reinforcement learning optimizer. The final network weights of the reinforcement learning optimizer are denoted as... The vector of candidate optimal process parameters for this round is obtained, denoted as... .
[0100] Step 5, in the When the round communication ends, the domain Upload two types of parameters to the server: (i) candidate agent model weights (ii) Network weights of the reinforcement learning optimizer The server is based on the domain. Number of samples With mixing coefficient In each domain Generate aggregation results separately:
[0101]
[0102] in, Represents the communication domain in round t. Aggregate proxy model weights and aggregate optimizer weights The aggregated weight parameter vector is formed. Representation domain The proxy model weights obtained after the t-th round of communication Network weights of reinforcement learning optimizer The resulting weight parameter vector , Representation domain Candidate agent model weights obtained after the t-th round of communication Network weights of reinforcement learning optimizer The resulting weight parameter vector , Let be the mixing coefficient of the t-th round of communication, which follows a uniform distribution from 0 to 1 and is used to characterize the strength of knowledge aggregation; For domain The number of samples, For domain The number of samples. The server will and Distribute to domain This is used for initializing the next communication round.
[0103] Step 6: Because the surrogate model is continuously updated between communication rounds, the reinforcement learning environment has time-varying characteristics. To improve training stability, this embodiment is based on the validation mean squared error. An instability score is calculated, and the number of environment interaction steps, learning rate, and exploration intensity of the reinforcement learning optimizer are adaptively adjusted accordingly. When the surrogate model instability score is high, the number of environment interaction steps, learning rate, or exploration intensity of the reinforcement learning optimizer is increased to enhance the optimizer's search capability in unstable surrogate model environments. When the surrogate model instability score is low, the number of environment interaction steps, learning rate, or exploration intensity of the reinforcement learning optimizer is decreased, allowing the optimizer to switch to refined optimization in stable surrogate model environments. This achieves an adaptive balance between the exploration capability and convergence accuracy of the reinforcement learning optimizer during the dynamic evolution of the surrogate model.
[0104] (1) Calculate the rate of change and smoothed rate of change between adjacent rounds based on the verification mean square error:
[0105]
[0106] in, For domain In the In round-robin communication, the rate of change of the verification mean squared error between adjacent rounds is used to measure the magnitude of change in the verification mean squared error of the current round relative to the previous round. For domain In the Mean square error of verification in round-robin communication; For domain In the The smoothed rate of change of verification error in round-robin communication, after exponential moving average processing, is used to reduce the impact of single-round error fluctuations on stability judgment. Representation domain In the The smoothed rate of change of verification error after exponential moving average processing in round-robin communication; The coefficient of the exponential moving average. The larger the value, the more emphasis is placed on historical trends. The smaller the value, the more attention is paid to changes in the current round.
[0107] (2) Calculate the instability score of the surrogate model based on the smoothed rate of change of the verification error and the mixing coefficient:
[0108]
[0109] in, For domain In the Stability score of the surrogate model in round-robin communication. A higher stability score indicates a smaller change in the mean square error of verification and a more stable energy consumption surrogate model; conversely, a lower score indicates that the energy consumption surrogate model fluctuates significantly between adjacent communication rounds. For domain In the The aggregation stability score in round-robin communication. Weighting the two together yields:
[0110]
[0111] in, For domain In the Instability score in round-robin communication The larger the value, the more unstable the environment, and the more exploration and step size should be increased; conversely, the smaller the value, the more refined the exploration should be. and For weight hyperparameters.
[0112] (3) Hyperparameter adaptive adjustment (example rule, upper / lower limits can be set according to actual needs):
[0113]
[0114] in, To enhance the learning optimizer's environmental interaction steps, and Set the upper and lower limits for the interaction step budget; To enhance the learning rate of the learning optimizer, and These are the upper and lower limits of the learning rate; To enhance the exploration intensity of the learning optimizer, and To explore the upper and lower limits of intensity.
[0115] By incorporating stability scores into the “surrogate model-optimizer” collaborative training framework, the instability of the reinforcement learning interaction environment caused by the continuous updates of the surrogate model between communication rounds can be quantified. Furthermore, the instability scores are adaptively correlated with the number of interaction steps, learning rate, and exploration intensity of the reinforcement learning optimizer.
[0116] Step 7: To improve cross-domain generalization reliability, the server performs cross-validation on candidate proxy models. In round-robin communication, the server will... Candidate Proxy Model Weights Broadcast to domain removal Other domains ;by domain Use its local validation set The model was validated, and the mean squared error of cross-validation was calculated:
[0117]
[0118] in, Indicates the first In round-robin communication, domain Candidate agent model In the domain The mean squared error of cross-validation obtained on the validation set; For domain The validation set; For domain The number of validation set samples; This represents actual energy consumption. (Various domains) The mean squared error of cross-validation is fed back to the server. The server then calculates the cross-domain weighted generalization error based on this.
[0119]
[0120] in, For the t-th round communication domain Cross-domain weighted generalization error, Indicates the first In round-robin communication, domain Candidate agent model In the domain The mean squared error of cross-validation obtained on the validation set.
[0121] After communication reaches the stopping condition in round T, the server is maintained. The historical sequence, and determine the historically optimal communication round:
[0122]
[0123] in: Representation domain The historical rounds that minimize cross-domain generalization error are considered as domain weights for the corresponding surrogate model weights. The historical best proxy model weights.
[0124] Step 8: After communication reaches the stopping condition T rounds, the domain... Collect the candidate optimal process parameter vector sequence obtained from each communication. .domain Use the historically best agency model (rounds) Corresponding historical best proxy model weights The sequence is back-evaluated, and the solution with the lowest predicted energy consumption is selected as the final output.
[0125]
[0126] in, For domain The final output is the optimal process parameter vector. In actual deployment, the domain can... It is denormalized back to its original dimensions and sent to the production control system for setting parameters such as current, concentration, temperature and liquid flow rate.
[0127] In this embodiment, considering that the energy consumption proxy models of each domain have not yet fully converged and the prediction error is relatively large in the early stage of cross-domain communication, some candidate process parameter vectors may obtain abnormally low predicted energy consumption values due to the error of the proxy model. If directly used for production control, it will lead to the problem that the actual energy consumption cannot be reproduced. The present invention adopts a backtracking evaluation and cumulative optimal decision-making mechanism: the server performs cross-validation on the candidate proxy models of each round based on the cross-domain weighted generalization error and determines the historical optimal proxy model, and distributes the historical optimal proxy model to the corresponding domain; each domain uniformly substitutes the candidate optimal process parameter vector sequence generated in each communication round into the historical optimal proxy model for re-evaluation. Under the same high-precision and strong generalization evaluation benchmark, all candidate solutions are compared horizontally, and the process parameter vector with the lowest predicted energy consumption is selected as the final output, thereby effectively eliminating the non-true low energy consumption candidate solutions caused by the early model error and improving the reproducibility and reliability of the output parameters.
[0128] This invention provides a cross-domain energy consumption optimization method for copper electrolysis based on the collaboration of a surrogate model and an optimizer. In each round of communication: each domain trains an energy consumption surrogate model on its local private dataset after unified boundary normalization, updating the surrogate model weights with the objective of minimizing the mean square error between predicted and actual energy consumption; a local reinforcement learning optimizer is constructed using the updated energy consumption surrogate model as the interaction environment, and continuous actions are performed to optimize within the local sub-feasible domain obtained by boundary mapping to obtain candidate optimal process parameter vectors, while simultaneously updating the optimizer network weights; each domain uploads the candidate surrogate model weights and the optimizer network weights to the server, which then performs optimization. A cross-domain aggregation strategy incorporating hybrid coefficients is adopted to weight and aggregate the surrogate model weights and optimizer network weights separately and distribute them to achieve cross-domain optimization strategy sharing. Furthermore, the model instability score constructed based on the validation mean square error sequence is used to adaptively adjust the hyperparameters of reinforcement learning interaction steps, learning rate, and exploration level. Finally, through cross-validation of cross-domain weighted generalization error and historical optimal screening, as well as backtracking and re-evaluation of candidate parameter sequences, a low-energy-consumption optimal process parameter vector that can be directly used for production control is determined. This improves the generalization ability, convergence stability, and collaborative optimization efficiency in cross-domain scenarios without sharing the original data.
[0129] In one specific embodiment, this embodiment is verified using on-site data from a copper electrolysis manufacturing company, as detailed below:
[0130] Using cell voltage as the energy consumption output, five process parameters (current, copper concentration, acid concentration, temperature, and liquid flow rate) are adjusted to construct an input feature vector, realizing a 5-dimensional input to 1-dimensional output energy consumption proxy model. The environment is set into two domains, divided according to the electrolytic cell number: samples from electrolytic cells 1–24 are assigned to domain 0, and samples from electrolytic cells 25–48 are assigned to domain 1, to characterize cross-domain distribution characteristics; the server aggregates the minimum / maximum values of features in the training sets of each domain to obtain a unified maximum and minimum value, and performs unified normalization processing on the corresponding feature columns of the dataset.
[0131] For training the energy consumption proxy model, the global communication rounds were set to 175 rounds, with 5 local iterations per domain per round and a local batch size of 128. Data was partitioned within each domain into training / validation / testing layers at a ratio of 0.6 / 0.2 / 0.2, and 5-fold cross-validation was used to improve evaluation reliability. The energy consumption proxy model adopted a CNN architecture containing convolutional and fully connected layers: the first fully connected layer fc1: 7→32 (ReLU), followed by two 1D convolutional layers conv1: 1→16, k=3, pad=1, conv2: 16→32, k=3, pad=1 (both containing ReLU and MaxPool), flattened, then passed through fc2: 32×30→128 (ReLU+Dropout), fc3: 128→32 (ReLU), and finally fc4: 32→1 to output the predicted slot voltage value; the loss function used was mean squared error (MSE). The local training optimizer used Adam with a learning rate of 2×10⁻⁶. -6 .
[0132] In terms of downstream process parameter optimization, reinforcement learning optimizers such as SAC, PPO, DDPG, TD3, TRPO (Trust Region Policy Optimization), AC, and GRPO are used to collaboratively optimize process parameters during communication training iterations. The final output is the combination of process parameters that minimizes the tank voltage, in order to verify the optimization performance and cross-domain generalization ability of the proposed framework under real industrial data.
[0133] like Figure 2 As shown, the generalization difference between independently trained models and the surrogate models of this method under the centralized framework is illustrated. While independently trained models (dashed lines) show improved performance initially, R² decreases with increasing training duration. 2The scores show significant differences, exhibiting obvious underfitting in domain 0 and obvious overfitting in domain 1. This is attributed to the excessive reliance of a single device on local private features under non-independent and identically distributed data. In contrast, our method, through a cross-domain aggregation strategy, allows the cross-domain distribution features of domain 1 to help correct the optimization direction of domain 0, effectively mitigating the risk of a single model getting trapped in local extrema. This ensures that the global aggregated model (solid line) remains highly stable even in the later stages of communication. This result confirms the superiority of selecting the globally optimal surrogate model based on cross-domain weighted generalization error: this mechanism can dynamically filter overfitting noise generated during the local optimization process, locking in and retaining the parameter state with the strongest generalization ability as a robust physical mapping benchmark, thereby ensuring that the method possesses excellent cross-domain generalization ability and long-term optimization stability under complex industrial conditions.
[0134] Table 1 presents the performance of the proposed method, comparing it with reinforcement learning optimizers within a centralized framework (including SAC, PPO, DDPG, TD3, TRPO, AC, and GRPO), traditional optimization algorithms (including BO (Bayesian Optimization), GA (Genetic Algorithm), PSO (Particle Swarm Optimization), SA (Simulated Annealing), SGD (Stochastic Gradient Descent), and WOA (Whale Optimization Algorithm)), and the FBO (Federated Bayesian Optimization) algorithm within a federated learning framework. The FBO algorithm utilizes a Gaussian process (GP) as a surrogate model and federates the mean and covariance of the GP to achieve cross-domain optimization. Experiments compare the obtained optimal process parameters against the optimal surrogate model, using five-fold cross-validation to ensure the reliability of the results and compare their performance across different domains. For convenience, this method is represented by RLCDO (Reinforcement Learning-Based Cross-Domain Optimizer) to distinguish various reinforcement learning optimizers, including AC-CDO (Actor-Critic Cross-Domain Optimization Algorithm), DDPG-CDO (Deep Deterministic Policy Gradient Cross-Domain Optimization Algorithm), GRPO-CDO (Group Relative Policy Cross-Domain Optimization Algorithm), SAC-CDO (Soft Actor-Critic Cross-Domain Optimization Algorithm), and TD3-CDO (Dual Delay Deep Deterministic Policy Gradient Cross-Domain Optimization Algorithm). TOA-Average represents the mean of the TOA algorithm, RL-Average represents the mean of the RL algorithm, and RLCDO-Average represents the mean of the RLCDO algorithm. Experimental results are shown in Table 1, demonstrating that RLCDO exhibits significant advantages in handling cross-domain data. Overall, the RLCDO algorithm performs best in the slot voltage optimization task, with average TOA values of 3.3464±0.0718V (domain 0) and 3.2693±0.0763V (domain 1), while its average TOA values for RL (reinforcement learning) are 3.3296±0.0667V (domain 0) and 3.26815±0.0764V (domain 1).In comparison, RLCDO's average values in domain 0 and domain 1 are 3.2680±0.0551V and 3.1819±0.0821V, respectively, which are significantly lower than the average values of TOA and RL. This indicates that RLCDO is more effective at reducing voltage during the optimization process. Among them, the AC algorithm achieved even lower voltage values of 3.2646V and 3.1773V.
[0135] Table 1 shows the comparison results between the present invention and various optimizers.
[0136]
[0137] To further analyze the performance improvement of the optimization algorithms, the maximum value of the optimizer in each domain was used as a benchmark, and the result of each optimizer was compared with this benchmark value. A bar chart was created to represent the improvement of each algorithm relative to the baseline, such as... Figure 3 As shown, RLCDO in domain 0 ( Figure 3 (a) and domain 1 ( Figure 3 Both (b) and (c) showed good performance, with an overall performance improvement of over 0.1V, but some differences still exist between them. The TOA and RL optimization algorithms performed relatively poorly, highlighting the limitations of centralized frameworks in handling cross-domain data. In contrast, FBO and RLCDO can effectively share optimization experience across domains, demonstrating strong optimization capabilities. However, the FBO algorithm still has certain limitations; it only shares the mean and covariance parameters in the Gaussian process (GP) surrogate model without fully considering the optimization needs between different domains, resulting in performance inferior to RLCDO.
[0138] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration, characterized in that, include: Each domain collects multi-dimensional process parameter vectors and corresponding actual energy consumption during the copper electrolysis production process to form a local dataset; Each domain calculates the local boundary of each dimension of the process parameters in the local dataset. The server aggregates the local boundaries of each domain. Each domain normalizes the process parameter vector based on the global feasible domain boundary after boundary aggregation and maps the local boundary to the normalized space to obtain the local sub-feasible domain boundary. In each round of communication, each domain trains and updates the energy consumption proxy model on the local dataset to obtain the weights of the candidate proxy models; Each domain uses the updated energy consumption proxy model as the interaction environment, constructs a reinforcement learning optimizer, and optimizes process parameters within the boundary of the local sub-feasible domain to obtain the candidate optimal process parameter vector under this round of communication, and updates the network weights of the reinforcement learning optimizer simultaneously. The server performs weighted aggregation of the candidate agent model weights and reinforcement learning optimizer network weights for each domain, generates aggregated agent model weights and aggregated optimizer weights for each domain, and sends them to the corresponding domain as initialization parameters for the next communication round. The server performs cross-validation and historical best-in-class selection on candidate proxy models for each domain, and then distributes the determined historical best-in-class proxy models to the corresponding domains. Each domain substitutes the candidate optimal process parameter vectors from previous communications into the historical optimal proxy model for backtracking evaluation, and selects the process parameter vector with the lowest predicted energy consumption as the optimal process parameter vector for final application in production control.
2. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, The energy consumption proxy model is constructed using a deep neural network, and the mean squared error loss function is minimized in each domain using the Adam optimization algorithm.
3. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, Each domain trains and updates its energy consumption proxy model on the local dataset to obtain candidate proxy model weights, including: Each domain uses the aggregated proxy model weights issued by the previous round of communication server as the initial values for local training, and executes training on the training set. The local iterative update aims to minimize the mean square error between predicted and actual energy consumption, thereby updating the surrogate model weights and obtaining the candidate surrogate model weights.
4. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, The reinforcement learning optimizer models the process parameter optimization process as a Markov decision process, where the state is a vector of process parameters and the action is the adjustment amount of each dimension of the process parameters, and the action satisfies the following conditions: ; After an action is performed, the state transition is performed according to the following rules: ; in, For domain No. The state of the interactive step, For domain No. The state of the interactive step, For projection operators, This represents the normalized boundary of the local subfeasible region. This is the scaling factor for the action step size. For domain No. The action of interactive step, For the dimensions of process parameters; The reward function consists of a main reward, an improvement reward, and a diversity reward weighted according to preset reward weights. The main reward is calculated based on the relative difference between the current predicted energy consumption and the lowest predicted energy consumption of the current communication round, using a hyperbolic tangent function. The improvement reward is determined based on whether the predicted energy consumption of the current interaction step is lower than the predicted energy consumption of the previous interaction step. The diversity reward is calculated based on the deviation between each action component in a preset historical window and the mean of the corresponding window.
5. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, The server performs weighted aggregation of the candidate agent model weights and reinforcement learning optimizer network weights for each domain, generating aggregated agent model weights and aggregated optimizer weights for each domain, expressed by the formula: in, Indicates the first Round communication domain Aggregate proxy model weights and aggregate optimizer weights The aggregated weight parameter vector is formed. Representation domain In the The proxy model weights obtained after round-robin communication Network weights of reinforcement learning optimizer The resulting weight parameter vector , Representation domain In the The proxy model weights obtained after round-robin communication Network weights of reinforcement learning optimizer The resulting weight parameter vector ; For the first Hybrid coefficients of round-robin communication For domain The number of samples, For domain The number of samples.
6. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, The server performs cross-validation and historical best-in-class selection on candidate proxy models for each domain, and distributes the determined historical best-in-class proxy models to the corresponding domains, including: The server will have a domain Candidate proxy model weights broadcast to the domain Other domains ; By domain Use local validation sets for the domain The candidate agent model weights are validated, and the domain is calculated. Candidate agent models in the domain Cross-validation error under different data distributions; Based on all other domain pairs The cross-validation error of the candidate surrogate model weights is calculated in the computational domain. Cross-domain weighted generalization error; Domain based on all communication rounds The cross-domain weighted generalization error is calculated, and the communication round with the smallest cross-domain weighted generalization error is selected as the domain. The historical best round is used to determine the weights of the candidate proxy models corresponding to the historical best round as the domain. The historical best proxy model weights.
7. The method for cross-domain energy consumption optimization in copper electrolysis based on surrogate model and optimizer collaboration as described in claim 1, characterized in that, Also includes: The server obtains the mean squared error of validation for each domain in the current communication round and the previous communication round, calculates the rate of change of validation error between adjacent rounds based on the magnitude of change between the two, and performs an exponential moving average on the rate of change between adjacent rounds to obtain the smoothed rate of change of validation error. This is then weighted with the mixing coefficient to determine the degree of fluctuation of the surrogate model as a reinforcement learning interaction environment, and calculates the instability score of the surrogate model. The server performs hyperparameter adaptive tuning of the reinforcement learning optimizer based on the agent model instability score. The hyperparameter adaptive tuning includes at least adjusting the number of environmental interaction steps, learning rate, and exploration intensity of the reinforcement learning optimizer.
Citation Information
Patent Citations
Energy conservation and emission reduction scheme generation method and system based on digital factory
CN120259021A
Fusion plasma loop voltage control method based on reinforcement learning
US20260196367A1