A method for determining the adjustable power range of an aluminum electrolytic cell
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2025-11-12
- Publication Date
- 2026-06-30
AI Technical Summary
Traditional aluminum electrolytic cells have limited adjustment capabilities and cannot quickly respond to changes in electricity prices or grid demand. Existing control methods lack physical models and data support, and are costly, making them difficult to promote quickly among different aluminum plants.
A multiphysics model was established using CFD simulation, and a dual-agent model and a multi-agent reinforcement learning system were constructed. The simulation model simulated the temperature, voltage, and concentration distribution, and combined with safety and economic function models, the power adjustment range of the aluminum electrolysis cell was optimized.
It achieves precise control of the power adjustment range of aluminum electrolytic cells, ensuring safety and economy, significantly widening the adjustable range, flexibly responding to dynamic working conditions, and improving operational safety margin and production efficiency.
Smart Images

Figure CN121351695B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aluminum electrolysis technology, and in particular to a method for determining the adjustable range of power in aluminum electrolysis cells based on CFD simulation and multi-agent reinforcement learning dual-agent game optimization. Background Technology
[0002] Traditional aluminum electrolysis cells typically operate with a constant current, relying on strict thermoelectric balance to maintain cell stability. Their adjustability is very limited; any significant fluctuation in current input or excessively high temperature can disrupt the thermal balance within the cell, easily leading to over-melting or thickening of the furnace lining, which in turn can cause aluminum solidification and damage to the cell lining, posing safety risks. To avoid such problems, current production systems often strictly limit the fluctuation range of single-cell power, and the adjustment process requires planning and execution several hours in advance, making it difficult to quickly respond to changes in electricity prices or grid demand.
[0003] Existing technologies attempt to improve the regulation capability of electrolyzers, for example by adding zoned air cooling, heat exchange devices, or intelligent control systems, to achieve short-term operation within a ±20% power range and maintain temperature balance through feedback control. However, these methods have the following problems:
[0004] 1) The heat-electricity-mass transport processes within the electrolytic cell are highly coupled, making it difficult to accurately determine the safe operating range at different power levels based on experience. Existing control methods are mostly based on fixed rules or manual adjustments, lacking power upper and lower limit calculation methods supported by physical models and data.
[0005] 2) Existing solutions for improving regulation capabilities often require the installation of additional equipment (such as zoned cooling, forced heat exchangers, etc.), which are costly, difficult to deploy, and highly dependent on the specific electrolytic cell structure, making it impossible to quickly promote them among different aluminum plants.
[0006] 3) Most current methods for determining the power adjustment range of electrolytic cells use fixed thresholds, static constraint optimization, or single calculation results, which cannot continuously learn and adjust the upper and lower limits of power, making it difficult to cope with dynamic operating conditions such as electricity price fluctuations, raw material changes, and equipment aging.
[0007] In summary, most of these methods rely on additional hardware investment, resulting in high costs and complex systems, making them unsuitable for large-scale deployment. Furthermore, the relevant control strategies are largely based on empirical rules and lack theoretical support, making it impossible to flexibly determine the safe and adjustable power range under different raw material properties and production objectives. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides a method for determining the adjustable power range of an aluminum electrolytic cell.
[0009] The technical solution of the present invention is: a method for determining the adjustable range of power in an aluminum electrolysis cell, comprising the following steps:
[0010] S1) Establish a multiphysics CFD simulation model for the electrolyzer, and use the CFD simulation model to solve the distribution characteristics of temperature, voltage and concentration of the electrolyzer under different input power P as a training dataset;
[0011] S2) Construct a dual-agent model that includes a tank safety function model and a production economy function model, and use the training dataset to train supervised learning to train the dual-agent model;
[0012] S3) Construct a multi-agent reinforcement learning system to solve for the minimum derating power of the aluminum electrolytic cell. With maximum upload power .
[0013] Preferably, in step S1), the simulation objective of the CFD simulation model is to simulate the temperature distribution, electric field distribution, heat generation and loss of the electrolytic cell under different input electric power P conditions, and to determine whether a thermal equilibrium state has been reached.
[0014] Preferably, in step S2), the cell safety function model is used to evaluate the thermal safety margin of the electrolytic cell under different input power P; the production economy function model is used to predict the output and unit energy consumption under different power.
[0015] Preferably, in step S2), the tank safety function model... The criterion for determining whether the tank meets the safety boundaries for heat, electricity, and concentration is expressed as follows:
[0016] ; (11)
[0017] In the formula, This is the temperature deviation term; This is the voltage deviation term; This is the alumina concentration deviation term; , , These are the weighting coefficients for the corresponding deviation terms; , , The allowable deviation threshold for the corresponding deviation item.
[0018] Preferably, in step S2), the production economy function model is described. The larger the value, the better the economic benefits. (This refers to the production economics function model.) Represented as:
[0019] ; (13)
[0020] In the formula, , Target output and target unit power consumption, respectively; , This is a weighting parameter used to reflect the relative importance of output and energy consumption; The actual output under power P; This represents the actual unit power consumption under power P.
[0021] Preferably, in step S3), the multi-agent reinforcement learning system includes a safety agent A and an economic agent B, wherein the safety agent A is used to determine whether a given power P violates the thermal safety boundary, and the economic agent B is used to select the power P with the greatest economic benefit within the safety constraints.
[0022] Preferably, in step S3), the security intelligent agent... Input state space for: Action output for: Reward function for: ;
[0023] The aforementioned economic intelligent agent Input state space for: Action output for: Reward function for: ;
[0024] in, , The outputs of the security agent and the economic agent, respectively, are predefined discrete sets. The specific action executed is determined by the policy network. According to the status The output action probability is used to determine this.
[0025] Preferably, in step S3), the security intelligent agent... By minimizing As a reward, namely Guide the agent to gradually approach the lower power limit. ;
[0026] The aforementioned economic intelligent agent By maximizing the reward function, i.e. Guide the intelligent agent to gradually approach the power limit. .
[0027] Preferably, in step S3), from the safe power point... We begin by searching for power points that satisfy the following conditions:
[0028] The security agent satisfies the following constraints: ;
[0029] Economic agency achieves stable optimality: ;
[0030] If the conditions are met, it is considered that a power extremum on one side has been found, and the power point that satisfies the conditions is recorded, i.e.:
[0031] ;
[0032] Select the extreme value interval from them. The final output is an adjustable power range for the electrolytic cell that is within a safe range and has good economic performance.
[0033] The beneficial effects of this invention are as follows:
[0034] 1. This invention accurately depicts the thermal field, electric field, and concentration distribution within the aluminum electrolytic cell using CFD simulation. Combined with a safety proxy model, it real-time assesses whether key indicators such as temperature, voltage, and concentration exceed limits, ensuring that the adjustment range remains within safe boundaries. It can quantitatively assess the thermal balance state under different power levels, promptly identifying and preventing dangerous conditions such as furnace wall melting and aluminum molten solidification, significantly improving operational safety margins. When the power test approaches the safety limit, the system automatically penalizes the over-limit behavior and reverts to a lower limit, ensuring that the electrolytic cell is not burned through and the lining does not overheat, resulting in higher safety.
[0035] 2. This invention takes into account both aluminum liquid production and energy consumption targets, incorporates production revenue and energy costs into the optimization consideration, evaluates the output and unit power consumption under different power levels through an economic proxy model, and sets dynamic weights with electricity price and aluminum ingot price as parameters to achieve quantitative measurement of economic benefits.
[0036] 3. The reinforcement learning agent of this invention autonomously seeks the power point with the maximum output / lowest energy consumption under the premise of ensuring safety, ensuring that the total output is not lower than the planned requirements and the unit power consumption does not exceed the standard;
[0037] 4. This invention significantly expands the adjustable range of the upper and lower limits of the power of aluminum electrolytic cells, allowing for a greater range of load increases and decreases based on grid demand while maintaining process stability. This invention uses dual-agent game theory to intelligently search for the minimum load reduction power and maximum load increase power of the electrolytic cell, significantly expanding the available adjustment range. When the grid is in a trough, the system can safely increase power to approach the upper limit for increased production, and quickly reduce power to the lower limit to reduce load and peak load during peak periods, truly transforming the aluminum electrolytic cell into a flexible adjustment resource. Attached Figure Description
[0038] Figure 1 This is a flowchart illustrating the method of the present invention. Detailed Implementation
[0039] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:
[0040] like Figure 1 As shown in the figure, this embodiment provides a method for determining the adjustable range of power in an aluminum electrolysis cell, including the following steps:
[0041] S1) Establish a multiphysics CFD simulation model for electrolyzers, and use the CFD simulation model to simulate the distribution characteristics of temperature, voltage and concentration of electrolyzers under different input power P as a training dataset;
[0042] The CFD simulation model simultaneously considers electric field distribution, Joule heat generation, reaction heat release, and heat and mass transfer processes. The simulation region of the CFD simulation model covers the following structural levels within the electrolyzer:
[0043] • Upper anode and air gap region;
[0044] • The molten electrolyte layer is mainly a mixture of cryolite and alumina;
[0045] • Liquid aluminum layer;
[0046] • Carbon cathode block and lower furnace lining;
[0047] • Surrounding furnace shell and insulation layer, used for heat dissipation simulation;
[0048] • Top interface with air contact.
[0049] The CFD simulation model aims to simulate the temperature distribution, electric field distribution, heat generation and dissipation of an electrolytic cell under different input power P conditions, and determine whether a thermal equilibrium state can be reached. Specifically, it includes the following steps:
[0050] S11) Establish the electric field solution model and calculate the Joule heat source based on the input power P. ;
[0051] In the electrolytic cell, current is introduced from the anode, passes through the electrolyte and molten aluminum, and reaches the cathode. Within the electrolytic cell, Ohm's law and Poisson's equation are satisfied; therefore, the electric field solution model is expressed as:
[0052] (1)
[0053] in, Let be the potential distribution function. These are spatial coordinate variables, representing the horizontal, vertical, and lateral position coordinates in the three-dimensional computational domain, used to describe the spatial position of any point inside the electrolytic cell; Represented as the gradient operator, it is used to find the spatial derivative of a scalar field (such as a potential distribution), and its function is to calculate the spatial rate of change of electric field intensity or temperature gradient, etc. The conductivity;
[0054] Boundary conditions: Current density applied to the anode busbar Cathode grounding, potential distribution function The sidewalls are insulated boundaries. ; Represented as a unit normal vector, it is used to describe the direction of the boundary surface; in boundary conditions This indicates an insulating boundary, meaning the current density is zero in the normal direction of the boundary. denoted as current density.
[0055] Calculation of Joule heat source based on current density ,Right now:
[0056] ; (2)
[0057] in, Represented as an electric field intensity vector, It is defined as the negative value of the potential gradient.
[0058] S12) Calculate the heat released by the electrolysis reaction based on the reaction rate. ,Right now:
[0059] ; (3)
[0060] in, The heat of reaction is expressed as the heat of decomposition of a single mole of alumina. For the reaction rate, For Ferrari's constant, The magnitude of the current; The volume of the control volume, i.e., the volume of the CFD mesh cells, is used to convert the total heat into local heat source intensity.
[0061] S13) Establish heat conduction and convection models and calculate temperature field distribution. ;
[0062] Heat in the electrolytic cell is mainly transferred through conduction and natural convection. Therefore, the heat conduction and convection model is expressed as follows:
[0063] ; (4)
[0064] in, Temperature field distribution; The thermal conductivity of each region; This is the heat released during the electrolysis reaction.
[0065] S14) Establish the boundary heat transfer model and calculate the total heat loss. ;
[0066] In CFD simulation models, the boundary heat transfer model is used to describe the heat exchange mechanism between the electrolyzer and its external environment, and is a key factor in determining whether thermal equilibrium can be reached. The heat input of the electrolyzer mainly comes from the heat released by the electrolysis reaction. and Joule heat source items Heat output is mainly achieved through heat transfer losses at the top, side walls, and bottom. Thermal boundary conditions need to be set for each boundary surface, as follows:
[0067] Heat transfer boundary conditions at the top of the electrolytic cell:
[0068] (5)
[0069] ;
[0070] in, This indicates the heat loss at the top of the electrolytic cell; The coefficient of heat transfer by natural convection at the top of the electrolytic cell reflects the air cooling effect; The ambient air temperature at the top of the electrolytic cell; This refers to the temperature of the outer surface of the top of the electrolytic cell; Surface emissivity represents the ability of a material's surface to absorb and emit infrared radiation; It is the Stefan-Boltzmann constant; It is thermal radiation heat flow; Indicates the top; Represented as a temperature distribution function, it represents any spatial coordinate within the computational domain of the electrolytic cell. Temperature value at that location; This represents the integral over the area of the top, sidewall, or bottom of the electrolytic cell;
[0071] The heat transfer boundary condition equation at the top of the electrolytic cell indicates that the heat loss at the top is accomplished through two mechanisms: natural cooling by the air and infrared thermal radiation from the surface of the electrolyte or covering material. Specifically, the left side represents the energy on the top surface of the electrolytic cell, and the right side represents the energy dissipated to the environment by convection and radiation. The equality of the two is the boundary condition for energy conservation.
[0072] Heat transfer boundary conditions of the sidewall of the electrolytic cell:
[0073] ; (6)
[0074] ;
[0075] in, This indicates the heat loss from the sidewalls of the electrolytic cell; Indicates the side view; The heat transfer coefficient of the sidewall of the electrolytic cell; This refers to the surface temperature of the sidewall of the electrolytic cell. The ambient air temperature of the sidewall of the electrolytic cell;
[0076] Heat transfer boundary conditions at the bottom of the electrolytic cell:
[0077] ; (7) ;
[0078] in, This indicates the heat loss at the bottom of the electrolytic cell; Indicates the bottom; Indicates the heat transfer coefficient at the bottom of the electrolytic cell; This refers to the temperature of the lower surface of the electrolytic cell. This refers to the air temperature at the bottom of the electrolytic cell.
[0079] The total heat loss mentioned above Represented as:
[0080] ; (8)
[0081] This represents the total heat loss through the top cover, side walls, and bottom.
[0082] S15) Determine the input power based on total heat absorption and total heat loss. Does thermal equilibrium exist?
[0083] ; (9)
[0084] ; (10)
[0085] in, The heat of reaction is the sum of the heat of reaction and the endothermic heat generated by the formation of molten aluminum;
[0086] This embodiment uses a set input power. Under these conditions, temperature field If the input power remains stable within the numerical convergence range, it is considered to be the input power. It can achieve thermal equilibrium and save relevant data: ;in, This indicates the power input. Below, the state vector of the electrolytic cell is a set of key physical quantities that describe the current operating state of the system; Indicates the first One input power value; Indicates the temperature deviation term; Indicates the voltage deviation term; This indicates the alumina concentration deviation term.
[0087] S2) Construct a dual-proxy model that includes a tank safety function model and a production economy function model, and train the dual-proxy model using the training dataset;
[0088] In this embodiment, a dual-proxy model is constructed using machine learning methods. The cell safety function model is used to evaluate the thermal safety margin of the electrolytic cell under different input power P; the production economy function model is used to predict the output and unit energy consumption under different power.
[0089] In this embodiment, the tank safety function model The criterion is whether the electrolytic cell body meets the safety boundaries for heat, electricity, and concentration. This indicates that the electrolytic cell is operating within a safe range at the current power P, and the aforementioned cell safety function model... The expression is:
[0090] ; (11)
[0091] In the formula, This is the temperature deviation term; This is the voltage deviation term; This is the alumina concentration deviation term; , , These are the weighting coefficients for the corresponding deviation terms; , , The allowable deviation threshold for the corresponding deviation item.
[0092] in:
[0093] The temperature deviation item is:
[0094] ; (11-1)
[0095] In the formula, This represents the ideal melt temperature at input power P. , These represent the maximum and minimum temperature thresholds that are allowed for operation of the electrolytic cell, respectively.
[0096] The voltage deviation term is:
[0097] (11-2)
[0098] In the formula, This represents the ideal voltage for electrolytic aluminum at input power P. These represent the maximum and minimum allowable voltage thresholds for the operation of the electrolytic cell, respectively.
[0099] Alumina concentration deviation item:
[0100] ; (11-3)
[0101] In the formula, This indicates the concentration of aluminum in the molten alumina raw material when the input power is P; These represent the concentration thresholds.
[0102] The allowable deviation thresholds are respectively expressed as:
[0103] ; (11-4)
[0104] In this embodiment, the judgment matrix is solved. The eigenvector corresponding to the largest eigenvalue is used to obtain the weight coefficients. ,Right now:
[0105] ; (11-5)
[0106] In the formula, Indicates the transpose operation; This represents the vector normalization operation, making... ; The matrix represents the judgment matrix. The eigenvector corresponding to the largest eigenvalue is used to reflect the relative importance of each safety indicator in the comprehensive safety evaluation function.
[0107] In this embodiment, a judgment matrix is constructed by comparing the importance of the temperature deviation term, voltage deviation term, and electrolyte concentration deviation term pairwise. ,Right now:
[0108] ; (11-6)
[0109] In the formula, This indicates the relative importance of temperature factors compared to voltage factors; This indicates the relative importance of temperature factors compared to concentration factors; This indicates the importance of the voltage factor relative to the concentration factor.
[0110] In this embodiment, to ensure the rationality of the judgment matrix, the judgment matrix is... Perform a consistency ratio test, i.e.:
[0111] ; (11-7)
[0112] In the formula, It is the largest eigenvalue; It is a random consistency indicator; 3 represents the order of the judgment matrix C; The Consistency Index measures the consistency of the judgment matrix. With a deviation that is completely consistent, when When, determine the matrix The consistency is acceptable, and the weighting results are valid.
[0113] In this embodiment, a supervised learning method is used to model the tank safety function. Training is conducted, where the inputs are temperature, voltage, and concentration deviations at different power levels. The output is the tank safety function model. The comprehensive security function value, namely:
[0114] ;(12)
[0115] In the formula, This represents a security agent model trained based on machine learning algorithms, used to establish power state features. With tank safety function model The mapping relationship between them was established, and the model was trained through supervised learning to accurately predict the tank safety index under different power inputs.
[0116] In this embodiment, the production economy function model is described. The larger the value, the better the economic benefits. (This refers to the production economics function model.) Represented as:
[0117] ; (13)
[0118] In the formula, , These are the target output and the target unit power consumption, respectively. , This is a weighting parameter used to reflect the relative importance of output and energy consumption; The actual output under power P; This represents the actual unit power consumption under power P.
[0119] Among them, input power Actual output Defined as:
[0120] ; (13-1)
[0121] In the formula, Current efficiency varies depending on factors such as power and temperature field; The molar mass of aluminum; For the number of electrons, Represents Ferrari's constant; For slot current; This refers to the scheduling cycle time.
[0122] power Actual unit power consumption for:
[0123] ; (13-2)
[0124] in, Input power Actual output; This refers to the scheduling cycle time.
[0125] In this embodiment, the weighting parameter , ; Electricity price is based on time of day; The unit marginal profit; This refers to the unit price of aluminum ingots; This refers to the variable cost per unit other than electricity.
[0126] Normalization to 1 indicates that the economic benefits of increased production are used as the benchmark. This reflects the economic penalty coefficient per unit of electricity consumption; when electricity prices rise or the target per unit of electricity consumption increases... As electricity prices increase, the model automatically pays more attention to energy consumption constraints; when electricity prices decrease... The model tends to increase output when the output is reduced.
[0127] By substituting equations (13-1) and (13-2) into equation (13), we obtain the production economy function model. The expression is:
[0128] ; (13-3)
[0129] Further analysis reveals:
[0130] ; (13-4)
[0131] As can be seen from the above equation, the production economy function model It is affected by voltage, power, current efficiency, electricity price, aluminum ingot unit price, marginal profit, and target output. When output... and unit power consumption They just met expectations. , At that time, the lower limit of the economy .
[0132] This embodiment also employs supervised learning fitting. The input data includes: voltage, power, current efficiency, electricity price, aluminum ingot unit price, marginal profit, and target output. These data are obtained through the CFD simulation model in step S1) and manually set indicators. The output is:
[0133] ;(14)
[0134] In the formula, This represents an economic agent model trained based on machine learning algorithms, used to establish power. Current efficiency ,Voltage Electricity price Unit marginal profit and target output The mapping relationship between the value of the comprehensive economic function and the value of the electrolyzer is obtained through supervised learning, which can accurately evaluate the economic operating efficiency of the electrolyzer under different power inputs.
[0135] S3) Construct a multi-agent reinforcement learning system to solve for the minimum derating power of the aluminum electrolytic cell. With maximum upload power ;
[0136] In this embodiment, the multi-agent reinforcement learning system includes a security agent. and economic intelligent agents The security intelligent agent mentioned above The economic agent is used to determine whether a given power P violates the thermal safety boundary and to select the power P with the greatest economic benefit within the safety constraints.
[0137] The aforementioned security intelligent agent Input state space for: Action output for: Reward function for: ;
[0138] The aforementioned economic intelligent agent Input state space for: Action output for: Reward function for: ;
[0139] in, , The outputs of the security agent and the economic agent, respectively, are predefined discrete sets. The specific action executed is determined by the policy network. According to the status The output action probability is used to determine this.
[0140] In this embodiment, the multi-agent reinforcement learning system updates power and state through environmental feedback and state update mechanisms, namely:
[0141] ; (15)
[0142] ; (16)
[0143] ; (17)
[0144] ; (18)
[0145] ; (19)
[0146] In the formula, express Power at any given moment; The input power is respectively Temperature deviation, voltage deviation, and alumina concentration deviation; The unit marginal profit; for Time-of-use electricity pricing; For input power is Current efficiency at that time; For input power is Ideal voltage of the electrolytic cell at that time; Indicates power input Below, based on the tank safety function model The predicted safety assessment function value reflects whether the electrolyzer's thermal stability, voltage, and concentration are within a safe range under the current power input. The smaller the value, the safer the system. Indicates power input Below, from the production economy function model The predicted economic function value reflects the combined benefits of output and unit power consumption at that power level. The higher the value, the greater the economic benefits.
[0147] In this embodiment, the security intelligent agent Evaluation function of tank safety status With the core, through minimization As a reward, namely ;Right now The smaller the better, meaning the smaller the better for safety. Therefore, a negative sign is added to penalize unsafe overreaching behavior caused by excessive power, guiding the agent to gradually approach the lower power limit. .
[0148] The aforementioned economic intelligent agent Production economic comprehensive function With this as the core, by maximizing the reward function, that is... In other words, the larger the better, rewarding higher returns and guiding intelligent agents to gradually approach their power limit. .
[0149] In this embodiment, a policy training mechanism based on the Proximal Policy Optimization (PPO) method is used to train a multi-agent reinforcement learning system, with each agent independently training a policy network. AND-valued function network The training process is as follows:
[0150] Construct the policy objective function, namely:
[0151]
[0152] ; (20)
[0153] In the formula, Let the objective function be the policy objective function; The expression represents the expected value of the sample at each time step during the reinforcement learning training process, which is equivalent to the average value over all time steps in a complete training sequence. This represents the total number of time steps within a training sequence (episode), used to calculate the average policy reward of the agent's interaction with the environment; Indicates the current policy network Parameters; It represents the strategy ratio, which is the ratio of the probability of choosing an action under the new strategy to that under the old strategy; The dominance function indicates whether the current power point is better than expected; Indicates will The strategy ratio is limited to Within this range, avoid excessively large strategy updates that could lead to learning instability; This is a constraint coefficient for strategy changes;
[0154] During the training process, continuously make Maximizing the strategy means making it more biased towards actions that yield high rewards, but updates shouldn't be too frequent. This ensures both stability and efficiency. Therefore, the following approach is adopted: .
[0155] In this embodiment, the policy network Indicates a given state The probability distribution of the actions (such as increasing or decreasing power) that it chooses in this state.
[0156] In this embodiment, the strategy ratio Represented as:
[0157] ; (twenty one)
[0158] Current strategy and old strategies In the same state Next, select an action. The probability ratio; if A value greater than 1 indicates that the policy is more inclined to perform the action; PPO will cut this ratio to limit the learning amplitude in order to maintain policy continuity.
[0159] In this embodiment, the advantage function The advantage function is used to measure whether the current action is better than the previous strategy. Represented as:
[0160] ; (twenty two)
[0161] in, .
[0162] In this embodiment, to improve computational efficiency, the reinforcement learning environment adopts a non-discounted reward accumulation method, wherein the accumulated reward... Defined as:
[0163] ; (twenty three)
[0164] ; (twenty four)
[0165] in, , respectively security intelligent agent and economic intelligent agents The cumulative returns.
[0166] In this embodiment, the value function network A deep neural network-based model is used for learning and fitting to estimate the expected cumulative reward in the current state. Value function network. Based on system state (safety intelligent agent) The system takes the current power, voltage difference, temperature difference, and concentration difference of the electrolyzer as input and outputs the predicted total future reward value. Its training objective is to minimize the difference between the model output and the actual cumulative reward. Mean Square Error (MSE) Loss Function The specific form is as follows:
[0167] ; (25)
[0168] In the formula, It represents the set of parameters of a value function network, including trainable parameters such as weights and biases in the network.
[0169] The network parameters are optimized using a gradient descent algorithm based on backpropagation (such as the Adam optimizer). Iterative updates enable the value function network to accurately reflect long-term return expectations under different operating conditions, thereby providing a high-quality valuation basis for the calculation of the advantage function and strategy optimization.
[0170] For example, for secure intelligent agents Its advantage function is as follows:
[0171] ; (26)
[0172] Its cumulative return Safety function model of the tank body The tank safety function model is used to control temperature, voltage, or alumina concentration when there are large fluctuations. It will rise, and the accumulated returns will follow. If it decreases, the action will be penalized; conversely, the strategy will learn in that direction.
[0173] During training, if the boundary is exceeded ( Or it may reach convergence. The value function network output is a neural network trained by the agent itself; its goal is to predict: from the current state... The difference between the expected cumulative reward value that can be experienced in the future and the expected value of the cumulative reward is the advantage function of PPO.
[0174] The total loss function of the multi-agent reinforcement learning system for:
[0175]
[0176] ;(27)
[0177] in, The network return prediction error term is a value function. This is the entropy regularization term; , These are the weighting coefficients; Let be the entropy of the policy distribution, representing randomness.
[0178] In this embodiment, the power boundary search is performed as follows:
[0179] By starting from the safe power point Initially, the current state is estimated using a dual-proxy model: ;
[0180] The agent performs actions and explores the power space, that is:
[0181] ;
[0182] ;
[0183] After executing the action, update the current power as follows:
[0184] ;
[0185] Use a proxy model to evaluate whether the boundary has been exceeded in real time and provide penalty feedback:
[0186] ( );
[0187] If the threshold is exceeded, it is considered that the safety boundary has been violated, the agent receives a negative reward, and the system automatically reverses the power action to punish the excessive exploration.
[0188] Determine whether convergence has been achieved based on the following conditions:
[0189] The security agent satisfies the following constraints: ;
[0190] Economic agency achieves stable optimality: .
[0191] If the conditions are met, it is considered that a power extremum on one side has been found, and the power point that satisfies the conditions is recorded, i.e.:
[0192] ;
[0193] Select the extreme value interval from them. The final output is an adjustable power range for the electrolytic cell that is within a safe range and has good economic performance.
[0194] The embodiments and descriptions above are merely illustrative of the principles and preferred embodiments of the present invention. Various changes and modifications may be made to the present invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed.
Claims
1. A method for determining the adjustable power range of an aluminum electrolytic cell, characterized in that, Includes the following steps: S1) Establish a multiphysics CFD simulation model for the electrolyzer, and use the CFD simulation model to solve the distribution characteristics of temperature, voltage and concentration of the electrolyzer under different input power P as a training dataset; S2) Construct a dual-agent model that includes a tank safety function model and a production economy function model, and use the training dataset to train supervised learning to train the dual-agent model; In step S2), the tank safety function model is used to evaluate the thermal safety margin of the electrolytic cell under different input power P; the production economy function model is used to predict the output and unit energy consumption under different power. In step S2), the tank safety function model The criterion for judgment is whether the tank meets the safety boundaries for heat, electricity, and concentration. The expression is as follows: ; (11) In the formula, This is the temperature deviation term; This is the voltage deviation term; This is the alumina concentration deviation term; , , These are the weighting coefficients for the corresponding deviation terms; , , The permissible deviation threshold for the corresponding deviation item; In step S2), the production economy function model is described. The larger the value, the better the economic benefits. (This refers to the production economics function model.) Represented as: ; (13) In the formula, , Target output and target unit power consumption, respectively; , This is a weighting parameter used to reflect the relative importance of output and energy consumption; The actual output under power P; This represents the actual unit power consumption under power P; S3) Construct a multi-agent reinforcement learning system to solve for the minimum derating power of the aluminum electrolytic cell. With maximum upload power ; In step S3), the multi-agent reinforcement learning system includes a safety agent A and an economic agent B. The safety agent A is used to determine whether a given power P violates the thermal safety boundary, and the economic agent B is used to select the power P with the greatest economic benefit within the safety constraints.
2. The method for determining the adjustable power range of an aluminum electrolytic cell according to claim 1, characterized in that: In step S1), the simulation objective of the CFD simulation model is to simulate the temperature distribution, electric field distribution, heat generation and loss of the electrolytic cell under different input electric power P conditions, and to determine whether a thermal equilibrium state has been reached.
3. The method for determining the adjustable power range of an aluminum electrolytic cell according to claim 1, characterized in that: In step S3), the security intelligent agent Input state space for: Action output for: Reward function for: ; The aforementioned economic intelligent agent Input state space for: Action output for: Reward function for: ; in, , These represent the outputs of security agent A and economic agent B, respectively. The aforementioned security intelligent agent By minimizing As a reward, namely Guide the agent to gradually approach the lower power limit. ; The aforementioned economic intelligent agent By maximizing the reward function, i.e. Guide the intelligent agent to gradually approach the power limit. .
4. The method for determining the adjustable power range of an aluminum electrolytic cell according to claim 3, characterized in that: In step S3), the multi-agent reinforcement learning system updates its power and state through an environmental feedback and state update mechanism, namely: ; (15) ; (16) ; (17) ; (18) ; (19) In the formula, express Power at any given moment; The input power is respectively Temperature deviation, voltage deviation, and alumina concentration deviation; The unit marginal profit; for Time-of-use electricity pricing; For input power is Current efficiency at that time; For input power is The ideal voltage for electrolytic aluminum; Indicates power input Below, based on the tank safety function model The predicted safety assessment function value; Indicates power input Below, from the production economy function model The predicted economic function value.
5. The method for determining the adjustable power range of an aluminum electrolytic cell according to claim 4, characterized in that: In step S3), the total loss function of the multi-agent reinforcement learning system for: ; (27) in, Let the objective function be the policy objective function; The mean squared error (MSE) loss function; The set of parameters representing a value function network; For policy networks; Representation of value function networks; This is the current state; The network return prediction error term is a value function. This is the entropy regularization term; , These are the weighting coefficients; Let be the entropy of the policy distribution, representing randomness; For cumulative returns.
6. The method for determining the adjustable power range of an aluminum electrolytic cell according to claim 5, characterized in that: In step S3), from the safe power point We begin by searching for power points that satisfy the following conditions: The security agent satisfies the following constraints: ; Economic agency achieves stable optimality: ; If the conditions are met, it is considered that a power extremum on one side has been found, and the power point that satisfies the conditions is recorded, i.e.: ; Select the extreme value interval from them. .
Citation Information
Patent Citations
Digital intelligent management and control platform for aluminum electrolysis production
CN112501655A
Unit combination method and device taking electrolytic aluminum heat conduction characteristics into consideration
CN114201868A