Multi-agent collaborative management method for energy storage power station based on deep reinforcement learning
By employing a multi-agent collaborative management method based on deep reinforcement learning, the problems of untimely coordination among subsystems and errors in battery state estimation in energy storage power stations are solved, enabling efficient, safe, and adaptive dynamic control of energy storage power stations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN TIER TECHNOLOGY CO LTD
- Filing Date
- 2025-11-24
- Publication Date
- 2026-05-12
AI Technical Summary
In existing energy storage power stations, the coordination between various subsystems is not timely, and the battery status estimation error is difficult to be fed back in a timely manner, which leads to the problems of strategy lag and equipment operation imbalance.
The method for multi-agent collaborative management of energy storage power stations based on deep reinforcement learning collects and preprocesses collaborative sensing data of energy storage, constructs a multi-control unit architecture, dynamically selects action strategies, constructs disturbance risk assessment values, constructs triplet samples of state, action and reward, and performs forward inference and backward correction on the action value assessment model to achieve a closed loop of action strategy optimization.
It enables information sharing and strategy interaction among subsystems, improves the overall system coordination efficiency, enhances the adaptive capability and security of high-frequency dynamic control, and solves the problem of insufficient real-time coordination and autonomous decision-making response among multiple control units in existing technologies.
Smart Images

Figure CN121192798B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, specifically to a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning. Background Technology
[0002] With the continuous increase in the penetration rate of new energy power generation, energy storage power stations are playing an increasingly prominent role in the power system, becoming an important means to ensure grid frequency regulation, peak shaving and valley filling, and improve power quality. Because energy storage systems contain multiple complex subsystems, their operating states exhibit strong coupling, time-varying, and nonlinear characteristics. Traditional rule-based or single-optimization-objective-based control strategies are insufficient to meet the optimal management needs of energy storage systems under multiple scenarios and dynamic disturbances. Therefore, how to achieve collaborative control and adaptive optimization of multiple control units in energy storage power stations has become a key research direction for intelligent power dispatch and intelligent energy storage operation and maintenance.
[0003] For example, invention publication CN116596028A discloses an Informer-based method for managing the power consumption of energy storage power stations, including a cloud server and multiple energy storage power station subsystems. The cloud server is equipped with an Informer network model, which uses data collected and uploaded by the energy storage power station subsystems for network training, outputting predicted SOC values for the energy storage power station under different temperatures and operating conditions. The energy storage power station subsystems download corresponding model weight files from the cloud server based on their own operating environment conditions. Simultaneously, the cloud server continuously receives data uploaded by the energy storage power station subsystems, performs incremental learning, and continuously optimizes the predicted temperature range with a 5°C step gradient. This method can improve the problem of excessively long training time due to insufficient hardware computing power of the subsystems, and can automatically update the weight files in a timely manner after changes in external conditions, eliminating the need for manual updates and improving prediction accuracy.
[0004] For example, the invention with publication number CN119378636A provides a parallel training method for multiple control units in a power system based on swarm intelligence optimization, including: constructing a system simulation model based on the novel power system to be controlled; constructing an initial multiple control unit based on the system simulation model; initializing particle swarm parameters and initializing the initial multiple control units, and several training control units; performing distributed parallel training on the several training control units until all training control units are trained, obtaining the current round of multiple control units; obtaining the fitness of each control unit in the current round of control unit group and updating the particle swarm parameters, and then updating the basic parameters of each control unit in the current round of control unit group to update the several training control units; repeating the distributed parallel training on the several training control units until the preset training conditions are met, obtaining the control unit group; and controlling the novel power system to be controlled based on the control unit group.
[0005] However, although the aforementioned technical solutions have made some progress in energy storage management prediction and multi-control unit training, they still lack the ability for real-time coordination, autonomous decision-making response, and adaptation under disturbance conditions among multiple control units, making it difficult to meet the optimization control requirements of energy storage power plants in high-frequency dynamic regulation. In addition, existing methods mostly rely on fixed model architectures and centralized inference, lacking a dynamic trade-off mechanism for the operational risks of energy storage systems, and cannot achieve effective coordination among control units under non-ideal operating conditions.
[0006] Therefore, in order to address the above problems, there is an urgent need for a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning. Summary of the Invention
[0007] Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning. This method solves the problems of untimely coordination among subsystems and difficulty in timely feedback of battery state estimation errors in existing energy storage power stations, which leads to policy lag and equipment operational imbalance.
[0009] Technical solution
[0010] To achieve the above objectives, this invention provides the following technical solution: a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning, comprising: S1, collecting energy storage collaborative sensing data, and performing time alignment, anomaly removal, standardization, and normalization processing on the energy storage collaborative sensing data to obtain preprocessed energy storage collaborative sensing data; S2, inputting the preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, evaluating the degree of operational coordination deviation of the energy storage power station, dynamically selecting the action strategy of each control unit, and, based on the evaluation results and action strategy, assigning each... The control unit constructs an action value assessment model; S3, extracts energy storage collaborative sensing data and operational collaborative deviation assessment results, constructs disturbance risk assessment values, and extracts state vectors and action vectors, constructs state, action, and reward triplet samples, performs forward inference and backward correction on the action value assessment model, and iteratively optimizes the action value assessment model round by round; S4, loads control targets for each control unit and deploys it to the simulated operation environment, executes action decisions and state interaction feedback; after each action decision is executed, assesses the degree of deviation between the current operating state and the action strategy, determines whether to trigger the strategy correction mechanism, and realizes the action strategy optimization closed loop.
[0011] Furthermore, the specific steps for collecting energy storage collaborative sensing data are as follows: Collect energy storage collaborative sensing data during the operation of the energy storage power station. The energy storage collaborative sensing data includes battery cell voltage, battery cell temperature, battery pack current, battery cycle count, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging and discharging power.
[0012] Furthermore, the specific steps for performing time alignment, anomaly removal, standardization, and normalization on the energy storage collaborative sensing data to obtain preprocessed energy storage collaborative sensing data are as follows: The energy storage collaborative sensing data is processed using an interpolation completion method based on unified sampling period alignment to align data from different acquisition sources to a consistent time series; the energy storage collaborative sensing data is screened using an anomaly identification method based on sliding window differential detection and physical range constraints to identify and remove data records with instantaneous mutations or out-of-bounds errors; the energy storage collaborative sensing data is numerically transformed using a maximum-minimum standardization method based on data type grouping to unify the numerical scale and distribution characteristics of different channels; and the energy storage collaborative sensing data is uniformly mapped using a symmetric interval normalization method to map all numerical data to a unified numerical range.
[0013] Further, the specific steps for inputting the preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, and evaluating the degree of operational collaborative deviation of the energy storage power station are as follows: The preprocessed energy storage collaborative sensing data is input into a multi-control unit architecture, which is divided into three control units according to function: battery management control unit, energy management control unit, and data acquisition and monitoring control unit, which respectively perform battery state estimation, power scheduling, and equipment operation monitoring tasks; the average voltage and average temperature of all batteries are calculated; the voltage of the i-th battery cell is subtracted from the average voltage of all batteries and then divided by the average voltage of all batteries; the comparison value is squared to obtain the voltage deviation factor; The temperature deviation factor is obtained by subtracting the average temperature of all batteries from the temperature of the i-th cell and then dividing by the average temperature of all batteries. The square of this value is then calculated. The temperature deviation factor is added to the voltage deviation factor to obtain the cell state fluctuation factor. The load strength coefficient is obtained by dividing the absolute value of the battery pack current by the rated current of the battery. The grid disturbance coefficient is obtained by subtracting the grid rated frequency from the grid frequency and then dividing the absolute value by the grid rated frequency. The operating disturbance amplification factor is obtained by multiplying the load strength coefficient by the grid disturbance coefficient. The operating disturbance amplification factor is obtained by taking the square root of the cell state fluctuation factor and multiplying it by the operating disturbance amplification factor. The cell operating coordination deviation term is obtained by averaging the cell operating coordination deviation terms of all cells to obtain the operating coordination deviation evaluation value.
[0014] Furthermore, the specific steps for dynamically selecting the action strategies of each control unit and constructing an action value assessment model for each control unit based on the evaluation results and action strategies are as follows: Real-time comparison of the operational coordination deviation evaluation value and the coordination deviation threshold, and dynamic switching of the action strategy selection logic: When the operational coordination deviation evaluation value is less than or equal to the coordination deviation threshold, the current energy storage power station is determined to be in a coordinated operating state, and each control unit selects an action strategy with power efficiency as the objective; when the operational coordination deviation evaluation value is greater than the coordination deviation threshold, the current energy storage power station is determined to be out of balance, and each control unit switches to an action strategy with heat distribution equilibrium and state pressure difference mitigation as the objectives; the operational coordination deviation evaluation value and energy storage coordination sensing data are jointly constructed into a state vector, and the action strategy is used as the action vector. Based on a deep neural network, an action value assessment model is constructed for each control unit.
[0015] Further, the specific steps for extracting energy storage collaborative sensing data and operational collaborative deviation assessment results to construct disturbance risk assessment values are as follows: A fixed-length sliding time window is set as the training period. Within each training period, the energy storage collaborative sensing data sequence and operational collaborative deviation assessment value sequence are extracted. The maximum and minimum charging / discharging power, maximum and minimum battery cell temperature are obtained, and the average battery temperature, average battery pack current, and average operational collaborative deviation assessment value are calculated. The maximum charging / discharging power minus the minimum charging / discharging power is divided by the battery's rated power, and the ratio is incremented by one and the natural logarithm is taken to obtain the power fluctuation coefficient. The maximum battery cell temperature minus the minimum battery cell temperature is divided by the average battery temperature to obtain the temperature difference distribution coefficient. The absolute value of the average battery pack current is divided by the battery's rated current to obtain the load intensity coefficient. The average operational collaborative deviation assessment value is incremented by one to obtain the collaborative disturbance correction coefficient. The power fluctuation coefficient, temperature difference distribution coefficient, load intensity coefficient, and collaborative disturbance correction coefficient are multiplied sequentially to obtain the disturbance risk assessment value.
[0016] Further, the state vector and action vector are extracted, and state, action, and reward triplet samples are constructed. Forward inference and backward correction are performed on the action value assessment model, and the specific steps for iteratively optimizing the action value assessment model round by round are as follows: Using the disturbance risk assessment value as the reward signal, state and action vectors are extracted, and state, action, and reward triplet samples are constructed and stored in the experience sample pool; a cyclic sample extraction strategy is adopted to extract state, action, and reward triplet samples from the experience sample pool in batches. Using the state vector as input, forward inference is performed based on the current action strategy to output the action assessment value; the deviation between the action assessment value and the disturbance risk assessment value is calculated to obtain the loss; using the loss as the training basis, backward correction operations are performed on the action value assessment model to update the action value assessment model and complete the model training for the current training cycle; after each training cycle, the next iteration begins, and the action value assessment model is continuously trained until convergence.
[0017] Furthermore, after loading the control targets for each control unit and deploying it to the simulated operating environment, the specific steps for executing action decisions and state interaction feedback are as follows: Based on the trained action value evaluation model, the scheduling behavior targets are configured for the control units: the battery management control unit is loaded with a control strategy aimed at individual cell thermal control and state balance; the energy management control unit is loaded with a power allocation strategy aimed at balancing economic benefits and battery life; and the data acquisition and monitoring control unit is loaded with an operation monitoring strategy centered on operating boundary constraints. After completing the strategy loading and scope configuration, the control units are deployed to the simulated operating environment to execute action decisions and state interaction feedback. The simulated operating environment includes battery cluster components, bidirectional converter components, and grid interface components.
[0018] Furthermore, the specific steps for evaluating the deviation between the current operating state and the action strategy after each action decision execution are as follows: After each action decision execution, extract the state vector within the current training cycle, perform inference based on the action value evaluation model, output the action evaluation value, and extract the battery cell voltage output value and battery cell temperature output value from the action evaluation value; subtract the corresponding voltage measurement value from the voltage output value of the i-th battery cell, and take the absolute value to obtain the voltage deviation value; subtract the corresponding temperature measurement value from the temperature output value of the i-th battery cell, and take the absolute value to obtain the temperature deviation value; add the voltage deviation value and temperature deviation value of each battery cell to obtain the state offset of the i-th battery cell; sum the state offset values of all battery cells and divide by the number of battery cells to obtain the average state offset; subtract the grid rated frequency from the grid frequency, and take the absolute value to obtain the frequency disturbance coefficient; multiply the average state offset by the frequency disturbance coefficient to obtain the state drift evaluation value.
[0019] Furthermore, the specific steps for determining whether to trigger the strategy correction mechanism and realize the closed loop of action strategy optimization are as follows: Real-time comparison of the state drift evaluation value and the state offset threshold. When the state drift evaluation value is less than the state offset threshold, the current operating state is determined to be normal, and the control unit maintains the existing action strategy. When the state drift evaluation value is greater than or equal to the state offset threshold, the current operating state is determined to deviate from the control target, triggering the strategy correction mechanism: using the state drift evaluation value as an adjustment signal, intervention and suppression are applied to the action evaluation value corresponding to the current action strategy, dynamically reconstructing the action value ranking, and outputting the corrected action evaluation value sequence. Based on the corrected action evaluation value sequence, the action selection logic is guided to adjust towards the state target direction, realizing the adaptive response and dynamic correction of the action strategy to the offset state.
[0020] Beneficial effects
[0021] The present invention has the following beneficial effects:
[0022] (1) This multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning divides the energy storage power station operation system into three control units: BMS, EMS, and SCADA. Based on the collaborative sensing data of energy storage, different control objectives and behavioral strategies are assigned to each control unit, thus constructing a multi-control unit architecture to realize information sharing and strategy interaction among the subsystems. Through the strategy optimization and action evaluation mechanism under the deep reinforcement learning framework, the existing information silo operation mode of each subsystem in the energy storage power station is broken, promoting the flow of state estimation results among the control units, so that local sensing results can quickly affect scheduling and safety control behavior, thereby improving the overall collaborative efficiency of the system.
[0023] (2) This multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning constructs a disturbance risk assessment value as a reward signal in reinforcement learning, comprehensively considering multiple influencing factors such as battery charging and discharging power fluctuations, individual cell temperature difference distribution, load intensity, grid frequency disturbances, and operational status deviations. This disturbance risk assessment value reflects the potential risks and dynamic pressures faced by the current operating state, which helps the control unit accurately identify sensitive points and vulnerabilities in operation during the learning process, improves the sensitivity of the training model to high-risk operating states, and thus realizes the rapid adaptation and early defense capabilities of the strategy under dynamic disturbance conditions, thereby improving the safety and stability of the energy storage power station.
[0024] (3) This multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning constructs a control unit action value evaluation model by utilizing the state-action-reward triplet mechanism in reinforcement learning. Through continuous forward reasoning and backward correction, the model achieves deep learning of the optimal action under various state scenarios. By real-time perception of the operating state and iterative training of the action evaluation model, the control unit can continuously improve the robustness and accuracy of strategy selection under complex operating conditions, forming a scheduling and control system with adaptive and self-correcting capabilities.
[0025] (4) This multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning introduces a state drift evaluation value to quantitatively assess the degree of state deviation after each action. When the deviation exceeds a preset threshold, the intervention and reconstruction mechanism of the action value model is triggered to dynamically adjust the action ranking and strategy output, guiding the action selection logic to shift towards the control target and achieving adaptive correction of the strategy. This can effectively solve problems such as strategy drift and model overfitting that occur during training, further improve the accuracy of action selection and the stability of operation control, and construct a complete strategy optimization closed loop. Attached Figure Description
[0026] Figure 1 The flowchart is for a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning.
[0027] Figure 2 A graph showing the change in perturbation risk during the training cycle of a multi-control unit;
[0028] Figure 3 This is a bar chart showing the state drift evaluation values. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Please see Figures 1-3This invention provides a technical solution: a multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning, comprising: S1, collecting energy storage collaborative sensing data, and performing time alignment, anomaly removal, standardization, and normalization processing on the energy storage collaborative sensing data to obtain preprocessed energy storage collaborative sensing data; S2, inputting the preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, evaluating the degree of operational coordination deviation of the energy storage power station, dynamically selecting the action strategy of each control unit, and, based on the evaluation results and action strategy, assigning a task to each control unit. S3: Construct an action value assessment model; extract energy storage collaborative sensing data and operation collaborative deviation assessment results, construct disturbance risk assessment values, and extract state vectors and action vectors. Construct state, action, and reward triplet samples, perform forward inference and backward correction on the action value assessment model, and iteratively optimize the action value assessment model round by round; S4: load the control targets for each control unit and deploy it to the simulated operation environment to execute action decisions and state interaction feedback; after each action decision is executed, assess the degree of deviation between the current operation state and the action strategy, determine whether to trigger the strategy correction mechanism, and realize the action strategy optimization closed loop.
[0031] Specifically, the steps for collecting energy storage collaborative sensing data are as follows: Collect energy storage collaborative sensing data during the operation of the energy storage power station. This data includes individual battery cell voltage, individual battery cell temperature, battery pack current, battery cycle count, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging / discharging power. Specifically, battery voltage data is obtained by collecting the voltage changes of each individual battery cell during operation; individual battery cell temperature data is obtained by monitoring the temperature values of surface or internal thermistors on the individual battery cells; the overall current output of the battery pack is measured in real time using current sensors installed at the bus or connection ports; the cumulative number of battery cycles is obtained by accumulating the counts of complete charge / discharge cycles; the ambient temperature around the equipment is obtained by deploying temperature sensors within the energy storage station's operating environment; the real-time AC grid voltage is obtained using a voltage monitoring device connected to the grid connection point; the current grid operating frequency is obtained using a synchronous measurement device; the transformer operating temperature is obtained by measuring the transformer winding temperature; and the real-time charging / discharging power is obtained by combining the voltage and current sampling results.
[0032] In this implementation plan, by accurately collecting data on individual battery cell voltage, individual battery cell temperature, battery pack current, ambient temperature, battery cycle count, grid voltage, grid frequency, transformer temperature, and charging / discharging power during the operation of the energy storage power station, comprehensive acquisition of collaborative sensing data for energy storage is achieved. This improves the accuracy of sensing the operating status of the energy storage power station, provides high-quality data support for subsequent action decision-making and value assessment model training, and enhances the real-time performance and reliability of the collaborative control strategy.
[0033] Specifically, the steps for obtaining preprocessed energy storage collaborative sensing data are as follows: First, the data is processed using an interpolation method based on a unified sampling period alignment to align data from different sources to a consistent time series, ensuring comparability of all monitoring data under the same time reference. Second, the data is screened using an anomaly identification method based on sliding window differential detection and physical range constraints. This involves considering the actual physical operating range and rate of change characteristics of individual battery cell voltage, individual battery cell temperature, battery pack current, battery cycle count, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging / discharging power. This process identifies and removes transient data and out-of-bounds data caused by abnormal fluctuations, measurement drift, and equipment failures, thereby improving the original data quality. Initial data quality is improved by performing numerical transformation on the energy storage collaborative sensing data using a max-min standardization method based on data type grouping. Independent standardization intervals are constructed for voltage, current, temperature, and frequency data to unify the numerical scale and distribution characteristics of individual battery voltage, individual battery temperature, battery pack current, battery cycle count, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging / discharging power before and after processing, eliminating the interference of different physical quantity units on model training. A symmetric interval normalization method is used to uniformly map the energy storage collaborative sensing data. Based on the normalization function, all standardized numerical data are mapped to a unified data interval, further improving the data's tractability and convergence in the deep reinforcement learning model. This provides a highly consistent and high-quality input data foundation for subsequent action strategy training and disturbance risk assessment value calculation.
[0034] In this implementation scheme, by performing unified sampling period alignment, sliding window differential detection and anomaly removal under physical range constraints, maximum and minimum standardization processing, and symmetric interval normalization mapping on the energy storage collaborative sensing data, the consistency of the data in the time dimension, the comparability in the numerical dimension, and the adaptability in the training dimension are significantly improved. This ensures that the input data has high precision, high stability, and high trainability, providing a solid data foundation for the construction of subsequent action value assessment models and the accurate calculation of disturbance risk assessment values, and improving the training efficiency and convergence performance of the multi-control unit collaborative strategy.
[0035] Specifically, the steps for inputting preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, and evaluating the degree of operational coordination deviation of the energy storage power station are as follows: The preprocessed energy storage collaborative sensing data is input into a multi-control unit architecture, functionally divided into three control units: the battery management control unit, the energy management control unit, and the data acquisition and monitoring control unit. These units respectively perform battery state estimation, power scheduling, and equipment operation monitoring tasks, ensuring efficient parallel processing of multi-source sensing data according to functional domains. The average voltage and temperature of all batteries are calculated, and global statistics are performed based on the voltage and temperature data of all individual batteries within the current sampling period. The voltage of the i-th individual battery is subtracted from the average voltage of all batteries, divided by the average voltage of all batteries, and the result is squared to obtain the voltage deviation factor, accurately quantifying the degree of state dispersion of the i-th individual battery at the voltage level. The temperature of the i-th individual battery is subtracted from the average temperature of all batteries, divided by the average temperature of all batteries, and the result is squared. The following steps are taken: First, a temperature deviation factor is obtained, reflecting the degree of abnormality in the thermal distribution of the i-th battery cell. Second, the voltage deviation factor is added to the temperature deviation factor to obtain the cell state fluctuation factor, which uniformly measures the comprehensive impact of multi-dimensional state parameters on cell stability. Third, the absolute value of the battery pack current is divided by the battery rated current to obtain the load intensity coefficient, representing the relative working intensity of the current system on the load side. Fourth, the absolute value of the grid frequency minus the grid rated frequency, divided by the grid rated frequency again, is obtained to obtain the grid disturbance coefficient, which reflects the stability fluctuation amplitude of the current grid operation. Fifth, the load intensity coefficient is multiplied by the grid disturbance coefficient to obtain the operation disturbance amplification factor, which characterizes the transmission effect of external grid disturbances at the load level from a coupling perspective. Sixth, the square root of the cell state fluctuation factor is multiplied by the operation disturbance amplification factor to obtain the cell operation coordination deviation term, which fully reflects the degree of operational deviation of the battery cell under the dual effects of state fluctuations and system disturbances. Finally, the average of the cell operation coordination deviation terms of all battery cells is used to obtain the operation coordination deviation evaluation value, which serves as a comprehensive quantitative indicator of the operational coordination of the current energy storage power station's multi-control unit.
[0036] The specific formula for calculating the operational coordination deviation assessment value is as follows:
[0037] ;
[0038] In the formula, This represents the operational coordination deviation assessment value. This represents the voltage of the i-th battery cell. This represents the average voltage of all batteries. This represents the temperature of the i-th battery cell. This represents the average temperature of all batteries. Indicates the battery pack current. Indicates the battery's rated current. Indicates the power grid frequency. Indicates the rated frequency of the power grid. This indicates the number of individual battery cells.
[0039] In this implementation scheme, preprocessed energy storage collaborative sensing data is input into a multi-control unit architecture, which is divided into a battery management control unit, an energy management control unit, and a data acquisition and monitoring control unit. This clarifies the functional affiliation of various data types and improves the processing efficiency of multi-dimensional sensing information. By constructing voltage deviation factors and temperature deviation factors, the scheme accurately reflects the degree of state deviation of individual battery cells in the voltage and temperature dimensions, improving the fine-grained accuracy of battery state estimation. By calculating the load intensity coefficient and grid disturbance coefficient, the scheme comprehensively reflects the current system operating load and grid fluctuation intensity, enhancing the ability to characterize operating disturbance characteristics. Based on the combined modeling of individual cell state fluctuation factors and operating disturbance amplification factors, the scheme achieves dynamic calculation of individual cell operating collaborative deviation terms, thereby obtaining an operating collaborative deviation evaluation value. This provides quantitative support for the switching of action strategies between multiple control units, effectively improving the agility and response accuracy of energy storage power station operation regulation.
[0040] Specifically, the dynamic selection of action strategies for each control unit, and the construction of an action value assessment model for each control unit based on the evaluation results and action strategies, involves the following steps: Real-time comparison of the operational coordination deviation assessment value and the coordination deviation threshold, dynamically switching the action strategy selection logic: When the operational coordination deviation assessment value is less than or equal to the coordination deviation threshold, the current energy storage station's operational state is determined to be coordinated. Each control unit selects an action strategy targeting power efficiency, calculated by comparing the energy storage system's charging and discharging efficiency with the battery pack's current fluctuation level. The action strategy maintains the consistency of individual battery cell voltage while controlling the rate of increase in battery cycle count to ensure operational stability. This stage also requires ensuring that the grid voltage remains within safe boundaries to avoid system failures caused by overvoltage or undervoltage. When the operational coordination deviation assessment value is greater than the coordination deviation threshold, the current energy storage station's operation is determined to be out of balance. Each control unit switches to an action strategy targeting thermal distribution equalization and mitigation of state pressure difference. Thermal distribution equalization is achieved through individual battery cell temperature control, while mitigation of state pressure difference relies on suppressing voltage fluctuations and reducing differences between battery state parameters. This stage emphasizes suppressing power surges and reducing the cumulative effect of operational disturbance amplification factors by dynamically responding to predictions of deviation trends based on individual battery cell operating states. The operational coordination deviation assessment value is spliced together with the energy storage coordination sensing data, including individual battery voltage, individual battery temperature, battery pack current, ambient temperature, battery cycle count, grid voltage, grid frequency, transformer temperature, and charging / discharging power, to jointly construct a state vector. The action strategies of each control unit after switching are encoded into action vectors and input into a deep neural network. Based on a supervised training mechanism, a structure-matched action value assessment model is constructed for each control unit to realize the quantitative modeling of action value under strategy differentiation.
[0041] In this implementation scheme, by comparing the operational coordination deviation assessment value with the coordination deviation threshold in real time, the action strategies of each control unit are dynamically switched, achieving adaptive selection of action strategies under different operating conditions. A strategy targeting power efficiency improves the consistency of battery cell voltage and suppresses the rate of increase in battery cycle count, ensuring operational stability. In cases of operational misalignment, a strategy targeting thermal distribution equalization and mitigation of state voltage differential is switched, enhancing the ability to regulate battery cell temperature and voltage fluctuations and effectively reducing the cumulative effect of operational disturbance amplification factors. Furthermore, by jointly constructing a state vector from the operational coordination deviation assessment value and energy storage collaborative sensing data, and constructing action vectors from the action strategies of each control unit, the matching accuracy and precision of the action value assessment model under strategy differentiation are further improved, achieving efficient modeling of multi-control unit action decisions.
[0042] Specifically, the steps for extracting energy storage collaborative sensing data and operational collaborative deviation assessment results to construct disturbance risk assessment values are as follows: Set a fixed-length sliding time window as the training period. In each training period, extract the energy storage collaborative sensing data sequence and the operational collaborative deviation assessment value sequence. Then, obtain the maximum and minimum charging and discharging power, the maximum and minimum battery cell temperature in the current training period from the energy storage collaborative sensing data sequence, and calculate the average battery temperature, the average battery pack current, and the average operational collaborative deviation assessment value. The power fluctuation coefficient is obtained by subtracting the minimum charge / discharge power from the maximum charge / discharge power, dividing the result by the battery's rated power, adding one to the ratio, and taking the natural logarithm. This coefficient measures the intensity of dynamic fluctuations in power output. The temperature difference distribution coefficient is obtained by subtracting the minimum battery cell temperature from the maximum battery cell temperature, dividing the result by the average battery temperature. This coefficient reflects the degree of temperature imbalance within the battery. The load intensity coefficient is obtained by dividing the absolute value of the average battery pack current by the battery's rated current. This coefficient characterizes the variation of the current load level relative to the design rating. The cooperative disturbance correction coefficient is obtained by adding one to the average operational coordination deviation assessment value. This coefficient is used to adjust the weighted impact of the overall disturbance intensity on the assessment results. The disturbance risk assessment value is obtained by multiplying the power fluctuation coefficient, temperature difference distribution coefficient, load intensity coefficient, and cooperative disturbance correction coefficient in sequence. This value serves as the reward signal source for subsequent control unit action value training.
[0043] The specific formula for calculating the disturbance risk assessment value is as follows:
[0044] ;
[0045] In the formula, This indicates the disturbance risk assessment value. Indicates the maximum charging and discharging power. This indicates the minimum charging and discharging power. Indicates the battery's rated power. This indicates the maximum temperature of a single battery cell. This indicates the minimum temperature of a single battery cell. This represents the average battery temperature. This represents the average current of the battery pack. Indicates the battery's rated current. This represents the mean of the operational coordination deviation assessment value.
[0046] In this embodiment, Table 1 is a disturbance risk assessment value data table, listing the key operating parameters and corresponding disturbance risk assessment values for five training cycles. Key operating parameters include maximum and minimum charge / discharge power, battery rated power, maximum and minimum battery cell temperature, average battery temperature, average battery pack current, battery rated current, and average operational coordination deviation assessment value. Specific data is explained as follows: In the first training cycle, the maximum charge / discharge power was 525, the minimum charge / discharge power was 455, the battery rated power was 500, the maximum battery cell temperature was 45.6, the minimum battery cell temperature was 31.5, the average battery temperature was 38.5, the average battery pack current was 215, the battery rated current was 230, the average operational coordination deviation assessment value was 0.15, and the corresponding disturbance risk assessment value was 0.052; In the second training cycle, the maximum charge / discharge power was 510, the minimum charge / discharge power was 455, the maximum charge / discharge power was 455, the minimum charge / discharge power was 455, the average ... The value is 460, the battery rated power is 500, the maximum battery cell temperature is 47.2, the minimum battery cell temperature is 32.8, the average battery temperature is 39.1, the average battery pack current is 225, the battery rated current is 230, the average operational coordination deviation assessment value is 0.18, and the corresponding disturbance risk assessment value is 0.041; in the third training cycle, the maximum charge / discharge power is 530, the minimum charge / discharge power is 452, the battery rated power is 500, the maximum battery cell temperature is 44.8, and the minimum battery cell temperature is... The minimum value was 30.2, the average battery temperature was 37.9, the average battery pack current was 218, the rated battery current was 230, the average operational coordination deviation assessment value was 0.14, and the corresponding disturbance risk assessment value was 0.060. In the 4th training cycle, the maximum charge / discharge power was 458, the minimum charge / discharge power was 500, the rated battery power was 500, the maximum battery cell temperature was 46.5, the minimum battery cell temperature was 33.1, the average battery temperature was 38.7, the average battery pack current was 220, and the rated battery current was... The current was 230, the average operational coordination deviation assessment value was 0.17, and the corresponding disturbance risk assessment value was 0.035; in the 5th training cycle, the maximum charge and discharge power was 520, the minimum charge and discharge power was 450, the battery rated power was 500, the maximum battery cell temperature was 45.9, the minimum battery cell temperature was 32.4, the average battery temperature was 38.2, the average battery pack current was 222, the battery rated current was 230, the average operational coordination deviation assessment value was 0.16, and the corresponding disturbance risk assessment value was 0.046.
[0047] Table 1. Disturbance Risk Assessment Data Table
[0048]
[0049] Figure 2The figure shows the changes in disturbance risk assessment values over five consecutive training cycles. As can be seen, the disturbance risk assessment values fluctuate across different cycles, influenced by a combination of factors including charging / discharging power fluctuations, battery temperature distribution, load intensity, and operational coordination deviations. Specifically, the disturbance risk assessment value reaches its highest level of 0.060 in the third training cycle, reflecting the highest degree of load fluctuation and state dispersion in the energy storage system during this cycle. The value is lowest in the fourth cycle, at only 0.035, indicating a relatively stable operating state. Figure 2 This intuitively reflects the sensitivity of the disturbance risk assessment value to the system's operating status, providing an effective feedback basis for control unit strategy training and verifying the feasibility and practical guiding value of using the disturbance risk assessment value as a reward signal in this invention.
[0050] In this implementation scheme, by extracting the energy storage collaborative sensing data sequence and the operational collaborative deviation evaluation value sequence based on a fixed sliding time window, a disturbance risk assessment value is constructed. This effectively integrates key thermoelectric parameters such as charging and discharging power, battery cell temperature, and battery pack current with operational collaborative deviation information. By using the power fluctuation coefficient, temperature difference distribution coefficient, load intensity coefficient, and collaborative disturbance correction coefficient for item-by-item calculation, the scheme accurately characterizes the comprehensive risk level of the electric-thermal-control interaction disturbance on the system's operational stability during the training period. This provides a highly robust and sensitive reward signal basis for the subsequent action value assessment model, improving the convergence efficiency and environmental adaptability of multi-control unit strategy learning.
[0051] Specifically, the steps for extracting state vectors and action vectors, constructing state, action, and reward triplet samples, and performing forward inference and backward correction on the action value assessment model to iteratively optimize the action value assessment model round by round are as follows: Using the disturbance risk assessment value as the reward signal, extract state vectors and action vectors, construct state, action, and reward triplet samples, and store them in an experience sample pool. The state vector includes battery cell voltage, battery cell temperature, battery pack current, battery cycle count, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging / discharging power from the energy storage collaborative sensing data. The action vector is a structured expression of the action strategy selected by the current control unit, used to represent the control output intention in the current state. The experience sample pool is a cache structure for storing historical interaction samples, with limited capacity and a first-in-first-out update mechanism. It can retain representative state-action-reward data samples during different training weeks, enhancing sample utilization efficiency, alleviating temporal correlation between samples, and improving the stability of model training. To improve generalization performance, a cyclic sampling strategy is adopted, combining the sample timestamp order with the fluctuation characteristics of disturbance risk assessment values for batch sampling. State, action, and reward triplet samples are selected from the experience sample pool to form training batches. Using the state vector as input, the deep neural network model inside the control unit performs forward inference based on the current action strategy, outputting action assessment values. The deviation between the output action assessment value and the disturbance risk assessment value is calculated as an error measure between the model prediction and expectation, further forming the loss amount used to optimize the objective function. Based on the loss amount and the optimizer parameter configuration, a backward correction operation is performed on the action value assessment model, updating the weight matrix and bias term parameters in the deep neural network, completing the model training for the current training cycle. After each training cycle, the next iteration begins. Through continuous and iterative model updates, the generalization ability of the action strategy to different state inputs is gradually improved, achieving stable convergence and optimal strategy response of the action value assessment model under complex energy storage power station operating conditions.
[0052] In this implementation scheme, structured state vectors, action vectors, and disturbance risk assessment values are constructed to form triplet samples of state, action, and reward. The caching mechanism of the experience sample pool enhances the temporal continuity of historical data and the coverage of sample diversity, improving the training efficiency of energy storage collaborative sensing data. Supported by a cyclical sample extraction strategy, high-frequency capture of dynamic state changes in energy storage power stations is achieved. A loss function is constructed based on the deviation between action assessment values and disturbance risk assessment values, ensuring that the action value assessment model can be stably updated and finely corrected under different collaborative deviation levels. Through a continuous forward inference and backward correction iterative mechanism, the model's response accuracy to energy storage collaborative sensing data is strengthened, enabling continuous optimization and dynamic adaptation of the action strategies of multiple control units in energy storage power stations, effectively supporting the evolution of intelligent control strategies under complex operating conditions.
[0053] Specifically, after loading control targets for each control unit and deploying them to the simulated operating environment, the specific steps for executing action decisions and state interaction feedback are as follows: Based on the trained and converged action value evaluation model, matching scheduling behavior targets are configured for various functional control units; among them, the battery management control unit is loaded with a control strategy aimed at individual cell thermal control and state balancing. This strategy achieves heat distribution uniformity control and voltage drop reduction by adjusting the difference between individual cell voltage and individual cell temperature; the energy management control unit is loaded with a power allocation strategy aimed at balancing economic benefits and battery life. This strategy dynamically adjusts the battery pack current allocation ratio based on the marginal benefit performance of battery cycle count and charge / discharge power at different time scales; the data acquisition and monitoring control unit is loaded with an operation monitoring strategy centered on operating boundary constraints. This strategy performs threshold judgment and strategy triggering decisions on the operating state by real-time monitoring of the changing trends of ambient temperature, grid voltage, grid frequency, and transformer temperature. After completing the strategy loading and scope configuration, the control unit is deployed to the simulation environment that constructs a complete physical topology and behavioral model to execute action decisions and state interaction feedback. The simulation environment includes a battery cluster component to simulate the thermoelectric coupling between the internal structures of the battery pack; a bidirectional converter component to simulate the power flow behavior between the grid side and the energy storage side; and a grid interface component to provide grid-side input and output boundary conditions with fluctuating characteristics, ensuring the effectiveness evaluation of the control unit strategy under multi-source disturbance conditions.
[0054] In this implementation scheme, differentiated control targets are loaded onto the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, and deployed in a simulated operating environment including battery cluster components, bidirectional converter components, and grid interface components. This enables precise execution of action strategies and interactive feedback of status during the operation of the energy storage system. Through the division of labor and synergy among control strategies, power allocation strategies, and operation monitoring strategies, the response efficiency and control accuracy of key energy storage collaborative sensing data such as battery cell voltage, battery cell temperature, battery pack current, battery cycle count, grid voltage, grid frequency, and ambient temperature are effectively improved. This enhances the generalization ability and control adaptability of the action value assessment model under multi-source operating condition disturbances, ensuring the transferability and robustness of the multi-control unit collaborative management mechanism in actual operation.
[0055] Specifically, after each action decision is executed, the steps for evaluating the deviation between the current operating state and the action strategy are as follows: After each action decision is executed, the state vector within the current training cycle is extracted. The state vector includes key energy storage collaborative sensing data such as battery cell voltage, battery cell temperature, battery pack current, ambient temperature, battery cycle count, grid voltage, grid frequency, transformer temperature, and charging / discharging power. Based on the trained action value evaluation model, inference operations are performed, and the action evaluation value is output while keeping the neural network structure parameters unchanged. The battery cell voltage output value and battery cell temperature output value are extracted from the action evaluation value. The voltage output value of the i-th battery cell is subtracted from the corresponding voltage measurement value collected in real time, and the absolute value is taken to obtain the voltage deviation value of the cell, ensuring... The system reflects instantaneous prediction errors; the absolute value of the temperature output value of the i-th battery cell is obtained by subtracting the corresponding temperature measurement value, which reflects the accuracy of thermal state control; the voltage deviation value and temperature deviation value of each battery cell are added together to form a state offset covering thermoelectric parameters, which is the state offset of the i-th battery cell; the sum of the state offsets of all battery cells is divided by the number of battery cells to obtain the average state offset in the current training period, which is used to characterize the deviation level between the global operating state and the model output; the absolute value of the difference between the current grid frequency and the grid rated frequency is obtained to obtain the frequency disturbance coefficient, which reflects the disturbance intensity of external grid dynamics on the operational stability of the energy storage system; the average state offset is multiplied by the frequency disturbance coefficient to obtain the state drift evaluation value.
[0056] The specific formula for calculating the state drift evaluation value is as follows:
[0057] ;
[0058] In the formula, This represents the state drift evaluation value. This represents the voltage output value of the i-th battery cell. This represents the measured voltage value of the i-th battery cell. This represents the temperature output value of the i-th battery cell. This represents the measured temperature value of the i-th battery cell. Indicates the number of individual battery cells. Indicates the power grid frequency. This indicates the rated frequency of the power grid.
[0059] In this embodiment, Table 2 is a data table of state drift assessment values, listing the average state drift, number of battery cells, grid frequency, grid rated frequency, and corresponding disturbance risk assessment values over five assessment periods. Specific data explanations are as follows: In period A1, the average state drift is 1.85, the number of battery cells is 8, the grid frequency is 50.08, the grid rated frequency is 50, and the state drift assessment value is 0.148; in period A2, the average state drift is 1.85, the number of battery cells is 8, the grid frequency is 50.20, the grid rated frequency is 50, and the state drift assessment value is 0.370; in period A3, the average state drift is 1.90, the number of battery cells... The average state drift is 0.095, the grid frequency is 50.05, the grid rated frequency is 50, and the state drift assessment value is 0.450. In cycle A4, the average state drift is 1.80, the number of battery cells is 8, the grid frequency is 50.25, the grid rated frequency is 50, and the state drift assessment value is 0.292.
[0060] Table 2. Data Table of State Drift Evaluation Values
[0061]
[0062] Figure 3 This chart compares the state drift assessment values across five evaluation periods. Each bar in the bar chart represents the state drift assessment value for the corresponding evaluation period, with specific values labeled above the bars for precise observation of the operating status in each period. The dashed line in the chart represents the state offset threshold, used to determine whether the current operating state deviates from the control target. The chart shows that the state drift assessment values for evaluation periods A2, A4, and A5 are all higher than the state offset threshold, indicating a significant deviation between the battery cell voltage and temperature output values and the actual measured values. Combined with the difference between the grid frequency and the grid's rated frequency, this may cause system operating state drift, requiring the triggering of an action strategy correction mechanism. In contrast, A1 and A3 are both below the state offset threshold, indicating relatively stable operating states. Figure 3 It intuitively reflects the trend of state deviation risk changing with the cycle, which helps guide the adaptive adjustment of subsequent action strategies.
[0063] In this implementation plan, by comparing key energy storage collaborative sensing data such as battery cell voltage, battery cell temperature, and grid frequency with the output results of the action value assessment model item by item, a state drift assessment value is constructed by multiplying the state offset by the frequency disturbance coefficient. This can accurately quantify the degree of deviation between the action strategy execution result and the current operating state, effectively improve the accuracy of strategy assessment and dynamic response capability, and enhance the control stability and timeliness of action strategy correction of energy storage power stations under fluctuating operating conditions.
[0064] Specifically, the steps for determining whether to trigger the strategy correction mechanism and achieve a closed loop for action strategy optimization are as follows: Real-time comparison of the state drift evaluation value and the state offset threshold. When the state drift evaluation value is less than the state offset threshold, the current operating state is determined to be normal. The control unit maintains the existing action strategy and continuously collects energy storage collaborative sensing data such as battery cell voltage, battery cell temperature, battery pack current, grid frequency, ambient temperature, and charging / discharging power within the current cycle for strategy stability monitoring. When the state drift evaluation value is greater than or equal to the state offset threshold, the current operating state is determined to deviate from the control target, triggering the strategy correction mechanism: using the state drift evaluation value within the current cycle as an adjustment signal, intervention and suppression are applied to the action evaluation value corresponding to the current action strategy. The regularization factor suppresses abnormal deviation behavior, dynamically reconstructs the action value ranking, removes high-risk action components from the action sequence, and outputs a corrected action evaluation value sequence. The regularization factor is a penalty weight coefficient within the model, used to limit abnormal surges in action evaluation values that deviate from the expected distribution range, suppress policy drift generated by the model in unstable states, and avoid non-target-oriented deviations in action selection results. Subsequently, based on the corrected action evaluation value sequence, the action selection logic is guided to adjust towards the state target direction, ensuring that the selected actions can effectively reduce the combined impact of voltage deviation, temperature deviation, and frequency disturbance coefficients, realizing the adaptive response and dynamic correction of the action strategy to the deviation state, and finally constructing a stable closed-loop optimized control path.
[0065] In this implementation scheme, a dynamic comparison mechanism between state drift evaluation values and state offset thresholds is introduced to achieve real-time monitoring and correction of the action strategy's effectiveness, significantly improving the intelligent response capability of the energy storage power station under complex operating conditions. Based on energy storage collaborative sensing data, combined with the penalty control effect of regularization factors, abnormal strategy outputs can be effectively constrained, unexpected action value fluctuations can be suppressed, and the action selection process can be ensured to have stability and goal orientation. Through intervention adjustment and reordering reconstruction of action evaluation values, the action strategy achieves adaptive response and dynamic closed-loop optimization to state offsets, enhancing the control robustness and strategy stability of the control unit under nonlinear disturbance conditions.
[0066] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0067] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning, characterized in that, Includes the following steps: S1 collects energy storage collaborative sensing data and performs time alignment, anomaly removal, standardization and normalization processing on the energy storage collaborative sensing data to obtain preprocessed energy storage collaborative sensing data; S2, input the preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit and data acquisition and monitoring control unit, evaluate the degree of operational collaborative deviation of the energy storage power station, dynamically select the action strategy of each control unit, and build an action value assessment model for each control unit based on the evaluation results and action strategy; The specific steps for inputting the preprocessed energy storage collaborative sensing data into the battery management control unit, energy management control unit, and data acquisition and monitoring control unit, and for evaluating the degree of operational collaborative deviation of the energy storage power station, are as follows: The pre-processed energy storage collaborative sensing data is input into the multi-control unit architecture, which is divided into three control units according to function: battery management control unit, energy management control unit and data acquisition and monitoring control unit, which respectively perform battery status estimation, power scheduling and equipment operation monitoring tasks. Calculate the average voltage and average temperature of all batteries. Subtract the average voltage of all batteries from the voltage of the i-th battery cell, divide by the average voltage of all batteries, and square the result to obtain the voltage deviation factor. Subtract the average temperature of all batteries from the temperature of the i-th battery cell, divide by the average temperature of all batteries, and square the result to obtain the temperature deviation factor. Add the voltage deviation factor to the temperature deviation factor to obtain the cell state fluctuation factor. Divide the absolute value of the battery pack current by the rated battery current to obtain the load strength coefficient. Subtract the rated grid frequency from the grid frequency, take the absolute value, and divide the absolute value by the rated grid frequency to obtain the grid disturbance coefficient. Multiply the load strength coefficient by the grid disturbance coefficient to obtain the operation disturbance amplification factor. Take the square root of the cell state fluctuation factor and multiply by the operation disturbance amplification factor to obtain the cell operation coordination deviation term. Average the cell operation coordination deviation terms of all battery cells to obtain the operation coordination deviation evaluation value. The specific steps for dynamically selecting the action strategy of each control unit and constructing an action value assessment model for each control unit based on the evaluation results and the action strategy are as follows: The system compares the operational coordination deviation assessment value and the coordination deviation threshold in real time and dynamically switches the action strategy selection logic: when the operational coordination deviation assessment value is less than or equal to the coordination deviation threshold, the current energy storage power station is determined to be in a coordinated state, and each control unit selects an action strategy with power efficiency as the target; when the operational coordination deviation assessment value is greater than the coordination deviation threshold, the current energy storage power station is determined to be out of balance, and each control unit switches to an action strategy with heat distribution balance and pressure difference reduction as the target. The operational coordination deviation assessment value and energy storage coordination sensing data are jointly constructed into a state vector, and the action strategy is used as the action vector. An action value assessment model is constructed for each control unit based on a deep neural network. S3: Extract energy storage collaborative sensing data and operation collaborative deviation assessment results, construct disturbance risk assessment values, extract state vectors and action vectors, construct state, action and reward triplet samples, perform forward reasoning and backward correction on the action value assessment model, and iteratively optimize the action value assessment model round by round. The specific steps for extracting energy storage collaborative sensing data and operational collaborative deviation assessment results to construct disturbance risk assessment values are as follows: A fixed-length sliding time window is set as the training period. In each training period, the energy storage collaborative sensing data sequence and the operation collaborative deviation evaluation value sequence are extracted. The maximum and minimum charging and discharging power, the maximum and minimum battery cell temperature are obtained, and the average battery temperature, the average battery pack current, and the average operation collaborative deviation evaluation value are calculated. The power fluctuation coefficient is obtained by subtracting the minimum charge / discharge power from the maximum charge / discharge power, dividing the result by the battery's rated power, adding one to the ratio, and taking the natural logarithm. The temperature distribution coefficient is obtained by subtracting the minimum battery cell temperature from the maximum battery cell temperature, dividing the result by the average battery temperature. The load strength coefficient is obtained by dividing the absolute value of the average battery pack current by the battery's rated current. The cooperative disturbance correction coefficient is obtained by adding one to the average operational coordination deviation assessment value. The disturbance risk assessment value is obtained by multiplying the power fluctuation coefficient, temperature distribution coefficient, load strength coefficient, and cooperative disturbance correction coefficient in sequence. S4 loads control targets for each control unit and deploys them to the simulated operating environment to execute action decisions and status interaction feedback. After each action decision is executed, the degree of deviation between the current operating state and the action strategy is evaluated to determine whether the strategy correction mechanism should be triggered, thus realizing a closed loop for action strategy optimization.
2. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 1, characterized in that: The specific steps for collecting energy storage collaborative sensing data are as follows: Collect energy storage collaborative sensing data during the operation of the energy storage power station. The energy storage collaborative sensing data includes individual battery voltage, individual battery temperature, battery pack current, battery cycle number, ambient temperature, grid voltage, grid frequency, transformer temperature, and charging and discharging power.
3. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 1, characterized in that: The specific steps for performing time alignment, anomaly removal, standardization, and normalization on the energy storage collaborative sensing data to obtain preprocessed energy storage collaborative sensing data are as follows: The energy storage collaborative sensing data is processed by an interpolation completion method based on unified sampling period alignment, aligning data from different acquisition sources to a consistent time series; the energy storage collaborative sensing data is screened by an anomaly identification method based on sliding window differential detection and physical range constraints, identifying and removing data records with instantaneous mutations or out-of-bounds errors; the energy storage collaborative sensing data is numerically transformed by a maximum-minimum standardization method based on data type grouping, unifying the numerical scale and distribution characteristics of different channels; and the energy storage collaborative sensing data is uniformly mapped by a symmetric interval normalization method, mapping all numerical data to a unified numerical range.
4. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 1, characterized in that: The specific steps for extracting state and action vectors, constructing state, action, and reward triplet samples, performing forward inference and backward correction on the action value assessment model, and iteratively optimizing the action value assessment model round by round are as follows: The disturbance risk assessment value is used as the reward signal. State vectors and action vectors are extracted, and samples of state, action and reward triples are constructed and stored in the experience sample pool. A cyclic sample extraction strategy is adopted to extract state, action and reward triple samples from the experience sample pool in batches. The state vector is used as input, and forward reasoning is performed based on the current action strategy to output the action evaluation value. The deviation between the motion assessment value and the disturbance risk assessment value is calculated to obtain the loss amount; Using the loss as the training basis, a backward correction operation is performed on the action value evaluation model to update the action value evaluation model and complete the model training for the current training cycle. After each training cycle ends, the next iteration begins, and the action value evaluation model is continuously trained until convergence.
5. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 1, characterized in that: The specific steps for loading control targets into each control unit and deploying it to the simulated operating environment to execute action decisions and state interaction feedback are as follows: Based on the completed action value evaluation model, the scheduling behavior objectives are configured for the control unit: the battery management control unit is loaded with a control strategy aimed at individual cell thermal control and state balance; the energy management control unit is loaded with a power allocation strategy aimed at balancing economic benefits and battery life; and the data acquisition and monitoring control unit is loaded with an operation monitoring strategy centered on operational boundary constraints. After completing the strategy loading and scope configuration, the control unit is deployed to the simulation environment to execute action decisions and status interaction feedback; the simulation environment includes battery cluster components, bidirectional converter components and grid interface components.
6. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 1, characterized in that: The specific steps for evaluating the deviation between the current operating state and the action strategy after each action decision is executed are as follows: After each action decision is executed, the state vector within the current training cycle is extracted. Inference is performed based on the action value evaluation model, and the action evaluation value is output. The battery cell voltage output value and battery cell temperature output value are extracted from the action evaluation value. The voltage output value of the i-th battery cell is subtracted from the corresponding voltage measurement value, and the absolute value is taken to obtain the voltage deviation value. The temperature output value of the i-th battery cell is subtracted from the corresponding temperature measurement value, and the absolute value is taken to obtain the temperature deviation value. The voltage deviation value and temperature deviation value of each battery cell are added together to obtain the state offset of the i-th battery cell. The state offset values of all battery cells are summed and divided by the number of battery cells to obtain the average state offset value. The grid frequency is subtracted from the grid rated frequency, and the absolute value is taken to obtain the frequency disturbance coefficient. The average state offset value is multiplied by the frequency disturbance coefficient to obtain the state drift evaluation value.
7. The multi-agent collaborative management method for energy storage power stations based on deep reinforcement learning according to claim 6, characterized in that: The specific steps for determining whether to trigger the strategy correction mechanism and realizing the action strategy optimization closed loop are as follows: The state drift assessment value and the state offset threshold are compared in real time. When the state drift assessment value is less than the state offset threshold, the current operating state is determined to be normal and the control unit maintains the existing action strategy. When the state drift evaluation value is greater than or equal to the state offset threshold, it is determined that the current operating state deviates from the control target, triggering the strategy correction mechanism: using the state drift evaluation value as an adjustment signal, the action evaluation value corresponding to the current action strategy is intervened and suppressed, the action value ranking is dynamically reconstructed, and the corrected action evaluation value sequence is output; and based on the corrected action evaluation value sequence, the action selection logic is guided to adjust towards the state target direction, realizing the adaptive response and dynamic correction of the action strategy to the offset state.