Photovoltaic energy storage joint optimization scheduling method and system based on deep reinforcement learning

By extracting energy storage health status features and constructing power fluctuation risk assessment indicators through deep reinforcement learning, the problem of insufficient health assessment of energy storage systems in existing photovoltaic-energy storage joint scheduling is solved, realizing efficient and flexible photovoltaic-energy storage joint optimization scheduling, and improving the system's operational safety and economy.

CN120875494AActive Publication Date: 2025-10-31CCE OASIS TECH CORP

Patent Information

Application Number
CN202511404369.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-10-31
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing photovoltaic and energy storage joint scheduling methods fail to fully consider the health status assessment of energy storage systems, fail to effectively handle multi-source heterogeneous time-series datasets, and struggle to cope with the volatility of photovoltaic power generation. This leads to accelerated wear and tear on energy storage devices, inflexible scheduling strategies, low computational efficiency, and difficulty in meeting real-time requirements.

Method used

A deep reinforcement learning-based approach is adopted to extract energy storage health status features through a dual encoder-decoder network and a recurrent graph convolutional network with a dynamic time-series graph structure. This enables the construction of energy storage safe operation range and power fluctuation risk assessment indicators, the generation of energy storage charging and discharging power time series, and the optimization of scheduling strategies.

Benefits of technology

It improves the lifespan and operational safety of energy storage systems, accurately assesses system fluctuation risks, enhances the search efficiency and quality of scheduling schemes, realizes efficient collaborative operation of photovoltaic energy storage systems, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875494A_ABST
    Figure CN120875494A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic energy storage joint optimization scheduling method and system based on deep reinforcement learning, and relates to the technical field of energy management, and the method comprises the steps: collecting multi-source heterogeneous time series data, extracting energy storage health state features through a dual codec network and a recurrence plot convolution network, dividing an energy storage power interval, and constructing a feasible region; based on the power difference value sequence, extracting fluctuation characteristics in a segmented manner, and constructing a risk evaluation index; and generating a plurality of search paths in the feasible region, obtaining an energy storage charging and discharging power time sequence through iterative search, and generating a scheduling instruction. According to the method, light storage joint optimization scheduling under the energy storage safety constraint is realized, and the photovoltaic consumption rate and the system economy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy management technology, and in particular to a photovoltaic energy storage joint optimization scheduling method and system based on deep reinforcement learning. Background Technology

[0002] With the increasing proportion of renewable energy in the power system, photovoltaic (PV) power generation is widely used due to its clean and environmentally friendly characteristics. However, the intermittency and volatility of PV power generation pose challenges to the safe and stable operation of the power grid. Energy storage systems, as an important regulation tool, can effectively smooth fluctuations in PV output and improve the stability and reliability of the power grid. The coordinated operation of PV and energy storage has become a key technology in new power systems, and reasonable joint dispatch of PV and energy storage is of great significance for improving the system's economy and reliability.

[0003] In the field of joint optimization scheduling of photovoltaic and energy storage, traditional methods mainly rely on deterministic mathematical models and simple rule-based strategies. These methods are typically based on historical data and empirical rules, employing optimization algorithms such as linear programming and dynamic programming for solution. With the increasing complexity of power systems and the rise of uncertainties, scheduling methods based on machine learning, particularly deep reinforcement learning, have gained increasing attention in recent years. These methods can learn optimal strategies through interaction with the environment, exhibiting strong adaptability and generalization capabilities.

[0004] However, existing photovoltaic and energy storage joint scheduling methods still have shortcomings. Existing methods generally ignore the health status assessment of energy storage systems and fail to fully consider the aging characteristics and lifespan degradation factors of energy storage batteries. This may lead to scheduling strategies accelerating the wear and tear of energy storage equipment and reducing the lifespan of energy storage systems. Traditional scheduling algorithms are difficult to effectively handle multi-source heterogeneous time-series datasets, do not accurately characterize the volatility of photovoltaic power generation, lack risk assessment mechanisms, and cannot respond flexibly to different power fluctuation scenarios, reducing the system's adaptability to extreme situations. Existing optimization algorithms are prone to getting trapped in local optima during the search process and have low computational efficiency, making it difficult to meet the real-time scheduling requirements of large-scale systems. They also have limited global optimization capabilities under complex constraints, affecting the economy and feasibility of scheduling schemes. Summary of the Invention

[0005] This invention provides a photovoltaic energy storage joint optimization scheduling method and system based on deep reinforcement learning, which can solve the problems in the prior art.

[0006] A first aspect of this invention provides a photovoltaic energy storage joint optimization scheduling method based on deep reinforcement learning, comprising: Collect photovoltaic power generation data, energy storage operation status data, and grid load data to form a multi-source heterogeneous time-series dataset; Based on a multi-source heterogeneous time-series dataset, energy storage features are extracted through a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, generating an energy storage health status feature vector. The comprehensive power feature value and capacity threshold are calculated based on the energy storage health status feature vector, and the safe operating range of energy storage is determined through a two-dimensional coordinate system. Based on the safe operating range of energy storage, the energy storage power range is divided, and combined with the physical constraints of energy storage, a feasible domain for energy storage charging and discharging power is constructed. Calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; segment the difference sequence and extract the power fluctuation feature vector; construct a power fluctuation risk assessment index based on the power fluctuation feature vector and correct the probability distribution parameters. Multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, the advantageous search path and the disadvantageous search path are divided. The energy storage charging and discharging power time series is obtained through iterative search and identification of local optimal solutions. A sequence of energy storage scheduling instructions is generated based on the energy storage charging and discharging power timing.

[0007] In one optional embodiment, based on a multi-source heterogeneous time-series dataset, energy storage features are extracted using a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure to generate an energy storage health status feature vector, including: The multi-source heterogeneous time-series dataset is input into the dual encoder-decoder network structure. The initial feature matrix is ​​obtained by extracting time-series features and fusing state features, which contains time-series information and state information of energy storage operation. A dynamic time-series graph structure is constructed based on the initial feature matrix. The nodes in the dynamic time-series graph structure correspond to the energy storage operation state at different times. The weights between the nodes in the dynamic time-series graph structure are calculated to obtain the time-series correlation weight matrix through the thermodynamic constraint adaptive attention mechanism. A recurrent graph convolutional network is constructed based on the temporal correlation weight matrix. The convolution kernel parameters in the recurrent graph convolutional network are updated by thermodynamically constrained recurrent units. The initial hidden state of the recurrent units is set according to the thermodynamic constraints. Graph convolution operation is performed to obtain thermodynamic temporal features. The thermodynamic temporal features are then fused with the initial feature matrix through residual connections to obtain fused enhanced features. The fused enhanced features are decomposed into multi-scale features to obtain a multi-scale feature representation that includes global operational features, local fluctuation features, and trend features. The multi-scale feature representation is input into a feature mapping network and reduced to a low-dimensional manifold space through nonlinear projection transformation to obtain the energy storage health status feature vector.

[0008] In one optional embodiment, the comprehensive power characteristic value and capacity threshold are calculated based on the energy storage health status characteristic vector, and the safe operating range of the energy storage is determined through a two-dimensional coordinate system, including: The energy storage health status feature vector is decomposed into power feature components and capacity feature components through nonlinear mapping. The power feature components include charging power components and discharging power components. The maximum charging power and maximum discharging power are determined by multiplying the charging power component and the discharging power component by the coupling coefficients of the energy storage internal resistance and temperature, respectively, and the comprehensive power characteristic value is calculated. The capacity feature component is multiplied by the weighting coefficients of the number of iterations and the depth of iterations, and then summed to obtain the capacity threshold. A two-dimensional coordinate system is established with the comprehensive power feature value as the vertical axis and the capacity threshold as the horizontal axis. Multiple feature points are selected as regional center points in the two-dimensional coordinate system. A regional division map is constructed based on the regional center points. The distance from any point in the two-dimensional coordinate system to each regional center point is calculated, and the region corresponding to the nearest regional center point is taken as the region to which the corresponding point belongs. A classifier is constructed using a kernel function to determine the classification boundaries of different regions. Based on the membership degree calculation, the comprehensive power feature value and the capacity threshold are mapped to different regions of the two-dimensional coordinate system. The classification boundaries of the different regions determine the safe operating range of energy storage.

[0009] In one optional embodiment, the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time-series dataset is calculated; the difference sequence is segmented and a power fluctuation feature vector is extracted; and a risk assessment index is constructed based on the power fluctuation feature vector to correct the probability distribution parameters, including: Calculate the difference sequence between photovoltaic power generation data and grid load data; The power fluctuation trend characteristics of the difference sequence are calculated. A fluctuation feature vector is constructed based on the power rise rate, fall rate, and duration. The fluctuation transition point is identified using a double threshold comparison method. The difference sequence is divided into multiple feature subsequences with the fluctuation transition point as the boundary. The first-order difference and second-order difference of each feature subsequence are calculated to construct a difference feature matrix. The difference feature matrix is ​​subjected to singular value decomposition to obtain a feature vector group. The feature vector group is dynamically weighted based on the energy storage health status feature vector to obtain a power fluctuation feature vector. The energy storage power interval is divided according to the direction and magnitude of the power fluctuation feature vector. The difference data within the energy storage power range is statistically analyzed, the data distribution characteristic parameters are calculated, and the power fluctuation characteristic vector is used as a correction factor to adjust the distribution characteristic parameters to obtain the corrected probability distribution function. Calculate the statistical characteristics of the modified probability distribution function, determine the risk assessment threshold by combining the energy storage charging and discharging response characteristics, and generate power fluctuation risk assessment indicators. The power fluctuation feature vector is updated based on real-time operating data, and the power fluctuation risk assessment index is dynamically adjusted.

[0010] In one optional embodiment, identifying fluctuation transition points using a dual threshold comparison method includes: Determine the instantaneous power change rate based on power time-series data; The instantaneous power change rate is accumulated and summed within the sliding time window to obtain the cumulative power change. The duration for which the sign of the instantaneous power change rate remains the same is recorded to obtain the fluctuation persistence index. Based on the statistical distribution characteristics of power time series data and the charging and discharging power limit of energy storage equipment, the upper and lower thresholds of power change rate are determined. Based on the capacity constraint of energy storage equipment and the duration of the sliding time window, the threshold of cumulative power change is determined. Determine whether the instantaneous power change rate exceeds the upper and lower thresholds of the power change rate, and whether the cumulative power change exceeds the threshold of the cumulative power change. Within the time period that satisfies the double threshold conditions, find the extreme point of the instantaneous power change rate and record it as a candidate conversion point. Calculate the time interval and power change amplitude of adjacent candidate switching points, set a time threshold based on the fluctuation persistence index, set an amplitude threshold based on the power change rate threshold, and delete candidate switching points whose time interval is less than the time threshold and whose power change amplitude is less than the amplitude threshold; The candidate transition points that have been screened are determined as power fluctuation transition points, and the power time series data is segmented using the power fluctuation transition points as boundaries.

[0011] In one optional embodiment, multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on a power fluctuation risk assessment index, advantageous and disadvantageous search paths are divided. The energy storage charging and discharging power time series is obtained through iterative search and identification of local optimal solutions, including: A uniformly distributed initial search point set is constructed within the feasible domain of energy storage charging and discharging power, and multiple parallel search paths are generated based on the initial search point set; Calculate the power fluctuation risk assessment index and its changing trend for each parallel search path, and determine the advantageous and disadvantageous search paths based on the dynamic weighted average. For each parallel search path, perform an iterative search, extract the search step size and search direction of the advantageous search path as feature parameters, and update the search step size and search direction of the disadvantageous search path. Record the change in evaluation value of the parallel search path within a preset number of consecutive iterations. When the change in evaluation value is less than a preset floating threshold, the current search point is recorded as a local optimum. A new search starting point is generated by randomly perturbing the neighborhood of the current search point. The search is performed again based on the new search starting point. The set of local optima that satisfy the feasible domain constraint of the energy storage charging and discharging power is taken as a candidate solution. The solution with the best evaluation value among the candidate solutions is selected to generate the energy storage charging and discharging power timing sequence, and the energy storage charging and discharging power timing sequence is adjusted online based on real-time operating data.

[0012] In one optional embodiment, calculating the power fluctuation risk assessment index and its changing trend for each parallel search path, and determining the advantageous and disadvantageous search paths based on the dynamic weighted average, includes: The power fluctuation risk assessment index of multiple parallel search paths is obtained as the evaluation value; The rate of change is obtained by calculating the difference in the evaluation values ​​within a preset time window, and the acceleration of change is obtained by calculating the difference in the rate of change. The path score is obtained by weighting the evaluation value and the rate of change, summing the values, and multiplying the sum by the acceleration. The dynamic weighted average of the path score is calculated by adjusting the weighting coefficients of the evaluation value and the rate of change based on the real-time status. Parallel search paths with path scores less than or equal to the dynamic weighted average are identified as inferior search paths, while parallel search paths with path scores greater than the dynamic weighted average are identified as superior search paths.

[0013] A second aspect of this invention provides a photovoltaic energy storage joint optimization scheduling system based on deep reinforcement learning, comprising: The first unit is used to collect photovoltaic power generation data, energy storage operation status data and grid load data to form a multi-source heterogeneous time series dataset. The second unit is used to extract energy storage features based on multi-source heterogeneous time-series datasets, through a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, and generate an energy storage health status feature vector; based on the energy storage health status feature vector, the comprehensive power feature value and capacity threshold are calculated, and the safe operating range of energy storage is determined through a two-dimensional coordinate system. The third unit is used to divide the energy storage power range according to the safe operating range of energy storage, and to construct the feasible domain of energy storage charging and discharging power in combination with the physical constraints of energy storage. The fourth unit is used to calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; the difference sequence is segmented and power fluctuation feature vectors are extracted; and power fluctuation risk assessment indicators are constructed based on the power fluctuation feature vectors to correct the probability distribution parameters. The fifth unit is used to generate multiple parallel search paths within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, it divides the search paths into advantageous and disadvantageous ones, and obtains the energy storage charging and discharging power time series through iterative search and identification of local optimal solutions. The sixth unit is used to generate an energy storage scheduling instruction sequence based on the energy storage charging and discharging power timing.

[0014] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0016] In this embodiment of the invention, by collecting multi-source heterogeneous time-series datasets and combining them with a dual codec network and a recursive graph convolutional network with a dynamic time-series graph structure, the health status characteristics of energy storage can be accurately extracted, the safe operating range of energy storage can be effectively determined, and the service life and operational safety of the energy storage system can be improved. An innovative power fluctuation risk assessment index is constructed. Through segmented analysis of the difference sequence and extraction of power fluctuation feature vectors, the system fluctuation risk can be assessed more accurately, providing a reliable basis for energy storage scheduling decisions. An optimization strategy using multiple parallel search paths, combined with a mechanism for dividing advantageous and disadvantageous paths, significantly improves the search efficiency and result quality of the scheduling scheme, achieving efficient collaborative operation of the photovoltaic energy storage system and reducing system operating costs. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the photovoltaic energy storage joint optimization scheduling method based on deep reinforcement learning, as described in an embodiment of the present invention. Figure 2 Adaptive iterative optimization flowchart for energy storage health status feature extraction; Figure 3 This is a data flow diagram for assessing the power fluctuation risk of photovoltaic energy storage. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0020] Figure 1 This is a flowchart illustrating the photovoltaic energy storage joint optimization scheduling method based on deep reinforcement learning, as described in an embodiment of the present invention. Figure 1 As shown, the method includes: Collect photovoltaic power generation data, energy storage operation status data, and grid load data to form a multi-source heterogeneous time-series dataset; Based on a multi-source heterogeneous time-series dataset, energy storage features are extracted through a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, generating an energy storage health status feature vector. The comprehensive power feature value and capacity threshold are calculated based on the energy storage health status feature vector, and the safe operating range of energy storage is determined through a two-dimensional coordinate system. Based on the safe operating range of energy storage, the energy storage power range is divided, and combined with the physical constraints of energy storage, a feasible domain for energy storage charging and discharging power is constructed. Calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; segment the difference sequence and extract the power fluctuation feature vector; construct a power fluctuation risk assessment index based on the power fluctuation feature vector and correct the probability distribution parameters. Multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, the advantageous search path and the disadvantageous search path are divided. The energy storage charging and discharging power time series is obtained through iterative search and identification of local optimal solutions. A sequence of energy storage scheduling instructions is generated based on the energy storage charging and discharging power timing.

[0021] In one optional implementation, based on a multi-source heterogeneous time-series dataset, energy storage features are extracted using a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure to generate an energy storage health status feature vector, including: The multi-source heterogeneous time-series dataset is input into the dual encoder-decoder network structure. The initial feature matrix is ​​obtained by extracting time-series features and fusing state features, which contains time-series information and state information of energy storage operation. A dynamic time-series graph structure is constructed based on the initial feature matrix. The nodes in the dynamic time-series graph structure correspond to the energy storage operation state at different times. The weights between the nodes in the dynamic time-series graph structure are calculated to obtain the time-series correlation weight matrix through the thermodynamic constraint adaptive attention mechanism. A recurrent graph convolutional network is constructed based on the temporal correlation weight matrix. The convolution kernel parameters in the recurrent graph convolutional network are updated by thermodynamically constrained recurrent units. The initial hidden state of the recurrent units is set according to the thermodynamic constraints. Graph convolution operation is performed to obtain thermodynamic temporal features. The thermodynamic temporal features are then fused with the initial feature matrix through residual connections to obtain fused enhanced features. The fused enhanced features are decomposed into multi-scale features to obtain a multi-scale feature representation that includes global operational features, local fluctuation features, and trend features. The multi-scale feature representation is input into a feature mapping network and reduced to a low-dimensional manifold space through nonlinear projection transformation to obtain the energy storage health status feature vector.

[0022] In one specific implementation, multi-source heterogeneous time-series data is first input into a dual encoder-decoder network structure for processing. This network receives heterogeneous data including energy storage system operating parameters, photovoltaic output power, and grid load demand, and performs preliminary processing on this data through a time-series feature extraction module. The time-series feature extraction module employs a combination of convolutional and recurrent layers. The convolutional layer uses a 3×3 kernel size with a stride of 1 to capture local temporal dependencies; the recurrent layer uses gated recurrent units with a hidden state dimension of 128 to learn long-term sequence patterns. The state feature fusion module uses a multi-head attention mechanism with 8 heads and an attention dimension of 64, adaptively weighting and fusing features from different sources to obtain an initial feature matrix of dimension 256×32, where 256 represents the time step and 32 represents the feature dimension.

[0023] When constructing a dynamic time-series graph structure based on the initial feature matrix, the energy storage operation state at each time step is treated as a node in the graph, with a node feature dimension of 32. The connection weights between nodes are calculated using a thermodynamically constrained adaptive attention mechanism. This mechanism first calculates the similarity between the features of any two nodes at any given time step, using cosine similarity as a metric, with similarity values ​​ranging from -1 to 1. Then, a thermodynamic constraint factor is introduced. This factor is set according to the thermodynamic characteristics of the energy storage system, considering parameters such as the rate of change of energy storage temperature and the rate of entropy increase, with a constraint factor value ranging from 0 to 2. The similarity is multiplied by the thermodynamic constraint factor, and then normalized using the softmax function to obtain the final time-series association weight matrix, with a matrix dimension of 256×256, representing the association strength between 256 time steps.

[0024] The recurrent graph convolutional network is constructed based on a temporal correlation weight matrix. The network consists of three graph convolutional layers, each with output feature dimensions of 64, 128, and 64, respectively. The graph convolutional kernel parameters are updated by thermodynamically constrained recurrent units, which employ a long short-term memory (LSM) network structure with a memory unit size of 128. The initial hidden state of each recurrent unit is set according to thermodynamic constraints, with an initial temperature of 25 degrees Celsius and an initial entropy value set to 85% of the standard value. During graph convolution operations, the features of each node are aggregated using a weighted summation method, with weights derived from the temporal correlation weight matrix. After the graph convolution operation, thermodynamic temporal features are obtained, with a dimension of 256×64. The thermodynamic time-series features are fused with the initial feature matrix through residual connection. Specifically, the thermodynamic time-series features are transformed into the same dimension of 256×32 as the initial feature matrix through linear projection, then element-wise addition is performed, and then processed by an activation function. The activation function selected is leakyReLU with a negative slope of 0.2, finally resulting in a fused enhanced feature with a dimension of 256×32.

[0025] The fusion enhancement features are decomposed into multi-scale features using the discrete wavelet transform method with the db4 wavelet basis and a decomposition level of 3. The decomposition yields global operational features, local fluctuation features, and trend features. Global operational features reflect the overall operational status of the energy storage system, with a dimension of 256×16; local fluctuation features capture short-term fluctuations, with a dimension of 256×8; and trend features represent long-term trends, with a dimension of 256×8. These three types of features combine to form a multi-scale feature representation with a total dimension of 256×32.

[0026] The multi-scale feature representation is input into a feature mapping network and subjected to a nonlinear projection transformation. The feature mapping network consists of three fully connected layers with 128, 64, and 32 neurons in each layer, and the activation function is the tanh function. Through this nonlinear transformation, the high-dimensional features are reduced to a low-dimensional manifold space, resulting in a 256×16 feature vector representing the energy storage health status.

[0027] In practical applications, taking a photovoltaic energy storage station as an example, 30 consecutive days of operational data were collected, including sampling points every 5 minutes, totaling 8640 time points. The data included battery pack voltage (range 360-420 volts), current (range -100 to 100 amperes), temperature (range 15-45 degrees Celsius), state of charge (range 10%-90%), and photovoltaic output power (range 0-1000 kilowatts). After inputting this data into the proposed method, the extracted initial feature matrix reflects the performance changes of the battery under different charge-discharge cycles. The constructed dynamic time series graph clearly shows the increased state transition probability under high-temperature conditions (above 38 degrees Celsius) and rapid charge-discharge conditions (current change rate exceeding 20 amperes / minute). The recursive graph convolutional network captured an abnormal pattern where the battery internal resistance increased by 12% after three consecutive days of high-load operation (discharge depth exceeding 60%). Multi-scale decomposition showed that on cloudy days with unstable sunlight, the local fluctuation characteristic value of the energy storage system increased by 35%, and the system regulation frequency increased. The final health status feature vector indicates that the energy storage system is in good health at 85%, with an estimated remaining lifespan of about 7 years. It is recommended to charge and discharge the system when the peak-valley price difference is greater than 0.4 yuan / kWh to extend battery life and improve economic efficiency.

[0028] like Figure 2 As shown, the adaptive iterative optimization flowchart for extracting energy storage health status features is illustrated.

[0029] In one optional implementation, the comprehensive power characteristic value and capacity threshold are calculated based on the energy storage health status characteristic vector, and the safe operating range of the energy storage is determined through a two-dimensional coordinate system, including: The energy storage health status feature vector is decomposed into power feature components and capacity feature components through nonlinear mapping. The power feature components include charging power components and discharging power components. The maximum charging power and maximum discharging power are determined by multiplying the charging power component and the discharging power component by the coupling coefficients of the energy storage internal resistance and temperature, respectively, and the comprehensive power characteristic value is calculated. The capacity feature component is multiplied by the weighting coefficients of the number of iterations and the depth of iterations, and then summed to obtain the capacity threshold. A two-dimensional coordinate system is established with the comprehensive power feature value as the vertical axis and the capacity threshold as the horizontal axis. Multiple feature points are selected as regional center points in the two-dimensional coordinate system. A regional division map is constructed based on the regional center points. The distance from any point in the two-dimensional coordinate system to each regional center point is calculated, and the region corresponding to the nearest regional center point is taken as the region to which the corresponding point belongs. A classifier is constructed using a kernel function to determine the classification boundaries of different regions. Based on the membership degree calculation, the comprehensive power feature value and the capacity threshold are mapped to different regions of the two-dimensional coordinate system. The classification boundaries of the different regions determine the safe operating range of energy storage.

[0030] In one specific implementation, the health state feature vector of the energy storage device can be obtained through monitoring. This feature vector contains multiple parameters such as voltage, current, temperature, and internal resistance. Nonlinear mapping is applied to this feature vector, and singular value decomposition (SVD) can be used to decompose it into power feature components and capacity feature components. For example, for a health state feature vector containing 25 parameters, the power feature component obtained after nonlinear mapping can include two parts: a charging power component and a discharging power component. Assuming the charging power component is 3.8 units and the discharging power component is 4.2 units, these values ​​reflect the charging and discharging capabilities of the energy storage device in its current state.

[0031] When determining the maximum charging and discharging power, it is crucial to consider the internal resistance and temperature effects of the energy storage device. The internal resistance coupling coefficient can be set to 0.85, and the temperature coupling coefficient can be set to 0.92 at 25°C. Multiplying the charging power component of 3.8 units by these coupling coefficients, the maximum charging power is calculated to be 3.8 × 0.85 × 0.92 = 2.97 units. Similarly, the discharging power component of 4.2 units, after the same calculation, yields a maximum discharging power of 3.30 units. The comprehensive power characteristic value can be obtained by averaging the two values, i.e., (2.97 + 3.30) / 2 = 3.14 units. This value represents the comprehensive power capability index of the energy storage device under the current conditions.

[0032] The processing of capacity characteristic components requires calculation in conjunction with the number of cycles and cycle depth. Assuming the capacity characteristic component is 85 units, the weighting coefficient for the number of cycles is set to 0.008, and the weighting coefficient for cycle depth is 0.015, when the energy storage device has undergone 200 cycles with an average cycle depth of 0.6, the calculated capacity threshold is 85 × (1 - 0.008 × 200 - 0.015 × 0.6) = 69.89 units. This value represents the capacity state threshold of the energy storage device under current operating conditions and is an important indicator for assessing the lifespan of energy storage.

[0033] When establishing a two-dimensional coordinate system, the capacity threshold is used as the horizontal axis, and the comprehensive power characteristic value is used as the vertical axis. Within this coordinate system, several representative characteristic points are selected as the region center points, such as points (70, 3.2), (60, 2.8), (50, 2.5), (40, 2.0), and (30, 1.5). These points can be determined through historical data analysis or expert experience, representing the characteristic performance of energy storage devices under different health conditions.

[0034] The region division uses a distance calculation method. For any point (x, y) in the two-dimensional coordinate system, the Euclidean distance from it to the center point of each region is calculated. For example, the distance from point (65, 3.0) to the center point (70, 3.2) is [(65-70)].2 +(3.0-3.2) 2 ] 1 / 2 =5.04 units. The distance to other center points can be calculated using the same method. This point belongs to the region where the nearest regional center point is located.

[0035] When constructing a classifier using kernel functions, radial basis functions (RBF) can be chosen. Assuming a RDF with parameter γ = 0.1 is used, for a boundary point (x, y), the kernel function value from its origin to the region center point (xi, yi) can be calculated using exp(-0.1 × [(x - xi)]). 2 +(y-yi) 2 The kernel function values ​​of point (x, y) can be calculated by comparing the kernel function values ​​of different regions. For example, the kernel function values ​​from point (55, 2.7) to the center point (60, 2.8) and the center point (50, 2.5) are 0.72 and 0.65 respectively, therefore this point belongs to the region corresponding to the first center point.

[0036] In the process of boundary determination, the concept of membership degree is introduced to enhance the rationality of classification. For any point (x, y), its membership degree to a region can be obtained through normalization. Assuming that the distances from the point (62, 2.9) to the center points of each region are 8.06, 2.24, 12.37, 22.47, and 32.57 units respectively, the membership degrees after normalization are 0.13, 0.57, 0.10, 0.06, and 0.04 respectively. Therefore, this point belongs to the second region, with a membership degree of 0.57.

[0037] Through the above analysis, the classification boundaries of different regions can be determined in a two-dimensional coordinate system, thereby defining the safe operating range of energy storage. For example, regions with a membership degree greater than 0.5 can be defined as safe operating regions, regions between 0.3 and 0.5 as warning regions, and regions less than 0.3 as danger regions. For an energy storage device with a comprehensive power characteristic value of 3.14 units and a capacity threshold of 69.89 units, its calculated membership degree is 0.68, placing it within the safe operating region, indicating that the energy storage device has the capability to continue operating safely.

[0038] In this embodiment, the energy storage health status feature vector can be monitored in real time. The comprehensive power characteristic value and capacity threshold are obtained through the aforementioned calculation method, and the current status point is located in a two-dimensional coordinate system. The safe operating status of the energy storage device is determined based on its location, providing a basis for operation and maintenance decisions. When the status point approaches the boundary of a warning zone or danger zone, an early warning message is issued to remind management personnel to take appropriate measures to ensure the continuous safe operation of the energy storage system.

[0039] In one optional implementation, the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time-series dataset is calculated; the difference sequence is segmented and a power fluctuation feature vector is extracted; and a risk assessment index is constructed based on the power fluctuation feature vector to correct the probability distribution parameters, including: Calculate the difference sequence between photovoltaic power generation data and grid load data; The power fluctuation trend characteristics of the difference sequence are calculated. A fluctuation feature vector is constructed based on the power rise rate, fall rate, and duration. The fluctuation transition point is identified using a double threshold comparison method. The difference sequence is divided into multiple feature subsequences with the fluctuation transition point as the boundary. The first-order difference and second-order difference of each feature subsequence are calculated to construct a difference feature matrix. The difference feature matrix is ​​subjected to singular value decomposition to obtain a feature vector group. The feature vector group is dynamically weighted based on the energy storage health status feature vector to obtain a power fluctuation feature vector. The energy storage power interval is divided according to the direction and magnitude of the power fluctuation feature vector. The difference data within the energy storage power range is statistically analyzed, the data distribution characteristic parameters are calculated, and the power fluctuation characteristic vector is used as a correction factor to adjust the distribution characteristic parameters to obtain the corrected probability distribution function. Calculate the statistical characteristics of the modified probability distribution function, determine the risk assessment threshold by combining the energy storage charging and discharging response characteristics, and generate power fluctuation risk assessment indicators. The power fluctuation feature vector is updated based on real-time operating data, and the power fluctuation risk assessment index is dynamically adjusted.

[0040] In one specific implementation, calculating the difference sequence between photovoltaic power generation data and grid load data is a fundamental step in implementing the photovoltaic energy storage joint optimization scheduling method based on deep reinforcement learning. By setting a sampling interval of 5 minutes, real-time power generation data of the photovoltaic power station and regional grid load data are obtained. Taking a certain day as an example, the photovoltaic power station's power generation at 10:00 AM is 2.35 MW, while the regional grid load at the same time is 3.75 MW. The calculated difference is -1.40 MW, indicating that the grid needs an additional 1.40 MW of power support at this time. Calculating the difference sequence using this method for the entire day's data yields a sequence containing 288 sampling points.

[0041] When calculating the power fluctuation trend characteristics of the difference series, it is necessary to identify the fluctuation transition points. A fluctuation transition point is a location in the difference series where the slope changes significantly, identified using a dual-threshold comparison method. Specifically, the rising slope threshold is set to 0.15 MW / min, the falling slope threshold to -0.12 MW / min, and the duration threshold to 15 minutes. When three consecutive sampling points are detected with slopes greater than the rising slope threshold or less than the falling slope threshold, and the duration exceeds the duration threshold, the last sampling point meeting the conditions is marked as a fluctuation transition point. For example, if, between 12:30 and 12:45, slopes of 0.16, 0.18, and 0.17 MW / min are detected consecutively, with a duration of 15 minutes, then 12:45 is marked as a fluctuation transition point. In this way, a total of 24 fluctuation transition points were identified in the entire day's difference series.

[0042] Using the fluctuation transition point as the boundary, the difference sequence is divided into 25 feature subsequences. For each feature subsequence, the first-order and second-order differences are calculated to construct a difference feature matrix. Taking the third feature subsequence as an example, this sequence contains 22 sampling points. The difference between every two adjacent points is calculated to obtain a first-order difference sequence with a length of 21; then, the difference between adjacent elements in the first-order difference sequence is calculated to obtain a second-order difference sequence with a length of 20. The first-order and second-order difference sequences are combined into a difference feature matrix with a size of 2×20.

[0043] Singular value decomposition (SVD) is performed on the difference feature matrix to obtain a set of feature vectors. Taking the third feature subsequence as an example, the SVD yields two main feature vectors: [0.85, 0.53] and [0.53, -0.85], with corresponding singular values ​​of 12.7 and 3.2. Dynamic weight allocation is performed on the feature vector set based on the health status feature vector of the energy storage system. The energy storage health status feature vector is calculated by monitoring parameters such as the real-time state of charge, cycle count, and temperature of the energy storage system. For example, the current health status feature vector of the energy storage system is [0.92, 0.78, 0.85], indicating that the energy storage system is in a good operating state. The health status feature vector and the feature vector set are multiplied by an inner product to obtain weight coefficients of 0.73 and 0.27, respectively. The feature vectors are then weighted and synthesized according to the weight coefficients to obtain the power fluctuation feature vector [0.76, 0.21]. The direction and magnitude of this vector represent the dominant direction and intensity of the power fluctuation, respectively, with a calculated magnitude of 0.79.

[0044] Based on the power fluctuation characteristic vector, the energy storage power range is divided into five intervals: maximum discharge interval (-2.5MW, -1.5MW), medium discharge interval (-1.5MW, -0.5MW), balance interval (-0.5MW, 0.5MW), medium charging interval (0.5MW, 1.5MW), and maximum charging interval (1.5MW, 2.5MW). The interval division takes into account the direction and magnitude of the power fluctuation characteristic vector. The vector direction determines the offset direction of the interval, and the magnitude determines the width adjustment coefficient of the interval.

[0045] Statistical analysis was performed on the difference data within each energy storage power range to calculate the data distribution characteristic parameters. Taking the medium charging range as an example, there are 53 sampling points in this range, with a calculated mean of 0.95MW, a standard deviation of 0.26MW, a skewness of 0.12, and a kurtosis of 2.86. The power fluctuation characteristic vector was used as a correction factor to adjust the distribution characteristic parameters. Specifically, the magnitude of the power fluctuation characteristic vector was used as the correction coefficient for the standard deviation, and the vector direction angle was used as the basis for correcting the skewness. The corrected standard deviation is 0.26 × 0.79 = 0.21MW, and the skewness is adjusted to 0.12 × (1 + 0.21 / 0.76) = 0.15. Based on the corrected characteristic parameters, a corrected probability distribution function was constructed using the kernel density estimation method.

[0046] The statistical characteristics of the corrected probability distribution function are calculated, including the expected value, variance, quintiles, and 95th percentiles. The calculated statistical characteristics of the corrected probability distribution function for the medium charging range are: expected value 0.93MW, variance 0.044, quintiles 0.59MW, and 95th percentile 1.27MW. Based on the charge and discharge response characteristics of the energy storage system, risk assessment thresholds are determined. The energy storage system has a charging response rate of 0.3MW / min, a discharging response rate of 0.35MW / min, a maximum continuous charging time of 120 minutes, and a maximum continuous discharging time of 90 minutes. Based on these parameters, an upper threshold of 1.3MW and a lower threshold of 0.55MW are set for risk assessment in the medium charging range. A risk warning is triggered when the real-time difference exceeds the threshold range.

[0047] Based on the risk assessment threshold and the corrected probability distribution function, a power fluctuation risk assessment index is generated. This index includes two dimensions: fluctuation risk probability and fluctuation impact intensity. The fluctuation risk probability is calculated as the probability that the difference sequence exceeds the threshold range. The fluctuation risk probability for the medium charging range is 0.13, indicating a 13% probability of power fluctuations exceeding the threshold. The fluctuation impact intensity is calculated as the average deviation when exceeding the threshold. The fluctuation impact intensity for the medium charging range is 0.22MW, indicating an average deviation of 0.22MW when exceeding the threshold.

[0048] The power fluctuation feature vector is updated based on real-time operational data to dynamically adjust the power fluctuation risk assessment indicators. Real-time data is processed using a sliding window approach, with a window length of 60 minutes and a sliding step of 5 minutes. After each window slide, the power fluctuation feature vector within the current window is recalculated and subjected to an exponentially weighted average with historical feature vectors, with the updated weight set to 0.3. The updated power fluctuation feature vector is used to adjust the risk assessment threshold and correct the probability distribution function, thus achieving dynamic adjustment of the risk assessment indicators. For example, if the magnitude of the power fluctuation feature vector increases by more than 20% for three consecutive windows, the sensitivity of the risk assessment threshold is increased accordingly, the trigger threshold is lowered, and an early risk warning is issued.

[0049] like Figure 3 As shown, the data flow diagram for photovoltaic energy storage power fluctuation risk assessment is displayed.

[0050] In one optional implementation, identifying fluctuation transition points using a dual threshold comparison method includes: Determine the instantaneous power change rate based on power time-series data; The instantaneous power change rate is accumulated and summed within the sliding time window to obtain the cumulative power change. The duration for which the sign of the instantaneous power change rate remains the same is recorded to obtain the fluctuation persistence index. Based on the statistical distribution characteristics of power time series data and the charging and discharging power limit of energy storage equipment, the upper and lower thresholds of power change rate are determined. Based on the capacity constraint of energy storage equipment and the duration of the sliding time window, the threshold of cumulative power change is determined. Determine whether the instantaneous power change rate exceeds the upper and lower thresholds of the power change rate, and whether the cumulative power change exceeds the threshold of the cumulative power change. Within the time period that satisfies the double threshold conditions, find the extreme point of the instantaneous power change rate and record it as a candidate conversion point. Calculate the time interval and power change amplitude of adjacent candidate switching points, set a time threshold based on the fluctuation persistence index, set an amplitude threshold based on the power change rate threshold, and delete candidate switching points whose time interval is less than the time threshold and whose power change amplitude is less than the amplitude threshold; The candidate transition points that have been screened are determined as power fluctuation transition points, and the power time series data is segmented using the power fluctuation transition points as boundaries.

[0051] In one specific implementation, power time-series data is collected. Taking a photovoltaic power station as an example, the sampling period is 1 second, and power data is continuously collected for 24 hours, totaling 86,400 data points. The collected power data is in kilowatts (kW), with values ​​ranging from 0 to 5000 kW. The data shows that power fluctuations are more pronounced during peak daytime power generation.

[0052] Based on the collected power time-series data, the instantaneous power change rate is calculated. The instantaneous power change rate is defined as the ratio of the difference in power values ​​between adjacent time points to the time interval. For example, if the power at time t is 4200 kW and the power at time t+1 is 4250 kW, and the sampling interval is 1 second, then the instantaneous power change rate at time t is (4250-4200) / 1 = 50 kW / s. This calculation is performed on all 86,400 data points, resulting in 86,399 instantaneous power change rate values.

[0053] To assess the power change trend over a period of time, a sliding time window is used to accumulate and sum the instantaneous power change rates. In this embodiment, the sliding window length is set to 30 seconds. For any time t, the sum of the instantaneous power change rates from t to t+29 is calculated to obtain the cumulative power change within the window. For example, if the instantaneous power change rates from t to t+29 are 50, 45, 30, -10, 15... kW / s, and their sum is 2000 kW, then the cumulative power change at time t is 2000 kW. Simultaneously, the duration for which the instantaneous power change rate sign remains the same is recorded. For example, if the instantaneous power change rate is positive for 10 consecutive sampling points, the fluctuation duration index is 10 seconds.

[0054] The power change rate threshold was determined based on the statistical characteristics of power time-series data. Analysis of the 24-hour power change rate distribution revealed that 95% of the instantaneous power change rate fell within the range of -100kW / s to 100kW / s. Considering the energy storage device's charging and discharging power limit of 1000kW, an upper threshold for the power change rate was set at 100kW / s, and a lower threshold at -100kW / s. Based on the energy storage device's capacity constraint (e.g., 2000kWh) and the sliding window duration (30 seconds), a cumulative power change threshold of 1500kW was set. This means that if the cumulative power change exceeds 1500kW within a 30-second window, charging and discharging regulation of the energy storage device will be considered.

[0055] The system then enters a dual-threshold judgment phase, examining the power time-series data point by point. For each time t, it checks whether the instantaneous power change rate exceeds the set upper and lower thresholds (-100kW / s to 100kW / s), and simultaneously checks whether the cumulative power change within the sliding window starting from that time exceeds 1500kW. When both conditions are met, the extreme point of the instantaneous power change rate is searched within that time period. For example, if the dual-threshold conditions are met within the time period from t to t+40, and the instantaneous power change rate reaches a local maximum of 130kW / s at time t+15, then time t+15 is recorded as a candidate transition point.

[0056] The initially identified candidate conversion points are screened. The time interval and power change amplitude between adjacent candidate conversion points are calculated. Based on the statistical fluctuation persistence index, a time threshold of 20 seconds is set, meaning that at least 20 seconds should pass between valid conversion points with power fluctuations. Simultaneously, an amplitude threshold of 500kW is set based on the power change rate threshold, meaning that the power change between adjacent conversion points should reach at least 500kW to be considered a valid conversion. For example, if the interval between candidate conversion points A and B is only 10 seconds and the power difference is 300kW, the one with the smaller power change rate is deleted.

[0057] The candidate transition points, after screening, were determined as the final power fluctuation transition points. In the actual case, 127 power fluctuation transition points were identified from the original 86,400 data points using a dual-threshold comparison method. These transition points clearly marked the alternation positions of power rise and fall segments, and each transition point recorded a timestamp and the corresponding power value.

[0058] Power time-series data is segmented using defined power fluctuation transition points as boundaries. For example, if two adjacent transition points are located at t1=3600 seconds and t2=3750 seconds respectively, then the 150 data points between t1 and t2 are divided into one power segment. After segmenting all data, 128 power segments are obtained (one more than the number of transition points). The data within each power segment exhibits a relatively consistent trend, which can be used for subsequent energy storage dispatch decisions or renewable energy power forecasting.

[0059] This dual-threshold comparison method can effectively identify key transition points in power fluctuations, avoiding misjudgments caused by relying solely on power change rate or cumulative amount, improving the accuracy of power fluctuation characteristic analysis, and providing a reliable basis for the optimized control of energy storage systems.

[0060] In one optional implementation, multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on a power fluctuation risk assessment index, advantageous and disadvantageous search paths are divided. The energy storage charging and discharging power time series is obtained through iterative search and identification of local optimal solutions, including: A uniformly distributed initial search point set is constructed within the feasible domain of energy storage charging and discharging power, and multiple parallel search paths are generated based on the initial search point set; Calculate the power fluctuation risk assessment index and its changing trend for each parallel search path, and determine the advantageous and disadvantageous search paths based on the dynamic weighted average. For each parallel search path, perform an iterative search, extract the search step size and search direction of the advantageous search path as feature parameters, and update the search step size and search direction of the disadvantageous search path. Record the change in evaluation value of the parallel search path within a preset number of consecutive iterations. When the change in evaluation value is less than a preset floating threshold, the current search point is recorded as a local optimum. A new search starting point is generated by randomly perturbing the neighborhood of the current search point. The search is performed again based on the new search starting point. The set of local optima that satisfy the feasible domain constraint of the energy storage charging and discharging power is taken as a candidate solution. The solution with the best evaluation value among the candidate solutions is selected to generate the energy storage charging and discharging power timing sequence, and the energy storage charging and discharging power timing sequence is adjusted online based on real-time operating data.

[0061] Within the feasible region of energy storage charging and discharging power, a uniformly distributed set of initial search points is constructed. The feasible region is assumed to consist of the upper and lower limits of the energy storage system's power and energy, for example, a power range of [-5MW, 5MW] and an energy range of [0, 10MWh]. To construct the initial search point set, the feasible region is divided according to a preset interval. For day-ahead scheduling at 96 time points, 20 initial search points can be generated within the feasible region. Each search point is represented as a 96-dimensional vector, representing the charging and discharging power over the 96 time periods throughout the day. These initial points can be obtained using the Latin hypercube sampling method to ensure the uniformity of the point distribution.

[0062] Based on the initial set of search points, multiple parallel search paths are generated. Each initial search point serves as the starting point of a search path, and the search process will explore feasible regions in different directions from these starting points. For example, for 20 initial search points, 20 parallel search paths will be generated. Each path evolves independently, but can share and exchange information during the search process.

[0063] A power fluctuation risk assessment index is calculated for each parallel search path. This index comprehensively considers the smoothness of energy storage charging and discharging power, the effect of grid power fluctuation suppression, and energy storage utilization efficiency. Specifically, the assessment index can be expressed as a weighted combination of the rate of change of energy storage charging and discharging power over a continuous period, the standard deviation of grid power fluctuation, and energy storage utilization efficiency. For example, if the charging and discharging power time series of a certain path at the initial point is P=[1.2, 1.5, -0.8, -1.2, ...] MW, its power fluctuation risk assessment index value is calculated to be 0.85.

[0064] Based on the magnitude and trend of power fluctuation risk assessment index values, a dynamic weighted average method is used to determine advantageous and disadvantageous search paths. Paths with assessment index values ​​lower than the average of all paths are defined as advantageous paths, while those higher than the average are defined as disadvantageous paths. For example, among 20 paths, the assessment index values ​​are [0.85, 0.92, 0.79, 1.05, ...], with a calculated average of 0.88. Paths with index values ​​of 0.85 and 0.79 are marked as advantageous paths, while the path with an index value of 1.05 is marked as a disadvantageous path.

[0065] Iterative search is performed on all parallel search paths. In each iteration, the search step size and search direction of the dominant search path are extracted as feature parameters. For example, the current search step size of a dominant path is 0.2MW, and the search direction vector is D=[0.1, -0.05, 0.08, ...]. These feature parameters are used to update the search behavior of the subordinate search path, causing the subordinate path to move closer to the dominant region. Specifically, the update method is: new step size of subordinate path = original step size of subordinate path × 0.5 + average step size of dominant path × 0.5, new direction of subordinate path = original direction of subordinate path × 0.7 + average direction of dominant path × 0.3.

[0066] During the iteration process, the change in evaluation value of each parallel search path is recorded within a preset number of consecutive iterations. When the change in evaluation value of a path within 10 consecutive iterations is less than a preset floating threshold of 0.01, the current search point is recorded as a local optimum. For example, if the evaluation values ​​of a path in iterations 50-60 are [0.782, 0.781, 0.780, 0.779, 0.779, 0.778, 0.778, 0.778, 0.777, 0.777], with a maximum change of 0.005, which is less than the threshold of 0.01, then the search point in iteration 60 is recorded as a local optimum.

[0067] For paths that find a local optimum, a random perturbation is applied within the neighborhood of the current search point to generate a new search starting point. The perturbation amplitude is a random value within ±5% of the current search point value. For example, if the charging / discharging power corresponding to the current local optimum is 2.5MW, then the new search starting point might be 2.5 + 2.5 × 5% × random(-1, 1) = 2.6MW. The search restarts based on the new starting point to avoid getting trapped in local optima.

[0068] After multiple iterations and re-searches, all local optimal solutions that satisfy the feasible domain constraint of energy storage charging and discharging power are selected as the candidate solution set. The constraints include upper and lower limits of energy storage power, upper and lower limits of energy, and state of charge range. For example, the power constraint is [-5MW, 5MW], the energy constraint is [0, 10MWh], and the state of charge range is [20%, 90%].

[0069] From the candidate solution set, the solution with the best evaluation value is selected as the final energy storage charge-discharge power timing sequence. For example, if there are 5 locally optimal solutions in the candidate solution set with evaluation values ​​of [0.77, 0.82, 0.79, 0.85, 0.80], then the solution with an evaluation value of 0.77 is selected as the final solution, and the corresponding charge-discharge power timing sequence is P_final = [1.2, 1.5, -0.8, -1.2, ...] MW.

[0070] During real-time operation, the charging and discharging power timing of the energy storage system is adjusted online based on real-time operational data. When the actual operational deviation exceeds a preset threshold (e.g., 5%), the charging and discharging power timing for future periods is recalculated using the rolling optimization method, with the current system state as the initial condition and the aforementioned multi-path search method. For example, if the real-time measured power is 1.3MW and the planned power is 1.2MW, the deviation is 8.3%, exceeding the 5% threshold, triggering online adjustment. The updated power timing is P_adjusted = [1.3, 1.6, -0.7, -1.1, ...]MW.

[0071] The above methods can effectively generate energy storage charging and discharging power timing that takes into account power fluctuation risks, thereby improving the operating efficiency of energy storage systems and grid stability.

[0072] In one optional implementation, calculating the power fluctuation risk assessment index and its changing trend for each parallel search path, and determining the advantageous and disadvantageous search paths based on the dynamic weighted average, includes: The power fluctuation risk assessment index of multiple parallel search paths is obtained as the evaluation value; The rate of change is obtained by calculating the difference in the evaluation values ​​within a preset time window, and the acceleration of change is obtained by calculating the difference in the rate of change. The path score is obtained by weighting the evaluation value and the rate of change, summing the values, and multiplying the sum by the acceleration. The dynamic weighted average of the path score is calculated by adjusting the weighting coefficients of the evaluation value and the rate of change based on the real-time status. Parallel search paths with path scores less than or equal to the dynamic weighted average are identified as inferior search paths, while parallel search paths with path scores greater than the dynamic weighted average are identified as superior search paths.

[0073] When calculating the power fluctuation risk assessment index for multiple parallel search paths, it is necessary to obtain the power fluctuation risk assessment index for each parallel search path as an evaluation value. These evaluation indicators typically include parameters such as the amplitude, frequency, and waveform characteristics of power fluctuations. For example, for three different parallel search paths, power fluctuation data can be collected over a period of time, resulting in an evaluation value of 0.85 for path 1, 0.72 for path 2, and 0.91 for path 3. These evaluation values ​​reflect the stability performance of each path; higher values ​​indicate lower power fluctuation risk and more stable system operation.

[0074] The changes in evaluation values ​​are calculated within a preset time window. Assuming a 10-minute time window with data collected every minute, the evaluation value sequence for path 1 at time points t1 to t10 is [0.85, 0.83, 0.86, 0.89, 0.87, 0.84, 0.82, 0.85, 0.88, 0.89]. The difference in evaluation values ​​between adjacent time points is calculated, yielding a rate of change sequence of [-0.02, 0.03, 0.03, -0.02, -0.03, -0.02, 0.03, 0.03, 0.01]. Further calculation of the rate of change difference yields an acceleration sequence of [0.05, 0, -0.05, -0.01, 0.01, 0.05, 0, -0.02]. These rate of change and acceleration data reflect the dynamic changes in the path's evaluation values.

[0075] Based on the calculated evaluation values ​​and rates of change, the overall performance of each path is calculated using a weighted summation method. The initial weighting coefficient for the evaluation value is set to 0.6, and the initial weighting coefficient for the rate of change is set to 0.4. Taking time t10 as an example, the evaluation value of path 1 is 0.89, the rate of change is 0.01, and the acceleration is -0.02. The path score for path 1 is calculated as: (0.89 × 0.6 + 0.01 × 0.4) × (-0.02) = -0.0108. Similarly, the path score for path 2 is calculated to be 0.0135, and the path score for path 3 is 0.0089. The path scores reflect the overall performance of each search path, taking into account the current evaluation value, the trend of change, and the acceleration characteristics of the change.

[0076] The weighting coefficients are dynamically adjusted based on the real-time operating status. When the system is stable, more emphasis is placed on the current evaluation value; the weight of the evaluation value can be increased to 0.75, while the weight of the rate of change can be decreased to 0.25. When the system is experiencing significant fluctuations, more attention is paid to the trend of change; the weight of the evaluation value can be decreased to 0.45, while the weight of the rate of change can be increased to 0.55. Assuming the system detects significant fluctuations, after adjusting the weights, the path score for path 1 is recalculated as: (0.89 × 0.45 + 0.01 × 0.55) × (-0.02) = -0.0089. Similarly, the path score for path 2 is calculated to be 0.0142, and the path score for path 3 is 0.0075.

[0077] To comprehensively consider path scores across multiple time points, a dynamically weighted average of the path scores is calculated. Within a time window, different weights are assigned to the path scores at each time point, with scores closer to the current time having higher weights. The weights for t1 to t10 are set as [0.05, 0.05, 0.06, 0.07, 0.08, 0.10, 0.12, 0.15, 0.16, 0.16]. The calculated dynamically weighted average is 0.0056 for path 1, 0.0102 for path 2, and 0.0083 for path 3. The dynamically weighted average reflects the overall performance of the search path throughout the entire time window.

[0078] Based on the comparison between path scores and dynamic weighted averages, dominant and suboptimal search paths are determined. The average of the dynamic weighted averages of the three paths is calculated, yielding a threshold of 0.0080. Path 1, with a score of 0.0056, is below the threshold and is therefore classified as a suboptimal search path; Path 2, with a score of 0.0142, is above the threshold and is classified as a dominant search path; Path 3, with a score of 0.0075, is below the threshold and is classified as a suboptimal search path. Resource allocation to the dominant search path 2 will be further strengthened, while resource investment in suboptimal search paths 1 and 3 will be reduced to improve overall search efficiency.

[0079] The determination of superior and inferior paths is updated periodically. Every preset update cycle, such as 30 minutes, the above calculation process is repeated to obtain the latest path scores and dynamic weighted averages, thus updating the determination of superior and inferior paths. For example, if after one update cycle, the dynamic performance of path 1 improves, and its new path score is 0.0124, higher than the updated threshold of 0.0095, then path 1 changes from a inferior search path to an advantageous search path. This dynamic update mechanism ensures timely response to changes in the performance of each search path, optimizing resource allocation strategies.

[0080] By calculating and analyzing the above power fluctuation risk assessment indicators, we can effectively identify advantageous search paths with stability and development potential, while eliminating inferior search paths with poor performance, thereby achieving a reasonable allocation of parallel search resources and improving utilization efficiency.

[0081] The photovoltaic energy storage joint optimization scheduling system based on deep reinforcement learning, as described in this embodiment of the invention, includes: The first unit is used to collect photovoltaic power generation data, energy storage operation status data and grid load data to form a multi-source heterogeneous time series dataset. The second unit is used to extract energy storage features based on multi-source heterogeneous time-series datasets, through a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, and generate an energy storage health status feature vector; based on the energy storage health status feature vector, the comprehensive power feature value and capacity threshold are calculated, and the safe operating range of energy storage is determined through a two-dimensional coordinate system. The third unit is used to divide the energy storage power range according to the safe operating range of energy storage, and to construct the feasible domain of energy storage charging and discharging power in combination with the physical constraints of energy storage. The fourth unit is used to calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; the difference sequence is segmented and power fluctuation feature vectors are extracted; and power fluctuation risk assessment indicators are constructed based on the power fluctuation feature vectors to correct the probability distribution parameters. The fifth unit is used to generate multiple parallel search paths within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, it divides the search paths into advantageous and disadvantageous ones, and obtains the energy storage charging and discharging power time series through iterative search and identification of local optimal solutions. The sixth unit is used to generate an energy storage scheduling instruction sequence based on the energy storage charging and discharging power timing.

[0082] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0083] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0084] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A photovoltaic energy storage joint optimization scheduling method based on deep reinforcement learning, characterized in that, include: Collect photovoltaic power generation data, energy storage operation status data, and grid load data to form a multi-source heterogeneous time-series dataset; Based on multi-source heterogeneous time-series datasets, energy storage features are extracted using a recurrent graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, generating an energy storage health status feature vector; according to The comprehensive power characteristic value and capacity threshold are obtained by calculating the energy storage health status feature vector, and the safe operating range of energy storage is determined by the two-dimensional coordinate system. Based on the safe operating range of energy storage, the energy storage power range is divided, and combined with the physical constraints of energy storage, a feasible domain for energy storage charging and discharging power is constructed. Calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; segment the difference sequence and extract the power fluctuation feature vector; construct a power fluctuation risk assessment index based on the power fluctuation feature vector and correct the probability distribution parameters. Multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, the advantageous search path and the disadvantageous search path are divided. The energy storage charging and discharging power time series is obtained through iterative search and identification of local optimal solutions. A sequence of energy storage scheduling instructions is generated based on the energy storage charging and discharging power timing.

2. The method according to claim 1, characterized in that, Based on a multi-source heterogeneous time-series dataset, energy storage features are extracted using a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure. The resulting energy storage health status feature vector includes: The multi-source heterogeneous time-series dataset is input into the dual encoder-decoder network structure. The initial feature matrix is ​​obtained by extracting time-series features and fusing state features, which contains time-series information and state information of energy storage operation. A dynamic time-series graph structure is constructed based on the initial feature matrix. The nodes in the dynamic time-series graph structure correspond to the energy storage operation state at different times. The weights between the nodes in the dynamic time-series graph structure are calculated to obtain the time-series correlation weight matrix through the thermodynamic constraint adaptive attention mechanism. A recurrent graph convolutional network is constructed based on the temporal correlation weight matrix. The convolution kernel parameters in the recurrent graph convolutional network are updated by thermodynamically constrained recurrent units. The initial hidden state of the recurrent units is set according to the thermodynamic constraints. Graph convolution operation is performed to obtain thermodynamic temporal features. The thermodynamic temporal features are then fused with the initial feature matrix through residual connections to obtain fused enhanced features. The fused enhanced features are decomposed into multi-scale features to obtain a multi-scale feature representation that includes global operational features, local fluctuation features, and trend features. The multi-scale feature representation is input into a feature mapping network and reduced to a low-dimensional manifold space through nonlinear projection transformation to obtain the energy storage health status feature vector.

3. The method according to claim 1, characterized in that, Based on the energy storage health status feature vector, the comprehensive power characteristic value and capacity threshold are calculated, and the safe operating range of energy storage is determined using a two-dimensional coordinate system, including: The energy storage health status feature vector is decomposed into power feature components and capacity feature components through nonlinear mapping. The power feature components include charging power components and discharging power components. The maximum charging power and maximum discharging power are determined by multiplying the charging power component and the discharging power component by the coupling coefficients of the energy storage internal resistance and temperature, respectively, and the comprehensive power characteristic value is calculated. The capacity feature component is multiplied by the weighting coefficients of the number of iterations and the depth of iterations, and then summed to obtain the capacity threshold. A two-dimensional coordinate system is established with the comprehensive power feature value as the vertical axis and the capacity threshold as the horizontal axis. Multiple feature points are selected as regional center points in the two-dimensional coordinate system. A regional division map is constructed based on the regional center points. The distance from any point in the two-dimensional coordinate system to each regional center point is calculated, and the region corresponding to the nearest regional center point is taken as the region to which the corresponding point belongs. A classifier is constructed using a kernel function to determine the classification boundaries of different regions. Based on the membership degree calculation, the comprehensive power feature value and the capacity threshold are mapped to different regions of the two-dimensional coordinate system. The classification boundaries of the different regions determine the safe operating range of energy storage.

4. The method according to claim 1, characterized in that, Calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time-series dataset; segment the difference sequence and extract power fluctuation feature vectors; construct risk assessment indicators based on the power fluctuation feature vectors after correcting the probability distribution parameters, including: Calculate the difference sequence between photovoltaic power generation data and grid load data; The power fluctuation trend characteristics of the difference sequence are calculated. A fluctuation feature vector is constructed based on the power rise rate, fall rate, and duration. The fluctuation transition point is identified using a double threshold comparison method. The difference sequence is divided into multiple feature subsequences with the fluctuation transition point as the boundary. The first-order difference and second-order difference of each feature subsequence are calculated to construct a difference feature matrix. The difference feature matrix is ​​subjected to singular value decomposition to obtain a feature vector group. The feature vector group is dynamically weighted based on the energy storage health status feature vector to obtain a power fluctuation feature vector. The energy storage power interval is divided according to the direction and magnitude of the power fluctuation feature vector. The difference data within the energy storage power range is statistically analyzed, the data distribution characteristic parameters are calculated, and the power fluctuation characteristic vector is used as a correction factor to adjust the distribution characteristic parameters to obtain the corrected probability distribution function. Calculate the statistical characteristics of the modified probability distribution function, determine the risk assessment threshold by combining the energy storage charging and discharging response characteristics, and generate power fluctuation risk assessment indicators. The power fluctuation feature vector is updated based on real-time operating data, and the power fluctuation risk assessment index is dynamically adjusted.

5. The method according to claim 4, characterized in that, Identifying fluctuation transition points using the dual threshold comparison method includes: Determine the instantaneous power change rate based on power time-series data; The instantaneous power change rate is accumulated and summed within the sliding time window to obtain the cumulative power change. The duration for which the sign of the instantaneous power change rate remains the same is recorded to obtain the fluctuation persistence index. Based on the statistical distribution characteristics of power time series data and the charging and discharging power limit of energy storage equipment, the upper and lower thresholds of power change rate are determined. Based on the capacity constraint of energy storage equipment and the duration of the sliding time window, the threshold of cumulative power change is determined. Determine whether the instantaneous power change rate exceeds the upper and lower thresholds of the power change rate, and whether the cumulative power change exceeds the threshold of the cumulative power change. Within the time period that satisfies the double threshold conditions, find the extreme point of the instantaneous power change rate and record it as a candidate conversion point. Calculate the time interval and power change amplitude of adjacent candidate switching points, set a time threshold based on the fluctuation persistence index, set an amplitude threshold based on the power change rate threshold, and delete candidate switching points whose time interval is less than the time threshold and whose power change amplitude is less than the amplitude threshold; The candidate transition points that have been screened are determined as power fluctuation transition points, and the power time series data is segmented using the power fluctuation transition points as boundaries.

6. The method according to claim 1, characterized in that, Multiple parallel search paths are generated within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, advantageous and disadvantageous search paths are divided. Through iterative search and identification of local optimal solutions, the energy storage charging and discharging power time series is obtained, including: A uniformly distributed initial search point set is constructed within the feasible domain of energy storage charging and discharging power, and multiple parallel search paths are generated based on the initial search point set; Calculate the power fluctuation risk assessment index and its changing trend for each parallel search path, and determine the advantageous and disadvantageous search paths based on the dynamic weighted average. For each parallel search path, perform an iterative search, extract the search step size and search direction of the advantageous search path as feature parameters, and update the search step size and search direction of the disadvantageous search path. Record the change in evaluation value of the parallel search path within a preset number of consecutive iterations. When the change in evaluation value is less than a preset floating threshold, the current search point is recorded as a local optimal solution. A new search starting point is generated by randomly perturbing the neighborhood of the current search point. The search is performed again based on the new search starting point. The set of local optimal solutions that satisfy the feasible domain constraint of the energy storage charging and discharging power is taken as a candidate solution. The solution with the best evaluation value among the candidate solutions is selected to generate the energy storage charging and discharging power timing sequence, and the energy storage charging and discharging power timing sequence is adjusted online based on real-time operating data.

7. The method according to claim 6, characterized in that, Calculate the power fluctuation risk assessment index and its changing trend for each parallel search path, and determine the advantageous and disadvantageous search paths based on the dynamic weighted average, including: The power fluctuation risk assessment index of multiple parallel search paths is obtained as the evaluation value; The rate of change is obtained by calculating the difference in the evaluation values ​​within a preset time window, and the acceleration of change is obtained by calculating the difference in the rate of change. The path score is obtained by weighting the evaluation value and the rate of change, summing the values, and multiplying the sum by the acceleration. The dynamic weighted average of the path score is calculated by adjusting the weighting coefficients of the evaluation value and the rate of change based on the real-time status. Parallel search paths with path scores less than or equal to the dynamic weighted average are identified as inferior search paths, while parallel search paths with path scores greater than the dynamic weighted average are identified as superior search paths.

8. A photovoltaic energy storage joint optimization scheduling system based on deep reinforcement learning, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to collect photovoltaic power generation data, energy storage operation status data and grid load data to form a multi-source heterogeneous time series dataset. The second unit is used to extract energy storage features based on multi-source heterogeneous time-series datasets, through a recursive graph convolutional network with a dual encoder-decoder network structure and a dynamic time-series graph structure, and generate an energy storage health status feature vector. The comprehensive power characteristic value and capacity threshold are calculated based on the energy storage health status characteristic vector, and the safe operating range of energy storage is determined by a two-dimensional coordinate system. The third unit is used to divide the energy storage power range according to the safe operating range of energy storage, and to construct the feasible domain of energy storage charging and discharging power in combination with the physical constraints of energy storage. The fourth unit is used to calculate the difference sequence between photovoltaic power generation data and grid load data in a multi-source heterogeneous time series dataset; the difference sequence is segmented and power fluctuation feature vectors are extracted; and power fluctuation risk assessment indicators are constructed based on the power fluctuation feature vectors to correct the probability distribution parameters. The fifth unit is used to generate multiple parallel search paths within the feasible domain of energy storage charging and discharging power. Based on the power fluctuation risk assessment index, it divides the search paths into advantageous and disadvantageous ones, and obtains the energy storage charging and discharging power time series through iterative search and identification of local optimal solutions. The sixth unit is used to generate an energy storage scheduling instruction sequence based on the energy storage charging and discharging power timing.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Distributed power supply optimization scheduling method and system based on demand side response

    CN120090295A

  • Traffic flow prediction method based on dynamic graph convolution circulation network

    CN120279714A

  • Distributed photovoltaic and energy storage combined planning method based on deep learning

    CN120471207A

  • Photovoltaic energy storage system power scheduling optimization method based on deep reinforcement learning

    CN120582254A

Cited By

  • Megawatt photovoltaic and energy storage integrated efficient energy management system

    CN121308120A

  • Method and system for evaluating risk of wind-solar-energy-storage combined power generation system

    CN122198673A