A method and system for site selection and capacity determination of energy storage in a substation considering uncertainty
Through deep reinforcement learning and improved particle swarm optimization algorithm combined with conditional generation adversarial network, the problem of uncertainty in distributed energy storage configuration is solved, the most preferred location and capacity of the energy storage system in the station area is realized, and the economic and security of the system is improved.
Patent Information
- Application Number
- CN202510854692.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing distributed energy storage configuration methods do not fully consider uncertainty factors, resulting in poor robustness and insufficient adaptability. The optimization algorithm is prone to local optimality, making it difficult to obtain the overall optimal solution of the system, and lacks the ability to deal with emergencies.
Deep reinforcement learning and improved particle swarm optimization algorithm are used to optimize collaboratively, and a conditional generation adversarial network is combined to simulate multiple uncertain scenarios, and the operational risk of the power system is quantitatively evaluated, and the objective function is constructed to solve the most preferred addressing and capacity scheme.
The economy, safety and operational adaptability of the energy storage system are significantly improved. Through the dynamic balance and robustness of the global search capabilities, local optimal traps are avoided, and the optimization results take into account both operational loss and system risk control.
Smart Images

Figure CN120372874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to power system planning and operation, and in particular to a method and system for site selection and sizing of energy storage in a substation considering uncertainty. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the widespread integration of renewable energy sources such as distributed photovoltaic and wind power, modern power distribution substations face increasingly severe challenges in operational stability and economic efficiency. Distributed energy storage, with its rapid response and flexible regulation, plays a key role in smoothing power fluctuations, improving voltage quality, shifting peaks and filling valleys, and reducing network losses. Therefore, scientific and rational site selection and sizing of distributed energy storage are crucial for optimizing substation operations and promoting the integration of renewable energy.
[0004] The inventors found in their research that the current mainstream distributed energy storage configuration methods mostly use deterministic models, which do not fully consider the impact of uncertain factors such as the randomness of distributed power output, load fluctuations, and extreme weather, resulting in poor robustness and insufficient adaptability of the configuration scheme in actual operation. At the same time, some optimization algorithms are prone to falling into local optimality, making it difficult to obtain the overall optimal solution for the system, and the assessment and control mechanism for operational risks is not sound, lacking the ability to cope with sudden working conditions, making it difficult to ensure the safe and stable operation of the system. Taking the traditional particle swarm optimization algorithm as an example, this algorithm is widely used to solve the problem of site selection and sizing of distributed energy storage due to its simple structure and easy implementation. However, the particle swarm algorithm is essentially a heuristic random search method with strong local search capabilities, but limited global search capabilities. When faced with problems with high objective function dimensions, complex feasible domains, or multiple local optimal solutions, the particle swarm in the particle swarm algorithm is prone to falling into local optimality, especially when the search speed slows down in the later stages of the algorithm iteration, making it difficult to jump out of the local extreme point, and the energy storage configuration scheme finally obtained may not be optimal. In addition, particle swarm optimization is usually based on static or simplified models, which makes it difficult to provide a robust and risk-controlled solution for the system. Summary of the Invention
[0005] To address the above-mentioned issues, the present invention proposes a method and system for energy storage site selection and sizing in a substation area that considers uncertainty. This method employs Deep Reinforcement Learning (DRL) and an improved Particle Swarm Optimization (PSO) algorithm for collaborative optimization, combines a Conditional Generative Adversarial Network (CGAN) to simulate various uncertainty scenarios, and utilizes Conditional Value at Risk (CVaR) to quantitatively assess power system operational risks, effectively improving the optimization algorithm's global search capabilities and the robustness of the solution. This method achieves optimal site selection and sizing for distributed energy storage in substations, significantly reducing network losses and enhancing the system's economic efficiency and reliability.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] One or more embodiments provide a method for site selection and sizing of energy storage in a substation considering uncertainty, including the following steps:
[0008] Obtain substation power data and environmental condition information and perform pre-processing;
[0009] Based on the acquired random disturbance noise and pre-processed environmental condition information, multiple uncertainty scenarios are generated, including distributed power output and load demand scenarios;
[0010] Clustering is performed on the obtained uncertainty scenarios;
[0011] Calculate network loss based on clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area;
[0012] An improved particle swarm optimization algorithm is combined with a deep reinforcement learning strategy. During the iteration process of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards. The objective function is solved to obtain the optimal location and capacity solution.
[0013] One or more embodiments provide a system for determining the site selection and capacity of energy storage in a substation considering uncertainty, including:
[0014] A data acquisition module is configured to acquire power data and environmental condition information of the substation area and perform preprocessing;
[0015] A scenario generation module is configured to generate multiple uncertainty scenarios, including distributed power output and load demand scenarios, based on the acquired random disturbance noise and pre-processed environmental condition information;
[0016] A clustering module is configured to perform clustering on the obtained uncertainty scenarios;
[0017] The target construction module is configured to calculate the network loss based on the clustered uncertainty scenarios, and to construct the objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area;
[0018] The solution module is configured to adopt a method of collaborating with an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration process of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards to solve the objective function and obtain the optimal site location and capacity solution.
[0019] One or more embodiments provide a system for selecting and sizing energy storage in a substation considering uncertainty, including: a data acquisition device and a processor;
[0020] A data acquisition device for acquiring power data and environmental condition information of the substation area;
[0021] The processor is configured to execute the steps of the above-mentioned method for selecting and sizing the energy storage site in a substation considering uncertainty.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] The method of the present invention is significantly superior to traditional deterministic optimization methods in terms of uncertainty modeling, and can effectively cover boundary scenarios caused by extreme weather and load fluctuations. The amount of data is compressed by scene clustering, which improves the optimization efficiency. The particle swarm optimization framework integrated with deep reinforcement learning achieves a dynamic balance between local and global search capabilities, avoids falling into local optimality, improves the stability and convergence quality of the algorithm in high-dimensional, non-convex optimization space, and uses deep reinforcement learning strategies to evaluate the "reward" level of current particle behavior and adjust the fitness function of PSO in real time. When the particle swarm as a whole falls into a state of convergence stagnation, DRL will impose penalties on inferior solutions based on the reward function to guide particles out of local trap areas. The optimization results take into account both operating losses and system risk control to obtain a more robust energy storage site selection and sizing solution. Overall, the method of the present invention can significantly improve the economy, safety and operational adaptability of the substation energy storage system.
[0024] The advantages of the present invention and its additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0026] Figure 1 This is a flow chart of a method for selecting and sizing energy storage in a substation considering uncertainty according to Example 1 of the present invention;
[0027] Figure 2 2 is a schematic diagram of the structure of the CGAN model according to Example 1 of the present invention;
[0028] Figure 3 is a flow chart of a method for coordinating an improved particle swarm optimization algorithm (PSO) with a deep reinforcement learning (DRL) strategy according to Example 1 of the present invention; DETAILED DESCRIPTION
[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0030] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0031] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. It should be noted that, in the absence of conflict, the various embodiments of the present invention and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.
[0032] Example 1
[0033] In the technical solutions disclosed in one or more embodiments, Figures 1 to 3 As shown in FIG, a method for selecting and sizing energy storage sites in a substation considering uncertainty includes the following steps:
[0034] Step 1: Obtain the power data and environmental condition information of the substation area and perform preprocessing;
[0035] Step 2: Based on the acquired random disturbance noise and pre-processed environmental condition information, multiple uncertainty scenarios are generated, including distributed power output and load demand scenarios;
[0036] Step 3: Cluster the obtained uncertainty scenarios;
[0037] Step 4: Calculate network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area;
[0038] Step 5: Using an improved particle swarm optimization (PSO) algorithm and a deep reinforcement learning (DRL) strategy, during the iteration of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning (DRL) strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards. The objective function is solved to obtain the optimal location and capacity solution.
[0039] The method in this embodiment is significantly superior to traditional deterministic optimization methods in terms of uncertainty modeling, and can effectively cover boundary scenarios caused by extreme weather and load fluctuations. By clustering scenarios, the amount of data is compressed, thereby improving the optimization efficiency. The particle swarm optimization framework integrated with deep reinforcement learning achieves a dynamic balance between local and global search capabilities, avoids falling into local optimality, improves the stability and convergence quality of the algorithm in high-dimensional, non-convex optimization space, and uses deep reinforcement learning strategies to evaluate the "reward" level of current particle behavior, and adjusts the fitness function of PSO in real time. When the particle swarm as a whole falls into a state of convergence stagnation, DRL will impose penalties on inferior solutions based on the reward function, guiding particles to jump out of the local trap area. The optimization results take into account both operating losses and system risk control, and obtain a more robust energy storage site selection and sizing scheme. Overall, the method of the present invention can significantly improve the economy, safety and operational adaptability of the substation energy storage system.
[0040] In step 1, the power data of the substation area may include distributed power generation output data, line parameters (such as resistance and reactance), node voltage data, load demand data, etc.
[0041] Environmental condition information may include corresponding area time data, geographical condition data, meteorological data and other condition information;
[0042] In step 1, data collection and preprocessing can utilize the power system's data acquisition and supervision (SCADA) system to collect distributed power output data, line parameters (such as resistance and reactance), node voltage data, and load demand data for a set time period in the substation, such as the past year; at the same time, obtain information on the substation's time, geography, meteorology, and other conditions.
[0043] Optional data preprocessing methods include:
[0044] Step 11: Clean the acquired data, remove obvious outliers, and remove noise and outliers in the data; for missing values, use spline interpolation to fill them.
[0045] Step 12: Use the Z-score normalization method to normalize the data and map the data to the [0,1] interval to improve the training stability and convergence speed of the model. The specific formula is:
[0046] ;
[0047] in, 、 is the feature mean and variance.
[0048] Furthermore, for the time data, time feature encoding is performed, time is encoded into angles, and then mapped through trigonometric functions (sine / cosine functions) to obtain the encoded time features;
[0049] Specifically, map the 24 hours of a day to the angle of a circle:
[0050] ;
[0051] The sine and cosine values of the angle are expressed as:
[0052] ;
[0053] Map a 365-day year to degrees:
[0054] ;
[0055] The sine and cosine values of the angle are expressed as:
[0056] ;
[0057] The time is decomposed into sine / cosine signals through Fourier transform. The specific formula is as follows:
[0058] ;
[0059] The Fourier transform can decompose any periodic signal into the sum of sine / cosine signals of different frequencies. For the time series f(t), its Fourier series is expressed as:
[0060] ;
[0061] In the substation energy storage scenario of this solution, 、 、 are coefficients respectively, and only the fundamental component with k=1 is retained, that is 、 It can effectively characterize major periodic features such as day and night, seasons, etc., and higher-order harmonics can be ignored to reduce the dimension.
[0062] The operation of the power grid has significant periodic characteristics. Traditional timestamps (such as Unix time) as continuous variables are difficult to effectively represent periodicity. Timestamps (such as hours and days of the year) are converted into periodic features to capture the diurnal and seasonal patterns of power load and distributed power output, that is, mapping 24 hours a day to angles. , represented by its sine and cosine values, to avoid the sudden changes caused by directly using numerical values (such as the proximity of 23 o'clock and 0 o'clock); similarly, the 365 days of a year are mapped to angles , characterizing seasonal variations. This not only avoids the model's misjudgment of temporal continuity due to linear encoding, but also preserves periodic symmetry, improving its ability to capture periodic fluctuations.
[0063] This embodiment overcomes the problem of traditional numerical input methods disrupting the periodic structure of time by periodically encoding time data. This effectively avoids encoding abrupt changes, such as large numerical differences between adjacent periods like 23:00 and 0:00, or between 365 days and the first day, and prevents the model from misjudging temporal continuity. This encoding method preserves periodic characteristics such as diurnal and seasonal characteristics, helping the model more accurately identify cyclical fluctuations in power load and distributed generation output, significantly improving the model's ability to perceive cyclical changes and its prediction accuracy.
[0064] Furthermore, the preprocessing method for meteorological data is as follows: performing correlation analysis, calculating the correlation between distributed energy output and meteorological data, and eliminating weakly correlated meteorological factors;
[0065] Optionally, the Pearson correlation coefficient can be used to calculate the correlation between distributed energy output and meteorological data. The calculation formula is as follows:
[0066] ;
[0067] in, Represents distributed power generation and meteorological factors The Pearson correlation coefficient ranges from [-1, 1]; is the time series data of distributed generation (DG) output power, The time series data of meteorological factors must be synchronized; otherwise, the correlation coefficient will be distorted. Linear interpolation can be used to solve the missing data. for and The covariance of 、 is the standard deviation of the two.
[0068] Furthermore, a correlation threshold can be set, and the calculated correlation coefficient is compared with the set threshold to screen key meteorological factors; optionally, the correlation threshold is not less than 0.5;
[0069] The output of distributed power generation is significantly affected by meteorological conditions. In this embodiment, the key meteorological factors are selected by the Pearson correlation coefficient. By removing weakly correlated factors and reducing data dimensionality, the team was able to prevent irrelevant features from interfering with subsequent model training. This ensured that the scenarios generated by the subsequent conditional generative adversarial network (CGAN) complied with physical laws.
[0070] In step 2, the uncertainty scenario can be generated by the conditional generative adversarial network (CGAN), which takes the pre-processed environmental condition information as input and generates the uncertainty scenario of the distributed power output and load demand through the generator;
[0071] Specifically, a CGAN model is constructed, and the generator and discriminator are trained using the Adam optimizer. During training, the generator and discriminator are trained alternately for a set number of rounds, such as 1000. After the CGAN model is trained, the generator is used to generate multiple distributed power output and load demand scenarios.
[0072] In one feasible implementation, the constructed CGAN model includes a generator and a discriminator. The generator adopts a multi-layer perceptron (MLP) structure, including 3 hidden layers, each with 128 neurons. The discriminator also adopts an MLP structure, including 2 hidden layers, each with 64 neurons. During the training process, Figure 2 As shown in Figure 2, the input of the generator is a random noise vector and preprocessed environmental condition information, and the output is a simulated distributed power output and load demand scenario. The input of the discriminator is the real scene data and the scene data generated by the generator, and the output is a judgment on the authenticity of the input data.
[0073] Optionally, use the Adam optimizer to alternately train the generator and discriminator, so that the generator can generate uncertain scenes with a distribution similar to that of real scene data. The alternating training method is to fix the generator G and train the discriminator D to improve its discrimination ability; fix the discriminator D and train the generator G to make the generated scenes more realistic.
[0074] Specifically, the loss function for CGAN model training is designed as:
[0075] For the discriminator, cross entropy loss is used to measure its classification error for real and fake data.
[0076] For the generator, an adversarial loss function is used to improve the ability of generated data to "fool" the discriminator.
[0077] In this example, a conditional generative adversarial network (CGAN) model is constructed by incorporating information such as the time, geographic location, and meteorological conditions of the substation as generation conditions. This makes the generated distributed power output and load demand scenarios more targeted and controllable, significantly improving the matching of simulation data to the actual operating environment. CGAN can generate diverse and representative scenarios under different temporal, spatial, and meteorological conditions, effectively capturing the uncertainties of distributed energy output and load demand. This provides a high-quality and diverse input data foundation for subsequent energy storage optimization, thereby enhancing the robustness of the overall model and the reliability of the optimization results.
[0078] In step 3, the K-means++ algorithm is used to cluster the obtained uncertain scenes, including the following steps:
[0079] Step 31: Use the obtained uncertainty scenario as a data point and randomly select a data point as the first cluster center;
[0080] Step 32: Calculate the distance from each data point to the nearest cluster center and select the point with the largest distance as the new center;
[0081] Step 33: Repeat step 32 until the specified number of cluster centers are selected.
[0082] Step 34: Perform traditional K-means clustering based on the obtained cluster centers, optimize the intra-class variance, and obtain the clustered scene;
[0083] Step 35: For the clustered scenarios, select nodes with voltage fluctuation rates and load rates higher than set values as candidate locations for distributed energy storage devices;
[0084] According to the above algorithm, based on the voltage fluctuation rate and load rate, K-means++ clustering is used to select high-fluctuation and high-load nodes as candidate installation locations, which can be used for the subsequent particle swarm algorithm population construction;
[0085] The voltage fluctuation rate is: ;
[0086] The load rate is: ;
[0087] in, represents the voltage fluctuation rate of node i, N is the number of samples in the time dimension. If 96 moments in a day are selected, with one point every 15 minutes, then N=96; is the voltage value of node i at time t, is the average voltage of node i at N moments;
[0088] is the load rate of node i, is the maximum active load of node i during the observation period, is the rated active load capacity of node i.
[0089] In this example, the generated scenarios are clustered using the K-means++ algorithm to reduce the number of redundant scenarios and lower computational complexity, while preserving key features and improving optimization efficiency. A large number of scenarios generated by the CGAN model are fed into the K-means++ algorithm, where cluster analysis reduces redundant scenarios. This example optimizes the selection of initial cluster centers to avoid the randomness inherent in traditional K-means and improve clustering quality. The clustered scenarios serve as input for subsequent optimization, preserving the original data distribution while reducing computational complexity.
[0090] Step 4: For the clustered uncertainty scenarios, establish an objective function that considers network loss and conditional loss at risk (CVaR), as follows:
[0091] ;
[0092] in, is the total number of scenes after clustering, For the The network loss in the substation area under the scenario and are weight coefficients, representing the importance of network loss and conditional risk loss (CVaR) in the objective function. CVaR represents the conditional risk loss for all scenarios;
[0093] Specifically, for each clustered scenario, the grid loss of the substation is calculated based on circuit theory and power flow calculation methods. The classic forward-backward substitution method is used for power flow calculation to obtain the active power loss and reactive power loss of each line, and then the total grid loss of the entire substation is calculated. The specific formula for grid loss is as follows:
[0094] ;
[0095] in, and is the active / reactive power of line ij, is the node voltage, is the resistance on line ij; i and j represent nodes on the line, and the line between the two nodes is marked as line ij;
[0096] Specifically, the method for determining conditional risk loss is as follows: calculate the network loss values under all scenarios, sort them in ascending order, set a confidence level, determine a network loss threshold, and calculate the average network loss value for scenarios exceeding the network loss threshold as the conditional risk loss value (CVaR value). The confidence level can be set to no less than 95%.
[0097] The objective function of this embodiment sets a conditional risk loss term to collect tail risks in the uncertainty scenarios generated by CGAN and quantify the potential losses caused by extreme scenarios. Extreme scenarios may include sudden drops in wind power / photovoltaic output and surges in load. The specific formula for conditional risk loss CVaR is:
[0098] ;
[0099] in, To set the confidence level The maximum possible network loss under Indicates exceeding The loss part, that is, the additional loss in extreme cases, is the probability of scenario s occurring, Indicates the total number of scenes.
[0100] The PSO algorithm is an optimization algorithm based on swarm intelligence. It searches for optimal solutions by simulating the foraging behavior of flocks of birds or schools of fish. In the PSO algorithm, each particle represents a potential solution. As particles move through the search space, they update their speed and position based on their own historical optimal position and the historical optimal position of the group.
[0101] In this embodiment, in order to prevent the PSO algorithm from falling into a local optimal solution and to improve the algorithm's search capability and convergence speed, the improved particle swarm optimization algorithm (PSO) adopts an adaptive inertia weight, which is reduced in sequence according to the number of iterations. In this way, in the early stage of the algorithm, a larger inertia weight is set, so that particles can perform a global search in a larger search space; in the later stage of the algorithm, a smaller inertia weight is used, so that particles can perform a fine search in a local area.
[0102] Specifically, the adaptive inertia weight is calculated as follows:
[0103] ;
[0104] in, and are the maximum and minimum values of the inertia weight, is the maximum number of iterations, is the current iteration number.
[0105] Furthermore, in the improved particle swarm optimization algorithm, the mutation operation uses Gaussian mutation to add a random variable that obeys Gaussian distribution to the position value of the particle in the mutation dimension.
[0106] After a particle updates its position, certain dimensions of the particle are mutated with a certain probability, introducing new search directions and increasing population diversity. This embodiment uses Gaussian mutation to add random noise (mean 0, standard deviation 1) that follows a Gaussian distribution to the particle position. When a particle falls into a local optimal solution for a certain node combination, Gaussian mutation perturbs the energy storage node or capacity dimension with a certain probability (e.g., mutation probability pm = 0.1) to generate a new site selection and capacity solution. This prevents premature population maturation due to homogeneity, allowing the particle to escape the current local optimal area and search for the optimal solution again after the update.
[0107] Step 5: For the clustered scenarios, an improved particle swarm optimization algorithm (PSO) and a deep reinforcement learning (DRL) strategy are used to solve the objective function and obtain the optimal location and capacity solution, which includes the following steps:
[0108] Step 51: Initialize the particle swarm and randomly generate the initial position and velocity of each particle. The position of each particle represents a site selection and capacity determination scheme for distributed energy storage.
[0109] Step 52: Based on the current particle set, a deep reinforcement learning (DRL) strategy is used to calculate a three-dimensional reward function RL including economy, risk, and constraints.
[0110] Step 53: Calculate the fitness value of each particle based on the calculated reward function RL and the reward and penalty terms obtained by the constraint judgment result;
[0111] Step 54: Based on the obtained fitness value, update the individual optimal position and the group optimal position of each particle;
[0112] Step 55: Update the inertia weight according to the reward function RL and the number of iterations, and update the particle's speed and position;
[0113] Step 56: Perform a Gaussian mutation operation on the updated particles to obtain updated particles as the current particle set;
[0114] Step 57: Determine whether the termination condition is met. If so, output the optimal solution. Otherwise, return to step 52 for the next round of iteration.
[0115] The termination condition in step 57 may be reaching the maximum number of iterations or convergence of the objective function value;
[0116] To solve dynamic decision-making problems in uncertain scenarios, this example introduces deep reinforcement learning (DRL) to build a policy optimization layer. This transforms the site selection and capacity determination problem into a sequential decision-making problem and uses an RL agent to learn the optimal policy in uncertain scenarios. The details are as follows:
[0117] Step 52: Based on the current particle set, a deep reinforcement learning (DRL) strategy is used to calculate the three-dimensional reward function RL that includes economy, risk, and constraints, including the following steps:
[0118] Step 52.1: Based on the current particle set, extract state space parameters to describe the corresponding power grid operating environment under the current particle set scenario;
[0119] Optionally, state space parameters include: integrating time characteristics (sine / cosine coding of hours / seasons), meteorological factors (the top two factors screened by the Pearson correlation coefficient), grid status (voltage deviation, network loss rate) and node characteristics (load fluctuation rate, historical limit-exceeding times) into a 12-dimensional state vector to achieve accurate characterization of the substation operating environment.
[0120] Step 52.2: Based on the current particle set, define the action space parameters including energy storage site selection and capacity configuration;
[0121] Optionally, the energy storage site selection and capacity determination decision can be discretized into 50 basic actions (5 candidate nodes × 10 capacity levels), supporting single-node and multi-node (≤3) combination configurations, and input into the policy network through one-hot encoding;
[0122] Step 52.3: Construct a reward function: Construct a three-dimensional reward function that includes economics, risk, and constraints, and explicitly quantify the multi-objective balance. The formula is as follows:
[0123] ;
[0124] Among them, α, β, and γ are the weights of economic, risky, and constraint rewards respectively;
[0125] Network loss reward for:
[0126] ;
[0127] in, is the total number of clustered scenes, is the probability of sub-scenario s, and the network loss of each scenario is obtained through power flow calculation;
[0128] Risk Reward Items for:
[0129] ;
[0130] Where Ω is the set of super-threshold scenarios;
[0131] Constraint reward items for:
[0132] ;
[0133] In this embodiment, CVaR is introduced to quantitatively assess system risk. Network losses and tail risks are comprehensively considered during the optimization process. This allows the resulting solution to reduce network losses while effectively controlling the system's operational risks under extreme conditions, thereby improving system reliability and safety.
[0134] Step 52.4: Use the deep reinforcement learning (DRL) policy network to output the probability distribution of the action space parameters corresponding to the state space parameters, including the probability of energy storage site selection and capacity parameters;
[0135] Step 52.5: Using the action probabilities output by the policy network and environmental feedback (network loss, risk, constraints) as input, estimate the expected reward through the value network of deep reinforcement learning. The expected reward is the cumulative reward value based on the weighted sum of the three-dimensional reward function, referred to as RL reward in this embodiment.
[0136] It should be noted that the policy network outputs actions based on state space parameters and provides the probability of executing each action. This embodiment, after one iteration of the particle swarm algorithm, can obtain a new particle swarm, that is, a particle swarm with known actions (energy storage site selection and capacity determination). This embodiment calculates the probability of the action output by the policy network through the policy network, and then uses this probability value as the value network to calculate the reward, providing a reward value for the next iteration of the particle swarm to guide the iterative optimization direction of the particle swarm.
[0137] Furthermore, the proximal policy optimization algorithm (PPO) is used to train the policy network and value network. The policy network outputs the location probability distribution and constant-volume Gaussian parameters, and the value network estimates the scene value. Training stability is improved through experience replay (such as 100,000 trajectories), gradient clipping (such as a threshold of ±0.5) and entropy regularization (such as a coefficient of 0.01).
[0138] Step 53: Calculate the fitness value of each particle based on the calculated reward function RL and the reward and penalty terms obtained by the constraint judgment result;
[0139] The RL average reward is embedded in the PSO fitness function. The specific formula of the PSO fitness function is as follows:
[0140] ;
[0141] in, 、 denote the average rewards of network loss and risk respectively, represents the average reward under the corresponding particle solution, which is obtained through step 52.
[0142] Furthermore, in step 55, the inertia weight is dynamically adjusted according to the RL reward, and the adjustment coefficient is added to the adaptive inertia weight of the current iteration. When the RL reward is higher than the set value, the adaptive inertia weight is multiplied by the first adjustment coefficient. , so that the inertia weight is less than 1 to accelerate convergence; when the RL reward is not higher than the set value, increase the set second adjustment coefficient , to increase the inertia weight so that the inertia weight is greater than 1 to enhance global exploration, where the first adjustment coefficient Less than the second adjustment coefficient .
[0143] The updated inertia weight is calculated as:
[0144] ;
[0145] in, ;
[0146] When the inertia weight w is larger, the particles rely more on the historical speed and have a wider exploration range; when the inertia weight w is smaller, the particles rely more on themselves / global optimality and have stronger local convergence ability; by dynamically adjusting the inertia weight using the number of iterations and rewards, the performance of the improved particle swarm algorithm is greatly improved.
[0147] In step 53, the constraints include power balance constraints, node voltage constraints, and energy storage power and capacity constraints. Each time a particle's fitness value is calculated, the site selection and sizing solution represented by the particle is identified to determine whether it satisfies these constraints. If the constraints are not met, the particle's fitness value is penalized, and the reward value RL is adjusted to gradually eliminate it during the iteration process.
[0148] Specifically, the constraints are as follows:
[0149] (1) Power balance constraint: In each scenario, the active power and reactive power of the substation should satisfy the power balance equation, that is:
[0150] ;
[0151] in, is the total number of nodes in the area, and Respectively The active and reactive output of distributed power sources at each node, and Respectively The load active and reactive demands of each node, and are the active power and reactive power of distributed energy storage, and They are active network loss and reactive network loss in the substation area respectively.
[0152] (2) Node voltage constraint: The voltage amplitude of each node should be within the allowable range, that is:
[0153] ;
[0154] in, For the The voltage amplitude of each node, and Respectively The minimum and maximum values of the node voltage amplitude.
[0155] (3) Energy storage power and capacity constraints: The charging and discharging power and capacity of distributed energy storage should meet the following constraints:
[0156] ;
[0157] in, is the maximum charge and discharge power of distributed energy storage, is the remaining capacity of distributed energy storage, is the maximum capacity of distributed energy storage.
[0158] In this embodiment, the PSO algorithm is combined with DRL to retain PSO's global search capability. RL is used to optimize the local decision-making process. Through the reinforcement learning strategy optimization layer, the energy storage configuration scheme can automatically adapt to extreme scenarios such as sudden load increases and sudden power output drops. The three-dimensional reward function in the multi-objective balance optimization can explicitly quantify the conflicting relationships between network losses, risks, and constraints. By tuning the weight parameters, it can flexibly adapt to different planning objectives such as "economic priority" (α=0.8) or "safety priority" (γ=0.3), improving the practicality of the scheme.
[0159] Example 2
[0160] Based on Example 1, this embodiment provides a system for selecting and sizing energy storage sites in a substation area taking uncertainty into consideration, including:
[0161] A data acquisition module is configured to acquire power data and environmental condition information of the substation area and perform preprocessing;
[0162] A scenario generation module is configured to generate multiple uncertainty scenarios, including distributed power output and load demand scenarios, based on the acquired random disturbance noise and pre-processed environmental condition information;
[0163] A clustering module is configured to perform clustering on the obtained uncertainty scenarios;
[0164] The target construction module is configured to calculate the network loss based on the clustered uncertainty scenarios, and to construct the objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area;
[0165] The solution module is configured to adopt a method of collaborating with an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration process of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards to solve the objective function and obtain the optimal site location and capacity solution.
[0166] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.
[0167] Example 3
[0168] Based on Example 1, this embodiment provides a system for selecting and sizing energy storage in a substation considering uncertainty, including: a data acquisition device and a processor;
[0169] Data acquisition device, used to collect power data and environmental conditions information of the substation area;
[0170] The processor is configured to execute the steps of a method for selecting and sizing a substation energy storage site considering uncertainty as described in Example 1.
[0171] The data acquisition device may include an electric power parameter acquisition device and an environmental information acquisition device;
[0172] Power parameter collection devices are used to collect power data such as load, voltage, current, active power, and reactive power in real time or periodically. This data is collected using the power system's Supervisory Control and Data Acquisition (SCADA) system, which can include smart meters (AMI meters), micro power monitoring terminals, and power analyzers.
[0173] Environmental information collection devices are used to obtain information about the substation's climate, topography, geographic coordinates, altitude, and other factors for assessing the suitability and cost impact of energy storage deployment. These include, but are not limited to, temperature and humidity sensors, air pressure sensors, and GPS positioning modules. Meteorological or weather data can also be obtained from meteorological websites.
[0174] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0175] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A method for site selection and capacity determination of energy storage in a substation considering uncertainty, characterized in that: The steps include: Obtain substation power data and environmental condition information and perform pre-processing; Based on the acquired random disturbance noise and pre-processed environmental condition information, multiple uncertainty scenarios are generated, including distributed power output and load demand scenarios; Clustering is performed on the obtained uncertainty scenarios; Calculate network loss based on clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area; An improved particle swarm optimization algorithm is used in conjunction with a deep reinforcement learning strategy. During the iteration of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards. The objective function is solved to obtain the optimal location and capacity solution. For the clustered scenarios, an improved particle swarm optimization algorithm is used in conjunction with a deep reinforcement learning strategy to solve the objective function and obtain the optimal location and capacity solution, which includes the following steps: Step 51: Initialize the particle swarm and randomly generate the initial position and velocity of each particle. The position of each particle represents a site selection and capacity determination scheme for distributed energy storage. Step 52: Based on the current particle set, a deep reinforcement learning strategy is used to calculate a three-dimensional reward function RL that includes economy, risk, and constraints. Step 53: Calculate the fitness value of each particle based on the calculated reward function RL and the reward and penalty terms obtained by the constraint judgment result; Step 54: Based on the obtained fitness value, update the individual optimal position and the group optimal position of each particle; Step 55: Update the inertia weight according to the reward function RL and the number of iterations, and update the particle's speed and position; Step 56: Perform a Gaussian mutation operation on the updated particles to obtain updated particles as the current particle set; Step 57: Determine whether the termination condition is met. If so, output the optimal solution. Otherwise, return to step 52 for the next round of iteration. Based on the current particle set, a deep reinforcement learning strategy is used to calculate the three-dimensional reward function RL that includes economy, risk, and constraints. The steps are as follows: Based on the current particle set, the state space parameters are extracted to describe the corresponding power grid operation environment under the current particle set scenario; Based on the current particle set, action space parameters are extracted, including energy storage site selection and capacity configuration; Construct a three-dimensional reward function that includes economy, risk, and constraints. The formula is as follows: ; Among them, α, β, and γ are the weights of economic, risky, and constraint rewards respectively; Network loss reward for: ; in, is the total number of clustering scenes, is the probability of sub-scenario s, and the network loss of each scenario is obtained through power flow calculation. For the The network loss in each scenario; Risk Reward Items for: ; Where Ω is the set of super-threshold scenarios; To set the confidence level The maximum possible network loss under Constraint reward items for: ; Through the deep reinforcement learning policy network, the probability distribution of action space parameters corresponding to the state space parameters are output, including the probability of energy storage site selection and capacity parameters; The output of the policy network is used as the input of the value network of deep reinforcement learning to obtain the estimated true value as the RL reward; The inertia weight is dynamically adjusted according to the RL reward, and an adjustment coefficient is added to the adaptive inertia weight of the current iteration. When the RL reward is higher than the set value, the adaptive inertia weight is multiplied by the first adjustment coefficient to make the inertia weight less than 1 to accelerate convergence; when the RL reward is not higher than the set value, the set second adjustment coefficient is added to increase the inertia weight so that the inertia weight is greater than 1 to enhance global exploration. The first adjustment coefficient is smaller than the second adjustment coefficient.
2. The method for selecting and sizing energy storage sites in a substation considering uncertainty according to claim 1, characterized in that: For the time data in the environmental condition information, the time is encoded into angles, and then mapped through trigonometric functions to obtain the encoded time features; The preprocessing method for meteorological data is: perform correlation analysis, use the Pearson correlation coefficient to calculate the correlation between distributed energy output and meteorological data, and eliminate weakly correlated meteorological factors.
3. The method for selecting and sizing energy storage sites in a substation considering uncertainty according to claim 1, characterized in that: For the obtained uncertain scenes, the K-means++ algorithm is used to cluster the scenes, which includes the following steps: Step 31: Use the obtained uncertainty scenario as a data point and randomly select a data point as the first cluster center; Step 32: Calculate the distance from each data point to the nearest cluster center and select the point with the largest distance as the new center. Step 33: Repeat step 32 until the specified number of cluster centers are selected. Step 34: Perform traditional K-means clustering based on the obtained cluster centers, optimize the intra-class variance, and obtain the clustered scene.
4. The method for selecting and sizing energy storage sites in a substation area considering uncertainty according to claim 1, characterized in that: The improved particle swarm optimization algorithm adopts adaptive inertia weight and reduces the inertia weight in turn according to the number of iterations.
5. The method for selecting and sizing energy storage sites in a substation considering uncertainty according to claim 1, characterized in that: In the improved particle swarm optimization algorithm, the mutation operation adopts Gaussian mutation, and a random variable obeying Gaussian distribution is added to the position value of the particle in the mutation dimension.
6. A system for site selection and capacity determination of a substation energy storage system taking into account uncertainty based on the method for site selection and capacity determination of a substation energy storage system taking into account uncertainty according to any one of claims 1 to 5, characterized in that: include: A data acquisition module is configured to acquire power data and environmental condition information of the substation area and perform preprocessing; A scenario generation module is configured to generate multiple uncertainty scenarios, including distributed power output and load demand scenarios, based on the acquired random disturbance noise and pre-processed environmental condition information; A clustering module is configured to perform clustering on the obtained uncertainty scenarios; The target construction module is configured to calculate the network loss based on the clustered uncertainty scenarios, and to construct the objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area; The solution module is configured to adopt a method of collaborating with an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration process of the particle swarm optimization algorithm, rewards are calculated based on the deep reinforcement learning strategy. The fitness function of the particle swarm optimization algorithm is dynamically adjusted based on the calculated rewards to solve the objective function and obtain the optimal site location and capacity solution.
7. A system for selecting and sizing energy storage in a substation area considering uncertainty, characterized in that: include: Data acquisition device and processor; Data acquisition device, used to collect power data and environmental conditions information of the substation area; The processor is configured to execute the steps of a method for selecting and sizing substation energy storage sites considering uncertainty as described in any one of claims 1 to 5.
Citation Information
Patent Citations
PG-W-PSO method of strategy gradient improved particle swarm based on reinforcement learning
CN116451737A
Prediction and reinforcement learning multi-target-based source network load electricity storage quantity distribution method
CN119382128A