Distribution area energy storage locating and sizing method and system considering uncertainty

Through deep reinforcement learning and improved particle swarm optimization algorithm combined with conditional generation adversarial networks, the problem of insufficient robustness and adaptability in distributed energy storage configuration is solved, the most preferred location and capacity of energy storage in the station area is realized, and the economy and safety of the power system are improved.

CN120372874AActive Publication Date: 2025-07-25SHANGHAI HUAQI NEW ENERGY TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510854692.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing distributed energy storage configuration methods do not fully consider uncertainty factors, resulting in poor robustness and insufficient adaptability. Optimization algorithms are prone to local optimality, difficult to obtain system optimal solutions, lack the ability to deal with emergencies, and difficult to ensure the safe and stable operation of the power system.

Method used

Introduce deep reinforcement learning and improved particle swarm optimization algorithm collaborative optimization, combine conditional generation adversarial networks to simulate multiple uncertain scenarios, and use conditional risks to evaluate the operating risks of the power system, and build an objective function to achieve the most preferred address and fixed capacity configuration.

Benefits of technology

It significantly improves the economy and reliability of the energy storage system. By dynamically balancing local and global search capabilities, avoiding falling into local optimization, improving optimization efficiency and robustness, and effectively coping with extreme meteorological and load fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372874A_ABST
    Figure CN120372874A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power system planning and operation, and provides a zone area energy storage locating and sizing method and system considering uncertainty, which introduces deep reinforcement learning and improved particle swarm optimization algorithm for collaborative optimization, and combines a conditional generative adversarial network to simulate various uncertain scenes. And quantitative evaluation is carried out on the operation risk of the power system by using the conditional risk, so that the global search capability of the optimization algorithm and the robustness of the solution are effectively improved. According to the method, the optimal site selection and constant volume configuration of the distributed energy storage in the transformer area are realized, the network loss is remarkably reduced, and the economical efficiency and reliability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to power system planning and operation. Specifically, it relates to a method and system for siting and sizing of distribution area energy storage considering uncertainty. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the large-scale access of new energy sources such as distributed photovoltaic and wind power, modern distribution areas are facing increasingly severe challenges in operation stability and economy. Due to its rapid response and flexible regulation characteristics, distributed energy storage plays a key role in suppressing power fluctuations, improving voltage quality, peak shaving and valley filling, and reducing network losses. Therefore, scientifically and reasonably carrying out the siting and sizing of distributed energy storage is of great significance for realizing the optimization of distribution area operation and promoting the consumption of new energy.

[0004] The inventor found in the research that most of the current mainstream distributed energy storage configuration methods adopt deterministic models and do not fully consider the influence of uncertain factors such as the randomness of distributed power output, load volatility, and extreme weather, resulting in poor robustness and insufficient adaptability of the configured schemes in actual operation. At the same time, some optimization algorithms are prone to falling into local optima, it is difficult to obtain the overall optimal solution of the system, and the evaluation and control mechanisms for operation risks are not perfect, lacking the ability to cope with sudden working conditions, and it is difficult to ensure the safe and stable operation of the system. Taking the traditional particle swarm optimization algorithm as an example, this algorithm is widely used to solve the problem of distributed energy storage siting and sizing because of its simple structure and easy implementation. However, the particle swarm algorithm essentially belongs to a heuristic random search method, with strong local search ability but limited global search ability. When facing problems with high-dimensional objective functions, complex feasible regions or containing multiple local optimal solutions, the particle swarm in the particle swarm algorithm is easy to fall into local optima, especially in the later stage of algorithm iteration, the search speed slows down, resulting in difficulty in jumping out of the local extreme point, and the finally obtained energy storage configuration scheme may not be optimal. In addition, the particle swarm algorithm is usually optimized based on static or simplified models, and it is difficult to provide a solution with strong robustness and risk control for the system. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method and system for siting and sizing of distribution network energy storage considering uncertainty. By introducing the collaborative optimization of Deep Reinforcement Learning (DRL for short) and improved Particle Swarm Optimization (PSO for short), combined with Conditional Generative Adversarial Network (CGAN for short) to simulate multiple uncertainty scenarios, and using Conditional Value at Risk (CVaR for short) to quantitatively evaluate the operation risk of the power system, the global search ability of the optimization algorithm and the robustness of the solution are effectively improved. This method realizes the optimal siting and sizing configuration of distributed energy storage in the distribution network area, significantly reduces network losses, and enhances the economy and reliability of the system.

[0006] To achieve the above object, the present invention adopts the following technical solutions: One or more embodiments provide a method for siting and sizing of distribution network energy storage considering uncertainty, including the following steps: Obtain the distribution network area power data and environmental condition information, and perform preprocessing; Based on the obtained random perturbation noise and preprocessed environmental condition information, generate multiple uncertainty scenarios, including distributed power generation output and load demand scenarios; Cluster the obtained uncertainty scenarios; Calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire distribution network area; Adopt a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning strategy, and dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal siting and sizing scheme.

[0007] One or more embodiments provide a system for siting and sizing of distribution network energy storage considering uncertainty, including: A data acquisition module, configured to obtain the distribution network area power data and environmental condition information, and perform preprocessing; A scenario generation module, configured to generate multiple uncertainty scenarios, including distributed power generation output and load demand scenarios, based on the obtained random perturbation noise and preprocessed environmental condition information; A clustering module, configured to cluster the obtained uncertainty scenarios; The target construction module is configured to calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area; The solution module is configured to use a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning strategy, dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal site selection and capacity determination scheme.

[0008] One or more embodiments provide a substation area energy storage site selection and capacity determination system considering uncertainty, including: a data acquisition device and a processor; The data acquisition device is used to collect substation area power data and environmental condition information of the substation area; The processor is configured to execute the steps of the above-mentioned substation area energy storage site selection and capacity determination method considering uncertainty.

[0009] Compared with the prior art, the beneficial effects of the present invention are: The method of the present invention is significantly superior to traditional deterministic optimization methods in terms of uncertainty modeling, and can effectively cover the boundary scenarios brought by extreme weather and load fluctuations. By clustering scenarios to compress the data volume, the optimization efficiency is improved. The particle swarm optimization framework integrated with deep reinforcement learning realizes the dynamic balance of local and global search capabilities, avoids falling into local optima, and improves the stability and convergence quality of the algorithm in high-dimensional and non-convex optimization spaces. Use the deep reinforcement learning strategy to evaluate the "reward" level of the current particle behavior, and adjust the fitness function of PSO in real time. When the entire particle swarm falls into a convergence stagnation state, DRL will impose a penalty on the inferior solutions according to the reward function, guiding the particles to jump out of the local trap area. The optimization results take into account both operating losses and system risk control, and obtain a more robust energy storage site selection and capacity determination scheme. Overall, the method of the present invention can significantly improve the economy, safety and operation adaptability of the substation area energy storage system.

[0010] The advantages of the present invention and the advantages of additional aspects will be described in detail in the following specific embodiments. Description of the Drawings

[0011] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.

[0012] Figure 1 It is a flowchart of the substation area energy storage site selection and capacity determination method considering uncertainty in Embodiment 1 of the present invention; Figure 2 It is a schematic structural diagram of the CGAN model in Embodiment 1 of the present invention; Figure 3 It is a flowchart of a method for the collaboration of the improved particle swarm optimization algorithm (PSO) and the deep reinforcement learning (DRL) strategy in Embodiment 1 of the present invention; Specific implementation manners The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0013] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0014] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features in the present invention can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.

[0015] Embodiment 1 In the technical solutions disclosed in one or more embodiments, as Figures 1 to 3 shown, a method for siting and sizing of distribution network energy storage considering uncertainties includes the following steps: Step 1: Obtain the power data and environmental condition information of the distribution network area and perform preprocessing; Step 2: Generate multiple uncertainty scenarios based on the obtained random perturbation noise and preprocessed environmental condition information, including distributed power generation output and load demand scenarios; Step 3: Cluster the obtained uncertainty scenarios; Step 4: Calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire distribution network area; Step 5: Adopt a method for the collaboration of the improved particle swarm optimization algorithm (PSO) and the deep reinforcement learning (DRL) strategy. During the iteration process of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning (DRL) strategy, dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal siting and sizing scheme; The method in this embodiment is significantly superior to traditional deterministic optimization methods in terms of uncertainty modeling, and can effectively cover the boundary scenarios brought by extreme meteorology and load fluctuations. By clustering scenarios to compress the data volume, the optimization efficiency is improved. The particle swarm optimization framework integrated with deep reinforcement learning realizes the dynamic balance of local and global search capabilities, avoids falling into local optima, and improves the stability and convergence quality of the algorithm in high-dimensional and non-convex optimization spaces. The fitness function of PSO is adjusted in real time by using the deep reinforcement learning strategy to evaluate the "reward" level of the current particle behavior. When the entire particle swarm falls into a convergence stagnation state, DRL will impose penalties on inferior solutions according to the reward function to guide the particles to jump out of the local trap area. The optimization results take into account both operation losses and system risk control, and obtain a more robust energy storage siting and sizing scheme. Overall, the method of the present invention can significantly improve the economy, safety and operation adaptability of the distribution network energy storage system.

[0016] In step 1, the distribution network power data may include distributed power generation output data, line parameters (such as resistance, reactance), node voltage data, load demand data, etc. The environmental condition information may include condition information such as corresponding distribution network time data, geographical condition data, meteorological data, etc. In step 1, for data collection and preprocessing, the data collection and monitoring (abbreviated as SCADA) system of the power system can be used to collect distributed power generation output data, line parameters (such as resistance, reactance), node voltage data and load demand data, etc. within a set time period of the distribution network, such as the past year; at the same time, obtain condition information such as the time, geography, and meteorology of the distribution network.

[0017] Optionally, the data preprocessing method includes: Step 11: Clean the obtained data, remove obvious outliers, and remove noise and outliers in the data; for missing values, use the spline interpolation method to fill them.

[0018] Step 12: Normalize the data using the Z-score normalization method, and map the data to the [0,1] interval to improve the training stability and convergence speed of the model. The specific formula is: ; Among them, 、 are the feature mean and variance.

[0019] Furthermore, for the time data, perform time feature encoding, encode the time into an angle, and then map it through trigonometric functions (sine / cosine functions) to obtain the encoded time feature. Specifically, map 24 hours of a day to the circumferential angle: ; Represented by the sine and cosine values of the mapped angle as: ; Map 365 days of a year to an angle: ; Represented by the sine and cosine values of the mapped angle as: ; Decompose time into sine / cosine signals through Fourier transform, and the specific formula is as follows: ; Fourier transform can decompose any periodic signal into the sum of sine / cosine signals of different frequencies. For the time series f(t), its Fourier series is expressed as: ; In the scenario of distribution network energy storage in this solution, , , are coefficients respectively. Only keep the fundamental wave component with k = 1, that is , can effectively characterize the main periodic features such as day and night, seasons, etc. Higher-order harmonics can be ignored to reduce the dimension.

[0020] The operation of the power grid has significant periodic characteristics. Traditional timestamps (such as Unix time) are difficult to effectively characterize periodicity as continuous variables. Convert timestamps (such as hours, days of the year) into periodic features to capture the day-night and seasonal patterns of power load and distributed power generation output, that is, map 24 hours of a day to an angle , represented by its sine and cosine values, to avoid the mutation caused by directly using numerical values (such as the adjacency between 23:00 and 0:00); similarly, map 365 days of a year to an angle , to characterize seasonal changes. Doing so can not only avoid the model misjudging the time continuity caused by linear coding; but also retain the periodic symmetry and improve the ability to capture periodic fluctuations.

[0021] In this embodiment, by performing periodic coding on time data, the problem that the traditional numerical input method destroys the time period structure is overcome, and the coding mutations such as the large numerical difference between adjacent times of 23:00 and 0:00, 365 days and the 1st day are effectively avoided, and the model misjudging the time continuity is avoided. This coding method retains periodic features such as day and night and seasons, helps the model to more accurately identify the periodic fluctuation laws of power load and distributed power generation output, and significantly improves the model's perception ability and prediction accuracy for periodic changes.

[0022] Furthermore, the preprocessing method for meteorological data is: conducting correlation analysis, calculating the correlation between distributed energy output and meteorological data, and eliminating weakly correlated meteorological factors; Optionally, the Pearson correlation coefficient can be used to calculate the correlation between distributed energy output and meteorological data. The calculation formula is as follows: ; in, Represents distributed power and meteorological factors The Pearson correlation coefficient ranges from [-1, 1]; is the time series data of the output power of distributed generation (DG), is the time series data of meteorological factors; the data timestamps of the two must be synchronized, otherwise the correlation coefficient will be distorted. Linear interpolation can be used to solve the missing data; for and The covariance of , is the standard deviation of the two.

[0023] Furthermore, a correlation threshold may be set, and the calculated correlation coefficient may be compared with the set threshold to screen key meteorological factors; optionally, the correlation threshold may be not less than 0.5; The output of distributed power generation is significantly affected by meteorological conditions. In this embodiment, the key meteorological factors are selected by the Pearson correlation coefficient and retained. The meteorological factors are removed, and weakly correlated factors are eliminated to reduce the data dimension and avoid irrelevant features from interfering with subsequent model training. This ensures that the scenarios generated by the subsequent conditional generative adversarial network (CGAN) conform to the laws of physics.

[0024] In step 2, the uncertainty scenario can be generated through the conditional generative adversarial network (CGAN), taking the preprocessed environmental condition information as input, and the uncertainty scenario of the distributed power output and load demand through the generator; Specifically, a CGAN model is constructed, and the generator and discriminator are trained using the Adam optimizer. During the training process, the generator and the discriminator are trained alternately for a set number of rounds, such as 1000. After the CGAN model is trained, the generator is used to generate multiple distributed power output and load demand scenarios.

[0025] In one achievable implementation, the constructed CGAN model includes a generator and a discriminator. The generator adopts a multi-layer perceptron (MLP) structure, including 3 hidden layers, each with 128 neurons. The discriminator also adopts an MLP structure, including 2 hidden layers, each with 64 neurons. During the training process, Figure 2As shown, the input of the generator is a random noise vector and the preprocessed environmental condition information, and the output is the simulated distributed power generation output and load demand scenarios. The input of the discriminator is the real scenario data and the scenario data generated by the generator, and the output is the judgment on the authenticity of the input data.

[0026] Optionally, the Adam optimizer is used to alternately train the generator and the discriminator, so that the generator can generate uncertainty scenarios similar to the real scenario data distribution. The alternating training method is as follows: fix the generator G and train the discriminator D to improve its discrimination ability; fix the discriminator D and train the generator G to make the generated scenarios more realistic. Specifically, the loss function design for the CGAN model training is as follows: For the discriminator, the cross-entropy loss is used to measure its classification error for real and forged data.

[0027] For the generator, an adversarial loss function is adopted to improve the ability of the generated data to "deceive" the discriminator.

[0028] In this embodiment, by introducing the time, geographical location, and meteorological conditions of the distribution area as the generation condition information, a conditional generative adversarial network (CGAN) model is constructed, so that the generated distributed power generation output and load demand scenarios have higher pertinence and controllability, significantly improving the matching degree of the simulated data to the actual operation environment. CGAN can generate representative and diverse scenarios under different spatio-temporal and meteorological conditions, effectively depicting the uncertainty characteristics of distributed energy output and load demand, providing a high-quality and diverse input data basis for subsequent energy storage optimization configuration, thereby improving the robustness of the overall model and the reliability of the optimization results.

[0029] In step 3, for the obtained uncertainty scenarios, the K-means++ algorithm is used for scenario clustering, including the following steps: Step 31: Take the obtained uncertainty scenarios as data points and randomly select a data point as the first clustering center. Step 32: Calculate the distance from each data point to the nearest clustering center, and select the point with the largest distance as the new center. Step 33: Repeat step 32 until the specified number of clustering centers are selected. Step 34: Based on the obtained clustering centers, perform traditional K-means clustering to optimize the within-class variance and obtain the clustered scenarios. Step 35: For the clustered scenarios, select the nodes with voltage volatility and load rate higher than the set values as the candidate locations for distributed energy storage devices. According to the above algorithm, based on the voltage volatility and load rate, the K-means++ clustering is used to select high-volatility and high-load nodes as candidate installation locations, which can be used for the construction of the population of the subsequent particle swarm algorithm; The voltage volatility is: ; The load rate is: ; Among them, represents the voltage volatility of node i, N is the number of sampling in the time dimension. If 96 moments within 1 day are selected, with 1 point every 15 minutes, then N = 96; is the voltage value of node i at the t-th moment, is the average voltage value of node i at N moments; is the load rate of node i, is the maximum active power load of node i during the observation period, is the rated active power load capacity of node i.

[0030] In this embodiment, the K-means++ algorithm is used to cluster the generated scenarios, reducing the number of redundant scenarios, reducing the computational complexity, while retaining the key features and improving the optimization efficiency; a large number of scenarios generated by the CGAN model are input into the K-means++ algorithm, and the redundant scenarios are reduced through cluster analysis. In this embodiment, by optimizing the selection of the initial cluster center, the randomness problem of the traditional K-means is avoided, and the clustering quality is improved. The clustered scenarios are used as the input for subsequent optimization, which not only retains the characteristics of the original data distribution but also reduces the computational complexity.

[0031] Step 4. For the clustered uncertain scenarios, establish an objective function considering network loss and conditional value-at-risk (CVaR) as follows: ; Among them, is the total number of clustered scenarios, is the network loss of the distribution transformer area under the -th scenario, and are weight coefficients, representing the importance of network loss and conditional value-at-risk (CVaR) in the objective function respectively. CVaR represents the conditional value-at-risk of all scenarios; Specifically, for each scenario generated after clustering, according to circuit theory and power flow calculation methods, calculate the network loss of the distribution transformer area. The classical forward-backward substitution method is used for power flow calculation to obtain the active power loss and reactive power loss of each line, and then calculate the total network loss of the entire distribution transformer area. The specific formula for network loss is as follows: ; Among them, and is the active / reactive power of line ij, is the node voltage, is the resistance on line ij; i and j represent the nodes on the line, and the line between two nodes is marked as line ij; Specifically, the method for determining conditional risk loss is as follows: calculate the network loss values under all scenarios, sort them in ascending order, set the confidence level, determine the network loss threshold, and calculate the average network loss of the scenarios exceeding the network loss threshold as the conditional risk loss value (CVaR value); among them, the confidence level can be set to not less than 95%; In the objective function of this embodiment, a conditional risk loss term is set, which can collect the tail risk in the uncertainty scenarios generated by CGAN and quantify the potential loss caused by extreme scenarios; among them, extreme scenarios can include scenarios such as sudden drops in wind power / photovoltaic power output and sudden increases in load; the conditional risk loss CVaR, the specific formula is: ; where, is the maximum possible network loss under the set confidence level ; represents the loss part exceeding , that is, the additional loss in extreme cases, is the probability of the occurrence of scenario s, represents the total number of scenarios.

[0032] The PSO algorithm is an optimization algorithm based on swarm intelligence, which finds the optimal solution by simulating the foraging behavior of bird flocks or fish schools. In the PSO algorithm, each particle represents a potential solution, and the particles move in the search space and update their velocities and positions according to their own historical optimal positions and the historical optimal positions of the group.

[0033] In this embodiment, in order to prevent the PSO algorithm from falling into a local optimal solution and improve the search ability and convergence speed of the algorithm, the improved particle swarm optimization algorithm (PSO) adopts an adaptive inertia weight, and the inertia weight is sequentially decreased according to the number of iterations, so that at the initial stage of the algorithm, a set larger inertia weight is adopted to enable the particles to perform global search in a larger search space; at the later stage of the algorithm, a smaller inertia weight is adopted to enable the particles to perform fine search in the local area; Specifically, the adaptive inertia weight, the calculation formula is as follows: ; where, and are respectively the maximum and minimum values of the inertia weight, is the maximum number of iterations, is the current number of iterations.

[0034] Further, in the improved particle swarm optimization algorithm, the mutation operation uses Gaussian mutation to add a random variable subject to a Gaussian distribution to the position value of the particles in the mutated dimension.

[0035] After the particles update their positions, a mutation operation is performed on some dimensions of the particles with a certain probability to introduce new search directions and increase the diversity of the population. In this embodiment, Gaussian mutation is used to add random noise subject to a Gaussian distribution (mean of 0 and standard deviation of 1) to the particle positions. When the particles fall into the local optimal solution of a certain node combination, Gaussian mutation perturbs the energy storage nodes or capacity dimensions with a certain probability (such as with a mutation probability pm = 0.1) to generate a brand-new site selection and capacity determination scheme, avoiding premature convergence of the population due to homogenization, so that the particles can jump out of the current local optimal region after update and search for the optimal solution again.

[0036] Step 5: For the clustered scenarios, use a method that collaborates the improved particle swarm optimization algorithm (PSO) and the deep reinforcement learning (DRL) strategy to solve the objective function to obtain the optimal site selection and capacity determination scheme, including the following steps: Step 51: Initialize the particle swarm, randomly generate the initial positions and velocities of each particle, and the position of each particle represents a site selection and capacity determination scheme for distributed energy storage; Step 52: Based on the current particle set, use the deep reinforcement learning (DRL) strategy to calculate the three-dimensional reward function RL including economy, risk, and constraint; Step 53: Based on the calculated reward function RL and the reward and penalty terms obtained by judging the results through the constraint conditions, calculate the fitness value of each particle; Step 54: Based on the obtained fitness values, update the individual optimal positions and the global optimal position of each particle; Step 55: Update the inertia weight according to the reward function RL and the number of iterations, and update the velocities and positions of the particles; Step 56: Perform Gaussian mutation operation on the updated particles to obtain the updated particles as the current particle set; Step 57: Judge whether the termination condition is satisfied. If it is satisfied, output the optimal solution; otherwise, return to Step 52 for the next round of iteration.

[0037] Among them, the termination condition in Step 57 can be reaching the maximum number of iterations or the convergence of the objective function value; To solve the dynamic decision-making problem under uncertain scenarios, this embodiment introduces deep reinforcement learning (DRL) to construct a policy optimization layer, transforms the site selection and capacity determination problem into a sequential decision-making problem, and uses the RL agent to learn the optimal policy in uncertain scenarios, which is specifically described as follows: Step 52: Based on the current particle set, adopt a deep reinforcement learning (DRL) strategy to calculate a three-dimensional reward function RL that includes economy, risk, and constraints, which includes the following steps: Step 52.1: Based on the current particle set, extract state space parameters to describe the corresponding power grid operation environment in the current particle set scenario; Optionally, the state space parameters include: integrating time features (sine / cosine encoding of hours / seasons), meteorological factors (the top 2 factors screened by Pearson correlation coefficient), power grid status (voltage deviation, line loss rate), and node features (load volatility, historical overlimit times) into a 12-dimensional state vector to achieve accurate characterization of the operation environment of the substation area.

[0038] Step 52.2: Based on the current particle set, define action space parameters including energy storage location selection and capacity configuration; Optionally, the energy storage location selection and capacity determination decision can be discretized into 50 basic actions (5 candidate nodes × 10 capacity levels), supporting single-node and multi-node (≤3 nodes) combined configuration, and input into the policy network through one-hot encoding; Step 52.3: Construct a reward function: Construct a three-dimensional reward function that includes economy, risk, and constraints, and explicitly quantify the multi-objective balance. The formula is as follows: ; Among them, α, β, and γ are the weights of the economy, risk, and constraint rewards respectively; Line loss reward term is: ; Among them, is the total number of clustering scenarios, is the probability of sub-scenario s, and the line loss of each scenario is obtained through power flow calculation; Risk reward term is: ; Among them, Ω is the set of scenarios exceeding the threshold; Constraint reward term is: ; In this embodiment, CVaR is introduced to quantitatively evaluate the system risk. During the optimization process, line loss and tail risk are comprehensively considered, so that the obtained solution can effectively control the operation risk of the system in extreme cases while reducing the line loss, improving the reliability and security of the system; Step 52.4: The policy network of deep reinforcement learning (DRL) outputs the probability distribution of the corresponding action space parameters in the state space parameters, including the probability of energy storage location selection and capacity determination parameters. Step 52.5: Taking the action probability output by the policy network and the environmental feedback (network loss, risk, constraints) as inputs, the value network of deep reinforcement learning estimates the expected return. The expected return is the cumulative reward value weighted and summed based on the three-dimensional reward function, which is simply referred to as the RL reward in this embodiment.

[0039] It should be noted that the policy network outputs actions based on the state space parameters and gives the probability of each action execution. After one iteration of the particle swarm algorithm in this embodiment, a new particle swarm can be obtained, that is, the known actions (energy storage location selection and capacity determination scheme). In this embodiment, the policy network calculates the probability of the actions output by the policy network, so as to calculate the reward based on this probability value by the value network, and provide the reward value for the next iteration of the particle swarm to guide the iterative optimization direction of the particle swarm. Furthermore, the proximal policy optimization algorithm (PPO) is used to train the policy network and the value network. The policy network outputs the location selection probability distribution and the capacity determination Gaussian parameters, and the value network estimates the scenario value. The training stability is improved through experience replay (such as 100,000 trajectories), gradient clipping (such as the threshold ±0.5), and entropy regularization (such as the coefficient 0.01).

[0040] Step 53: Based on the calculated reward function RL and the reward and penalty terms obtained by judging the results through the constraint conditions, calculate the fitness value of each particle. Embed the RL average reward in the PSO fitness function. The specific formula of the PSO fitness function is as follows: ; where and respectively represent the average rewards of network loss and risk, represents the average reward under the corresponding particle scheme, which is obtained through Step 52.

[0041] Furthermore, in Step 55, the inertia weight is dynamically adjusted according to the RL reward. An adjustment coefficient is added to the adaptive inertia weight of the current iteration. When the RL reward is higher than the set value, the adaptive inertia weight is multiplied by the first adjustment coefficient so that the inertia weight is less than 1 to accelerate convergence; when the RL reward is not higher than the set value, the set second adjustment coefficient is increased to increase the inertia weight, making the inertia weight greater than 1 to enhance global exploration. Among them, the first adjustment coefficient is less than the second adjustment coefficient

[0042] The calculation formula for the updated inertia weight is as follows: ; where, ; When the inertia weight w is larger, the particle is more dependent on the historical velocity and has a wide exploration range; when the inertia weight w is smaller, the particle is more dependent on its own / global optimum and has a strong local convergence ability; by dynamically adjusting the inertia weight using the number of iterations and rewards, the performance of the improved particle swarm optimization algorithm is greatly improved.

[0043] In step 53, the constraint conditions include power balance constraint, node voltage constraint, and energy storage power and capacity constraint. When calculating the fitness value of each particle, it is identified whether the siting and sizing scheme represented by the particle satisfies the power balance constraint, node voltage constraint, and energy storage power and capacity constraint. If the constraint conditions are not met, the fitness value of the particle is penalized, and the size of the reward value RL is adjusted so that it is gradually eliminated during the iteration process.

[0044] Specifically, the constraint conditions are as follows: (1) Power balance constraint: In each scenario, the active power and reactive power of the distribution area should satisfy the power balance equation, that is: ; where, is the total number of nodes in the distribution area, and are the active power output and reactive power output of the distributed power source at the th node respectively, and are the active power demand and reactive power demand of the load at the th node respectively, and are the active power and reactive power of the distributed energy storage respectively, and are the active power loss and reactive power loss of the distribution area respectively.

[0045] (2) Node voltage constraint: The voltage amplitude of each node should be within the allowable range, that is: ; where, is the voltage amplitude of the th node, and are the minimum and maximum values of the voltage amplitude of the th node respectively.

[0046] (3) Energy storage power and capacity constraint: The charging and discharging power and capacity of the distributed energy storage should satisfy the following constraint conditions: ; Among them, is the maximum charge-discharge power of the distributed energy storage, is the remaining capacity of the distributed energy storage, is the maximum capacity of the distributed energy storage.

[0047] In this embodiment, the PSO algorithm is combined with DRL, retaining the global search ability of PSO, using RL to optimize the local decision-making process, and through the reinforcement learning policy optimization layer, the energy storage configuration scheme can automatically adapt to extreme scenarios such as sudden increases in load and sudden drops in power generation. The three-dimensional reward function in the multi-objective balance optimization can explicitly quantify the conflict relationship between network loss, risk and constraints. By tuning the weight parameters, different planning objectives such as "economy first" (α = 0.8) or "safety first" (γ = 0.3) can be flexibly adapted, improving the practicality of the scheme.

[0048] Embodiment 2 Based on Embodiment 1, this embodiment provides a system for locating and sizing a distribution network energy storage considering uncertainty, including: A data acquisition module, configured to acquire distribution network power data and environmental condition information, and perform preprocessing; A scenario generation module, configured to generate multiple uncertainty scenarios based on the acquired random perturbation noise and preprocessed environmental condition information, including distributed power generation output and load demand scenarios; A clustering module, configured to cluster the obtained uncertainty scenarios; An objective construction module, configured to calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the minimum total network loss and conditional risk loss of the entire distribution network as the objective; A solution module, configured to use a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning strategy, and dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal location and sizing scheme.

[0049] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1, and its specific implementation process is the same, so it will not be repeated here.

[0050] Embodiment 3 Based on Embodiment 1, this embodiment provides a system for locating and sizing a distribution network energy storage considering uncertainty, including: a data collection device and a processor; The data collection device is used to collect distribution network power data and the environmental condition information of the distribution network; The processor is configured to execute the steps of a method for locating and sizing a distribution network energy storage considering uncertainty described in Embodiment 1.

[0051] Among them, the data acquisition device may include a power parameter acquisition device and an environmental information acquisition device; The power parameter acquisition device is used to collect power data such as the power load, voltage, current, active power, and reactive power of the transformer substation area in real time or periodically. The collection is carried out by using the data acquisition and monitoring (abbreviated as SCADA) system of the power system. The SCADA system may include smart meters (AMI meters), micro power monitoring terminals, power analyzers, etc.

[0052] The environmental information acquisition device is used to obtain information such as the climate, terrain, geographical coordinates, and altitude of the environment where the transformer substation area is located, and is used to evaluate the suitability of energy storage construction and the cost impact. It includes, but is not limited to, temperature and humidity sensors, barometric pressure sensors, GPS positioning modules, etc., and meteorological or weather data can also be obtained through meteorological websites.

[0053] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0054] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for locating and sizing energy storage in a distribution transformer area considering uncertainty, characterized in that It includes the following steps: Obtain the power data of the power distribution area and environmental condition information, and perform preprocessing; Based on the obtained random disturbance noise and the preprocessed environmental condition information, generate multiple uncertainty scenarios, including distributed power generation output and load demand scenarios; Cluster the obtained uncertainty scenarios; Calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire power distribution area; Adopt a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration process of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning strategy, and dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal site selection and capacity determination plan.

2. The method for locating and sizing the energy storage in the substation area considering uncertainty as claimed in claim 1, wherein: For the time data in the environmental condition information, encode the time as an angle, and then perform mapping through trigonometric functions to obtain the encoded time feature; The preprocessing method for meteorological data is: perform correlation analysis, use the Pearson correlation coefficient to calculate the correlation between distributed energy output and meteorological data, and eliminate weakly correlated meteorological factors.

3. A method for determining the site selection and capacity of energy storage in a power distribution area considering uncertainty according to claim 1, characterized in that: For the obtained uncertainty scenarios, use the K-means++ algorithm for scenario clustering, including the following steps: Step 31: Take the obtained uncertainty scenarios as data points, and randomly select a data point as the first clustering center; Step 32: Calculate the distance from each data point to the nearest clustering center, and select the point with the largest distance as the new center; Step 33: Repeat step 32 until the specified number of clustering centers is selected; Step 34: Perform traditional K-means clustering based on the obtained clustering centers to optimize the within-class variance and obtain the clustered scenarios.

4. The method for locating and sizing the distribution network energy storage considering uncertainty according to claim 1, wherein: The improved particle swarm optimization algorithm uses an adaptive inertia weight, and the inertia weight decreases sequentially according to the number of iterations.

5. A method for determining the location and capacity of distribution network energy storage considering uncertainty as described in claim 1, characterized in that: In the improved particle swarm optimization algorithm, the mutation operation uses Gaussian mutation, and a random variable obeying a Gaussian distribution is added to the position value of the particles in the mutated dimension.

6. The method for siting and sizing of distribution network energy storage considering uncertainty according to claim 1, characterized in that: For the clustered scenarios, adopt a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy to solve the objective function to obtain the optimal site selection and capacity determination plan, including the following steps: Step 51: Initialize the particle swarm, randomly generate the initial position and velocity of each particle, and the position of each particle represents a site selection and capacity determination plan for distributed energy storage; Step 52: Based on the current particle set, use the deep reinforcement learning strategy to calculate a three-dimensional reward function RL including economy, risk, and constraints; Step 53: Calculate the fitness value of each particle based on the calculated reward function RL and the reward and penalty terms obtained by judging the results through the constraint conditions; Step 54: Based on the obtained fitness value, update the individual optimal position and the global optimal position of each particle; Step 55: Update the inertia weight according to the reward function RL and the number of iterations, and update the velocity and position of the particles; Step 56: Perform Gaussian mutation operation on the updated particles to obtain the updated particles as the current particle set; Step 57: Determine whether the termination condition is met. If it is met, output the optimal solution; otherwise, return to Step 52 for the next round of iteration.

7. The method for siting and sizing of distribution network energy storage considering uncertainty according to claim 6, characterized in that: Based on the current particle set, adopt a deep reinforcement learning strategy to calculate a three-dimensional reward function RL that includes economy, risk, and constraint, including the following steps: Based on the current particle set, extract state space parameters to describe the corresponding power grid operating environment in the current particle set scenario; Based on the current particle set, extract action space parameters including energy storage location selection and capacity configuration; Construct a three-dimensional reward function that includes economy, risk, and constraint; Through the policy network of deep reinforcement learning, output the probability distribution of the corresponding action space parameters in the state space parameters, including the probabilities of energy storage location selection and capacity determination parameters; Use the output of the policy network as the input of the value network of deep reinforcement learning to obtain the estimated true value as the RL reward.

8. A method for locating and sizing energy storage in a distribution transformer area considering uncertainty as described in claim 6, characterized in that: Dynamically adjust the inertia weight according to the RL reward, and add an adjustment coefficient to the adaptive inertia weight of the current iteration. When the RL reward is higher than the set value, multiply the adaptive inertia weight by the first adjustment coefficient so that the inertia weight is less than 1 to accelerate convergence; when the RL reward is not higher than the set value, increase the set second adjustment coefficient to increase the inertia weight so that the inertia weight is greater than 1 to enhance global exploration, where the first adjustment coefficient is less than the second adjustment coefficient.

9. A system for locating and sizing energy storage in a distribution transformer area considering uncertainty, characterized in that, Include: A data acquisition module configured to acquire power data and environmental condition information of the substation area and perform preprocessing; A scenario generation module configured to generate multiple uncertainty scenarios based on the acquired random perturbation noise and preprocessed environmental condition information, including distributed power generation output and load demand scenarios; A clustering module configured to cluster the obtained uncertainty scenarios; An objective construction module configured to calculate the network loss based on the clustered uncertainty scenarios, and construct an objective function with the goal of minimizing the total network loss and conditional risk loss of the entire substation area; A solution module configured to adopt a method that combines an improved particle swarm optimization algorithm and a deep reinforcement learning strategy. During the iteration of the particle swarm optimization algorithm, calculate the reward based on the deep reinforcement learning strategy, dynamically adjust the fitness function of the particle swarm optimization algorithm based on the calculated reward, and solve the objective function to obtain the optimal location and capacity determination scheme.

10. A system for locating and sizing energy storage in a distribution transformer area considering uncertainty, characterized in that, Include: A data acquisition device and a processor; The data acquisition device is used to acquire power data and environmental condition information of the substation area; The processor is configured to execute the steps of a method for energy storage location and capacity determination in a substation area considering uncertainty according to any one of claims 1-8.

Citation Information

Patent Citations

  • PG-W-PSO method of strategy gradient improved particle swarm based on reinforcement learning

    CN116451737A

  • Prediction and reinforcement learning multi-target-based source network load electricity storage quantity distribution method

    CN119382128A

  • Power distribution network reliability cooperative control method and system under distributed power supply access

    CN119448301A

  • Memory polynomial based digital predistorter

    US9160280B1