Method and system for optimizing operation mode of power distribution network for distributed photovoltaic absorption
Through the distribution network operation mode optimization method for distributed photovoltaic absorption, the distribution network topology is optimized using reinforcement learning and classification prediction models, and the problem of inefficient calculation efficiency of traditional algorithms is solved, and efficient photovoltaic absorption and stable operation are achieved.
Patent Information
- Application Number
- CN202510629934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional algorithms have low computational efficiency and cannot perform real-time optimization and calculation, which makes it difficult for the distribution network to achieve efficient absorption and stable operation when facing the volatility and randomness of distributed photovoltaics.
The distribution network operation mode optimization method is adopted for distributed photovoltaic absorption. By collecting system data, using scenario classification model and photovoltaic power prediction model, inputting reinforcement learning decision model, we obtain the distribution network topology reconstruction strategy, and through the security layer correction strategy, we ensure that the system safety constraints are met.
It has achieved that the distribution network can improve the photovoltaic absorption capacity and improve the economic and safety of system operation without adding energy storage equipment, and can meet system safety constraints in real-time high-speed reconstruction and optimization.
Smart Images

Figure CN120150263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system operation optimization, and in particular to a method and system for optimizing the operation mode of a distribution network for distributed photovoltaic accommodation. Background Art
[0002] With the explosive growth of distributed photovoltaics, the distribution network has become an important channel for distributed photovoltaic grid connection, and a large number of distributed photovoltaics have been successively connected to the end of the distribution network line. During high photovoltaic output periods, the sudden increase in photovoltaic output will raise the grid connection point voltage, directly affecting the power consumption quality of surrounding users, and even causing distributed photovoltaics to disconnect from the grid due to excessive grid connection point voltage. At the same time, due to the fact that the distributed photovoltaic power generation power in high-penetration areas is greatly affected by weather, the intermittent large fluctuations also have a certain probability of causing large fluctuations in the voltage of the low-voltage distribution bus. The overall accommodation capacity of the distribution network shows deficiencies, and it is urgent to newly build a large number of regulating resources such as new energy storage, but the high economic cost is not conducive to the large-scale sustainable development of photovoltaics. And the operation reconstruction technology of the distribution network is one of the important measures to solve this problem. Generally, in order to ensure the quality and safety of power supply, the distribution network is designed to be able to operate in different ways at the beginning of the grid structure design. Therefore, without adding additional energy storage equipment, the accommodation capacity of photovoltaics can be improved by operating the line standby switch to transform the grid structure.
[0003] Compared with the traditional distribution network system, when the new power system performs topological reconstruction, it is extremely easy to appear problems such as line overload and voltage violation; secondly, although adjusting the grid through the line switch can directly control the power flow distribution and node voltage of the system, it also exacerbates the complexity of solving the power grid dispatching operation problem. In addition, compared with the dispatching problem on a daily scale, the operation reconstruction of the distribution network has higher requirements for calculation timeliness. The output of new energy represented by photovoltaics is volatile and random, and the impact of the system grid on the operation index is closely related to the output of new energy and photovoltaics at the corresponding moment. Generating a reconstruction strategy in a timely and rapid manner at the corresponding time period can greatly improve the economic efficiency of system operation, while most of the current technologies usually have difficulty achieving the required calculation speed without occupying a large amount of computing power. Summary of the Invention
[0004] Aiming at the problem that the traditional algorithm has low calculation efficiency and cannot perform real-time optimization calculation, the present invention provides a method and system for optimizing the operation mode of a distribution network for distributed photovoltaic accommodation, aiming to give an efficient distribution network operation scheduling strategy with the goal of promoting distributed photovoltaic accommodation.
[0005] In order to achieve the above invention purpose, the technical solution of the present invention is as follows: According to the first aspect of the present invention, a method for optimizing the operation mode of a distribution network for distributed photovoltaic accommodation is proposed. The method includes the following steps: Step 1: Collect system data, including power data with timestamps, system structure data, and meteorological data; the system power data includes photovoltaic power data and load data of each node. Step 2: Use a scenario classification model to identify the scenario pattern corresponding to the current photovoltaic power, and call the photovoltaic power prediction model corresponding to the identified scenario pattern to obtain the photovoltaic power prediction value; the scenario classification model and the photovoltaic power prediction model are trained based on clustering technology by dividing the system power data using meteorological data. Step 3: Take the system power data, system structure data, and photovoltaic power prediction value of the current time period as the current system state and input them into a decision-making model based on reinforcement learning to obtain the distribution network topology reconstruction strategy under the current system state. Step 4: Monitor the distribution network topology reconstruction strategy, and iteratively correct the strategy that violates the system security constraints until the system security constraint strategy is met.
[0006] Further, in Step 1, the meteorological data includes temperature data, humidity data, and air pressure data of each area at each moment. The system structure data includes the topological structure of the power system, the states of line switches, and the impedance data of each line at each moment.
[0007] Further, in Step 2, the scenario classification model and the photovoltaic power prediction model are trained in the following way: Step 21: Divide the photovoltaic power data of each moment into the corresponding season according to the timestamp. Step 22: Select humidity, air pressure, and temperature from the meteorological data set as the characteristic variables for clustering analysis. Through the K-Means clustering algorithm, further divide the photovoltaic power data of each season into three weather patterns: sunny, cloudy, and rainy, so as to divide the photovoltaic power data into sub-power data sets under 12 types of scenario patterns. Step 23: Based on the division results, use the system data as the data source to train the scenario classification model and the photovoltaic power prediction model for each scenario pattern.
[0008] Further, Step 22 includes: Step 221: Randomly select 3 initial centroids, and complete the clustering division of the power data by continuously updating the centroids and assigning the meteorological data points corresponding to each photovoltaic power data to the nearest centroid according to the distance difference from each centroid; among them, each centroid is a three-dimensional vector composed of the average values of humidity, air pressure, and temperature. The calculation methods of the centroid and the distance are:
[0009] where is the meteorological data pointi The Euclidean distance from the centroid j ; H i , P i and T i represent the humidity value, air pressure value, and temperature value of the data point i respectively; k is the number of the category cluster to which the meteorological data point i belongs, N k is the k number of meteorological data points in the th category cluster, k is the new centroid coordinate; Cluster k represents the set of meteorological data points corresponding to each photovoltaic power data in the th category cluster; Step 222: Analyze the characteristics of the meteorological data in each category cluster, and determine and label the weather pattern corresponding to each category cluster according to the analysis results.
[0010] Furthermore, Step 23 includes: Step 231: Construct the loss functions of the scenario classification model and the photovoltaic power prediction model respectively, expressed as:
[0011] where N 1 and N 2 are the numbers of samples in the training sample sets corresponding to the scenario classification sub-model and the classification prediction sub-model respectively, k is the sample category; is the loss function of the scenario classification sub-model, is the true category label of the i th sample in the training sample set corresponding to the scenario classification sub-model, is the probability that the i th sample in the training sample set corresponding to the scenario classification sub-model belongs to the category k , is the loss function of the photovoltaic power prediction model, and are the true photovoltaic power value and the predicted photovoltaic power value corresponding to the i th sample in the training sample set corresponding to the classification prediction sub-model respectively.
[0012] Step 232: Minimize and With the goal of using cross - entropy loss for the scene classification model and mean absolute error for the photovoltaic power prediction model respectively, the Adam optimizer is called through gradient backpropagation to update the parameters, realizing the training of the scene classification sub - model and the classification prediction sub - model.
[0013] Furthermore, in step 3, the decision - making model based on reinforcement learning is trained as follows: Step 31, establish a virtual power system environment model, which is used to simulate the steady state of the distribution network through distributed power flow calculation based on the distribution network topology reconstruction strategy output by the decision - making model, and generate simulated system power data and system structure data in combination with the set photovoltaic power prediction value; Step 32, make the virtual power system environment model and the decision - making model interact, and train the neural network based on reinforcement learning with the goal of maximizing the reward function to obtain the decision - making model; the reward function is related to the photovoltaic accommodation ratio index, network loss index, switch loss index, whether the power flow calculation result converges, and whether the current distribution network topology reconstruction strategy meets the system security constraints.
[0014] Furthermore, step 31 includes: Step 311: Perform distributed power flow calculation through the following formula:
[0015] Where, and are the active power and reactive power of branch ij respectively, and are the active power and reactive power of branch jk respectively, and are the active and reactive powers of nodes i respectively, and are the squares of the voltage amplitudes at nodes i and j respectively, is the square of the current amplitude on branch ij respectively, and represent the sets of nodes and branches respectively, represents the branch i between nodes j and ij , and are the resistance and reactance of branch ij respectively, and represent the upper and lower bounds of the variable respectively; Step 312: Determine whether the distributed power flow solution converges. When it is determined that the power flow solution does not converge, a first negative penalty is fed back. r pen1 ; Step 313: Determine whether the current distribution network topology reconstruction strategy meets the system safe operation conditions. When the system safe operation conditions are violated, a second negative penalty is fed back. r pen2 .
[0016] Further, in step 32, the reward function is expressed as:
[0017] where is the absorption ratio of distributed PV in the system, , are the PV powers before and after optimization, is the line loss of the network, is the number of PV nodes, is the total number of system branches, is the basic reward of the reinforcement learning model, is the number of switch changes under the current policy, , , are proportionality coefficients used to adjust the weights of the PV absorption ratio index, line loss index, and switch loss index in the reward respectively. is a small quantity used to prevent reward overflow due to too small denominator or division by zero error reporting; is the total negative penalty, and r pen = r pen1+ r pen2 ; is the total reward function of the model.
[0018] Further, in step 32, the neural network based on reinforcement learning is updated using the policy gradient algorithm, and includes an actor network and a critic network. The actor network inputs the state data of the system and outputs the current distribution network reconstruction strategy. The critic network inputs the current data of the system and the reconstruction strategy given by the actor network and outputs the state-action value to guide the training parameters. The power system executes the optimization strategy given by the actor network, returns the reward parameters, and updates its own system state. The neural network based on reinforcement learning is expressed as:
[0019] where is the switch action of the j th branch, is the The distribution network topology reconstruction action at a moment, is the topological state of the system, is the total state of the system at the is the node load demand, is the predicted value of future photovoltaic power; S represents the state space, with a dimension of 1× N branch ; is the loss function of the critic network for updating the parameters of the critic network; is the system state at the next moment; 、 are the main critic network and the main actor network respectively; is the target critic network, is the target actor network, and 、 and have exactly the same structure, and obtain their parameters through soft parameter update after several interactions; s and a are the system state and the distribution network topology reconstruction action at each moment respectively; is the parameter of the main actor network, obtained by the main critic network through gradient ascent for the state-action value, is the parameter change amount, n is the total number of samples collected and used for training in the current training batch.
[0020] Furthermore, in step 4, the correction includes: Step 41, determine the photovoltaic power coefficient of all nodes through the following formula:
[0021] where, is the photovoltaic power coefficient of node ; Step 42: Prioritize the reduction of the node with the highest photovoltaic power coefficient through the following formula:
[0022] where, is the reduction coefficient for regulating the reduction ratio; and are the photovoltaic powers of the node with the highest photovoltaic power coefficient before and after the reduction respectively; Step 43: After the curtailment, update the status and perform power flow calculation again. If there are still violations, recalculate the PV power coefficient and perform curtailment again. Iterate until the distribution network topology reconstruction strategy given by the model meets the system security constraints.
[0023] According to the second aspect of the present invention, there is provided a distribution network operation mode optimization system for distributed PV accommodation using the method as described in the first aspect of the present invention. The system includes: A data acquisition module for acquiring system data, including system power data with time stamps, system structure data, and meteorological data; the system power data includes PV power data and load data of each node. A power prediction module for identifying the scenario pattern of the current PV power using a scenario classification model, and calling the corresponding PV power prediction model according to the identified scenario pattern to obtain the PV power prediction value; the scenario classification model and the PV power prediction model are obtained by training after dividing the system power data based on meteorological data using clustering technology. A distribution network reconstruction optimization module for taking the current system power data, system structure data, and PV power prediction value as the system state and inputting them into a decision-making model based on reinforcement learning to obtain the distribution network topology reconstruction strategy under the current system state. A decision correction module for monitoring the distribution network topology reconstruction strategy, and correcting the distribution network topology reconstruction strategy that violates the system security constraints until it meets the system security constraints, so as to obtain an optimized distribution network topology reconstruction strategy.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention introduces a classification model to call different prediction models for PV power under different scenario patterns for targeted prediction, providing a data basis for the optimization model. 2. Based on the reinforcement learning theory, the present invention constructs a virtual environment training decision-making model for the distribution network system, realizing real-time and high-speed reconstruction and optimization of the distribution network topology.
[0025] 3. The present invention constructs an additional security layer that can detect and correct decision instructions that violate the system security constraints. The decision-making model based on reinforcement learning of the present invention can dynamically adjust the power grid operation mode, further improving the PV accommodation capacity of the distribution network system and enhancing the security and economy of system operation.
[0026] In summary, the present invention identifies and solves the physical mechanism and power change trend of the distribution network by integrating classification technology, prediction technology, and reinforcement learning theory, and gives an efficient distribution network operation scheduling strategy aiming at promoting the accommodation of distributed photovoltaic power. Aiming at the problem that the strategy given by the neural network may violate the limits, a safety layer is constructed to ensure that all system safety constraints are met, improving the economy and safety of the operation of the distribution network with distributed photovoltaic power. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 FIG. is a schematic diagram of the general process of an optimization method for the operation mode of a distribution network for accommodating distributed photovoltaic power according to the present invention; Figure 2 FIG. is a schematic diagram of the specific model structure of an optimization method for the operation mode of a distribution network for accommodating distributed photovoltaic power according to the present invention; Figure 3 FIG. is a schematic diagram of the topological structure of the optimized distribution network under different time sections; Figure 4 FIG. is a schematic diagram of the structure of an optimization system for the operation mode of a distribution network for accommodating distributed photovoltaic power in an embodiment of the present invention.
[0028] FIGS. 5(a) to 5(c) are respectively the clustering result diagrams obtained by using the K-means clustering, AP clustering, and hierarchical clustering methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] In the first aspect of the present invention, an optimization method for the operation mode of a distribution network for accommodating distributed photovoltaic power is designed. As shown in FIGS. and, in one embodiment, the method includes the following steps: Figure 1 and 2 shown, in one embodiment, the method includes the following steps: Step 1: Collect system data, including system power data with time stamps, system structure data, and meteorological data; the system power data includes photovoltaic power data and load data of each node.
[0031] In this step, the meteorological data includes temperature data, humidity data, and air pressure data at each moment; the system structure data includes the topological structure of the current distribution network system, the status of line switches, and the impedance data of each line. The timestamp of the meteorological data corresponds to the power data and is used for subsequent clustering of power data into pattern scenarios. The load data and system structure data are used as decision-making information for power system scheduling. For the operation of the power system, no specific distinction is made between the user load and the distributed photovoltaic power. Therefore, the difference between the two, that is, the net load data, is used as the power data input of the power system. Except for the status of line switches, the grid structure data are inherent parameters of the system and will not change with the change of the scheduling strategy.
[0032] Step 2: Use the scenario classification model to identify the scenario pattern corresponding to the current photovoltaic power, and call the photovoltaic power prediction model corresponding to the identified scenario pattern to obtain the photovoltaic power prediction value. The scenario classification model and the photovoltaic power prediction model are trained based on the clustering technology by dividing the system power data through meteorological data.
[0033] Specifically, in Step 2, the scenario classification model and the photovoltaic power prediction model are trained in the following way: Step 21, divide the photovoltaic power data at each moment into the corresponding season according to the timestamp.
[0034] In this step, since the timestamp is recorded when the power data is collected, the power data can be divided into four seasons: spring, summer, autumn, and winter according to the timestamp.
[0035] Step 22, select humidity, air pressure, and temperature from the meteorological dataset as the characteristic variables for clustering analysis. Through the K-Means clustering algorithm, further divide the photovoltaic power data in each season into three weather patterns: sunny, cloudy, and rainy, so as to divide the photovoltaic power data into sub-power datasets under 12 types of scenario patterns.
[0036] Specifically, Step 22 includes: Step 221, randomly select 3 initial centroids, and complete the clustering division of the power data by continuously updating the centroids and allocating the meteorological data points corresponding to each photovoltaic power data to the nearest centroid according to the distance difference from each centroid. Among them, each centroid is a three-dimensional vector composed of the average values of humidity, air pressure, and temperature; the calculation process of the centroid and the distance is as follows:
[0037] Among them, is the meteorological data point i and the centroid j the Euclidean distance between them; H i 、 Pi and T i represent the humidity value, air pressure value, and temperature value of the data point respectively; i The humidity value, air pressure value, and temperature value of the data point; k is the number of the category cluster to which the meteorological data point belongs, N k is the k number of meteorological data points in the th category cluster, and the new centroid coordinate; Cluster k represents the k th category cluster, and the set of meteorological data points corresponding to each photovoltaic power data.
[0038] It should be noted that when performing clustering, setting each centroid as a three-dimensional vector composed of the average values of humidity, air pressure, and temperature is carefully considered. Specifically, for photovoltaic prediction, although irradiance is the key data that directly affects the power generation of a photovoltaic power station and traditional algorithms often use this data as a clustering feature, high-precision irradiance measurement equipment is expensive, requires regular maintenance and calibration, and can only represent data within a certain range with limited coverage. In contrast, it is easier to obtain humidity, air pressure, and temperature data. The measurement equipment is relatively popular, and meteorological data can be obtained through various public channels such as weather forecasts and online meteorological platforms. The data of a single measurement point can reflect the meteorological characteristics of a certain area with a large coverage. Therefore, different from traditional feature selection, the present invention trains a classification model through macroscopic meteorological data (i.e., humidity, air pressure, and temperature, etc.), so as to reduce the data acquisition cost and improve the prediction efficiency while ensuring the prediction accuracy.
[0039] Step 222: Analyze the characteristics of the meteorological data in each category cluster, and determine and mark the weather pattern corresponding to each category cluster according to the analysis results.
[0040] Among them, the weather pattern corresponding to each category cluster can be determined according to the characteristic values of the clustering center, combined with meteorological knowledge and experience. For example, if a certain cluster has the characteristics of high humidity, low temperature, and low air pressure, it can be defined as a "rainy weather pattern"; if a certain cluster has the characteristics of low humidity, high temperature, and high air pressure, it can be defined as a "sunny weather pattern"; if a certain cluster has the characteristics of relatively high humidity, relatively low temperature, and relatively low air pressure, it can be defined as a "cloudy weather pattern". It should be noted that the humidity, temperature, and air pressure characteristics in rainy, sunny, and cloudy days have certain universality and regularity, but they will also change due to various factors. Therefore, in practical applications, it is necessary to conduct comprehensive analysis and judgment in combination with specific meteorological observation data and meteorological principles.
[0041] Step 23: Based on the partitioning result, use the system data as the data source to train the scenario classification model and the photovoltaic power prediction model for each scenario pattern.
[0042] Generally speaking, this step trains the corresponding scenario photovoltaic power prediction models for different data sets respectively, trains the power scenario classification model using the entire system data set and classification label information, and determines the current scenario pattern based on the power sequence. Two types of models, namely the scenario classification model and the photovoltaic power prediction models for each scenario pattern, are both constructed using neural networks.
[0043] Specifically, Step 23 includes: Step 231: Construct the loss functions of the scenario classification model and the photovoltaic power prediction model respectively, expressed as:
[0044] Where, N 1 and N 2 are the numbers of samples in the training sample sets corresponding to the scenario classification sub-model and the classification prediction sub-model respectively, k is the sample category; is the loss function of the scenario classification sub-model, is the true category label of the i th sample in the training sample set corresponding to the scenario classification sub-model, is the probability that the i th sample in the training sample set corresponding to the scenario classification sub-model belongs to the category k , is the loss function of the photovoltaic power prediction model, and are the true photovoltaic power value and the predicted photovoltaic power value corresponding to the i th sample in the training sample set corresponding to the classification prediction sub-model respectively.
[0045] Step 232: With the goal of minimizing and , update the parameters of the scenario classification model and the photovoltaic power prediction model respectively by using cross-entropy loss and mean absolute error and calling the Adam optimizer through gradient backpropagation to realize the training of the scenario classification sub-model and the classification prediction sub-model.
[0046] In this way, the scenario pattern of the current power fluctuation can be identified through the scenario classification model according to the photovoltaic power sequence, and the photovoltaic power prediction model under the corresponding scenario pattern can be called to obtain the predicted photovoltaic power value of the distributed photovoltaic.
[0047] Further, methods such as cross-validation can be used to evaluate the stability and accuracy of the clustering results, and according to the verification results, the feature variable parameters can be adjusted to optimize the clustering results.
[0048] To further illustrate the advantages of photovoltaic power prediction combined with the K-means clustering method, Table 1 shows the improvement of the prediction accuracy index of the prediction model using K-means clustering compared with the ordinary prediction model that does not use K-means clustering and only uses timestamp classification.
[0049] Table 1
[0050] Among them, RMSE represents the root mean square error, MAE represents the mean absolute error, and SMAPE represents the symmetric mean absolute percentage error. It can be seen from Table 1 that after using K-means clustering, the prediction accuracy of the photovoltaic prediction model has been significantly improved.
[0051] In addition, it should be noted that in addition to the K-means clustering algorithm, other clustering algorithms can also be used, such as the AP clustering algorithm and the hierarchical clustering algorithm. Figures 5(a) to 5(c) respectively show the K-means clustering, AP clustering, and hierarchical clustering results in sequence, where the number of AP clustering categories is greater than 3. It can be seen from Figures 5(a) to 5(c) that after merging adjacent categories, the clustering results obtained by the three clustering algorithms are less different. However, among these three algorithms, the K-means clustering algorithm is simple and fast, and can efficiently aggregate large-scale feature data. Therefore, considering comprehensively, the K-means algorithm is the best choice among these three algorithms.
[0052] Step 3: Input the system power data, system structure data, and photovoltaic power prediction value of the current period as the current system state into the decision-making model based on reinforcement learning to obtain the distribution network topology reconstruction strategy under the current system state.
[0053] Specifically, Step 3 includes: Step 31, establish a virtual environment model of the power system, and the virtual environment model of the power system is used to simulate the steady state of the distribution network through distributed power flow calculation based on the distribution network topology reconstruction strategy output by the decision-making model, and generate simulated system power data and system structure data in combination with the set photovoltaic power prediction value.
[0054] Among them, by generating simulated system power data and system structure data through the virtual environment model of the power system, data samples for parameter training of the decision-making model can be obtained. Specifically, Step 31 includes: Step 311: Perform distributed power flow calculation through the following formula:
[0055] Among them, and are the active power and reactive power of branch ij respectively, and are the active power and reactive power of branch jk respectively, and are the active and reactive power of node i respectively, and are the squares of the voltage amplitudes at nodes i and j respectively, is the square of the current amplitude on branch ij , and represent the sets of nodes and branches respectively, represents the branch i between nodes j and ij , and are the resistance and reactance of branch ij respectively, and represent the upper and lower bounds of the variable respectively.
[0056] Step 312: Determine whether the distributed power flow solution result converges. When it is determined that the power flow solution result does not converge, a first negative penalty is fed back r pen1 .
[0057] Step 313: Determine whether the current distribution network topology reconstruction strategy meets the system safe operation conditions. When the system safe operation conditions are violated, a second negative penalty is fed back r pen2 .
[0058] Comparing with the method of using the AC power flow formula to solve by the Newton-Raphson method in the prior art, in this embodiment, a faster solution speed can be obtained by using the distributed power flow solution formula, so that the result can be generated quickly. And the reinforcement learning model has extremely high requirements for the amount of training data, usually interacting with the environment hundreds of thousands to millions of times. Therefore, adopting the distributed power flow solution method can significantly improve the interaction efficiency between the reinforcement learning model and the environment.
[0059] Step 32: Enable the power system virtual environment model and the decision-making model to interact, and train a neural network based on reinforcement learning with the goal of maximizing the reward function to obtain the decision-making model; the reward function is related to the PV accommodation ratio index of the distribution network operation, the network loss index, the switch loss index, whether the power flow calculation result converges, and whether the current distribution network topology reconstruction strategy meets the system security constraints.
[0060] The reward function determines the optimization direction of the decision-making model. In the present invention, by setting an appropriate reward function, that is, considering both the PV accommodation index, economic indexes (including the network loss index and the switch loss index), and safety indexes (i.e., whether the power flow converges and whether the topology reconstruction strategy meets the system security constraints), the decision given by the trained model can find a topology structure that can not only meet the safety constraints but also minimize the network loss and economic cost while accommodating as much PV as possible, which is of great significance for improving the operation efficiency and economy of the power system.
[0061] Specifically, the reward function of this embodiment is expressed as:
[0062] Where, is the accommodation ratio of the distributed PV in the system, and are the PV outputs before and after optimization, is the line loss of the network, is the number of PV nodes, is the total number of system branches, is the basic reward of the reinforcement learning model, is the number of switch changes under the current strategy, and and are proportionality coefficients used to adjust the weights of the PV accommodation ratio index, the network loss index, and the switch loss index in the reward respectively. is a small quantity used to prevent the reward from overflowing due to too small a denominator or reporting an error due to division by zero. is an additional negative penalty introduced when the current distribution network topology reconstruction strategy does not meet the system security constraints, is the total reward function of the model.
[0063] The neural network based on reinforcement learning is updated using the policy gradient algorithm, and the data structure adopts a triple architecture of action-state-reward. The neural network based on reinforcement learning consists of an actor network and a critic network. The actor network inputs the state data of the system and outputs the current distribution network reconstruction strategy. The critic network inputs the current data of the system and the reconstruction strategy given by the actor network and outputs the state-action value to guide the training parameters. In this process, the power system executes the optimization strategy given by the actor network, returns the reward parameter, and updates its own system state. The neural network based on reinforcement learning is expressed as:
[0064] where, is the switch action of the j th branch, is the distribution network topology reconstruction action at the th moment, is the topological state of the system, is the th moment of the total state of the system, is the node load demand, is the predicted value of future PV output; S represents the state space, with a dimension of 1× N branch ; is the loss function of the critic network for updating the parameters of the critic network, is the system state at the next moment; , are the main critic network and the main actor network respectively; is the target critic network of the model, is the target actor network of the model, and , and have exactly the same structure, and parameters are obtained by soft-updating the parameters after several interactions. s and a are the system state and the distribution network topology reconstruction action at each moment respectively; is the parameter of the main actor network, obtained by gradient ascent of the main critic network for the state-action value, is the parameter change amount, n is the total number of samples collected and used for training in the current training batch.
[0065] Step 4: Monitor the distribution network topology reconstruction strategy, correct the distribution network topology reconstruction strategy that violates the system security constraints until the system security constraints are met, and obtain the optimized distribution network topology reconstruction strategy.
[0066] Specifically, in step 4, the correction includes: Step 41: Determine the photovoltaic power coefficients of all nodes through the following formula:
[0067] where is the photovoltaic output coefficient of the node. The higher the proportion of indicates that the photovoltaic output of this node is more serious and needs to be preferentially reduced.
[0068] Step 42: Preferentially reduce the node with the highest photovoltaic power coefficient through the following formula:
[0069] where, is the reduction coefficient used to regulate the reduction ratio. and are the photovoltaic powers of the node with the highest photovoltaic power coefficient before and after reduction, respectively.
[0070] Step 43: After reduction, update the state and re-perform the power flow solution. If there are still violations, recalculate the output coefficient and perform reduction again, and iterate until the given distribution network topology reconstruction strategy meets the system safety constraints.
[0071] To verify the performance of the proposed method, the actual optimization effect of the model is evaluated based on the actual system data. Among them, the case study is carried out based on a 100-node distribution network system constructed from the local data of Lianyungang. The network topology diagram includes three standby lines and several distributed photovoltaic power sources. Among them, the distributed power sources and the user's power load demand are combined as the system input. The photovoltaic data is collected from the actual photovoltaic output in Lianyungang and then normalized and incorporated into the system data. Based on this dataset, the distribution network system is reconstructed and optimized in real time. The distribution network system is equipped with 14 tie switches, which are usually in the closed state. To reduce the network loss and stabilize the voltage, and further improve the photovoltaic accommodation capacity of the regional system, they are opened during operation to adjust the network topology. The experimental parameter configuration is as follows: the negative penalty for violating the safety constraint is -0.5, the negative penalty for the case of non-convergence is -1, and the discount return coefficient for DDPG training is 0.99. Under the above configuration, the proposed classification prediction model and decision model are both trained based on the real-world meteorological and power datasets.
[0072] Figure 3 Shows the schematic diagram of the topology structure of the power system distribution network at different time sections. Table 2 shows the results of various indicators for each model to reconstruct and optimize the operation mode of the distribution network. The effectiveness of the method is measured by comparing the number of switch operations, the average reward value, and the time used for scheduling decision calculation during its cycle.
[0073] Table 2 Optimization Results of Different Models
[0074] It can be observed that for different load demands and changes in PV output, the proposed method can accurately capture the distance weights between the load and the power source, distribute the user load by adjusting the tie switches of the controllable branches, and improve the operating conditions of the power grid with the least number of operations. For the reconfiguration scheduling results within one day, the average network loss is controlled from 3.57 MW to 3.34 MW, and the total network loss is reduced by about 6%. The voltage under the initial topology is concentrated in the range of 0.97 - 1.02 p.u. After reconfiguration optimization, the voltage distribution of each node at different times is more concentrated near the rated voltage. During the optimization period, the voltage fluctuation amplitude is significantly reduced, and the voltage deviation is reduced by about 3%. The local voltage rise during the peak period of PV output is significantly suppressed. This shows that the method proposed in this study can improve the power flow distribution of the distribution network system through small-step topology changes and further enhance the PV accommodation capacity of the system.
[0075] In the second aspect of the present invention, a distribution network operation mode optimization system for distributed PV accommodation is provided, which utilizes the method as described in the first aspect of the present invention. As Figure 4 , the system includes: A data acquisition module for collecting system data, including system power data with timestamps, system structure data, and meteorological data; the system power data includes PV power data and load data of each node; A power prediction module for identifying the scenario pattern of the current PV power using a scenario classification model and obtaining the PV power prediction value by calling the corresponding PV power prediction model according to the identified scenario pattern; the scenario classification model and the PV power prediction model are trained by dividing the system power data based on meteorological data through clustering technology; A distribution network reconfiguration optimization module for taking the current system power data, system structure data, and PV power prediction value as the system state and inputting them into a decision-making model based on reinforcement learning to obtain the distribution network topology reconfiguration strategy under the current system state; A decision correction module for monitoring the distribution network topology reconfiguration strategy and correcting the distribution network topology reconfiguration strategy that violates the system security constraints until the system security constraints are met, obtaining the optimized distribution network topology reconfiguration strategy.
[0076] In summary, compared with the prior art, the method of the present invention analyzes the operation mechanism of a distribution network with distributed photovoltaic power generation, constructs a reinforcement learning model for optimizing the operation mode of the distribution network for photovoltaic power consumption, introduces a classification mechanism and prediction technology to provide solid and effective optimization information data for decision-making, and embeds the artificial intelligence agent technology into the traditional operation optimization problem by designing a unique reward function and agent architecture, which can ensure that the dispatching optimization scheme fully meets the system safety constraint conditions, thereby guaranteeing the safety and stability of the operation of the active distribution network.
[0077] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0078] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0079] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a respective computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0080] Computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0081] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the knowledge scope of those of ordinary skill in the art.
Claims
1. A distribution network operation mode optimization method for distributed photovoltaic consumption, characterized in that: The following steps are involved: Step 1: Collect system data, including system power data with timestamps, system structure data and meteorological data; the system power data includes photovoltaic power data and load data of each node; Step 2: Use the scene classification model to identify the scene mode corresponding to the current photovoltaic power, and call the photovoltaic power prediction model corresponding to the identified scene mode to obtain the photovoltaic power prediction value; The scene classification model and the photovoltaic power prediction model are obtained by training after dividing the system power data through meteorological data based on clustering technology; Step 3: Input the system power data, system structure data and photovoltaic power forecast value of the current period as the current system state into the decision model based on reinforcement learning to obtain the distribution network topology reconstruction strategy under the current system state; Step 4: Monitor the distribution network topology reconstruction strategy, and modify the distribution network topology reconstruction strategy that violates the system safety constraints until the system safety constraints are met, thereby obtaining an optimized distribution network topology reconstruction strategy.
2. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 1, characterized in that: In step 1, the meteorological data includes the temperature data, humidity data and air pressure data of the substation area at each moment; the system structure data includes the topological structure of the power system, the state of the line switches and the impedance data of each line at each moment.
3. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 1, characterized in that: In step 2, the scene classification model and the photovoltaic power prediction model are trained in the following manner: Step 21, dividing the photovoltaic power data at each moment into corresponding seasons according to the timestamp; Step 22, humidity, air pressure and temperature are selected from the meteorological data set as characteristic variables for cluster analysis, and the photovoltaic power data of each season is further divided into three weather modes: sunny, cloudy and rainy, by using the K-Means clustering algorithm, so as to divide the photovoltaic power data into sub-power data sets under 12 scene modes; Step 23, based on the division result, the system data is used as a data source to train the scene classification model and the photovoltaic power prediction model under each scene mode.
4. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 3 is characterized in that: Step 22 includes: Step 221, randomly select three initial centroids, continuously update the centroids and assign the meteorological data points corresponding to each photovoltaic power data to the nearest centroid according to the distance difference from each centroid, so as to complete the clustering of the power data; wherein each centroid is a three-dimensional vector composed of the average values of humidity, air pressure and temperature; the centroid and distance are calculated as follows: in, For weather data points i With the centroid j The Euclidean distance between H i , P i and T i Represents the data points i Humidity, air pressure and temperature values; k is the number of the category cluster to which the meteorological data point belongs, N k For the k The number of meteorological data points in the class cluster, is the new centroid coordinate; Cluster k Indicates k The set of meteorological data points corresponding to each photovoltaic power data in the category cluster; Step 222: Analyze the characteristics of the meteorological data in each category cluster, and determine and mark the weather pattern corresponding to each category cluster according to the analysis results.
5. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 3 is characterized in that: Step 23 includes: Step 231: construct the loss functions of the scene classification model and the photovoltaic power prediction model respectively, expressed as: in, N 1 and N 2 are the number of samples in the training sample set corresponding to the scene classification sub-model and the classification prediction sub-model, respectively. k is the sample category; is the loss function of the scene classification sub-model, is the first i The true category labels of samples, is the first i samples belong to the category k The probability of is the loss function of the photovoltaic power prediction model, and They are the first i The actual value of photovoltaic power and the predicted value of photovoltaic power corresponding to the samples; Step 232, to minimize and As the goal, the Adam optimizer is called to update the parameters of the scene classification model and the photovoltaic power prediction model by using cross entropy loss and mean absolute error through gradient back propagation, so as to realize the training of the scene classification sub-model and the classification prediction sub-model.
6. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 1, characterized in that: In step 3, the decision model based on reinforcement learning is trained in the following way: Step 31, establishing a power system virtual environment model, wherein the power system virtual environment model is used to simulate the steady state of the distribution network through distributed power flow calculation based on the distribution network topology reconstruction strategy output by the decision model, and generate simulated system power data and system structure data in combination with the set photovoltaic power prediction value; Step 32, making the power system virtual environment model and the decision model interact, training a neural network based on reinforcement learning with the goal of maximizing the reward function, and obtaining the decision model; the reward function is related to the photovoltaic absorption ratio index, network loss index, switch loss index, whether the power flow solution result converges, and whether the current distribution network topology reconstruction strategy meets the system safety constraints.
7. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 6, characterized in that: Step 31 includes: Step 311: Perform distributed power flow calculation using the following formula: in, , Branch ij The active power and reactive power of , Branch jk The active power and reactive power of , Node j The active and reactive power, and The nodes are i and j The square of the voltage amplitude on For branch ij The square of the current amplitude, , denote the set of nodes and branches respectively, Representation Node i and j Between branches ij , , Branch ij The resistance and reactance, and Respectively represent the upper and lower bounds of the variable; Step 312: Determine whether the distributed power flow solution result converges, and feedback the first negative value penalty when it is determined that the power flow solution result does not converge. r pen1 ; Step 313: Determine whether the current distribution network topology reconstruction strategy meets the system safety operation conditions, and feedback the second negative value penalty when the system safety operation conditions are violated. r pen2 .
8. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 7, characterized in that: In step 32, the reward function is expressed as: in, is the absorption ratio of distributed photovoltaic system, , is the photovoltaic power before and after optimization, is the line loss of the network, is the number of photovoltaic nodes, is the total number of system branches, is the basic reward for the reinforcement learning model, is the number of switch changes under the current strategy, , , The proportional coefficients are used to adjust the weights of the photovoltaic absorption ratio index, network loss index and switch loss index in the reward; A small amount used to prevent the denominator from being too small and causing reward overflow or division by zero error; is the total negative penalty, and r pen = r pen1+ r pen2 ; is the total reward function of the model.
9. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 8, characterized in that: In step 32, the reinforcement learning-based neural network is updated using a policy gradient algorithm and includes an actor network and a critic network; the actor network inputs the state data of the system and outputs the current distribution network reconstruction strategy, the critic network inputs the current system data and the reconstruction strategy given by the actor network and outputs the state action value to guide the training parameters, the power system executes the optimization strategy given by the actor network to return the reward parameter and update its own system state; the reinforcement learning-based neural network is expressed as: in, For the j The switch action of each branch, For the The distribution network topology reconstruction action at each moment, is the topological state of the system, For the The total state of the system at time is the node load demand, is the predicted value of future photovoltaic power; S represents the state space with a dimension of 1× N branch ; is the loss function of the critic network, which is used to update the parameters of the critic network; is the system state at the next moment; , They are the main critic network and the main actor network respectively; For the target critic network, For the target actor network, and , and The structure is exactly the same. and The parameters of are obtained by performing soft parameter updates after several interactions; s and a They are the system status at each moment and the distribution network topology reconstruction action; The parameters of the main actor network are obtained by gradient ascent of the state-action value of the main critic network. is the parameter change, n It is the total number of samples collected and used for training in the current training batch.
10. The method for optimizing the operation mode of distribution network for distributed photovoltaic consumption according to claim 8, characterized in that: In step 4, the correction includes: Step 41: Determine the photovoltaic power factor of all nodes by the following formula: in, For Node The photovoltaic power coefficient; Step 42: The node with the highest photovoltaic power coefficient is preferentially reduced by the following formula: in, The reduction coefficient is used to adjust the reduction ratio; and They are the nodes with the highest photovoltaic power coefficient. PV power before and after curtailment; Step 43: After the reduction, the state is updated and the power flow is recalculated. If there is still a violation, the photovoltaic power coefficient is recalculated and reduced again. The cycle is iterated until the given distribution network topology reconstruction strategy meets the system safety constraints.
11. A distribution network operation mode optimization system for distributed photovoltaic consumption using the method according to any one of claims 1 to 10, characterized in that: include: A data acquisition module, used to collect system data, including system power data with time stamps, system structure data and meteorological data; The system power data includes photovoltaic power data and load data of each node; A power prediction module is used to identify the current photovoltaic power scene mode using a scene classification model, and call the corresponding photovoltaic power prediction model according to the identified scene mode to obtain the photovoltaic power prediction value; The scene classification model and the photovoltaic power prediction model are obtained by training after dividing the system power data through meteorological data based on clustering technology; The distribution network reconstruction optimization module is used to input the current system power data, system structure data and photovoltaic power prediction value as the system state into the decision model based on reinforcement learning to obtain the distribution network topology reconstruction strategy under the current system state; The decision correction module is used to monitor the distribution network topology reconstruction strategy and correct the distribution network topology reconstruction strategy that violates the system safety constraints until the system safety constraints are met, thereby obtaining an optimized distribution network topology reconstruction strategy.
Citation Information
Patent Citations
Hydropower station group monthly transaction plan electric quantity decomposition method considering power grid section constraints
CN110400232A
Distributed photovoltaic power generation power prediction system and method
CN117933531A
Photovoltaic short-term generation power combined prediction method, system, equipment and medium
CN118801343A
Distributed photovoltaic power prediction method and system based on machine learning
CN119853029A
Power distribution network reconstruction method and system based on topology security constraint and integrated reinforcement learning
CN119891150A