High-permeability distributed photovoltaic power distribution network reconstruction method and system

By constructing an expert policy trajectory library and state imitation reinforcement learning, the problem of unstable training process in high-penetration distributed photovoltaic distribution networks was solved, achieving rapid response and stable distribution network reconfiguration, and improving the system's carrying capacity and decision-making performance.

CN122000984APending Publication Date: 2026-05-08STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE
Filing Date
2025-11-25
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing reinforcement learning methods suffer from large fluctuations in the training process, slow convergence speed, and low sample efficiency in high-penetration distributed photovoltaic power distribution networks. They are unable to respond quickly to fluctuations in photovoltaic output and lack effective prior knowledge guidance, resulting in insufficient decision stability and feasibility, and failing to meet the needs of rapid reconstruction.

Method used

By constructing an expert policy trajectory library, combining state imitation reinforcement learning and traditional topology optimization algorithms, a continuous state space is designed, policy samples are generated and high-quality trajectories are selected, an agent model is trained, and the expert policy trajectory library is dynamically updated to improve the model's convergence speed and decision stability.

Benefits of technology

It enables rapid response to photovoltaic power output fluctuations, alleviates voltage over-limit and grid loss problems, improves the overall carrying capacity and decision-making stability of the system, and enhances its adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122000984A_ABST
    Figure CN122000984A_ABST
Patent Text Reader

Abstract

The invention discloses a high-permeability distributed photovoltaic power distribution network reconstruction method and system, and belongs to the technical field of power system scheduling, and the method comprises the steps: designing a continuous state space based on the system data of a distributed photovoltaic power distribution network, and generating a plurality of attribute state matrixes; based on a traditional distribution network topology optimization algorithm, generating a strategy sample by using an attribute state matrix, and screening high-quality strategy trajectories to construct an expert strategy trajectory library; based on a state simulation reinforcement learning algorithm, using an expert strategy trajectory library to train an agent model; inputting the current distribution network operation state data into the agent model to generate a real-time reconstruction strategy; evaluating a real-time reconstruction strategy based on a comprehensive index system, and dynamically updating an expert strategy trajectory library adopted during intelligent agent model training; according to the method, the expert strategy trajectory library is introduced, so that the reinforcement learning model can be quickly applied to a new scene, a relatively good decision-making effect is obtained, distributed photovoltaic fluctuation is responded in real time, and the network frame topology of the system is quickly and accurately adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for reconfiguring a high-penetration distributed photovoltaic power distribution network, belonging to the field of power system dispatching technology. Background Technology

[0002] With the large-scale integration of distributed photovoltaic (PV) power generation, the operation of distribution networks has undergone profound changes. Voltage exceedances and increased network losses are becoming increasingly prominent, necessitating dynamic topology reconfiguration to optimize power flow distribution and enhance system capacity. Traditional distribution network reconfiguration largely relies on mathematical optimization methods to find the optimal topology solution while satisfying radial constraints, switching operation limits, and power flow feasibility. In recent years, reinforcement learning has been introduced into this field due to its potential in complex decision-making problems, enabling autonomous decision-making by constructing an environment state and action space. Existing research has employed deep reinforcement learning models to generate distribution network reconfiguration strategies, utilizing neural networks to approximate policy or value functions to achieve responses to real-time operating conditions.

[0003] However, pure reinforcement learning methods face challenges in practical applications, including large fluctuations in the training process, slow convergence speed, and low sample efficiency. This is particularly true in high-dimensional, nonlinear power systems with stringent safety constraints, where ensuring decision stability and feasibility is difficult. Furthermore, due to the lack of effective prior knowledge guidance, agents require extensive exploration to obtain optimal policies, leading to high training costs and difficulties in rapid deployment to new scenarios. Existing methods typically rely solely on online learning mechanisms, failing to fully utilize historical high-quality policy trajectories as empirical guidance. This limits the model's generalization ability and real-time response performance, making it unsuitable for the rapid reconstruction requirements of frequently fluctuating high-penetration photovoltaic scenarios. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a high-penetration distributed photovoltaic power distribution network reconfiguration method and system. By introducing an expert policy trajectory library, reinforcement learning models can be quickly applied to new scenarios and achieve relatively good decision-making results. It can respond to distributed photovoltaic fluctuations in real time and quickly and accurately adjust the system network topology.

[0005] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0006] In a first aspect, the present invention provides a method for reconfiguring a high-penetration distributed photovoltaic distribution network, comprising:

[0007] The acquired current distribution network operation status data is input into a pre-trained intelligent agent model to generate a real-time reconstruction strategy;

[0008] The real-time reconstruction strategy is evaluated based on a comprehensive index system, and the expert strategy trajectory library used during the training of the intelligent agent model is dynamically updated.

[0009] The training method for the intelligent agent model includes:

[0010] Based on the acquired system data of the distributed photovoltaic power distribution network, a continuous state space is designed to generate multiple attribute state matrices;

[0011] Based on the traditional distribution network topology optimization algorithm, the attribute state matrix is ​​used to generate strategy samples, and high-quality strategy trajectories are selected to build an expert strategy trajectory library.

[0012] Based on the state imitation reinforcement learning algorithm, the agent model is trained using the aforementioned expert policy trajectory library;

[0013] The intelligent agent model takes the current attribute state matrix and the target attribute state matrix as inputs and outputs the distribution network topology reconfiguration strategy vector.

[0014] Furthermore, the system data includes system network topology data, line impedance data, node voltage data, node injected active power data, and reactive power data.

[0015] Furthermore, the method for acquiring system data of the distributed photovoltaic distribution network includes: collecting parameters of reactive power equipment, energy storage equipment, and transformers configured at each node in the system, as well as the voltage amplitude and voltage phase angle of each node.

[0016] Furthermore, the design of the continuous state space includes: using a multidimensional scaling analysis algorithm to transform the topological connection relationship of the distribution network system nodes into two-dimensional coordinates, wherein the Euclidean distance between the two-dimensional coordinates is as close as possible to the distance matrix calculated based on the line impedance magnitude. The calculation process is as follows:

[0017] ;

[0018] Among them, matrix This is the distance matrix between system nodes. It is in the squared form of the distance matrix. for An identity matrix of order 1, where the diagonal elements are 1s and the off-diagonal elements are 0s. It is a vector of all 1s. This is a centered matrix used to shift data to the origin. It is an inner product matrix used to reflect the inner product relationship between nodes. This is an eigenvalue matrix, where the eigenvalues ​​are arranged in descending order. The eigenvector matrix, This is a diagonal matrix composed of the two largest eigenvalues, whose eigenvalues ​​are arranged in descending order. These are the first two main eigenvector matrices. It is a two-dimensional coordinate matrix.

[0019] Furthermore, the method also includes: fixing the coordinates of node 0 at the origin and aligning the coordinates of node 1 to the vertical axis to unify the coordinate system, as shown in the following expression:

[0020] ;

[0021] in, These are the original coordinates of node 0. The coordinate matrix after translation. Let the vertical and horizontal coordinates of node 1 be... For rotation angle, It is the arctangent function in the fourth quadrant. These are the aligned system node coordinates.

[0022] Furthermore, the multiple attribute state matrices are obtained by downsampling and interpolation within a fixed window range, wherein the attribute matrices include a voltage attribute matrix, an active power attribute matrix, and a reactive power attribute matrix, and the coordinate calculation process for the attribute matrices is as follows:

[0023] ;

[0024] in, For the first The two-dimensional x and y coordinates of the node This represents the boundary range of the transformed coordinates, where l is the side length of the square window covering all coordinate points, and d is a redundant length used to avoid coordinate exceeding the limit in special cases. These represent the center coordinates and boundary coordinates of the square window, respectively. This indicates the number of sampling points on one side of the downsampling grid within the square window. Represents the coordinate sequence of grid points along the horizontal and vertical axes, using This represents the global grid point coordinate matrix.

[0025] Furthermore, the downsampling interpolation uses the radial basis function interpolation method. For sampling grid coordinate points without numerical representation, interpolation is performed based on the attributes under the original system node coordinates to generate a continuous attribute state matrix. The radial basis function interpolation process is as follows:

[0026] ;

[0027] in, For interpolation functions, Let be the radial basis function weight coefficients of the i-th node. For radial basis functions, This is a shape parameter used to control the decay rate of the function. Let be the Euclidean distance from point (x, y) to node . These are the coefficients of the terms in the linear polynomial. For the elements of the radial basis function matrix, , , These represent the polynomial condition matrix, the node attribute value vector, and the interpolation result of the grid points, respectively.

[0028] Furthermore, the generation of strategy samples using the attribute state matrix based on the traditional distribution network topology optimization algorithm includes: solving for the optimal reconfiguration strategy and generating a binary decision vector as a strategy sample, under the conditions of satisfying the system's power flow constraints, line switch operation frequency constraints, and radial topology constraints.

[0029] Furthermore, the strategy samples are spliced ​​together at a set time scale to form a real-time reconstructed strategy trajectory matrix within a day, and stored in the expert strategy trajectory library. Each strategy trajectory corresponds to three sets of attribute state matrices.

[0030] Furthermore, the state imitation reinforcement learning algorithm includes two sets of neural network models: an action network model and a judge network model;

[0031] The action network model takes as input three complete attribute state matrices at the current moment and a target voltage attribute state matrix from the expert policy trajectory library, and outputs a decision vector that transfers the current environmental state to the target voltage state.

[0032] The evaluation network model takes the decision vector and the complete attribute state matrix at the current moment as input, and outputs an estimated value of the system voltage attribute state matrix after the decision is executed.

[0033] Furthermore, the training of the intelligent agent model uses a convolutional neural network as a pre-feature extraction layer, and uses the L2 norm of the deviation between the system's target voltage state matrix and the estimated voltage state as the loss function to guide parameter optimization. The calculation process is as follows:

[0034] ;

[0035] in, These are the internal network parameters of the model. This refers to the observation state at adjacent time steps, and here it generally refers to the state matrix of the three types of attributes. For the predicted state transition action, To evaluate the model's fitting loss and ensure that the evaluation model f can accurately predict real actions. The change in the following environmental state, For forward consistency loss, the L2 norm is used to measure the deviation between the transition state and the actual state. It is a standard action loss function, which is used to mimic decision-making actions in an expert policy trajectory library through labeled training.

[0036] Furthermore, the step of inputting the acquired current distribution network operation status data into a pre-trained agent model to generate a real-time reconstruction strategy includes: matching the state variable most similar to the current state from the expert strategy trajectory library, and using the current state matrix and the next-time state matrix in the expert strategy trajectory as input to the agent model to generate a reconstruction strategy vector.

[0037] Furthermore, the expert policy trajectory library used during the dynamic updating of the intelligent agent model training includes: calculating the economic and safety benefits of the current system policy trajectory using a comprehensive index system; when the value reaches a threshold, the current policy trajectory is identified as a high-quality sample and updated to the expert policy trajectory library. The calculation formula is as follows:

[0038] ;

[0039] in, The reward function uses a comprehensive indicator system to quantify the merits of the current decision. This is the network loss coefficient. This refers to the system network loss after the action is performed. The economic coefficient for switch operation. This represents the number of switch actions.

[0040] Secondly, the present invention provides a high-penetration distributed photovoltaic distribution network reconfiguration system for implementing the high-penetration distributed photovoltaic distribution network reconfiguration method described in any one of the foregoing claims, comprising:

[0041] The state space design module is used to design a continuous state space and generate multiple attribute state matrices based on the acquired system data of the distributed photovoltaic distribution network.

[0042] The expert strategy trajectory library module is used to generate strategy samples based on the attribute state matrix using the traditional distribution network topology optimization algorithm, and to select high-quality strategy trajectories to build an expert strategy trajectory library.

[0043] The agent training module is used to train an agent model based on the state imitation reinforcement learning algorithm and using the expert policy trajectory library; wherein, the agent model takes the current attribute state matrix and the target attribute state matrix as input and outputs the distribution network topology reconstruction policy vector.

[0044] The reconstruction strategy calculation module is used to input the acquired current distribution network operation status data into the pre-trained intelligent agent model to generate a real-time reconstruction strategy;

[0045] The dynamic update module is used to evaluate the real-time reconstruction strategy based on a comprehensive index system and dynamically update the expert strategy trajectory library used during the training of the intelligent agent model.

[0046] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0047] Fourthly, the present invention provides an electronic device, comprising:

[0048] Memory, used to store computer programs / instructions;

[0049] A processor for executing the computer program / instructions to implement the steps of any of the methods described above.

[0050] Fifthly, the present invention provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0051] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0052] This invention provides a method and system for reconfiguring a high-penetration distributed photovoltaic (PV) distribution network. By constructing a continuous state space that integrates topology, impedance, and voltage information, it addresses the weakness in feature representation caused by state discretization in traditional reinforcement learning. By establishing an expert policy trajectory library and introducing a state imitation learning mechanism, it utilizes historical high-quality policy state trajectories as prior knowledge to guide agent training, significantly improving model convergence speed and decision stability. Furthermore, by continuously injecting new high-quality policies through a dynamic update mechanism, it enhances the system's long-term adaptability. This method enables rapid response to PV output fluctuations, effectively mitigating voltage exceedance and network loss issues while ensuring safety constraints, and improving the overall carrying capacity of the high-penetration distributed PV distribution network. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the model structure of the high-penetration distributed photovoltaic power distribution network reconfiguration method provided in the embodiments of the present invention;

[0054] Figure 2 This is a schematic diagram of network loss changes before and after reconstruction provided in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of voltage changes before and after reconstruction provided in an embodiment of the present invention. Detailed Implementation

[0056] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0057] Example 1: This example introduces a high-penetration distributed photovoltaic distribution network reconfiguration method, including:

[0058] The acquired current distribution network operation status data is input into a pre-trained intelligent agent model to generate a real-time reconstruction strategy;

[0059] The real-time reconstruction strategy is evaluated based on a comprehensive index system, and the expert strategy trajectory library used during the training of the intelligent agent model is dynamically updated.

[0060] The training method for the intelligent agent model includes:

[0061] Based on the acquired system data of the distributed photovoltaic power distribution network, a continuous state space is designed to generate multiple attribute state matrices;

[0062] Based on the traditional distribution network topology optimization algorithm, the attribute state matrix is ​​used to generate strategy samples, and high-quality strategy trajectories are selected to build an expert strategy trajectory library.

[0063] Based on the state imitation reinforcement learning algorithm, the agent model is trained using the aforementioned expert policy trajectory library;

[0064] The intelligent agent model takes the current attribute state matrix and the target attribute state matrix as inputs and outputs the distribution network topology reconfiguration strategy vector.

[0065] like Figure 1 As shown in the figure, the high-penetration distributed photovoltaic distribution network reconfiguration method provided in this embodiment involves the following steps in its application process:

[0066] Step 1: Based on the distributed photovoltaic distribution network system data and node voltage information collected by physical devices, a continuous reinforcement learning state space is designed by integrating existing system network topology information and line impedance information to construct a distribution network topology scheduling optimization environment. The distributed photovoltaic distribution network system data includes reactive power compensation equipment (such as SVC, SVG) configured at each node, energy storage system parameters (capacity, charging and discharging efficiency), transformer tap positions and turns ratios, and other control resource information. Node voltage information covers the voltage amplitude and phase angle measurements of each node, typically acquired in real time by SCADA systems or PMU devices. The system network topology information describes the connection relationships between buses, switches, and branches, forming a graph structure representation. Line impedance information includes the resistance and reactance parameters of each branch, used for accurate modeling of power transmission characteristics. After normalization, the above multi-source heterogeneous data is mapped to a continuous state representation of a unified dimension, avoiding the "curse of dimensionality" problem caused by traditional discretized states. This state space design not only preserves the physical meaning of the original power system but also possesses good differentiability, facilitating end-to-end learning of subsequent neural network models.

[0067] Step 2: Generate a massive number of strategy samples based on traditional distribution network topology optimization algorithms. Perform strategy calculations under massive distributed photovoltaic power output scenarios, and select high-quality strategy trajectories to construct an expert strategy trajectory library. The traditional distribution network topology optimization algorithms can be selected from Mixed Integer Linear Programming (MILP), Genetic Algorithm, Particle Swarm Optimization (PSO), or convex relaxation methods (such as SOCP). Under the premise of satisfying radial topology, power flow constraints, and switching operation limits, find the optimal or near-optimal reconfiguration scheme. Strategy calculations cover various operating scenarios such as typical days, extreme weather, and seasonal load changes, ensuring the strategy library has broad representativeness. Each strategy trajectory consists of a series of topology decision vectors at different time points and their corresponding state matrix sequences, forming a complete transition path of "state-action-next state".

[0068] Step 3: Train the internal parameters of the decision model based on the state imitation reinforcement learning algorithm, and generate a real-time distribution network topology reconfiguration strategy by inputting the current operating state of the distribution network. The state imitation reinforcement learning algorithm comprises two core components: an action network and a judge network (i.e., a forward dynamics model). The action network receives the current state matrix and the target state matrix as input and outputs an n-dimensional binary decision vector representing the on / off state of each tie switch (0 for closed, 1 for open). The judge network predicts the system's voltage response under a given action and evaluates its consistency with the actual state. During training, a convolutional neural network is used to extract spatial features from the state image, and state transition pairs from the aforementioned expert trajectory library are used as supervision signals to minimize the L2 distance between the predicted and actual states. The loss function considers both action accuracy and state transition consistency to improve model robustness. After training, the agent can complete inference within milliseconds, making it suitable for real-time control requirements.

[0069] Step 4: Construct a comprehensive indicator system to calculate the economic and safety benefits brought by the current system strategy trajectory, extract real-time system strategy data, and dynamically update the expert strategy trajectory library. The comprehensive indicator system quantifies strategy performance in a weighted combination, mainly including three indicators: network loss reduction rate, voltage deviation improvement degree, and switching operation cost. The specific strategy used in a given operation corresponds to... When a value is better than the average level of similar historical scenarios by a certain threshold (e.g., 20% higher), it is judged as a high-quality sample, and its complete state-action sequence is packaged and uploaded to the expert policy trajectory library. The update process supports incremental writing without interrupting system operation. As an optional implementation, a clustering algorithm can also be introduced to classify and archive newly added trajectories, improving subsequent retrieval efficiency. Through the above steps, this application achieves closed-loop accumulation and self-evolution of policy knowledge, enhances the system's adaptability to new operating modes, and extends the model's lifespan.

[0070] The following detailed implementation process of using the method of this invention for distributed photovoltaic (PV) distribution network reconfiguration is illustrated with specific examples. The data used in this invention originates from the IEEE 33-node system. The PV power data used are actual distributed PV output and load data from 2020, which have been normalized to adapt to the system's total capacity setting. The data sampling interval is 15 minutes, totaling 30,000 samples. Distributed PV systems are configured at nodes 3, 5, 7, 10, 16, 22, and 24. Based on the collected data, the specific implementation steps of the method of this invention are as follows:

[0071] (1) System data acquisition and discrete topology information aggregation processing;

[0072] First, the basic parameters and dynamic variables of the system need to be obtained by using the power measurement equipment and state estimation equipment installed in the distribution network system. The data of the distributed photovoltaic distribution network system includes the parameters of the reactive power equipment, energy storage equipment and transformers configured at each node in the system. The node voltage information includes the voltage amplitude and voltage phase angle of each node, taking into account the injected active and reactive power of the nodes after the distributed photovoltaic is connected. The 33-node system contains 37 switching lines used to adjust the grid structure and change the topology flow. In the optimization process, only the system topology connection relationship will be adjusted, and no additional line parameters will be added or changed. The actual decision is equivalent to selecting from a limited number of grid topologies.

[0073] Then, the reinforcement learning agent's parameters are trained through a distribution network topology scheduling optimization environment, guiding the reinforcement learning agent to optimize its own parameters and improve decision-making performance. During parameter training, the state needs to represent complete environmental information in numerical form. Traditional distribution network topology is represented by a node admittance matrix, which becomes discrete and distributed when the topology changes, hindering rapid training convergence. To improve training efficiency, the designed environment state is presented as three continuous two-dimensional attribute state matrices in a single time segment, integrating graph-structured system line topology data, node voltage data, node injected active power data, and node injected reactive power data. The nearest line impedance magnitude is used to characterize the distance attribute between system nodes, where the nearest line is calculated using the Dijkstra algorithm. Based on a multidimensional scaling analysis algorithm, the topological connection relationship of distribution network system nodes is transformed into two-dimensional coordinates. The Euclidean distance between these two-dimensional coordinates is made as close as possible to the given distance matrix. The calculation process is as follows:

[0074] ;

[0075] Among them, matrix This is the distance matrix between system nodes. The squared form of the distance matrix facilitates subsequent calculations. for An identity matrix of order 1, where the diagonal elements are 1s and the off-diagonal elements are 0s. It is a vector of all 1s. This is a centered matrix used to shift data to the origin. This is the inner product matrix, which reflects the inner product relationship between nodes. This is an eigenvalue matrix, where the eigenvalues ​​are arranged in descending order. The eigenvector matrix, This is a diagonal matrix composed of the two largest eigenvalues, whose eigenvalues ​​are arranged in descending order. These are the first two main eigenvector matrices. It is a two-dimensional coordinate matrix.

[0076] Finally, to ensure the consistency of the training process, the coordinates of node 0 are fixed at the origin, and the coordinates of node 1 are aligned to the vertical axis, thus converting the discrete line connection information into continuous two-dimensional coordinates:

[0077] ;

[0078] in, These are the original coordinates of node 0. The coordinate matrix after translation. Let the vertical and horizontal coordinates of node 1 be... For rotation angle, It is the arctangent function in the fourth quadrant. These are the aligned system node coordinates.

[0079] (2) Environmental quantity design for distribution network reconfiguration optimization;

[0080] First, based on the aforementioned two-dimensional node coordinates, the remaining power and voltage information are fused into a unified two-dimensional attribute state matrix. To ensure training stability, i.e., topological states of the same type should have similarity, the environmental state design method requires downsampling and interpolation within a fixed window range to form the attribute matrix. , , The selection of the fixed window range should cover the node coordinates of all system topology scenarios as much as possible. The calculation process of the downsampled attribute matrix coordinates is as follows:

[0081] ;

[0082] in, For the first The two-dimensional x and y coordinates of the node This represents the boundary range of the transformed coordinates, where l is the side length of the square window covering all coordinate points, and d is a redundant length used to avoid coordinate exceeding the limit in special cases. These represent the center coordinates and boundary coordinates of the square window, respectively. This indicates the number of sampling points on one side of the downsampling grid within the square window. Represents the coordinate sequence of grid points along the horizontal and vertical axes, using This represents the global grid point coordinate matrix.

[0083] Then, since the sampling points and coordinate points do not completely coincide, for sampling grid coordinate points without numerical representation, the point value is obtained by attribute interpolation under the original system node coordinates. The radial basis function interpolation process is as follows:

[0084] ;

[0085] in, For interpolation functions, Let be the radial basis function weight coefficients of the i-th node. For radial basis functions, a Gaussian function is used here. The shape parameter is used to control the decay rate of the function. Let be the Euclidean distance from point (x, y) to node . These are the coefficients of the terms in the linear polynomial. For the elements of the radial basis function matrix, , , These represent the polynomial condition matrix, the node attribute value vector, and the interpolation result of the grid points, respectively.

[0086] Finally, the reconstruction strategy proposed by the intelligent agent must satisfy the system's power flow constraints and the number of line switch operations. Simultaneously, the new network structure obtained by the system's line switch operations must strictly satisfy the radial topology constraint. The system's steady state is obtained by solving the power flow equations; the specific calculation formula is as follows:

[0087] ;

[0088] in, This represents the total number of times all switches operate within the scheduling cycle. Switching during the scheduling cycle Total number of actions; For the first Switch at any time Status, 0 indicates enabled, 1 indicates disabled; This represents the maximum number of all switching actions within the scheduling period. For switch The maximum number of times; , These represent the active power and reactive power injected into the nodes, respectively. , They are respectively Branch conductance and branch susceptance between nodes for At this moment Node voltage.

[0089] (3) Construction and updating of expert policy trajectory library and training of distribution network reconstruction agent based on state imitation reinforcement learning;

[0090] First, the samples in the expert strategy trajectory database are obtained by organizing historical decision data or by simulation using traditional optimization algorithms. The single-section solution result of the traditional optimization algorithm under constraints is an n-dimensional binary vector, where 0 represents on and 1 represents off. These 96 intraday topology decision vectors are concatenated at a 15-minute timescale to form the intraday real-time reconstruction strategy trajectory matrix for the distributed photovoltaic distribution network, and stored in the knowledge base. Each single-section decision vector under a strategy trajectory corresponds to... , , The three sets of attribute state matrices correspond to the optimal decision path. Therefore, the training objective of the agent is to guide the environment state to follow the state trajectory under the optimal decision, rather than directly learning how to solve the optimal decision. The goal is to quickly train a better distribution network reconstruction agent based on historical experience decision data.

[0091] Secondly, the agent is trained using a state imitation reinforcement learning algorithm, constructing two sets of neural network models: the action network model takes into account the current complete state matrix and the target state matrix, i.e., the state matrix at the current time point. , , Three sets of attribute state matrices output a decision vector that transitions the current environmental state to the target state. Since the active and reactive power state variables injected into the system cannot be changed through network reconstruction, the evaluation network model (forward dynamics model) only estimates the voltage state based on the input decision vector and state matrix. To effectively extract information from the state matrix, this invention uses a convolutional neural network as a pre-feature extraction layer, with the L2 norm of the deviation between the system's target voltage state matrix and the estimated voltage state used as the loss function to guide the training of the agent parameters. The calculation process is as follows:

[0092] ;

[0093] in, These are the internal network parameters of the model. This refers to the observation state at adjacent time steps, and here it generally refers to the state matrix of the three types of attributes. For the predicted state transition action, To evaluate the model's fitting loss and ensure that the evaluation model f can accurately predict real actions. The change in the following environmental state, For forward consistency loss, the L2 norm is used to measure the deviation between the transition state and the actual state. It is a standard action loss function, which is used to mimic decision-making actions in an expert policy trajectory library through labeled training.

[0094] Then, during actual decision-making, the state variable most similar to the current state will be matched from the state matrix in the expert policy trajectory library. The current state matrix and the state matrix of the next moment in the expert policy trajectory will be used as input to the agent action network model to generate the reconstructing policy vector of the distribution network topology in real time.

[0095] Finally, the model will extract real-time system strategy data, select high-quality samples to dynamically update the expert strategy trajectory library, enrich the state diversity in the expert strategy trajectory library, and use a comprehensive index system to calculate the economic and security benefits brought by the current system strategy trajectory. Once the value reaches a threshold, it is identified as a high-quality sample. The calculation formula is as follows:

[0096] ;

[0097] in, The reward function uses a comprehensive indicator system to quantify the merits of the current decision. This is the network loss coefficient. This refers to the system network loss after the action is performed. The economic coefficient for switch operation. This represents the number of switch actions.

[0098] To further evaluate the reconstruction performance of the proposed method in a 33-node system, traditional optimization algorithms were selected for comparison in terms of decision performance and computational efficiency. Table 1 shows the intraday topology decision given by the patented method.

[0099] Table 1. Intraday Topology Reconstruction Results

[0100]

[0101] As shown in the table, the reinforcement learning model successfully identified four different types of topologies, which are suitable for reconstruction at different time periods to meet system requirements. The neural network effectively captures the evolution of the network structure before and after each cycle by integrating switching states and system topology into the state and reward functions. During the reconstruction process, the model only operates a portion of the line switches, gradually transforming the system topology into the optimal configuration. This strategy significantly mitigates the negative impact of frequent switching operations on switching performance and system economic operation.

[0102] Figure 2The paper demonstrates the changing trend of system network losses after topology reconfiguration. Overall, network losses are lower during the day and increase at night. This is mainly because photovoltaic output is higher during the day, making energy storage regulation more significant. At this time, most of the electricity can be quickly consumed locally, while the load reaches its peak at night. Despite energy storage regulation, the overall load remains high, leading to a sharp increase in network losses. The algorithm proposed in this paper effectively reduces system network losses at different times by adjusting the system topology in a timely manner and rationally configuring transmission paths. The average network loss reduction reaches 23.8%, with the lowest reduction being 8.2%.

[0103] In environments with high photovoltaic penetration, voltage deviation becomes a key challenge for power distribution network optimization. Figure 3 The voltage distribution of each node under different topologies is shown. In the initial stage, the voltage values ​​are concentrated in the range of 0.95 to 1.02 per unit. After reconstruction, the voltage distribution of each node is closer to the rated voltage. The voltage fluctuation amplitude during the optimization period is significantly reduced, and the local voltage rise during the peak photovoltaic output is effectively suppressed, thereby improving the power quality of the distribution network. These subtle structural adjustments further enhance the system's ability to absorb photovoltaic power.

[0104] Table 2 Results of the intraday decision-making comparison experiment

[0105]

[0106] Comparing the data in Table 2 reveals that as the coupling coefficient of multi-time-period segments increases, the power system can more comprehensively consider the temporal variation characteristics of source-load interaction, thereby achieving a gradual reduction in overall cost. However, the MISOCP method only exhibits limited optimization performance when dealing with a small number of multi-time-period segments. As the problem size increases, the model's computation time increases significantly. In contrast, the decision agent built based on reinforcement learning is not limited by the size of the optimization problem and can provide an accurate and complete scheduling scheme within seconds, with its optimization cost and network loss only slightly higher than the optimal result of 24-hour time-period scheduling.

[0107] In summary, the method of this invention introduces state-mimicking reinforcement learning into the reconstruction optimization field of high-penetration distributed photovoltaic (PV) distribution networks. By making discrete topology information continuous, it forms an efficient and concise representation of the distribution network system state, making the model training more stable. Therefore, the optimization results of this invention can be used for grid dispatching and formulating safer and more reliable operating strategies for distributed PV access areas, thereby ensuring the safe and stable operation of the power system during large-scale distributed PV grid connection.

[0108] Example 2: This example provides a high-penetration distributed photovoltaic (PV) distribution network reconfiguration system for implementing the high-penetration distributed PV distribution network reconfiguration method described in any one of Examples 1, including:

[0109] The state space design module is used to design a continuous state space and generate multiple attribute state matrices based on the acquired system data of the distributed photovoltaic distribution network.

[0110] The expert strategy trajectory library module is used to generate strategy samples based on the attribute state matrix using the traditional distribution network topology optimization algorithm, and to select high-quality strategy trajectories to build an expert strategy trajectory library.

[0111] The agent training module is used to train an agent model based on the state imitation reinforcement learning algorithm and using the expert policy trajectory library; wherein, the agent model takes the current attribute state matrix and the target attribute state matrix as input and outputs the distribution network topology reconstruction policy vector.

[0112] The reconstruction strategy calculation module is used to input the acquired current distribution network operation status data into the pre-trained intelligent agent model to generate a real-time reconstruction strategy;

[0113] The dynamic update module is used to evaluate the real-time reconstruction strategy based on a comprehensive index system and dynamically update the expert strategy trajectory library used during the training of the intelligent agent model.

[0114] The specific functions of each module described above are explained in the relevant content of the method in Embodiment 1, and will not be repeated here.

[0115] Example 3: This example provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in Example 1.

[0116] Example 4: This example provides an electronic device, including:

[0117] Memory, used to store computer programs / instructions;

[0118] A processor for executing the computer program / instructions to implement the steps of any of the methods described in Embodiment 1.

[0119] Example 5: This example provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described in any one of Examples 1.

[0120] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

[0121] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0122] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure and not to limit its protection scope. Although this disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading this disclosure, they can still make various changes, modifications or equivalent substitutions to the specific implementation of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the pending claims.

Claims

1. A high-permeability distributed photovoltaic power distribution network reconstruction method, characterized in that, include: The acquired current distribution network operation status data is input into a pre-trained intelligent agent model to generate a real-time reconstruction strategy; The real-time reconstruction strategy is evaluated based on a comprehensive index system, and the expert strategy trajectory library used during the training of the intelligent agent model is dynamically updated. The training method for the intelligent agent model includes: Based on the acquired system data of the distributed photovoltaic power distribution network, a continuous state space is designed to generate multiple attribute state matrices; Based on the traditional distribution network topology optimization algorithm, the attribute state matrix is ​​used to generate strategy samples, and high-quality strategy trajectories are selected to build an expert strategy trajectory library. Based on the state imitation reinforcement learning algorithm, the agent model is trained using the aforementioned expert policy trajectory library; The intelligent agent model takes the current attribute state matrix and the target attribute state matrix as inputs and outputs the distribution network topology reconfiguration strategy vector.

2. The method of claim 1, wherein, The system data includes system network topology data, line impedance data, node voltage data, node injected active power data, and reactive power data.

3. The method of claim 1, wherein, The method for acquiring system data of the distributed photovoltaic power distribution network includes: collecting parameters of reactive power equipment, energy storage equipment, and transformers configured at each node in the system, as well as the voltage amplitude and voltage phase angle of each node.

4. The method of claim 3, wherein, The design of the continuous state space includes: using a multidimensional scaling analysis algorithm to transform the topological connection relationship of the distribution network system nodes into two-dimensional coordinates, wherein the Euclidean distance between the two-dimensional coordinates is as close as possible to the distance matrix calculated based on the line impedance modulus. The calculation process is as follows: ; Among them, matrix This is the distance matrix between system nodes. It is in the squared form of the distance matrix. for An identity matrix of order 1, where the diagonal elements are 1s and the off-diagonal elements are 0s. It is a vector of all 1s. This is a centered matrix used to shift data to the origin. It is an inner product matrix used to reflect the inner product relationship between nodes. This is an eigenvalue matrix, where the eigenvalues ​​are arranged in descending order. The eigenvector matrix, This is a diagonal matrix composed of the two largest eigenvalues, whose eigenvalues ​​are arranged in descending order. These are the first two main eigenvector matrices. It is a two-dimensional coordinate matrix.

5. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 4, characterized in that, The method further includes: fixing the coordinates of node 0 at the origin and aligning the coordinates of node 1 to the vertical axis to unify the coordinate system, as shown in the following expression: ; in, These are the original coordinates of node 0. The coordinate matrix after translation. Let the vertical and horizontal coordinates of node 1 be... For rotation angle, It is the arctangent function in the fourth quadrant. These are the aligned system node coordinates.

6. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 4, characterized in that, The multiple attribute state matrices are obtained by downsampling and interpolation within a fixed window range. These attribute matrices include a voltage attribute matrix, an active power attribute matrix, and a reactive power attribute matrix. The coordinate calculation process for these attribute matrices is as follows: ; in, For the first The two-dimensional x and y coordinates of the node This represents the boundary range of the transformed coordinates, where l is the side length of the square window covering all coordinate points, and d is a redundant length used to avoid coordinate exceeding the limit in special cases. These represent the center coordinates and boundary coordinates of the square window, respectively. This indicates the number of sampling points on one side of the downsampling grid within the square window. Represents the coordinate sequence of grid points along the horizontal and vertical axes, using This represents the global grid point coordinate matrix.

7. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 4, characterized in that, The downsampling interpolation uses the radial basis function interpolation method. For sampling grid coordinate points without numerical representation, interpolation is performed based on the attributes under the original system node coordinates to generate a continuous attribute state matrix. The radial basis function interpolation process is as follows: ; in, For interpolation functions, Let be the radial basis function weight coefficients of the i-th node. For radial basis functions, This is a shape parameter used to control the decay rate of the function. Let be the Euclidean distance from point (x, y) to node . These are the coefficients of the terms in the linear polynomial. For the elements of the radial basis function matrix, , , These represent the polynomial condition matrix, the node attribute value vector, and the interpolation result of the grid points, respectively.

8. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 1, characterized in that, The method for generating strategy samples using the attribute state matrix based on the traditional distribution network topology optimization algorithm includes: solving for the optimal reconfiguration strategy and generating a binary decision vector as the strategy sample, under the conditions of satisfying the system's power flow constraints, line switch operation frequency constraints, and radial topology constraints.

9. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 8, characterized in that, The strategy samples are spliced ​​together at a set time scale to form a real-time reconstructed strategy trajectory matrix within a day, and stored in the expert strategy trajectory library. Each strategy trajectory corresponds to three sets of attribute state matrices.

10. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 9, characterized in that, The state imitation reinforcement learning algorithm includes two sets of neural network models: an action network model and a judge network model; The action network model takes as input three complete attribute state matrices at the current moment and a target voltage attribute state matrix from the expert policy trajectory library, and outputs a decision vector that transfers the current environmental state to the target voltage state. The evaluation network model takes the decision vector and the complete attribute state matrix at the current moment as input, and outputs an estimated value of the system voltage attribute state matrix after the decision is executed.

11. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 10, characterized in that, The training of the intelligent agent model uses a convolutional neural network as the pre-feature extraction layer, and the L2 norm of the deviation between the system's target voltage state matrix and the estimated voltage state is used as the loss function to guide parameter optimization. The calculation process is as follows: ; in, These are the internal network parameters of the model. This refers to the observation state at adjacent time steps, and here it generally refers to the state matrix of the three types of attributes. For the predicted state transition action, To evaluate the model's fitting loss and ensure that the evaluation model f can accurately predict real actions. The change in the following environmental state, For forward consistency loss, the L2 norm is used to measure the deviation between the transition state and the actual state. It is a standard action loss function, which is used to mimic decision-making actions in an expert policy trajectory library through labeled training.

12. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 1, characterized in that, The step of inputting the acquired current distribution network operation status data into a pre-trained agent model to generate a real-time reconstruction strategy includes: matching the state quantity most similar to the current state from the expert strategy trajectory library, and using the current state matrix and the next-moment state matrix in the expert strategy trajectory as input to the agent model to generate a reconstruction strategy vector.

13. The high-penetration distributed photovoltaic distribution network reconfiguration method according to claim 1, characterized in that, The expert policy trajectory library used during the dynamic updating of the intelligent agent model training includes: calculating the economic and safety benefits of the current system policy trajectory using a comprehensive index system; when the value reaches a threshold, the current policy trajectory is identified as a high-quality sample and updated to the expert policy trajectory library; the calculation formula is as follows: ; in, The reward function uses a comprehensive indicator system to quantify the merits of the current decision. This is the network loss coefficient. This refers to the system network loss after the action is performed. The economic coefficient for switch operation. This represents the number of switch actions.

14. A high-penetration distributed photovoltaic distribution network reconfiguration system, used to implement the high-penetration distributed photovoltaic distribution network reconfiguration method according to any one of claims 1-13, characterized in that, include: The state space design module is used to design a continuous state space and generate multiple attribute state matrices based on the acquired system data of the distributed photovoltaic distribution network. The expert strategy trajectory library module is used to generate strategy samples based on the attribute state matrix using the traditional distribution network topology optimization algorithm, and to select high-quality strategy trajectories to build an expert strategy trajectory library. The agent training module is used to train an agent model based on the state imitation reinforcement learning algorithm and using the expert policy trajectory library; wherein, the agent model takes the current attribute state matrix and the target attribute state matrix as input and outputs the distribution network topology reconstruction policy vector. The reconstruction strategy calculation module is used to input the acquired current distribution network operation status data into the pre-trained intelligent agent model to generate a real-time reconstruction strategy; The dynamic update module is used to evaluate the real-time reconstruction strategy based on a comprehensive index system and dynamically update the expert strategy trajectory library used during the training of the intelligent agent model.