Dynamic networking fault protection method, device and equipment based on deep reinforcement learning
By applying a dynamic network fault protection method based on deep reinforcement learning in the distribution network, the problems of slow response speed, limited adjustment ability, insufficient intelligence and poor system coordination in the existing technology are solved, and rapid response and effective protection for sudden failures are achieved, and the intelligent and adaptive capabilities of the distribution network are improved.
Patent Information
- Application Number
- CN202510498610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems such as slow response speed, limited regulation capability, insufficient intelligence and poor system coordination in dynamic networking and fault protection of distribution networks, which are difficult to effectively prevent and control the spread of faults, affecting the reliability and stability of the power grid.
The dynamic networking fault protection method based on deep reinforcement learning is adopted. By obtaining the operating parameters of the distribution network, inputting it into the pre-trained dynamic networking decision model, and outputting the optimal networking adjustment decision to minimize fault expansion and distribution network losses. This model achieves dynamic adjustment of grid topology and rapid response to faults through training of multi-level protection strategies and experience playback pools.
It realizes rapid response and effective protection for sudden failures, improves the intelligence and adaptability of the distribution network, improves the operating efficiency and stability of the power grid, and reduces the risk of failure expansion.
Smart Images

Figure CN120016417A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a dynamic networking fault protection method, device and equipment based on deep reinforcement learning, belonging to the technical field of power transmission and distribution. Background Art
[0002] With the large-scale access of smart grids and renewable energy, the complexity and dynamics of distribution networks have increased significantly. Traditional distribution networks mainly rely on fixed topology structures and static fault protection mechanisms, such as circuit breakers and protection relays. These traditional methods have the defects of slow response speed and limited regulation ability when dealing with complex and changeable grid operation states. It is difficult to effectively prevent and control the spread of faults, which seriously affects the reliability and stability of the grid.
[0003] In the existing technology, dynamic networking technology optimizes the distribution of power flow by adjusting the topology of the distribution network in real time, thereby improving the operating efficiency and stability of the power grid. However, the existing dynamic networking methods are mostly based on heuristic algorithms or linear optimization models, which are difficult to handle the high nonlinearity and complexity of power grid operation. In addition, when faced with sudden failures, the decision-making speed and accuracy are insufficient, and it is impossible to achieve rapid response to failures and effective protection.
[0004] In addition, with the expansion of the scale and complexity of distribution networks, traditional fault protection strategies are difficult to meet the requirements of modern power grids for high reliability and high security. Existing protection mechanisms mostly rely on preset protection strategies and parameters, lack intelligence and adaptive capabilities, and cannot dynamically adjust to adapt to the real-time operating status of the power grid, resulting in unsatisfactory protection effects in complex situations with multiple faults and multiple variables, and may even cause large-scale power outages.
[0005] In summary, the existing technologies have the following major defects and deficiencies in dynamic networking and fault protection of distribution networks: Slow response speed: Traditional dynamic networking methods cannot achieve rapid response and real-time adjustment when faced with sudden failures, resulting in fault expansion and grid instability.
[0006] Limited regulation capability: Existing heuristic and linear optimization methods have difficulty in handling the high nonlinearity and complexity of the power grid, and the regulation effect is not ideal.
[0007] Lack of intelligence: Traditional fault protection strategies lack intelligence and adaptability, and are unable to dynamically adjust protection parameters and strategies based on real-time operating status.
[0008] Poor system coordination: Insufficient coordination and control mechanisms for multiple devices and multiple levels result in protection measures being unable to work together effectively in complex fault situations.
[0009] Therefore, there is an urgent need for a dynamic networking fault protection solution based on deep reinforcement learning, which can achieve rapid response and effective protection against sudden faults while ensuring the safe and stable operation of the power grid, and improve the intelligence and adaptability of the distribution network. Summary of the invention
[0010] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a dynamic networking fault protection method based on deep reinforcement learning.
[0011] In order to solve the above technical problems, the present invention is implemented by adopting the following technical solutions.
[0012] In a first aspect, the present invention discloses a dynamic networking fault protection method based on deep reinforcement learning, comprising: Obtain the operating parameters of the distribution network and input them into a pre-trained dynamic networking decision model based on deep reinforcement learning. With the purpose of minimizing the expansion of faults and the loss of distribution networks, the optimal networking adjustment decision is output. The parameters are input into a pre-trained dynamic networking decision model based on deep reinforcement learning. With the purpose of minimizing the expansion of faults and the loss of distribution networks, the optimal networking adjustment decision is output. The training process of the dynamic networking decision model based on deep reinforcement learning includes: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
[0013] Furthermore, the operating parameters of the distribution network include node voltage, branch current, branch active power, branch reactive power and switch status obtained after preprocessing.
[0014] Furthermore, the preprocessing includes: performing noise filtering, missing value processing, outlier detection and data normalization processing on the collected data in sequence.
[0015] Furthermore, the multi-level protection strategy based on nodes and lines dynamically networks the distribution network topology structure and adjusts the distribution network topology structure to the optimal distribution network topology structure, including: Collect fault events in the distribution network; Determine a fault location result according to the fault event, and generate an adjustment strategy for an optimal distribution network topology structure using a deep reinforcement learning model according to the fault location result; The distribution network topology structure is adjusted based on the adjustment strategy of the optimal distribution network topology structure.
[0016] Furthermore, after adjusting the topological structure of the distribution network, the power flow of the distribution network is recalculated, and the operation state of the power grid is optimized based on the power flow of the distribution network with the goal of ensuring the voltage stability of each node.
[0017] Furthermore, the use of batch samples to train the dynamic networking decision model based on deep reinforcement learning to obtain a trained dynamic networking decision model based on deep reinforcement learning includes: Step 1): The dynamic networking decision model based on deep reinforcement learning uses a deep deterministic policy gradient algorithm to initialize the policy network Actor weights , initialize the weight of the value network Critic , initialize the Actor weights of the target network , Critic's weight , , , initialize the experience replay pool D; Step 2): At each training step, based on the current policy network , select action , the corresponding formula is: ; In the formula, N t for noise; Step 3): Execute the action , observe the next state and rewards ,Will Stored in experience replay pool D; Step 4): Randomly extract batch samples from the experience replay pool D , where Respectively represent i Current state, action, reward, and next state in samples; Update the value network, the formula is: ; In the formula, y i is the updated value of the value network; c is the discount factor used to determine the current value of future rewards;Q The Q value output by the target Q network; Indicates in status According to the Actor weight of the target network Actions taken; Minimize the loss function, the formula is: ; In the formula, To minimize the value of the loss function, N is the total amount of training data; Update the policy network, the formula is: ; In the formula, is the updated value of the policy network; It is the gradient information of the Q function for the action under a specific strategy; The adjustment direction of the strategy is used to guide the strategy update to maximize the reward; Step 5): Update the target network parameters, the formula is: ; In the formula, is the soft update parameter; Step 6): Repeat steps 2) to 5) until the model converges and achieves the expected dynamic networking adjustment effect of the distribution network, and obtains a trained dynamic networking decision model based on deep reinforcement learning.
[0018] Furthermore, the calculation formula of the reward is: ; In the formula, R t represents the reward function, α is the weight coefficient of distribution network loss, β is the weight coefficient of the voltage deviation, or is the weight coefficient of the number of switching times, d is the weight coefficient of power supply reliability, L is the distribution network loss, △ V is the voltage deviation, C For the time period t The number of switching times of the internal switching device, W It is the quantitative value of power supply reliability; ; In the formula, R ij For line< i , j > resistance, I ij For line<i , j > current; ; In the formula, V set,i For Node i The set voltage value, V i ( t ) for t Time Node i Voltage value.
[0019] Furthermore, the method also includes optimizing the network adjustment decision, including the following steps: Regularly calculating the performance of the current dynamic networking decision model based on deep reinforcement learning according to the strategy evaluation index to determine the evaluation result, wherein the strategy evaluation index = {average loss reduction rate, average voltage deviation reduction, response time}; According to the evaluation results, adjust the weight coefficient in the reward function α , β , or , d , the learning direction of the optimization strategy; ; In the formula, Δ α , Δ β , Δ or , Δ d Adjust the step size for each parameter separately.
[0020] In a second aspect, the present invention discloses a dynamic networking fault protection device based on deep reinforcement learning, comprising: An acquisition module, used to obtain the operating parameters of the distribution network; The processing module is used to input the operating parameters of the distribution network into a pre-trained dynamic networking decision model based on deep reinforcement learning, so as to minimize the fault extension and distribution network loss, and output the optimal networking adjustment decision; The processing module comprises a training unit, wherein the training unit is used for: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
[0021] In a third aspect, the present invention discloses a computer device, comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method of the first aspect.
[0022] The beneficial effects achieved by the present invention are: First, the present invention adopts deep reinforcement learning technology to realize the intelligent and adaptive capabilities of dynamic networking strategies. Through the powerful learning ability of deep reinforcement learning technology, it can efficiently handle the complex nonlinear relationships in the operation of the distribution network, optimize networking decisions in real time, and improve the accuracy and response speed of fault protection.
[0023] Secondly, the proposed solution is highly real-time and flexible. Through the real-time monitoring and data acquisition system, the system can instantly obtain the operating status and fault information of the power grid, and combine the deep reinforcement learning model to quickly generate the optimal network adjustment plan, achieve rapid response and effective isolation of faults, prevent fault expansion, and ensure the stable operation of the power grid.
[0024] In addition, the present invention has good scalability and compatibility. The deep reinforcement learning model can be flexibly adjusted and expanded according to the scale and operation characteristics of the power grid, and can adapt to distribution networks of different scales and complexities. In addition, the dynamic networking control and execution module adopts a distributed architecture, which can be seamlessly integrated with existing power grid equipment and control systems, reducing system transformation costs and improving the overall compatibility of the system.
[0025] Finally, the present invention significantly improves the intelligence level and adaptive ability of the distribution network through intelligent dynamic networking and fault protection strategies. The system can automatically optimize the networking structure, quickly respond to faults, and ensure the safe and stable operation of the power grid in a complex and changeable power grid operation environment. It has broad application prospects and significant economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 is a system architecture diagram of the present invention; Figure 3 It is a specific workflow diagram of the dynamic networking decision model. DETAILED DESCRIPTION
[0027] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0028] Embodiment 1, as Figure 1 As shown, this embodiment introduces a dynamic networking fault protection method based on deep reinforcement learning, including: Obtain the operating parameters of the distribution network and input them into a pre-trained dynamic networking decision model based on deep reinforcement learning, with the goal of minimizing fault expansion and distribution network loss, and outputting the optimal networking adjustment decision; The training process of the dynamic networking decision model based on deep reinforcement learning includes: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
[0029] The operating parameters of the distribution network include node voltage, branch current, branch active power, branch reactive power and switch status obtained after preprocessing.
[0030] The preprocessing includes: performing noise filtering, missing value processing, outlier detection and data normalization processing on the collected data in sequence.
[0031] The multi-level protection strategy based on nodes and lines dynamically networks the distribution network topology structure and adjusts the distribution network topology structure to the optimal distribution network topology structure, including: Collect fault events in the distribution network; Determine a fault location result according to the fault event, and generate an adjustment strategy for an optimal distribution network topology structure using a deep reinforcement learning model according to the fault location result; The distribution network topology structure is adjusted based on the adjustment strategy of the optimal distribution network topology structure.
[0032] After adjusting the distribution network topology, the power flow of the distribution network is recalculated, and the grid operation state is optimized based on the power flow of the distribution network to ensure the voltage stability of each node.
[0033] The method of training the dynamic networking decision model based on deep reinforcement learning by using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning includes: Step 1): The dynamic networking decision model based on deep reinforcement learning uses a deep deterministic policy gradient algorithm to initialize the policy network Actor weights , initialize the weight of the value network Critic , initialize the Actor weights of the target network , Critic's weight , , , initialize the experience replay pool D; Step 2): At each training step, based on the current policy network , select action , the corresponding formula is: ; In the formula, N t for noise; Step 3): Execute the action , observe the next state and rewards ,Will Stored in experience replay pool D; Step 4): Randomly extract batch samples from the experience replay pool D , where Respectively represent i Current state, action, reward, and next state in samples; Update the value network, the formula is: ; In the formula, y i is the updated value of the value network; c is the discount factor used to determine the current value of future rewards; Q The Q value output by the target Q network; Indicates in status According to the Actor weight of the target network Actions taken; Minimize the loss function, the formula is: ; In the formula, To minimize the value of the loss function, N is the total amount of training data; Update the policy network, the formula is: ; In the formula, is the updated value of the policy network; It is the gradient information of the Q function for the action under a specific strategy; The adjustment direction of the strategy is used to guide the strategy update to maximize the reward; Step 5): Update the target network parameters, the formula is: ; In the formula, is the soft update parameter; Step 6): Repeat steps 2) to 5) until the model converges and achieves the expected dynamic networking adjustment effect of the distribution network, and obtains a trained dynamic networking decision model based on deep reinforcement learning.
[0034] The calculation formula for the reward is: ; In the formula, R t represents the reward function, α is the weight coefficient of distribution network loss, β is the weight coefficient of the voltage deviation, or is the weight coefficient of the number of switching times, d is the weight coefficient of power supply reliability, L is the distribution network loss, △ V is the voltage deviation, C For the time period t The number of switching times of the internal switching device, W is the quantitative value of power supply reliability (the proportion of loads that can still be powered normally under fault conditions); ; In the formula, R ij For line< i , j > resistance, I ij For line< i , j > current; ; In the formula, V set,i For Node i The set voltage value, V i ( t ) for t Time Node i Voltage value.
[0035] It also includes network adjustment decision optimization, including the following steps: Regularly calculating the performance of the current dynamic networking decision model based on deep reinforcement learning according to the strategy evaluation index to determine the evaluation result, wherein the strategy evaluation index = {average loss reduction rate, average voltage deviation reduction, response time}; According to the evaluation results, adjust the weight coefficient in the reward function α , β , or , d , the learning direction of the optimization strategy; ; In the formula, Δ α , Δ β , Δ or , Δ d Adjust the step size for each parameter separately.
[0036] Example 2, based on the same inventive concept as Example 1, this example introduces a dynamic networking fault protection method based on deep reinforcement learning, such as Figure 2 As shown, the hardware part of the system includes: intelligent sensors and filters deployed on key nodes and lines to monitor node and line parameters, a data acquisition system that receives data from sensors and performs preprocessing, a deep reinforcement learning model that analyzes data, evaluates status and generates decisions, a dynamic networking control module that receives decision-making instructions, intelligent switching devices that execute control instructions, and dynamic networking that achieves topology adjustment and fault isolation through dynamic networking.
[0037] like Figure 3 As shown, the specific working process is: Step 1: Initialize system parameters; Step 11: Define basic parameters of distribution network; First, set the basic parameters of the distribution network, including the upper and lower limits of node voltage, line capacity, maximum switching times of switchgear, and time delay. The specific parameter settings are as follows: The upper and lower limits of node voltage are set to 0.95pu to 1.05pu, that is, the voltage amplitude Vi of each node satisfies: ; Line capacity: Each line<i,j> The maximum active power transfer P max,ij and reactive power transfer Q max,ij They are defined as: ; Switching device parameters: The maximum switching times of each switching device is set to 10 times / hour, or set according to actual conditions, and the time delay does not exceed 1 second.
[0038] Step 12: Initialize deep reinforcement learning model parameters Neural network structure: Defines the neural network structure of the deep reinforcement learning model, including the number of nodes in the input layer, hidden layer, and output layer. Assume that the input layer includes n nodes, the hidden layer includes m neurons, and the output layer contains a action nodes.
[0039] Learning rate: Set the learning rate α ,For example α =0.001; Discount Factor: Set the discount factor c ,For example c =0.99; Exploration Rate: Set the exploration rate e , initial value e =1.0, gradually decrease to e min =0.01; Experience replay pool size: set to 100,000 experiences; The above parameters can be adjusted according to actual conditions.
[0040] Step 2: Real-time monitoring and data collection.
[0041] Step 21: Sensor deployment. Smart sensors are deployed at key nodes and lines of the distribution network to ensure comprehensive monitoring of the grid operation status. Sensor types include voltage sensors, current sensors, power factor sensors, etc.
[0042] Step 22: Data transmission. The real-time collected voltage, current, load and other operating data are transmitted to the central data processing center through a high-speed communication network (such as optical fiber, wireless communication, etc.). The data transmission adopts an encryption protocol to ensure the security and real-time nature of the data.
[0043] Step 23: Data preprocessing. Preprocess the collected data, including data cleaning, outlier detection and data normalization. Data cleaning: Remove noise and erroneous data generated during the collection process. The specific steps include: 1. Noise filtering: Use low-pass filter to remove high-frequency noise and retain the main signal of power grid operation; ; In the formula, X org ( t ) is the original signal, h ( t-t ) is the impulse response of the filter, X ( t ) is the signal after filtering.
[0044] , Missing value processing: interpolation processing is performed on missing data points. The commonly used method is linear interpolation: ; In the formula, X ( t ) is a missing value, t 1. t 2 are the time points before and after the missing value, X ( t 1) X ( t 2) are the data values at the corresponding time points.
[0045] Outlier detection: Detect and process outliers using statistical methods (such as Z-score), which is specifically defined as: ; Where X is the data point, is the mean, is the standard deviation, when When it is >3, it is considered an outlier and processed, for example, it is identified as a missing value and filled with linear interpolation or directly deleted.
[0046] Data normalization: Normalize the data to the range of [0,1] to improve the efficiency and stability of model training: ; In the formula, X min and X max are the minimum and maximum values in the data set, respectively.
[0047] Step 3: Construction and training of deep reinforcement learning model.
[0048] Step 31: Define the state space. The state space S t Contains all key operating parameters of the distribution network at time t, specifically defined as: ; Where V i (t) is the voltage of node i at time t, I l (t), P l (t), Q l (t) are the current, active power and reactive power of branch l at time t, Switch state (t) is the switch status at time t.
[0049] Step 32: Define the action space. Action space A t Contains all possible network adjustment operations of the distribution network at time t, which is specifically defined as: ; Among them, the Switch operation indicates whether a certain switch device needs to switch state (on or off); the load transfer instruction indicates that a certain load needs to be transferred from one node to another. The specific operation includes realizing power transmission from one node of the SOP to another node through the smart soft switch (SOP).
[0050] Step 33: Design a reward function. Reward function R t Taking into account multiple objectives of power grid operation, the specific design is as follows: ; In the formula, α is the weight coefficient of power grid loss, β is the weight coefficient of voltage deviation, or is the weight coefficient of the number of switching times, d is the weight coefficient of power supply reliability, which is specifically defined as: Grid losses: ; In the formula, R ij For line<i,j> resistance.
[0051] Voltage deviation: ; Where V set,i is the set voltage value of node i.
[0052] Switching times: The number of switching times of the switching device within the time period t.
[0053] Power supply reliability: It is based on the assessment of the continuity and stability of power supply of the power grid within a period of time t, which can be quantified by evaluating the number and duration of power outages.
[0054] Step 34: Select a deep reinforcement learning algorithm. The present invention adopts the Deep Deterministic Policy Gradient (DDPG) algorithm, which is suitable for continuous action space and high-dimensional state space, and can better meet the complexity and real-time requirements of dynamic networking of distribution networks. The DDPG algorithm combines deep neural networks and policy gradient methods to achieve policy optimization and value evaluation through the Actor-Critic architecture.
[0055] Step 35: Train the deep reinforcement learning model. Through multiple rounds of interaction with the distribution network simulation environment, the model continuously learns and optimizes the network adjustment strategy. The specific training steps include: Step 351: Initialize model parameters. Randomly initialize the policy network Actor weights , initialize the weight of the value network Critic , initialize the Actor weights of the target network , Critic's weight , , , initialize the experience replay pool D; Step 352: Sampling and Exploration. At each training step, based on the current policy network m(s|q), select an action: ; Where N t It is noise, which is used for strategy exploration and increases the exploratory nature of the strategy.
[0056] Step 353: Interact with the environment. Perform actions , observe the next state and rewards .Will Stored in experience replay pool D.
[0057] Step 354: Experience replay and gradient descent. Randomly extract small batches of samples from the experience replay pool D: .
[0058] Update value network: ; Minimize the loss function: ; Update policy network: ; Step 355: Target network soft update; Update the target network parameters: ; In the formula, t is the soft update parameter, usually t <<1, e.g. t =0.001.
[0059] Step 356: Iterative training. Repeat steps 352 to 355 until the model converges and the expected dynamic network adjustment effect of the distribution network is achieved.
[0060] Step 4: Dynamic networking decision and execution; Step 41: Decision generation. The deep reinforcement learning model receives the current state S t After that, the optimal networking adjustment action A is generated through the policy network t , the specific steps are as follows: Step 411: Input state. The current distribution network operation state St Input policy network m(s|q).
[0061] Step 412: Output action. The policy network outputs continuous actions , including the switching instructions of each switchgear and the load transfer instructions of SOP: ; Step 42: Send the control command. Converted into specific control instructions and sent to the intelligent switch devices in the distribution network through the communication network. The control instructions include: Switch switching command: indicates whether a switch device needs to switch state (on / off); Load transfer instruction: indicates that a load node needs to be transferred from one node to another node through a smart soft switch (SOP) to achieve power transfer.
[0062] Step 43: Implementation of the network adjustment strategy. After receiving the control command, the intelligent switch device performs the corresponding switch operation, dynamically adjusts the topology of the distribution network, and realizes the optimal distribution of power flow and isolation of fault areas. The specific implementation process is as follows: Step 431: Switch operation. According to the control instruction, the switch device performs an on or off operation to adjust the connection relationship of the power grid; ; In the formula, Switch state,i ( t +1) for switch i At the moment t +1 for the status.
[0063] Step 432: Load transfer operation: Use SOP to realize load transfer.
[0064] Step 433: Recalculate the power flow distribution. After adjusting the topology, recalculate the power flow distribution using the power flow calculation method to ensure that the voltage of each node is stable. The calculation method can use the commonly used Newton-Raphson method.
[0065] Step 5: Fault detection and diagnosis; Step 51: Fault detection. Use smart sensors and monitoring equipment to detect fault events in the distribution network in real time, such as short circuits, overloads, etc. Specific detection methods include: Step 511: Voltage anomaly detection. When the voltage of a node suddenly drops or rises beyond a set threshold, it is determined to be a voltage anomaly: ; In the formula, is the voltage deviation threshold.
[0066] Step 512: Current abnormality detection. When the current of a certain line exceeds its maximum capacity, it is determined to be a current abnormality: ; In the formula, I max,ij For branch<i,j> Maximum current capacity.
[0067] Step 52: Fault isolation. Based on the fault location results, use the network adjustment strategy generated by the deep reinforcement learning model (the model training process can be performed using a ready-made Python library) to quickly isolate the fault area. The specific steps are as follows: Switch disconnection: According to the fault location, determine the switch device that needs to be disconnected, perform the disconnection operation, and isolate the fault area.
[0068] Load transfer: Transfer the load in the area affected by the fault to other safe areas, or transmit power from other areas through SOP to supply important loads to ensure the continuity and stability of power supply.
[0069] Power flow optimization: After adjusting the topology, recalculate the power flow, optimize the grid operation status, and ensure the voltage stability of each node.
[0070] Step 6: Model optimization and self-learning.
[0071] Step 61: Data feedback. Feedback the actual operation data and network adjustment effects to the deep reinforcement learning model as new training samples. Specifically including: Network adjustment effect evaluation: Evaluate the operation status of the power grid after network adjustment, such as voltage stability, loss changes, etc. Effect evaluation index = {voltage stability score, loss reduction, power supply reliability index}; Feedback data storage: the current state S t 、Action A t , Reward R t , the next state S t+1 The training data is stored in the experience replay pool D: ; Step 62: Model update. Use feedback data to continuously optimize and self-learn the deep reinforcement learning model. The specific steps are as follows: Step 621: Batch training. Randomly extract small batches of samples from the experience replay pool D, perform model training, and update the parameters of the policy network and the value network.
[0072] Randomly draw batches of samples from D ; Step 622: Online learning. During the actual operation, the model continuously receives new experience data and performs online training to improve the model's decision-making ability and adaptability.
[0073] Step 63: Strategy optimization. Through continuous training and optimization, the model can continuously improve the dynamic networking adjustment strategy. The specific steps are as follows: Step 631: Strategy evaluation. Regularly evaluate the performance of the current strategy and analyze the performance of the strategy under fault conditions; Strategy evaluation index = {average loss reduction rate, average voltage deviation reduction, response time}.
[0074] Step 632: Strategy adjustment. Adjust the weight coefficient in the reward function according to the evaluation results. α , β , or , d , the learning direction of the optimization strategy; ; In the formula, Δ α , Δ β , Δ or , Δ d Adjust the step size for each parameter separately.
[0075] Embodiment 3 is based on the same inventive concept as Embodiment 1. This embodiment introduces a dynamic networking fault protection device based on deep reinforcement learning, including: An acquisition module, used to obtain the operating parameters of the distribution network; The processing module is used to input the operating parameters of the distribution network into a pre-trained dynamic networking decision model based on deep reinforcement learning, so as to minimize the fault extension and distribution network loss, and output the optimal networking adjustment decision; The processing module comprises a training unit, wherein the training unit is used for: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
[0076] Example 4 is based on the same inventive concept as Example 1. This example introduces a computer device, including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method described in Example 1.
[0077] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0078] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure one A process or multiple processes and / or boxes Figure one A device that provides the functions specified in a block or multiple blocks.
[0079] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure one A process or multiple processes and / or boxes Figure one A function specified in one or more boxes.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure one A process or multiple processes and / or boxes Figure one A step that specifies a function in one or more boxes.
[0081] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A dynamic networking fault protection method based on deep reinforcement learning, characterized in that: include: Obtain the operating parameters of the distribution network and input them into a pre-trained dynamic networking decision model based on deep reinforcement learning. With the purpose of minimizing the expansion of faults and the loss of distribution networks, the optimal networking adjustment decision is output. The parameters are input into a pre-trained dynamic networking decision model based on deep reinforcement learning. With the purpose of minimizing the expansion of faults and the loss of distribution networks, the optimal networking adjustment decision is output. The training process of the dynamic networking decision model based on deep reinforcement learning includes: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
2. The dynamic networking fault protection method based on deep reinforcement learning according to claim 1 is characterized in that: The operating parameters of the distribution network include node voltage, branch current, branch active power, branch reactive power and switch status obtained after preprocessing.
3. The dynamic networking fault protection method based on deep reinforcement learning according to claim 2 is characterized in that: The preprocessing includes: performing noise filtering, missing value processing, outlier detection and data normalization processing on the collected data in sequence.
4. The dynamic networking fault protection method based on deep reinforcement learning according to claim 1 is characterized in that: The multi-level protection strategy based on nodes and lines dynamically networks the distribution network topology structure and adjusts the distribution network topology structure to the optimal distribution network topology structure, including: Collect fault events in the distribution network; Determine a fault location result according to the fault event, and generate an adjustment strategy for an optimal distribution network topology structure using a deep reinforcement learning model according to the fault location result; The distribution network topology structure is adjusted based on the adjustment strategy of the optimal distribution network topology structure.
5. The dynamic networking fault protection method based on deep reinforcement learning according to claim 4 is characterized in that: After adjusting the distribution network topology, the power flow of the distribution network is recalculated, and the grid operation state is optimized based on the power flow of the distribution network to ensure the voltage stability of each node.
6. The dynamic networking fault protection method based on deep reinforcement learning according to claim 1, characterized in that: The method of training the dynamic networking decision model based on deep reinforcement learning by using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning includes: Step 1): The dynamic networking decision model based on deep reinforcement learning uses a deep deterministic policy gradient algorithm to initialize the policy network Actor weights , initialize the weight of the value network Critic , initialize the Actor weights of the target network , Critic's weight , , , initialize the experience replay pool D; Step 2): At each training step, based on the current policy network , select action , the corresponding formula is: ; In the formula, N t for noise; Step 3): Execute the action , observe the next state and rewards ,Will Stored in experience replay pool D; Step 4): Randomly extract batch samples from the experience replay pool D , where Respectively represent i Current state, action, reward, and next state in samples; Update the value network, the formula is: ; In the formula, y i is the updated value of the value network; γ is the discount factor used to determine the current value of future rewards; Q The Q value output by the target Q network; Indicates in status According to the Actor weight of the target network Actions taken; Minimize the loss function, the formula is: ; In the formula, To minimize the value of the loss function, N is the total amount of training data; Update the policy network, the formula is: ; In the formula, is the updated value of the policy network; It is the gradient information of the Q function for the action under a specific strategy; The adjustment direction of the strategy is used to guide the strategy update to maximize the reward; Step 5): Update the target network parameters, the formula is: ; In the formula, is the soft update parameter; Step 6): Repeat steps 2) to 5) until the model converges and achieves the expected dynamic networking adjustment effect of the distribution network, and obtains a trained dynamic networking decision model based on deep reinforcement learning.
7. The dynamic networking fault protection method based on deep reinforcement learning according to claim 6 is characterized in that: The calculation formula for the reward is: ; In the formula, R t represents the reward function, α is the weight coefficient of distribution network loss, β is the weight coefficient of the voltage deviation, η is the weight coefficient of the number of switching times, δ is the weight coefficient of power supply reliability, L is the distribution network loss, △ V is the voltage deviation, C For the time period t The number of switching times of the internal switching device, W It is the quantitative value of power supply reliability; ; In the formula, R ij For line< i , j > resistance, I ij For line< i , j > current; ; In the formula, V set,i For Node i The set voltage value, V i ( t ) for t Time Node i Voltage value.
8. The dynamic networking fault protection method based on deep reinforcement learning according to claim 7 is characterized in that: It also includes network adjustment decision optimization, including the following steps: Regularly calculating the performance of the current dynamic networking decision model based on deep reinforcement learning according to the strategy evaluation index to determine the evaluation result, wherein the strategy evaluation index = {average loss reduction rate, average voltage deviation reduction, response time}; According to the evaluation results, adjust the weight coefficient in the reward function α , β , η , δ , the learning direction of the optimization strategy; ; In the formula, Δ α , Δ β , Δ η , Δ δ Adjust the step size for each parameter separately.
9. A dynamic networking fault protection device based on deep reinforcement learning, characterized in that: include: An acquisition module, used to obtain the operating parameters of the distribution network; The processing module is used to input the operating parameters of the distribution network into a pre-trained dynamic networking decision model based on deep reinforcement learning, so as to minimize the fault extension and distribution network loss, and output the optimal networking adjustment decision; The processing module comprises a training unit, wherein the training unit is used for: Based on the multi-level protection strategy of nodes and lines, the distribution network topology is dynamically networked, the distribution network topology is adjusted to the optimal distribution network topology, and a dynamic networking decision model based on deep reinforcement learning is constructed according to the optimal distribution network topology and deep reinforcement learning algorithm; Randomly extract batch samples from a pre-built experience replay pool, each of which includes the state corresponding to the current networking adjustment decision of the distribution network, the execution action after the current networking adjustment decision interacts with the environment, the reward after the current networking adjustment decision interacts with the environment, and the next state after the current networking adjustment decision interacts with the environment; The dynamic networking decision model based on deep reinforcement learning is trained using batch samples to obtain a trained dynamic networking decision model based on deep reinforcement learning.
10. A computer device, characterized in that: include, One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described in claims 1 to 8.
Citation Information
Patent Citations
Recommendation model determination method and device and computer readable storage medium
CN116186376A
Fault processing method, device and system based on intelligent switch and medium
CN116488169A
Wind power optimization control system and method based on multi-target coupling enhancement
CN119084222A
Cited By
Power distribution area topology optimization generation method based on deep reinforcement learning
CN120414524A
A power distribution district topology optimization generation method based on deep reinforcement learning
CN120414524B