Power distribution network fault recovery method and system based on reinforcement learning
By modeling the power distribution network as a weighted directed graph model and combining quantum-classical hybrid architecture and quantum reinforcement learning algorithms, the problem that traditional fault recovery methods are difficult to cope with complex dynamic characteristics and uncertainties is solved, and efficient and reliable recovery of power distribution network failures is achieved.
Patent Information
- Application Number
- CN202411667138.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The traditional power distribution network fault recovery method is based on fixed rules and manual experience, and it is difficult to cope with the complex dynamic characteristics and uncertainties of the power distribution network, resulting in the inflexible and adaptable fault recovery.
Using a reinforcement learning method, the power distribution network is modeled as a weighted directed graph model, and a data transmission protocol and result reception analysis method of quantum-classical hybrid architecture are designed, combined with quantum reinforcement learning algorithms and strategy optimization, the power distribution network fault recovery is achieved.
It has achieved rapid, accurate and reliable recovery of power distribution network failures, and can effectively deal with the complexity and uncertainty brought about by the diversified development of power loads.
Smart Images

Figure CN119171440B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for power distribution network fault recovery based on reinforcement learning. Background Art
[0002] In the field of power distribution network fault restoration, traditional methods are mainly based on fixed rules and manual experience. First, by modeling the power distribution network, a simple circuit model is usually used to represent the topological structure and component characteristics of the distribution network. In this model, the distribution network is regarded as a linear circuit system composed of nodes, lines, and power supplies, and the connection relationship between nodes and line parameters are clearly determined. For fault detection, relay protection devices such as overcurrent protection and distance protection are usually used. These protection devices determine whether a fault occurs based on a preset current or voltage threshold. Once a fault is detected, the fault area is isolated by disconnecting a specific switch according to a pre-established fault isolation strategy. In the fault recovery stage, based on manually set rules and experience, appropriate recovery paths and operating switches are selected. For example, the power supply of important loads is restored first, or the power supply is gradually restored in a preset priority order. In the recovery process, only some simple factors, such as line capacity and voltage constraints, are usually considered, while the complex dynamic characteristics and uncertainties of the distribution network are ignored.
[0003] However, in the face of increasingly complex power systems, this traditional power distribution network fault recovery method based on fixed rules and manual experience ignores the complex dynamic characteristics and uncertainties of the distribution network, lacks flexibility and adaptability, and is unable to cope with the complexity and uncertainty of the distribution network brought about by the diversified development of power loads in the power distribution network, making it difficult to achieve efficient and reliable recovery of power distribution network faults. Summary of the invention
[0004] The embodiments of the present invention provide a power distribution network fault recovery method and system based on reinforcement learning, which can effectively cope with the complexity and uncertainty of the distribution network brought about by the diversified development of power loads in the power distribution network, thereby achieving efficient and reliable recovery of power distribution network faults.
[0005] An embodiment of the present invention provides a method for fault recovery of a power distribution network based on reinforcement learning, comprising the following steps:
[0006] Modeling the power distribution network, modeling the power distribution network as a weighted directed graph model;
[0007] Based on the weighted directed graph model of the power distribution network, a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery are designed; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation results of the quantum computing part are converted into control instructions or decision results; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations;
[0008] Based on the weighted directed graph model of the power distribution network, a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery are designed, and the training data of the power distribution network is used for training. Among them, the state and action in reinforcement learning are represented by quantum state, and the strategy network and value network of the quantum reinforcement learning algorithm are constructed and updated by using quantum gate operations, and the quantum reward function is designed.
[0009] When a fault in the power distribution network is detected, the state data of the power distribution network after the fault is obtained according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and the state data of the power distribution network after the fault is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part;
[0010] The quantum computing part uses the designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy of the power distribution network and feeds it back to the classical computing part. The classical computing part then converts the optimal fault recovery strategy into actual control instructions to execute the fault recovery work of the power distribution network.
[0011] Another embodiment of the present invention provides a power distribution network fault recovery system based on reinforcement learning, including:
[0012] A modeling module, used for modeling the power distribution network, and modeling the power distribution network as a weighted directed graph model;
[0013] The first design module is used to design a data transmission protocol and result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery based on a weighted directed graph model of the power distribution network; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation results of the quantum computing part are converted into control instructions or decision results; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations;
[0014] The second design module is used to design a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery based on a weighted directed graph model of the power distribution network, and use the training data of the power distribution network for training; in which, the state and action in reinforcement learning are represented by quantum state, and quantum gate operations are used to realize the construction and update of the strategy network and value network of the quantum reinforcement learning algorithm, and the quantum reward function is designed;
[0015] A computing module, which is used for, when a fault in the power distribution network is detected, obtaining the state data of the power distribution network after the fault according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and encoding the state data of the power distribution network after the fault in the power distribution network into the probability amplitude of quantum bits and transmitting it to the quantum computing part;
[0016] The fault recovery module is used to calculate the optimal fault recovery strategy of the power distribution network through the quantum computing part using the designed quantum reinforcement learning algorithm and feed it back to the classical computing part, and convert the optimal fault recovery strategy into actual control instructions through the classical computing part to perform fault recovery work of the power distribution network.
[0017] Another embodiment of the present invention provides a power distribution network fault recovery system based on reinforcement learning, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the power distribution network fault recovery method based on reinforcement learning described in the above-mentioned embodiment of the invention.
[0018] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0019] First, by modeling the power distribution network as a weighted directed graph model, an accurate distribution network structure and parameter basis is provided for the entire fault recovery method. On this basis, the quantum-classical hybrid architecture makes the advantages of classical computing and quantum computing complement each other. The classical computing part can efficiently process and encode the distribution network state data, and convert the quantum computing results into actual control instructions or decision results; the quantum computing part uses the mapping relationship between quantum bits and distribution network elements to accurately represent the distribution network state with quantum states, and simulates the evolution of the distribution network state through quantum gate operations, including fault propagation and equipment actions, so as to more accurately analyze the dynamic changes of the distribution network. At the same time, the quantum reinforcement learning algorithm and strategy optimization designed based on the weighted directed graph model use training data for training, use quantum states to represent the states and actions in reinforcement learning, build and update the strategy network and value network through quantum gate operations, and design quantum reward functions, so that the algorithm can learn the optimal fault recovery strategy. When a fault occurs in the distribution network, the classical computing part quickly and accurately obtains the distribution network state data after the fault based on the weighted directed graph model, and encodes it into the probability amplitude of quantum bits and transmits it to the quantum computing part. The quantum computing part uses the trained quantum reinforcement learning algorithm to quickly calculate the optimal fault recovery strategy and feeds it back to the classical computing part. The classical computing part converts it into actual control instructions to achieve precise operation of the distribution network equipment, thereby quickly and effectively recovering the power distribution network fault. From the above analysis, it can be seen that the embodiment of the present invention is based on the weighted directed graph model of the power distribution network, constructs a quantum-classical hybrid architecture, and realizes distribution network fault recovery through quantum reinforcement learning algorithm and strategy optimization. In this way, the optimal fault recovery strategy of the power distribution network can be determined quickly and accurately, and the complexity and uncertainty of the distribution network brought about by the diversified development of the power load of the power distribution network can be effectively dealt with, thereby achieving efficient and reliable recovery of the power distribution network fault. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of a method for fault recovery of a power distribution network based on reinforcement learning provided by an embodiment of the present invention;
[0021] Figure 2 It is a structural schematic diagram of a power distribution network fault recovery system based on reinforcement learning provided by an embodiment of the present invention;
[0022] Figure 3 It is a structural diagram of a power distribution network fault recovery system based on reinforcement learning provided by another embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] See also Figure 1 , is a flow chart of a method for restoring a power distribution network fault based on reinforcement learning provided by an embodiment of the present invention. The method for restoring a power distribution network fault based on reinforcement learning comprises the following steps:
[0025] S10, modeling the power distribution network, and modeling the power distribution network as a weighted directed graph model;
[0026] S11, based on the weighted directed graph model of the power distribution network, a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery are designed; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation results of the quantum computing part are converted into control instructions or decision results; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations;
[0027] S12, based on the weighted directed graph model of the power distribution network, design a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery, and use the training data of the power distribution network for training; in which, the quantum state is used to represent the state and action in reinforcement learning, and the quantum gate operation is used to realize the construction and update of the strategy network and value network of the quantum reinforcement learning algorithm, and the quantum reward function is designed;
[0028] S13, when a fault in the power distribution network is detected, the state data of the power distribution network after the fault is obtained according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and the state data of the power distribution network after the fault is encoded into a probability amplitude of a quantum bit and transmitted to the quantum computing part;
[0029] S14, the quantum computing part uses the designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy of the power distribution network and feeds it back to the classical computing part, and the classical computing part converts the optimal fault recovery strategy into actual control instructions to execute the fault recovery work of the power distribution network.
[0030] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0031] First, by modeling the power distribution network as a weighted directed graph model, an accurate distribution network structure and parameter basis is provided for the entire fault recovery method. On this basis, the quantum-classical hybrid architecture makes the advantages of classical computing and quantum computing complement each other. The classical computing part can efficiently process and encode the distribution network state data, and convert the quantum computing results into actual control instructions or decision results; the quantum computing part uses the mapping relationship between quantum bits and distribution network elements to accurately represent the distribution network state with quantum states, and simulates the evolution of the distribution network state through quantum gate operations, including fault propagation and equipment actions, so as to more accurately analyze the dynamic changes of the distribution network. At the same time, the quantum reinforcement learning algorithm and strategy optimization designed based on the weighted directed graph model use training data for training, use quantum states to represent the states and actions in reinforcement learning, build and update the strategy network and value network through quantum gate operations, and design quantum reward functions, so that the algorithm can learn the optimal fault recovery strategy. When a fault occurs in the distribution network, the classical computing part quickly and accurately obtains the distribution network state data after the fault based on the weighted directed graph model, and encodes it into the probability amplitude of quantum bits and transmits it to the quantum computing part. The quantum computing part uses the trained quantum reinforcement learning algorithm to quickly calculate the optimal fault recovery strategy and feeds it back to the classical computing part. The classical computing part converts it into actual control instructions to achieve precise operation of the distribution network equipment, thereby quickly and effectively recovering the power distribution network fault. From the above analysis, it can be seen that the embodiment of the present invention is based on the weighted directed graph model of the power distribution network, constructs a quantum-classical hybrid architecture, and realizes distribution network fault recovery through quantum reinforcement learning algorithm and strategy optimization. In this way, the optimal fault recovery strategy of the power distribution network can be determined quickly and accurately, and the complexity and uncertainty of the distribution network brought about by the diversified development of the power load of the power distribution network can be effectively dealt with, thereby achieving efficient and reliable recovery of the power distribution network fault.
[0032] It can be understood that in step S10, by modeling the distribution network as a weighted directed graph, the relationship and parameters of the nodes, lines and other elements in the distribution network are clarified, providing a basic framework for subsequent data processing, state representation and strategy formulation. Only by establishing an accurate distribution network model can the various states and behaviors of the distribution network be effectively analyzed and processed. In step S11, in the classical computing part, the distribution network state data is encoded as the probability amplitude of the quantum bit and transmitted to the quantum computing part, so that the quantum computing can process the distribution network information. At the same time, the classical computing part is also responsible for receiving the results of the quantum computing part and converting them into control instructions or decision results. This two-way data interaction enables the powerful computing power of quantum computing to be combined with the actual control capabilities of classical computing. The quantum computing part establishes a mapping relationship between quantum bits and distribution network elements, represents the distribution network state with quantum states, and simulates the evolution of the distribution network state through quantum gate operations, including fault propagation and equipment action. This enables quantum computing to model and analyze the complex dynamics of the distribution network at the quantum state level, providing a basis for subsequent reinforcement learning algorithms. In step S12, the state and action in reinforcement learning are represented by quantum state, the policy network and value network are constructed and updated by quantum gate operation, and the quantum reward function is designed, which is the core of realizing intelligent fault recovery strategy. The representation of quantum state can describe the complex state of the distribution network more comprehensively and accurately, and the quantum gate operation provides an efficient way to update the strategy and value. By training with training data, the quantum reinforcement learning algorithm can learn the optimal strategy under different distribution network states, so that decisions can be made quickly and accurately when facing actual faults. In step S13, this step realizes that when an actual fault occurs, the real-time distribution network state information is transmitted to the quantum computing part. Based on the previously established model and the trained quantum reinforcement learning algorithm, the quantum computing part can analyze and calculate according to the current fault state. In step S14, the quantum computing part calculates the optimal fault recovery strategy based on the received fault state data using the quantum reinforcement learning algorithm, and feeds it back to the classical computing part. The classical computing part converts the strategy into actual control instructions, realizes the operation of the distribution network equipment, and completes the fault recovery work.
[0033] In addition, the specific explanation of the application of the power distribution network modeling as a weighted directed graph model in step S10 in other steps is as follows:
[0034] In step S11, the classical computing part: when the distribution network state data processed by the classical computing is encoded into the probability amplitude of the quantum bit and transmitted to the quantum computing part, the modeling of the weighted directed graph provides a basis for determining the distribution network state data. For example, the load power of the node, the output power of the distributed power source, the resistance and reactance of the line, etc., are all important components of the distribution network state. Through the weighted directed graph model, the specific parameters of each node and line can be clarified, so as to accurately obtain and process these state data. When receiving the calculation results of the quantum computing part and converting them into control instructions or decision results, the weighted directed graph model helps to understand and interpret these results. Because the control instructions or decision results are usually related to switch operations, distributed power scheduling, etc., and these operations are closely related to the nodes and lines in the weighted directed graph. For example, according to the graph model, it can be determined which switches need to be closed or opened to achieve the optimal fault recovery strategy. Quantum computing part: The mapping relationship between quantum bits and distribution network elements is established based on the weighted directed graph. The voltage state of the node, the on-off state of the line, etc. correspond to the nodes and edges in the graph. By mapping quantum bits to distribution network elements, the state of the distribution network can be accurately represented by quantum states. For example, the position and connection relationship of a node in a weighted directed graph determines how its corresponding quantum bit is represented in the quantum state. The evolutionary simulation of the distribution network state, such as fault propagation and equipment action, is also based on weighted directed graphs. The propagation path of the fault in the graph is closely related to the connection relationship of the line, and the equipment action (such as switch operation) changes the structure of the graph. Simulating these processes through quantum gate operations actually reflects the changes in the state of the distribution network in the weighted directed graph at the quantum state level.
[0035] In step S12:
[0036] State and action representation: The quantum state representation of states and actions in reinforcement learning is based on a weighted directed graph model. The state of the distribution network, such as node voltage, line current, load power, etc., has a clear definition and position in the weighted directed graph. These state information is encoded into the quantum state, forming a comprehensive representation of the state of the distribution network. The representation of actions, such as the opening and closing operations of switches, the adjustment of the output power of distributed power sources, etc., are also related to the nodes and lines in the weighted directed graph. Different actions will change the structure and parameters of the graph, and thus affect the state of the distribution network. By encoding these actions, the quantum state enables the reinforcement learning algorithm to learn the impact of different actions on the state of the distribution network.
[0037] Construction and update of policy network and value network: The construction and update of policy network and value network rely on the accurate representation of the state and action of the distribution network, which is based on the weighted directed graph model. The update of policy network and value network through quantum gate operation is actually to adjust the network parameters according to the state change and action selection of the distribution network in the weighted directed graph to find the optimal strategy. For example, when calculating the probability of taking a certain action in the policy network, it is necessary to consider the current state of the distribution network in the weighted directed graph, including the voltage level of the node, the load condition of the line, etc. The value network's evaluation of the value of the distribution network state also needs to be based on various parameters and state information in the weighted directed graph.
[0038] Quantum reward function design: The quantum reward function comprehensively considers multiple factors of the distribution network, such as power outage losses, voltage quality, network losses, etc. The calculation of these factors is closely related to the node and line parameters in the weighted directed graph. For example, the calculation of power outage losses involves the load power of the node, which is determined in the weighted directed graph model. The evaluation of voltage quality depends on the voltage of the node, and the calculation of network losses is related to the resistance and current of the line, which are all parameters defined in the weighted directed graph. Through the quantum reward function, the reinforcement learning algorithm can give corresponding rewards or penalties based on the state and action selection of the distribution network in the weighted directed graph, thereby guiding the algorithm to find the optimal fault recovery strategy.
[0039] In step S13, when a fault is detected in the power distribution network, the classical computing part obtains the distribution network status data after the fault, such as the location of the fault node, the affected lines, load changes, etc., according to the weighted directed graph model, and encodes these data into the probability amplitude of quantum bits and transmits them to the quantum computing part. The weighted directed graph model provides a clear framework for determining the fault location and the scope of influence, so that the classical computing part can accurately extract and process the fault status data.
[0040] In step S14, the quantum computing part uses the designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy based on the distribution network structure and parameter information provided by the weighted directed graph model. When the classical computing part converts the optimal fault recovery strategy into actual control instructions, it also determines the specific switch operation, distributed power supply scheduling and other control actions based on the weighted directed graph model to achieve fault recovery of the distribution network.
[0041] As an improvement of the above embodiment, the power distribution network state data includes current, voltage and power; and the weighted directed graph model includes node information and line information of the power distribution network.
[0042] Specifically, in an embodiment of the present invention, current is an important component of the power distribution network status data. Current data can reflect the load conditions of each line in the distribution network. By measuring and collecting the current value in the line, the current magnitude and direction at different locations can be understood. These current data can be combined with the line information in the weighted directed graph model. For example, for each line, its current value at different times and the current change under different operating conditions (such as different load levels, distributed power supply output changes, etc.) can be recorded. In the event of a fault, the current will change significantly. By monitoring the sudden change or abnormality of the current, the occurrence of the fault can be quickly detected. For example, when the current on a line suddenly increases to exceed its rated value, it may mean that a short circuit fault has occurred in the line. At this time, the current data will be included in the distribution network status data after the fault as one of the important bases for fault diagnosis, and used for subsequent fault recovery strategy calculation. Voltage data reflects the potential of each node in the distribution network. It is not only closely related to the normal operation of power equipment, but also one of the key indicators for measuring the power supply quality of the distribution network. In the weighted directed graph model, the node information includes voltage-related parameters of each node, such as rated voltage, actual voltage, etc. Like current, voltage varies under different operating conditions. For example, during peak load periods, node voltage may drop, while when the output power of distributed power sources increases, node voltage may rise. When a fault occurs, the voltage changes more significantly. For example, near a short-circuit fault point, the voltage drops sharply. The changes in these voltage data will be recorded in detail and correlated with the node information in the weighted directed graph model, providing important information for the status assessment and fault recovery of the distribution network. Power includes active power and reactive power, which respectively reflect the transmission and consumption of energy in the distribution network and the flow of reactive power in the circuit. In the weighted directed graph model, the power information of the node is closely related to the operating status of the load and distributed power sources. For load nodes, the size of active power and reactive power reflects the power demand of the user or equipment connected to the node. For distributed power source nodes, the active power and reactive power they output determine their power supply capacity for the distribution network. After a fault occurs, the distribution of power will change. For example, some loads may lose power due to the fault, causing the power they consume to become zero; and distributed power sources may need to adjust their output power to support fault recovery. These changes in power data will serve as an important part of the distribution network status data, used to determine the distribution network operation status after the fault, and provide a basis for formulating fault recovery strategies.
[0043] In addition, node information: The nodes in the weighted directed graph model represent various electrical connection points in the power distribution network, such as substations, distribution boxes or user access points. Node information includes but is not limited to the node number, location, voltage level, rated voltage, connected load or distributed power type and capacity, etc. For example, a node may be marked as an important load node, the load power connected to it is large, and the power supply reliability requirements are high. During the fault recovery process, the recovery priority of the node may be increased. Alternatively, if a node is a distributed power access point, the range and characteristics of its output power will also be recorded in the node information, so as to reasonably schedule the output of the distributed power source in the fault recovery strategy.
[0044] In addition, line information: Line information covers the detailed parameters of the lines connecting various nodes, such as line resistance, reactance, conductance, susceptance, line length, line capacity, etc. These parameters determine the electrical characteristics and transmission capacity of the line. For example, if the resistance of a line is large, it will cause a large power loss when the current passes through it. When formulating a fault recovery strategy, it is necessary to consider the line loss factor and try to avoid transmitting too much power through lines with large losses. In addition, the line capacity also limits the maximum current or power passing through the line. If the power flow needs to be redistributed during the fault recovery process, it is necessary to ensure that the power transmitted by the line does not exceed its capacity limit to prevent the line from being overloaded.
[0045] For ease of understanding, here are some examples:
[0046] Assume that in an actual power distribution network, there is a weighted directed graph model consisting of multiple nodes and lines. It is a substation, node They are user access points and nodes in different areas. It is a distributed power access point. Under normal operation, the voltage, current and power data of each node are recorded. When a short circuit occurs, the current at the fault point increases sharply, and the node and The voltage will drop sharply. At this time, the current, voltage and power data after the fault are collected and combined with the node and line information in the weighted directed graph model. The classical computing part of the quantum-classical hybrid architecture will process the distribution network status data after the fault, encode it into the probability amplitude of quantum bits and transmit it to the quantum computing part. The quantum computing part uses quantum states to represent the state of the distribution network after the fault based on the mapping relationship between quantum bits and distribution network elements, and simulates fault propagation and possible equipment actions through quantum gate operations. Based on the weighted directed graph model and quantum reinforcement learning algorithm, the optimal fault recovery strategy is calculated. For example, a possible strategy is to disconnect the faulty line , adjust the distributed power supply output power, close some switches to redistribute power flow and restore important load nodes The classical computing part converts the optimal strategy calculated by the quantum computing part into actual control instructions, performs corresponding switching operations and distributed power supply power adjustment, and realizes fault recovery of the power distribution network. In this way, the embodiment of the present invention can make full use of the power distribution network state data and the weighted directed graph model, combined with the quantum reinforcement learning algorithm, to quickly and accurately realize the fault recovery of the power distribution network and improve the reliability and operation efficiency of the distribution network.
[0047] As an improvement of the above embodiment, the power distribution network is modeled as a weighted directed graph model, including:
[0048] Modeling the power distribution network as a weighted directed graph ;in, represents the set of nodes of the power distribution network, is the number of spare nodes, each node represents an electrical connection point in the power distribution network; Represents a set of lines, each line Connect two nodes and ; Each line Both have resistance Reactance , Conductivity and electrosusceptor Electrical parameters, which reflect the electrical characteristics of the line; nodes Active power for load and reactive power They represent the power demand of the users or devices connected to the node; distributed generation at the node The output power of and Indicates the power generated by distributed energy at the node; the switch state is represented by a binary variable express, It means the switch is closed and current can flow through the line; Indicates that the switch is disconnected and the line is in disconnected state;
[0049] For nodes The node voltage, using the complex form To indicate that is the real part of the node voltage, is the imaginary part of the node voltage; the relationship between the node voltage and the line current and power is obtained through the following calculation equation. , the power is expressed as: ;in, It is the line admittance, which reflects the line's ability to conduct current; Is a node The conjugate of the voltage, i.e. .
[0050] In this embodiment, by accurately modeling the power distribution network as a weighted directed graph model, various elements and parameters of the distribution network are fully covered, providing a solid foundation for the subsequent fault recovery strategy formulation. On the one hand, it is a refined modeling of the distribution network, which not only clarifies the basic elements such as nodes and lines, but also considers in detail the electrical parameters such as resistance, reactance, conductance, and susceptance of the line, as well as the load of the node, the output power of the distributed power supply, and the switch state, so that the model is closer to the complex situation of the actual distribution network; on the other hand, the node voltage is represented in complex form, and the relationship between the node voltage and the line current and power is established through a specific calculation equation, which can more accurately reflect the electrical characteristics and energy transmission laws of the distribution network, and provide a powerful tool for in-depth analysis of the distribution network state and fault conditions. Specifically, the power distribution network is constructed as a weighted directed graph model. Among them, the node set represents each electrical connection point in the distribution network, and each node has its specific meaning and function. Each line in the line set connects two nodes, and each line has parameters such as resistance, reactance, conductance, and susceptance that reflect its electrical characteristics. The load of the node is measured by active power and reactive power, which reflects the power demand of users or equipment. The output power of distributed power generation at the node is also clearly represented, and the switch state is described by binary variables to determine the on and off status of the line.
[0051] To make it easier to understand, here is an example:
[0052] Assume that there is a real power distribution network scenario. In this distribution network, there are multiple nodes, such as node It is a substation that connects multiple user access points (such as nodes etc.) and distributed power access points (such as nodes ).line Has a certain resistance Reactance , Conductivity and electrosusceptor .node The load active power is , the reactive power is , and the node The distributed power output active power is , the reactive power is .switch The initial state is closed, that is , current can flow through the line When the operation status of the distribution network needs to be analyzed, the node Assuming that the node The voltage is , and the node Other connected nodes (such as nodes ) is , the line admittance is , then according to the formula , we can calculate the node If at a certain moment, the line A fault occurred on the switch. May be disconnected, i.e. . At this point, by recalculating the relationship between the node voltage and the line current and power, the impact of the fault on the entire distribution network can be accurately evaluated. Moreover, based on this detailed model and calculation method, the subsequent quantum-classical hybrid architecture and quantum reinforcement learning algorithm can be used to formulate the optimal fault recovery strategy, such as adjusting the switching state of other lines, changing the output power of distributed power sources, etc., to restore the normal operation of the distribution network as soon as possible. In summary, this embodiment provides comprehensive and accurate basic information for fault recovery of the power distribution network through precise modeling and detailed calculations, which helps to achieve efficient and reliable fault recovery strategies.
[0053] As an improvement of the above embodiment, the weighted directed graph model based on the power distribution network is used to design a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by the classical computing is encoded as a probability amplitude of a quantum bit and transmitted to the quantum computing part, and the calculation result of the quantum computing part is converted into a control instruction or a decision result; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by a quantum state, and the evolution of the power distribution network state is simulated through quantum gate operations, including:
[0054] Based on the node information and line information of the weighted directed graph model of the power distribution network, a data transmission protocol and result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery are designed in the following ways:
[0055] a. Classical computing part:
[0056] Data transmission protocol design: The power distribution network state data processed by the classical computing part is transmitted to the quantum computing part in the form of the initialization state of the quantum bit; the normalized value of the node state data is encoded as the probability amplitude of the quantum bit; the normalized value of the node voltage is set to , then the corresponding quantum bit state is expressed as ; indicates that when When it approaches 0, the quantum bit is in The probability of the state is higher; when When it is close to 1, the quantum bit is in The probability of the state is higher;
[0057] Receiving and analyzing quantum computing results: The quantum state returned by the quantum computing part is measured through quantum measurement to obtain the measurement result; the classical computing part measures the quantum state multiple times and statistically analyzes the distribution of the measurement results; the quantum state is assumed to represent the probability of selecting an action, and the frequency of different actions being selected is obtained through multiple measurements; the action with the highest frequency is selected as the final decision result;
[0058] b. Quantum computing part:
[0059] The mapping relationship between quantum bits and power distribution network elements:
[0060] The voltage state of each node is mapped to a qubit; for example, the node The state of the power distribution network using quantum bits The ground state of the quantum bit is Indicates that the power distribution network status data is within the lower limit of the normal range, the excited state It indicates the upper limit of the power distribution network status data within the normal range, while the superposition state indicates different values of the power distribution network status data within the normal range;
[0061] The on-off state of the circuit is represented by the entangled state of two quantum bits; The on and off states of the quantum bit and It means that when the circuit is closed, the quantum state is expressed as ; When the line is disconnected, the quantum state is represented by ; By controlling the degree of entanglement between quantum bits, the on-off state of the circuit can be accurately represented;
[0062] c. Quantum state representation of the power distribution network state:
[0063] The state of the power distribution network Using quantum state It is the tensor product of multiple quantum bit states; for a The state of the power distribution network with nodes is expressed as ; The state of each quantum bit is represented by ,in and is a complex coefficient, satisfying ; Representation Node The probability amplitude of the power distribution network status data being within the predetermined interval within the normal range, Representation Node The probability amplitude that the power distribution network state data is within another predetermined interval within a normal range;
[0064] d. Quantum gate operation simulation of power distribution network state evolution:
[0065] Fault propagation simulation: When a fault occurs, the quantum NOT gate is used To indicate the disconnection operation of the switch; A short circuit occurs in the line where the corresponding quantum bit and Apply the quantum NOT gate, that is ,in and They are and The opposite state of the quantum bit is changed through the quantum NOT gate operation to simulate the disconnection of the line; for the propagation of the fault current, the quantum control-NOT gate is used to achieve it; set the node The fault current needs to propagate to the node , the node The quantum bit is used as the control bit, and the node The qubit is used as the target bit; if the control bit is in state, apply quantum NOT gate operation to the target bit to realize the propagation of fault current;
[0066] Equipment action simulation: When the agent decides to take corresponding actions, the actions include closing the switch, opening the switch, and adjusting the output power, which are realized through the corresponding quantum gate operations; among them, the adjustment of the output power of the distributed power supply is represented by changing the probability amplitude of the quantum bit. Assume that the output power of the distributed power supply can be range, when the output power is When , the corresponding quantum bit state is expressed as , thereby realizing the encoding of the output power of the distributed power supply by adjusting the probability amplitude of the quantum bit; for the switch operation, the quantum gate is used to change the quantum bit state corresponding to the switch state, where for the closed switch operation, the quantum register state corresponding to the switch is set to ; For disconnect switch operation, set to .
[0067] In this embodiment, based on the weighted directed graph model of the power distribution network, a data transmission protocol and result receiving and parsing method of a quantum-classical hybrid architecture are designed to achieve the collaborative work of classical computing and quantum computing, so as to efficiently handle the problem of power distribution network fault recovery. In terms of data transmission, the distribution network state data processed by classical computing is encoded into the probability amplitude of quantum bits for transmission, which realizes the seamless connection between classical data and quantum computing; in addition, through the clever mapping of quantum bits and distribution network elements, and the precise representation of the distribution network state by quantum states, the complex state and evolution process of the distribution network can be simulated more comprehensively and accurately; in addition, the use of quantum gate operations to simulate fault propagation and equipment action provides a more sophisticated analysis method for the formulation of fault recovery strategies.
[0068] Specifically, for the classical computing part:
[0069] Data transmission protocol design: The processed distribution network state data is transmitted to the quantum computing part in the form of the initialization state of the quantum bit. The normalized value of the node state data is encoded as the probability amplitude of the quantum bit. For example, the normalized value of the node voltage is , the corresponding quantum bit state is .
[0070] Receiving and analyzing quantum computing results: The quantum state returned by the quantum computing part is measured through quantum measurement, and the classical computing part measures the quantum state multiple times to statistically analyze the distribution of the measurement results. If the quantum state represents the probability of action selection, the frequency of different actions being selected is obtained through multiple measurements, and the action with the highest frequency is selected as the decision result.
[0071] For the quantum computing part:
[0072] The mapping relationship between quantum bits and power distribution network elements: The node voltage state is mapped to a quantum bit, the ground state represents the lower limit of the state data, the excited state represents the upper limit, and the superposition state represents different values within the normal range. For the on-off state of the line, the entangled state of two quantum bits is used to represent it, such as when the line is closed and disconnected, there are different quantum states.
[0073] Quantum state representation of power distribution network state: The distribution network state is represented by a quantum state, which is the tensor product of multiple quantum bit states. The state of the distribution network of nodes is expressed as , the state of each quantum bit is ,in and satisfy .
[0074] Quantum gate operation simulation of power distribution network state evolution:
[0075] Fault propagation simulation: When a fault occurs, a quantum NOT gate is used to represent the switch disconnection operation. If the line where the switch is located is short-circuited, a quantum NOT gate is applied to the corresponding quantum bit to simulate the line disconnection. For fault current propagation, a quantum control-NOT gate is used to implement it, and the node quantum bit is used as the control bit and the target bit for operation.
[0076] Equipment action simulation: The actions taken by the intelligent agent (such as opening and closing the switch, adjusting the output power) are realized through quantum gate operations. The output power of the distributed power supply is adjusted by changing the quantum bit probability amplitude, and the switch operation is realized by changing the quantum register state corresponding to the switch. The closed switch and the open switch have different quantum register states.
[0077] To make it easier to understand, here is an example:
[0078] Assume that in an electric power distribution network, there are multiple nodes and lines. The voltage state of the node needs to be transmitted to the quantum computing part for processing. First, in the classical computing part, The voltage state data is normalized to obtain the normalized value This is then encoded as the probability amplitude of the qubit , and transmit it to the quantum computing part. After receiving the quantum bit, the quantum computing part maps the quantum bit to the node according to the mapping relationship between the quantum bit and the distribution network element. Assume that the node The voltage of the circuit is within a certain value within the normal range, and the corresponding quantum bit is in a superposition state. The on-off state of the circuit is represented by the entangled state of two quantum bits. Initially in a closed state, its quantum state is a specific entangled state. When a short circuit occurs on the circuit, the classical computing part detects the fault and transmits the relevant information to the quantum computing part. The quantum computing part uses quantum NOT gate pairs to represent the circuit. The quantum bit is operated to simulate the switch disconnection operation. To Node The propagation of nodes The quantum bit is used as the control bit, and the node The quantum bit of the target bit is used to simulate the fault current propagation through quantum control-NOT gate operation. If the intelligent agent decides to adjust the output power of a distributed power source, such as the distributed power source located at the node , its output power can be When the output power is When a quantum bit changes, the probability amplitude of the quantum bit is changed to represent the change in output power. For switch operations, such as closing or opening a switch, the quantum register state corresponding to the switch state is changed through the corresponding quantum gate operation. Finally, the quantum computing part calculates the optimal fault recovery strategy and feeds it back to the classical computing part. The classical computing part converts it into actual control instructions, such as switch operation instructions or distributed power output power adjustment instructions, to achieve fault recovery of the distribution network. Through such a quantum-classical hybrid architecture data transmission protocol and result reception and analysis method, the advantages of classical computing and quantum computing can be fully utilized to handle the fault recovery problem of the power distribution network more accurately and efficiently.
[0079] As an improvement of the above embodiment, the weighted directed graph model based on the power distribution network is used to design a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery, and the training data of the power distribution network is used for training; wherein the state and action in reinforcement learning are represented by quantum state, and the construction and update of the strategy network and value network of the quantum reinforcement learning algorithm are realized by quantum gate operation, and the quantum reward function is designed, including:
[0080] Based on the node information and line information of the weighted directed graph model of the power distribution network, a quantum reinforcement learning algorithm and strategy optimization for fault recovery of the power distribution network are designed in the following way, where the quantum state of the quantum reinforcement learning algorithm represents the state and action in reinforcement learning:
[0081] a. Status indication:
[0082] The state of the power distribution network Using quantum state It is the tensor product of multiple quantum bit states; for a The state of the power distribution network with nodes is expressed as ; The state of each quantum bit is represented by ,in and is a complex coefficient, satisfying ; Set node The power distribution network status data has three different intervals within the normal range, corresponding to low, medium and high interval values; Representation Node The probability amplitude of the power distribution network status data being in the low range, Representation Node The probability amplitude of the power distribution network status data is in the high range, and It indicates the probability amplitude that the power distribution network status data is in the middle range;
[0083] b. Action representation:
[0084] action Expressed in quantum state; assuming the action space includes different actions, including closing the switch, opening the switch, and adjusting the output power; each action can be represented by a quantum register; for the switch operation action , using a quantum register consisting of multiple qubits Indicates that is the number of qubits required to represent the action; where , and set the corresponding quantum register state to , which means that when the quantum register is in this state, it represents a switch closing action; if you want to open the switch , set it to ; For the adjustment of the output power of the distributed power source, the different values of the output power are mapped to the different states of the quantum bit; Assume that the output power of the distributed power source is The range varies, and this range is divided into multiple discrete values, each of which corresponds to a specific quantum state; when the output power is When , the corresponding quantum state is expressed as ; When the output power is When , the corresponding quantum state is expressed as ;
[0085] c. Quantum gate operations to implement strategy updates and optimizations:
[0086] c1. Policy network:
[0087] Defining the Quantum Strategy Network ,in is the parameter of the quantum gate; the strategy network is used to adjust the current power distribution network status Calculate each action taken The probability of using a quantum revolving door To adjust the state of the quantum bit, The operation of the quantum rotating gate is expressed as: ; is an adjustable parameter. The value of , adjust the quantum bit to and The probability of the state is used to update the strategy. In the strategy network, a complex quantum state transformation is constructed by combining multiple quantum rotating gates. For a system composed of qubits, the strategy network is expressed as: ;in It is a unitary transformation composed of quantum rotating gates; the unitary transformation ensures that the evolution of the quantum state is reversible and maintains the normalized properties of the quantum state; then, by Take measurements to get the probabilities of taking different actions;
[0088] c2. Value network:
[0089] Building a quantum value network , used to evaluate the value of the current power distribution network status; are the parameters of the quantum gates in the value network; the Hamiltonian of the value network is expressed as , by the quantum state Applying the Hamiltonian and performing quantum measurements yields an estimate of the value, where the measured quantum state The expected value obtained is expressed as: ;
[0090] Hamiltonian Designed as a function of various operating indicators and incentive factors of the power distribution network ,in is the weight coefficient; It is used to measure the importance of the deviation of the power distribution network of the node from the rated voltage; Used to control the effect of line current on value; Reflects the impact of load power on value;
[0091] c3. Strategy update:
[0092] Use the policy gradient algorithm to update the parameters of the policy network ; The policy gradient is expressed as: ;in is the objective function of the policy, which measures the performance of the policy; is the action value function, which means that in state Take action the value of
[0093] The action-value function is estimated by the following Bellman formula: ; is the quantum reward function, which comprehensively considers various factors after the power distribution network takes action in the current state; is a discount factor that weighs the importance of future rewards against current rewards; Indicates that considering the state transition probability Finally, the expected value of the future state;
[0094] According to the calculation results of the strategy gradient, by adjusting the parameters of the quantum rotation gate To update the strategy, the gradient descent method is used to update the parameters: ,in is the learning rate, which determines the step size of each update; is the number of iterations; in each iteration, the policy gradient is calculated based on the current policy network and value network, and then the parameters are updated to gradually optimize the strategy;
[0095] c4. Quantum reward function design:
[0096] Quantum Reward Function Comprehensively consider multiple factors of the power distribution network:
[0097] Power outage losses: ,in Is a node Load power; It represents a node A binary variable indicating whether there is a power outage, Indicates that there is no power outage. Indicates a power outage; is the unit outage loss cost, which means that when a node loses power, a power outage loss will occur, and the size of the loss is proportional to the load power and the unit outage loss cost;
[0098] Voltage quality: ,in Is a node The actual voltage, is the rated voltage. This formula measures the sum of squared deviations of the node voltage from the rated voltage. The larger the deviation, the worse the voltage quality and the lower the reward value.
[0099] Network loss: ,in It is a line The current, is the line resistance; the network loss is proportional to the square of the line current and the resistance. The greater the loss, the lower the reward value;
[0100] The quantum reward function is expressed as ,in , , is the weight coefficient, which is used to balance the importance of different factors;
[0101] After designing the quantum reinforcement learning algorithm and strategy optimization method, the quantum reinforcement learning algorithm is trained and the strategy is optimized in the following ways:
[0102] Collect historical operation data of the power distribution network, including voltage, current and power in different seasons, time periods, load levels and fault conditions;
[0103] Clean the historical operation data of the power distribution network to remove outliers and noise;
[0104] Normalize the cleaned historical operation data and map it to interval;
[0105] The normalized historical running data is divided into a training set and a validation set; the training set is used to train the quantum reinforcement learning algorithm; the validation set is used to adjust the algorithm's hyperparameters, including the learning rate and quantum gate parameters;
[0106] Use the training set to train the designed quantum reinforcement learning algorithm, and use the validation set to adjust the hyperparameters of the quantum reinforcement learning algorithm and perform strategy optimization at the same time.
[0107] In this embodiment, based on the weighted directed graph model of the power distribution network, a quantum reinforcement learning algorithm and strategy optimization are designed, the state and action are represented by quantum state, the strategy network and value network are constructed and updated by quantum gate operation, and a quantum reward function is designed, and the training data of the power distribution network is used for training. The technical innovations include: the accurate representation of the state and action of the distribution network by quantum state can more comprehensively capture the complex characteristics of the distribution network; the construction and update of the strategy network and value network are realized through quantum gate operation, which provides a new method for strategy optimization; the quantum reward function that comprehensively considers multiple factors such as power outage loss, voltage quality and network loss is designed to make the strategy more targeted and effective; the rich historical operation data of the power distribution network is used for training and hyperparameter adjustment, which improves the adaptability and performance of the algorithm.
[0108] It is understandable that the training process of the designed quantum reinforcement learning algorithm can refer to the training method of the existing reinforcement learning algorithm, which will not be elaborated here.
[0109] As an improvement of the above embodiment, when a fault in the power distribution network is detected, the state data of the power distribution network after the fault is obtained according to the weighted directed graph model by the classical computing part of the quantum-classical hybrid architecture, and the state data of the power distribution network after the fault is encoded into a probability amplitude of a quantum bit and transmitted to the quantum computing part, including:
[0110] When the relevant sensors and monitoring equipment of the classical computing part detect that the power distribution network has failed, the power distribution network status data after the power distribution network failure is collected according to the weighted directed graph model;
[0111] The classical computing part encodes the state data of the power distribution network after the fault according to the pre-designed data transmission protocol, and after the encoding is completed, the classical computing part transmits the state data of the power distribution network in the form of quantum state to the quantum computing part;
[0112] The quantum computing part receives the power distribution network status data in the form of quantum state and uses it as input.
[0113] In this embodiment, when a fault occurs in the power distribution network, the classical computing part of the quantum-classical hybrid architecture obtains and processes the distribution network state data after the fault according to the weighted directed graph model, and encodes it into the probability amplitude of the quantum bit and transmits it to the quantum computing part for subsequent fault recovery strategy calculation. The technical innovation lies in making full use of the weighted directed graph model to accurately collect the state data after the fault, and realizes the efficient interaction between classical computing and quantum computing through a carefully designed data transmission protocol, so that quantum computing can obtain comprehensive and accurate fault information in a timely manner, laying the foundation for the rapid formulation of effective fault recovery strategies. Specifically, when the sensors and monitoring equipment of the classical computing part detect that a fault occurs in the distribution network, the state data of the distribution network after the fault will be collected according to the weighted directed graph model. These data include but are not limited to the voltage value of each node, the current value of the line, the load power, the output power of the distributed power supply, and the switch state. The classical computing part encodes these state data after the fault according to the pre-designed data transmission protocol. During the encoding process, these data are converted into the probability amplitude form of quantum bits. After the encoding is completed, the classical computing part will transmit the distribution network status data in the form of quantum state to the quantum computing part. After receiving this data, the quantum computing part will use it as input for subsequent calculations.
[0114] For ease of understanding, an example is given here: Assume that in an actual power distribution network, at a certain moment, the sensor of the classical computing part detects a short circuit fault in a line (for example, in an actual power distribution network, current transformers, voltage transformers, zero-sequence current transformers, and smart meters distributed on various nodes and lines continuously collect the operation data of the power system. These data are transmitted to the monitoring center through the communication network, and the software system of the monitoring center analyzes these data in real time. When it is found that the current of a certain line suddenly increases to exceed the rated value, or the voltage of a certain node drops to an abnormal range, the system will automatically determine that a fault may have occurred and issue an alarm. At the same time, through the comprehensive analysis of multiple parameters, the type and location of the fault can be further determined to provide accurate information for subsequent fault recovery work. Among them, the fault analysis method can refer to the existing technology). According to the weighted directed graph model, the classical computing part quickly collects the status data related to the fault, such as a sharp drop in the voltage of the nodes at both ends of the fault line, a sharp increase in the current of the fault line, and a change in the load power connected to the line. Then, according to the data transmission protocol, the classical computing part encodes these data into the probability amplitude of quantum bits. For example, the node voltage value is converted into a specific quantum bit probability amplitude after processing. After the encoding is completed, the classical computing part transmits the state data in the form of these quantum states to the quantum computing part. After receiving this data, the quantum computing part can calculate the optimal fault recovery strategy based on these input information and the previously trained quantum reinforcement learning algorithm and model, such as whether to disconnect certain switches, adjust the output power of which distributed power sources, etc., to achieve rapid fault recovery of the distribution network. Through this process, when a fault occurs in the distribution network, the fault state data can be quickly and accurately transmitted to the quantum computing part, giving full play to the advantages of the quantum-classical hybrid architecture and improving the efficiency and accuracy of fault recovery.
[0115] As an improvement of the above embodiment, the quantum computing part uses a designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy of the power distribution network and feeds it back to the classical computing part, and the optimal fault recovery strategy is converted into actual control instructions by the classical computing part to perform the fault recovery work of the power distribution network, including:
[0116] After receiving the power distribution network status data after the fault, the quantum computing part calculates the optimal fault recovery strategy of the power distribution network based on the node information and line information provided by the weighted directed graph model and the trained quantum reinforcement learning algorithm. The specific process is as follows:
[0117] According to the quantum strategy network Calculate the current power distribution network status Take each action The probability of; where the quantum strategy network through the quantum revolving door Adjust the state of the quantum bit, and the combination of multiple quantum rotating gates constructs a complex quantum state transformation, thereby calculating the probability distribution of the action;
[0118] Through the value network Evaluate the value of each state-action pair; where the Hamiltonian of the value network It contains various operating indicators and reward factors of the power distribution network, and obtains value estimates by applying Hamiltonian to quantum states and performing quantum measurements;
[0119] According to the quantum reward function , Discount Factor and state transition probability , calculate the action value function through the Bellman equation ; Among them, the quantum reward function comprehensively considers factors such as power outage losses, voltage quality and network losses. The discount factor is used to balance the importance of current and future rewards. The state transition probability reflects the possibility of transferring to the next state after taking a certain action from the current state.
[0120] Use the policy gradient algorithm to update the parameters of the policy network , and find the optimal action sequence, i.e. the optimal fault recovery strategy, through quantum gate operation; wherein the optimal fault recovery strategy includes the opening and closing operation of the switch and the adjustment of the power supply output power;
[0121] The quantum computing part feeds back the calculated optimal fault recovery strategy to the classical computing part. After receiving the strategy, the classical computing part converts it into actual control instructions. If the strategy requires closing a switch, , the classical calculation part will generate the corresponding switch closing instruction; if the strategy requires adjusting the output power of a distributed power source, the classical calculation part will calculate the specific power adjustment value and generate the corresponding control instruction.
[0122] In this embodiment, after the quantum computing part receives the state data of the power distribution network after the fault, it combines the node and line information provided by the weighted directed graph model, uses the trained quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy, and feeds back to the classical computing part for control command conversion. The technical innovation is that through the synergy of quantum strategy network, value network, quantum reward function, policy gradient algorithm and quantum gate operation, the accurate calculation of the optimal fault recovery strategy of the power distribution network is realized. The quantum strategy network uses quantum rotating gates to adjust the state of quantum bits to calculate the probability distribution of actions, the value network evaluates the value of state-action pairs through Hamiltonian, the quantum reward function integrates multiple factors, and the policy gradient algorithm updates the policy network parameters. The combination of these innovative methods can more effectively cope with the complexity and uncertainty of distribution network fault recovery. Specifically, after receiving the state data of the distribution network after the fault, the quantum computing part starts to calculate the optimal fault recovery strategy. First, according to the quantum strategy network, the state of the quantum bit is adjusted by the quantum rotating gate, and multiple quantum rotating gates are combined to construct a complex quantum state transformation, and then the probability distribution of taking each action under the current distribution network state is calculated. Next, the value network comes into play, and its Hamiltonian covers various operating indicators and reward factors of the distribution network. The Hamiltonian is applied to the quantum state and quantum measurements are performed to evaluate the value of each state-action pair. Then, according to the quantum reward function that comprehensively considers factors such as power outage losses, voltage quality, and network losses, combined with the discount factor and state transition probability, the action value function is calculated through the Bellman equation. Finally, the policy gradient algorithm is used to update the parameters of the policy network, and the optimal action sequence is determined through quantum gate operations, that is, the optimal fault recovery strategy including switch opening and closing operations, power supply output power adjustment, etc. The quantum computing part feeds this strategy back to the classical computing part, and the classical computing part generates corresponding control instructions based on the content of the strategy, such as switch closing instructions or specific distributed power supply output power adjustment values.
[0123] For ease of understanding, here is an example: suppose that a fault occurs in a certain area of an electric power distribution network, and the quantum computing part receives the state data after the fault. Based on the node information (such as node voltage, connected load, etc.) and line information (such as line resistance, current, etc.) of the area in the weighted directed graph model, the quantum strategy network calculates the probability of different actions (such as the closing or opening of certain switches, the increase or decrease of the output power of distributed power sources) through quantum rotating gates. The value network evaluates the value of each state-action pair according to the Hamiltonian, for example, considering the impact of the closing of a certain switch on the line current distribution and node voltage, as well as the degree of improvement in the overall operating indicators. The quantum reward function comprehensively considers factors such as power outage losses (such as losses caused by power outages at important load nodes), voltage quality (deviation of node voltage from rated voltage), and network losses (losses caused by line current and resistance). The action value function is calculated through the Bellman equation, and the policy gradient algorithm updates the policy network parameters. Finally, the optimal fault recovery strategy is determined through quantum gate operations, such as closing certain key switches to restore the power supply of important loads and appropriately adjusting the output power of distributed power sources. The quantum computing part feeds back this strategy to the classical computing part, which then generates specific control instructions, such as closing a switch to a specific position or adjusting the output power of a distributed power source to a specific value, and then sends these control instructions to the corresponding devices for execution to achieve fault recovery of the power distribution network. In this way, the optimal fault recovery strategy can be formulated and executed quickly and accurately, improving the reliability and recovery efficiency of the distribution network.
[0124] See also Figure 2 , is a schematic diagram of a structure of a power distribution network fault recovery system based on reinforcement learning provided by an embodiment of the present invention. The power distribution network fault recovery system based on reinforcement learning includes:
[0125] A modeling module 10, used for modeling the power distribution network, and modeling the power distribution network as a weighted directed graph model;
[0126] The first design module 11 is used to design a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery based on a weighted directed graph model of the power distribution network; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by the classical computing is encoded into the probability amplitude of the quantum bit and transmitted to the quantum computing part, and the calculation result of the quantum computing part is converted into a control instruction or a decision result; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between the quantum bit and the power distribution network element of the power distribution network is established, the power distribution network state of the power distribution network is represented by the quantum state, and the evolution of the power distribution network state is simulated through quantum gate operations;
[0127] The second design module 12 is used to design a quantum reinforcement learning algorithm and strategy optimization for fault recovery of the power distribution network based on a weighted directed graph model of the power distribution network, and use the training data of the power distribution network for training; wherein the state and action in the reinforcement learning are represented by quantum states, and the construction and update of the strategy network and value network of the quantum reinforcement learning algorithm are realized by using quantum gate operations, and a quantum reward function is designed;
[0128] The computing module 13 is used for, when a fault in the power distribution network is detected, obtaining the state data of the power distribution network after the fault according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and encoding the state data of the power distribution network after the fault in the power distribution network into the probability amplitude of quantum bits and transmitting it to the quantum computing part;
[0129] The fault recovery module 14 is used to calculate the optimal fault recovery strategy of the power distribution network by using the designed quantum reinforcement learning algorithm through the quantum computing part and feed it back to the classical computing part, and convert the optimal fault recovery strategy into actual control instructions through the classical computing part to execute the fault recovery work of the power distribution network.
[0130] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0131] The embodiment of the present invention is based on the weighted directed graph model of the power distribution network, constructs a quantum-classical hybrid architecture, and realizes distribution network fault recovery through quantum reinforcement learning algorithm and strategy optimization. In this way, the optimal fault recovery strategy of the power distribution network can be determined quickly and accurately, and the complexity and uncertainty of the distribution network brought about by the diversified development of power loads in the power distribution network can be effectively dealt with, thereby realizing efficient and reliable recovery of power distribution network faults.
[0132] It should be noted that the relevant scheme details of the power distribution network fault recovery system based on reinforcement learning in the embodiment of the present invention can correspond to the relevant embodiments of the power distribution network fault recovery method based on reinforcement learning mentioned above, and will not be repeated here.
[0133] See also Figure 3 , is a schematic diagram of a power distribution network fault recovery system based on reinforcement learning provided in one embodiment of the present invention. The power distribution network fault recovery device based on reinforcement learning of this embodiment includes: a processor 100, a memory 101, and a computer program stored in the memory 101 and executable on the processor 100, such as a power distribution network fault recovery program based on reinforcement learning. When the processor 100 executes the computer program, the steps in the above-mentioned various embodiments of the power distribution network fault recovery method based on reinforcement learning are implemented. Alternatively, when the processor 100 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0134] The power distribution network fault recovery device based on reinforcement learning can be a main control device for controlling the power distribution network, such as a desktop computer, a notebook, a PDA, and a cloud server. The power distribution network fault recovery device based on reinforcement learning may include, but is not limited to, a processor and a memory. The processor may be a central processing unit, or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The memory can be used to store the computer program and / or module, and the processor realizes the various functions of the power distribution network fault recovery device based on reinforcement learning by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory.
[0135] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for fault recovery of power distribution network based on reinforcement learning, characterized in that: include: Modeling the power distribution network, modeling the power distribution network as a weighted directed graph model; Based on the weighted directed graph model of the power distribution network, a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery are designed; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation results of the quantum computing part are converted into control instructions or decision results; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations; Based on the weighted directed graph model of the power distribution network, a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery are designed, and the training data of the power distribution network is used for training. Among them, the state and action in reinforcement learning are represented by quantum state, and the strategy network and value network of the quantum reinforcement learning algorithm are constructed and updated by using quantum gate operations, and the quantum reward function is designed. When a fault in the power distribution network is detected, the state data of the power distribution network after the fault is obtained according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and the state data of the power distribution network after the fault is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part; The quantum computing part uses the designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy of the power distribution network and feeds it back to the classical computing part. The classical computing part converts the optimal fault recovery strategy into actual control instructions to execute the fault recovery work of the power distribution network. When a fault in the power distribution network is detected, the state data of the power distribution network after the fault is obtained according to the weighted directed graph model by the classical computing part of the quantum-classical hybrid architecture, and the state data of the power distribution network after the fault is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, including: When the relevant sensors and monitoring equipment of the classical computing part detect that the power distribution network has failed, the power distribution network status data after the power distribution network failure is collected according to the weighted directed graph model; The classical computing part encodes the state data of the power distribution network after the fault according to the pre-designed data transmission protocol, and after the encoding is completed, the classical computing part transmits the state data of the power distribution network in the form of quantum state to the quantum computing part; After receiving the power distribution network status data in the form of quantum state, the quantum computing part uses it as input; The quantum computing part uses the designed quantum reinforcement learning algorithm to calculate the optimal fault recovery strategy of the power distribution network and feeds it back to the classical computing part, and converts the optimal fault recovery strategy into actual control instructions through the classical computing part to perform the fault recovery work of the power distribution network, including: After receiving the power distribution network status data after the fault, the quantum computing part calculates the optimal fault recovery strategy of the power distribution network based on the node information and line information provided by the weighted directed graph model and the trained quantum reinforcement learning algorithm. The specific process is as follows: The probability of taking each action |a> under the current power distribution network state |s> is calculated according to the quantum strategy network π(|a>|s>; θ); wherein the quantum strategy network is connected through the quantum revolving gate R y (θ i ) Adjust the state of the quantum bit, and the combination of multiple quantum rotating gates constructs a complex quantum state transformation, thereby calculating the probability distribution of the action; Through the value network Evaluate the value of each state-action pair; where the Hamiltonian of the value network It contains various operating indicators and reward factors of the power distribution network, and obtains value estimates by applying Hamiltonian to quantum states and performing quantum measurements; According to the quantum reward function R(|s>, |a>), the discount factor γ and the state transition probability P(|s′||s>, |a>), the action value function Q(|s>, |a>) is calculated by the Bellman equation; among them, the quantum reward function comprehensively considers the power outage loss, voltage quality and network loss factors, the discount factor is used to balance the importance of current and future rewards, and the state transition probability reflects the possibility of transferring to the next state after taking a certain action from the current state; Use the policy gradient algorithm to update the parameters θ of the policy network, and find the optimal action sequence, i.e., the optimal fault recovery strategy, through quantum gate operations. The optimal fault recovery strategy includes the opening and closing operations of the switch and the adjustment of the power supply output power. The quantum computing part feeds back the calculated optimal fault recovery strategy to the classical computing part. After receiving the strategy, the classical computing part converts it into actual control instructions. If the strategy requires closing a switch S ij , the classical calculation part will generate the corresponding switch closing instruction; if the strategy requires adjusting the output power of a distributed power source, the classical calculation part will calculate the specific power adjustment value and generate the corresponding control instruction.
2. The power distribution network fault recovery method based on reinforcement learning according to claim 1, characterized in that: The power distribution network status data includes current, voltage and power; the weighted directed graph model includes node information and line information of the power distribution network.
3. The power distribution network fault recovery method based on reinforcement learning as claimed in claim 2, characterized in that: The power distribution network is modeled as a weighted directed graph model, including: The power distribution network is modeled as a weighted directed graph model G = (V, E); where V = {υ1, υ2, ..., υ n } represents the node set of the power distribution network, n is the number of standby nodes, and each node represents an electrical connection point in the power distribution network; E = {e ij |i, j∈V} represents a set of lines, each line e ij =(υ i , j ) connects two nodes v i and j ; Each line e ij Both have resistance R ij Reactance X ij , conductivity G ij and electrosusceptor B ij The electrical parameters of the line reflect the electrical characteristics of the line; the load of node i is expressed by active power and reactive power They represent the power demand of the user or device connected to the node; the output power of the distributed generation at node i is and It represents the power generation of distributed energy at this node; the switch state is represented by the binary variable S ij ∈{0,1} means, S ij =1 means the switch is closed and current flows through the circuit; S ij =0 means the switch is off and the line is disconnected; For the node voltage at node i, use the complex form V i =V Ri +jV Ii To represent, where V Ri is the real part of the node voltage, V Ii is the imaginary part of the node voltage; the relationship between the node voltage and the line current and power is obtained through the following calculation equation. For node i, the power is expressed as: Among them, Y ij =G ij +jR ij It is the line admittance, which reflects the line's ability to conduct current; is the conjugate of the voltage at node j, that is 4. The power distribution network fault recovery method based on reinforcement learning as claimed in claim 3, characterized in that: The weighted directed graph model based on the power distribution network is used to design a data transmission protocol and a result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation result of the quantum computing part is converted into a control instruction or a decision result; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations, including: Based on the node information and line information of the weighted directed graph model of the power distribution network, a data transmission protocol and result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery are designed in the following ways: a. Classical computing part: Data transmission protocol design: The power distribution network state data processed by the classical computing part is transmitted to the quantum computing part in the form of the initialization state of the quantum bit; the normalized value of the node state data is encoded as the probability amplitude of the quantum bit; let the normalized value of the node voltage be x, then the corresponding quantum bit state is expressed as It means that when the difference between x and 0 is less than the preset threshold, the probability that the quantum bit is in the |0> state is high; when the difference between x and 1 is less than the preset threshold, the probability that the quantum bit is in the |1> state is high; Receiving and analyzing quantum computing results: The quantum state returned by the quantum computing part is measured through quantum measurement to obtain the measurement result; the classical computing part measures the quantum state multiple times and statistically analyzes the distribution of the measurement results; the quantum state is assumed to represent the probability of selecting an action, and the frequency of different actions being selected is obtained through multiple measurements; the action with the highest frequency is selected as the final decision result; b. Quantum computing part: The mapping relationship between quantum bits and power distribution network elements: The voltage state of each node is mapped to a quantum bit; the power distribution network state of node i is represented by quantum bit q i Indicates that the ground state of the quantum bit |0> represents the lower limit of the power distribution network state data within the normal range, the excited state |1> represents the upper limit of the power distribution network state data within the normal range, and the superposition state represents different values of the power distribution network state data within the normal range; The on-off state of the circuit is represented by the entangled state of two quantum bits; suppose the circuit e ij The on and off states of the quantum bit and It means that when the circuit is closed, the quantum state is expressed as When the line is disconnected, the quantum state is represented by By controlling the degree of entanglement between quantum bits, the on and off states of the circuit can be accurately represented; c. Quantum state representation of the power distribution network state: The state s of the power distribution network is represented by the quantum state |s>, which is the tensor product of multiple quantum bit states; for a power distribution network with n nodes, the state is represented as The state of each qubit is represented by |q i >=α i |0>+β i |1>, where α i and β i are complex coefficients, satisfying |α i | 2 +|β i | 2 =1;α i represents the probability amplitude of the power distribution network state data of node i being within the normal range of the predetermined interval, β i The probability amplitude indicating that the power distribution network state data of node i is within another predetermined interval within a normal range; d. Quantum gate operation simulation of power distribution network state evolution: Fault propagation simulation: When a fault occurs, a quantum NOT gate X is used to represent the disconnection operation of the switch; let the switch S ij A short circuit occurs in the line where the corresponding quantum bit and Apply the quantum NOT gate, that is in and They are and The opposite state of node i is thus changed through quantum NOT gate operation to simulate the disconnection of the line; for the propagation of fault current, quantum control-NOT gate is used to achieve it; suppose the fault current of node i is to be propagated to node j, the quantum bit of node i is used as the control bit, and the quantum bit of node j is used as the target bit; if the control bit is in the |1> state, the quantum NOT gate operation is applied to the target bit to achieve the propagation of fault current; Equipment action simulation: When the intelligent agent decides to take corresponding actions, the actions include closing the switch, opening the switch, and adjusting the output power, which are realized through the corresponding quantum gate operation; among them, the adjustment of the output power of the distributed power supply is represented by changing the probability amplitude of the quantum bit. Assume that the output power of the distributed power supply is [P min , P max ] range, when the output power is P, the corresponding quantum bit state is expressed as [P min , P max ], thereby encoding the output power of the distributed power supply by adjusting the probability amplitude of the quantum bit; for the switch operation, the quantum gate is used to change the quantum bit state corresponding to the switch state, where for the closed switch operation, the quantum register state corresponding to the switch is set to For disconnect switch operation, set to 5. The power distribution network fault recovery method based on reinforcement learning as claimed in claim 4, characterized in that: The weighted directed graph model based on the power distribution network is used to design a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery, and the training data of the power distribution network is used for training; wherein the state and action in reinforcement learning are represented by quantum state, and the strategy network and value network of the quantum reinforcement learning algorithm are constructed and updated by using quantum gate operation, and a quantum reward function is designed, including: Based on the node information and line information of the weighted directed graph model of the power distribution network, a quantum reinforcement learning algorithm and strategy optimization for fault recovery of the power distribution network are designed in the following way, where the quantum state of the quantum reinforcement learning algorithm represents the state and action in reinforcement learning: a. Status indication: The state s of the power distribution network is represented by the quantum state |s>, which is the tensor product of multiple quantum bit states; for a power distribution network with n nodes, the state is represented as The state of each qubit is represented by |q i >=α i |0>+β i |1>, where α i and β i are complex coefficients, satisfying |α i | 2 +|β i | 2 =1′; Assume that the power distribution network status data of node i has three different intervals within the normal range, corresponding to low, medium and high interval values respectively; α i Indicates the probability amplitude that the power distribution network state data of node i is in the low range, β i represents the probability amplitude of the power distribution network state data of node i being in the high range, and 1-|α i | 2 -|β i | 2 It indicates the probability amplitude that the power distribution network status data is in the middle range; b. Action representation: Action a is represented by a quantum state; suppose the action space includes m different actions, including closing the switch, opening the switch, and adjusting the output power; each action is represented by a quantum register; for the switch operation action a j , using a quantum register consisting of multiple qubits where k is the number of qubits required to represent the action; where if the switch S is closed ij , and set the corresponding quantum register state to Indicates that when the quantum register is in this state, it represents a switch closing action; if you want to open the switch S ij , set it to For the adjustment of the output power of the distributed power source, different values of the output power are mapped to different states of the quantum bit; assuming that the output power of the distributed power source is [P min , P max ] range, dividing this range into multiple discrete values, each of which corresponds to a specific quantum state; when the output power is P1, the corresponding quantum state is expressed as When the output power is P2, the corresponding quantum state is expressed as c. Quantum gate operations to implement strategy updates and optimizations: c1. Policy network: Define a quantum strategy network π(|a>|s>θ), where θ = {θ1, θ2, …, θ l } are the parameters of the quantum gate; the strategy network is used to calculate the probability of taking each action |a> according to the current power distribution network state |s>; the quantum rotation gate R is used y (θ i ) to adjust the state of the quantum bit. For the i-th quantum bit, the operation of the quantum rotation gate is expressed as: θ i is an adjustable parameter that can be adjusted by changing θ i The value of adjusts the probability of the quantum bit being in the |0> and |1> states, thereby updating the strategy; in the strategy network, complex quantum state transformations are constructed through the combination of multiple quantum rotating gates. For a system composed of n quantum bits, the strategy network is expressed as: |ψ policy >=U(θ)|s>; where is a unitary transformation composed of quantum rotating gates; the unitary transformation ensures that the evolution of the quantum state is reversible and maintains the normalized properties of the quantum state; then, by policy >Measure the probability of taking different actions; c2. Value network: Building a quantum value network Used to evaluate the value of the current power distribution network status; are the parameters of the quantum gates in the value network; the Hamiltonian of the value network is expressed as By applying the Hamiltonian to the quantum state |s> and performing quantum measurements, we obtain a value estimate, where the expected value of the measured quantum state |ψ> is expressed as: Hamiltonian Designed as a function of various operating indicators and incentive factors of the power distribution network in is the weight coefficient; It is used to measure the importance of the deviation of the power distribution network of the node from the rated voltage; Used to control the effect of line current on value; Reflects the impact of load power on value; c3. Strategy update: Use the policy gradient algorithm to update the parameters θ of the policy network; the policy gradient is expressed as: Where J(θ) is the objective function of the policy, which measures the performance of the policy; Q(|s>, |a>) is the action value function, which represents the value of taking action |a> in state |s>; The action-value function is estimated by the following Bellman formula: R(|s>, |a>) is the quantum reward function, which comprehensively considers various factors after the power distribution network takes action in the current state; γ is the discount factor, which is used to weigh the importance of future rewards and current rewards; It represents the expected value of the future state after considering the state transition probability P(|s′||s>, |a>); According to the calculation results of the policy gradient, by adjusting the parameter θ of the quantum rotation gate i To update the strategy, the gradient descent method is used to update the parameters: Where α is the learning rate, which determines the step size of each update; t is the number of iterations; in each iteration, the policy gradient is calculated based on the current policy network and value network, and then the parameters are updated to gradually optimize the strategy; c4. Quantum reward function design: The quantum reward function R(|s>, |a>) takes into account multiple factors of the power distribution network: Power outage losses: in is the load power of node i; S i is a binary variable indicating whether node i has power outage, S i =1 means there is no power outage, S i =0 means power outage; C outage is the unit outage loss cost, which means that when a node loses power, a power outage loss will occur, and the size of the loss is proportional to the load power and the unit outage loss cost; Voltage quality: Where V i is the actual voltage at node i, V ref is the rated voltage. This formula measures the sum of squared deviations of the node voltage from the rated voltage. The larger the deviation, the worse the voltage quality and the lower the reward value. Network loss: Among them I ij It is line e ij The current, R ij is the line resistance; the network loss is proportional to the square of the line current and the resistance. The greater the loss, the lower the reward value; The quantum reward function is expressed as R(|s>,|a>)=-ω1L outage -ω2V quality -ω3P loss , where ω1, ω2, and ω3 are weight coefficients used to balance the importance of different factors; After designing the quantum reinforcement learning algorithm and strategy optimization method, the quantum reinforcement learning algorithm is trained and the strategy is optimized in the following ways: Collect historical operation data of the power distribution network, including voltage, current and power in different seasons, time periods, load levels and fault conditions; Clean the historical operation data of the power distribution network to remove outliers and noise; Normalize the cleaned historical operation data and map it to the [0, 1] interval; The normalized historical running data is divided into a training set and a validation set; the training set is used to train the quantum reinforcement learning algorithm; the validation set is used to adjust the algorithm's hyperparameters, including the learning rate and quantum gate parameters; Use the training set to train the designed quantum reinforcement learning algorithm, and use the validation set to adjust the hyperparameters of the quantum reinforcement learning algorithm and perform strategy optimization at the same time.
6. A power distribution network fault recovery system based on reinforcement learning, characterized in that: The method for fault recovery of a power distribution network based on reinforcement learning according to any one of claims 1 to 5 is applied, comprising: A modeling module, used for modeling the power distribution network, and modeling the power distribution network as a weighted directed graph model; The first design module is used to design a data transmission protocol and result receiving and parsing method of a quantum-classical hybrid architecture for power distribution network fault recovery based on a weighted directed graph model of the power distribution network; wherein, in the classical computing part of the quantum-classical hybrid architecture, the power distribution network state data of the power distribution network processed by classical computing is encoded into the probability amplitude of quantum bits and transmitted to the quantum computing part, and the calculation results of the quantum computing part are converted into control instructions or decision results; in the quantum computing part of the quantum-classical hybrid architecture, a mapping relationship between quantum bits and power distribution network elements of the power distribution network is established, the power distribution network state of the power distribution network is represented by quantum states, and the evolution of the power distribution network state is simulated through quantum gate operations; The second design module is used to design a quantum reinforcement learning algorithm and strategy optimization for power distribution network fault recovery based on a weighted directed graph model of the power distribution network, and use the training data of the power distribution network for training; in which, the state and action in reinforcement learning are represented by quantum state, and quantum gate operations are used to realize the construction and update of the strategy network and value network of the quantum reinforcement learning algorithm, and the quantum reward function is designed; A computing module, which is used for, when a fault in the power distribution network is detected, obtaining the state data of the power distribution network after the fault according to the weighted directed graph model through the classical computing part of the quantum-classical hybrid architecture, and encoding the state data of the power distribution network after the fault in the power distribution network into the probability amplitude of quantum bits and transmitting it to the quantum computing part; The fault recovery module is used to calculate the optimal fault recovery strategy of the power distribution network through the quantum computing part using the designed quantum reinforcement learning algorithm and feed it back to the classical computing part, and convert the optimal fault recovery strategy into actual control instructions through the classical computing part to perform fault recovery work of the power distribution network.
7. A power distribution network fault recovery system based on reinforcement learning, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the power distribution network fault recovery method based on reinforcement learning as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Quantum and classic hybrid cloud platform and task execution method
CN112465146A
Power grid operation and maintenance diagnosis and overhaul device
CN113471876A