Power system PMU optimal configuration method, system, device and storage medium
Through reinforcement learning guidance search tree method, the PMU configuration sequence is optimized, and the problem of unreasonable PMU configuration in the power system is solved, and the complete obscurity and safety of the power system under limited PMU is achieved.
Patent Information
- Application Number
- CN202111387630.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The prior art is difficult to achieve complete observability of the power system under the case of limited PMUs, and lacks defense against cyber attacks, resulting in unreasonable PMU configuration.
Using a search tree method based on reinforcement learning, the reinforcement learning model is optimized through neural networks and PMU configuration, and the configuration sequence of PMU is optimized, and combined with the obscurity and security of the power system, we find the best PMU configuration solution.
In the case of insufficient PMU, more buses can be effectively protected, complete and observable of the power system, and improve the safety of the system to adapt to engineering reality.
Smart Images

Figure CN114358366B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical application field of power system measurement, and relates to a method, system, device and storage medium for optimizing the configuration of PMUs in a power system. Background Art
[0002] Currently, the power system is developing towards complex interconnection, and the state estimation of the power system highly depends on accurate and secure data. Traditional state estimation analyzes the data obtained by SCADA (Supervisory Control and Data Acquisition system), but its accuracy is low and it is vulnerable to cyberattacks, such as data integrity attacks. Therefore, PMUs (Phasor Measurement Units) are configured in the power system.
[0003] The PMU uses the GPS signal as the system clock signal, can synchronize the measurement data from different regions, and since the measurement data will be marked with GPS timestamps, the data provided by the PMU is difficult to be tampered with by malicious attackers. A single PMU can measure the node voltage of a bus and the measurement data of the branch current connected thereto in the power system. If PMUs are installed on each bus, obviously, the complete observability of the power system can be achieved. However, the cost of PMUs is high, and it is not realistic to comprehensively configure PMUs. Therefore, the problem of optimizing the configuration of PMUs with the constraint of achieving complete system observability and the goal of using fewer PMUs has been widely studied.
[0004] Currently, a large number of PMU optimization configuration methods have emerged, mainly divided into numerical algorithms and heuristic algorithms. However, few methods consider the PMU configuration for defending against cyberattacks. And most methods can only obtain the location of the PMU configuration. In fact, in the case of lack of funds, it is very difficult to configure PMUs at one time to achieve the complete observability of the system. Therefore, it is crucial to obtain a reasonable PMU configuration order to observe more buses with limited PMUs. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method, system, device and storage medium for optimizing the configuration of PMUs in a power system.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] In the first aspect of the present invention, a method for optimizing the configuration of PMUs in a power system includes the following steps:
[0008] S1: Obtain the initial state of PMU configuration in the power system. Using the initial state of PMU configuration as the root node, construct a search tree with PMU configuration states as nodes.
[0009] S2: Use the initial state of PMU configuration as the current state of PMU configuration.
[0010] S3: Obtain the node corresponding to the current state of PMU configuration in the search tree to get the current node. According to the current state of PMU configuration, through a preset neural network, select a child node from all the unexpanded child nodes of the current node in the search tree as the current action.
[0011] S4: According to the current state of PMU configuration and the current action, through a preset reinforcement learning model for PMU optimal configuration, obtain the updated state of PMU configuration and the reward of the current action, and combine the current state of PMU configuration, the current action, the reward of the current action, and the updated state of PMU configuration to obtain a training sample.
[0012] S5: Update the neural network parameters according to the training sample. Use the updated state of PMU configuration as the current state of PMU configuration, and repeat S3 - S4 until the updated state of PMU configuration reaches the preset end state of PMU configuration, and record the action sequence from the initial state of PMU configuration to the end state of PMU configuration.
[0013] S6: Repeat S2 - S5 until the preset number of repetitions or the action sequence is less than the preset length. Select the action sequence with the minimum length from all the action sequences to obtain the PMU optimal configuration scheme, and perform PMU optimal configuration for the power system according to the initial state of PMU configuration and the PMU optimal configuration scheme.
[0014] Optionally, the specific method for obtaining the initial state of PMU configuration in the power system is:
[0015] Select the bus with the least number of adjacent buses in the power system to obtain the vulnerable bus. Place the PMU on the adjacent bus of the vulnerable bus to obtain the initial state of PMU configuration in the power system.
[0016] Optionally, the specific method for selecting a child node from all the unexpanded child nodes of the current node in the search tree as the current action according to the current state of PMU configuration through a preset neural network is:
[0017] According to the current state of PMU configuration, through a preset neural network, obtain the predicted value of each child node among all the unexpanded child nodes of the current node in the search tree. Select the child node with the maximum predicted value from all the unexpanded child nodes of the current node in the search tree as the current action.
[0018] Optionally, when selecting the child node with the greatest predicted value as the current action from all the unexpanded child nodes of the current node in the search tree, the ε-greedy strategy is adopted for selection.
[0019] Optionally, the reward function of the PMU optimal configuration reinforcement learning model is:
[0020] r t+1 = (||s t+1 - s t ||0 - 1) × 10
[0021] where r t+1 is the current action reward, s t+1 is the PMU configuration update status, and s t is the current PMU configuration status.
[0022] Optionally, the neural network includes a target neural network and a main neural network. The specific method for updating the neural network parameters according to the training samples is as follows:
[0023] Based on the k-th training sample (s k , a k , r k , s k '), the target value y k is:
[0024]
[0025] where s k is the current PMU configuration status, a k is the current action, r k is the reward obtained for the current action, s k ' is the PMU configuration update status, q(s k ', a'; θ - ) is the target neural network function, a' is the action that maximizes the function q(s k ', a'; θ - ), and θ - is the target neural network parameter;
[0026] Based on the target value y k , the main neural network parameters are updated through the following formula:
[0027]
[0028] where θ t+1 is the updated main neural network parameter, θ t is the current main neural network parameter, τ is the update step size, L is the loss function, N is the number of training samples, is the gradient of the loss function L;
[0029] When the number of updates of the main neural network parameters reaches the preset number of times, assign the target neural network parameters to the main neural network parameters, and re-accumulate the number of updates of the main neural network parameters.
[0030] Optionally, the preset number of times is 200 times.
[0031] In the second aspect of the present invention, a power system PMU optimal configuration system is characterized by comprising:
[0032] An acquisition module, configured to acquire the initial state of the PMU configuration of the power system, and construct a search tree with the initial state of the PMU configuration as the root node and the PMU configuration state as the node;
[0033] An action determination module, configured to use the initial state of the PMU configuration as the current state of the PMU configuration, obtain the node corresponding to the current state of the PMU configuration in the search tree to obtain the current node; according to the current state of the PMU configuration, select a child node from all the unexpanded child nodes of the current node in the search tree as the current action through a preset neural network;
[0034] A reinforcement learning module, configured to obtain the updated state of the PMU configuration and the current action reward according to the current state of the PMU configuration and the current action through a preset PMU optimal configuration reinforcement learning model, and combine the current state of the PMU configuration, the current action, the current action reward, and the updated state of the PMU configuration to obtain a training sample;
[0035] A training module, configured to update the neural network parameters according to the training sample, use the updated state of the PMU configuration as the current state of the PMU configuration, repeatedly trigger the action determination module and the reinforcement learning module until the updated state of the PMU configuration reaches the preset end state of the PMU configuration, and record the action sequence from the initial state of the PMU configuration to the end state of the PMU configuration;
[0036] An optimal configuration module, configured to repeatedly trigger the action determination module, the reinforcement learning module, and the first repetition module until the preset number of repetitions or the action sequence is less than the preset length, select the action sequence with the minimum length from all the action sequences to obtain the PMU optimal configuration scheme, and perform power system PMU optimal configuration according to the initial state of the PMU configuration and the PMU optimal configuration scheme.
[0037] In the third aspect of the present invention, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned power system PMU optimal configuration method are implemented.
[0038] In the fourth aspect of the present invention, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned power system PMU optimal configuration method are implemented.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] In the power system PMU optimal configuration method of the present invention, by obtaining the initial state of the PMU configuration of the power system, then taking the initial state of the PMU configuration as the root node and the PMU configuration state as the nodes to construct a search tree, and further obtaining the node corresponding to the current state of the PMU configuration in the search tree to get the current node; according to the current state of the PMU configuration, through a preset neural network, select a child node from all the unexpanded child nodes of the current node in the search tree as the current action; then according to the current state of the PMU configuration and the current action, through a preset PMU optimal configuration reinforcement learning model, obtain the updated state of the PMU configuration and the current action reward, and combine the current state of the PMU configuration, the current action, the current action reward and the updated state of the PMU configuration to obtain a training sample, and update the neural network parameters according to the training sample; then repeat the above steps until the updated state of the PMU configuration reaches the preset end state of the PMU configuration, and record the action sequence from the initial state of the PMU configuration to the end state of the PMU configuration. On this basis, through several times of the above operations, several action sequences are obtained, and the action sequence with the minimum length is selected from all the action sequences to obtain the PMU optimal configuration scheme, and according to the initial state of the PMU configuration and the PMU optimal configuration scheme, the power system PMU is optimized and configured. On the basis of the initial state of the PMU configuration, in order to achieve the complete observability of the power system, a method based on reinforcement learning to guide the search tree is adopted to find the best PMU optimal configuration scheme, which comprehensively considers the observability and security of the power system, can guide the PMU optimal configuration in the case of insufficient PMUs, use limited PMUs to protect more buses, and fully consider the engineering practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow block diagram of the power system PMU optimal configuration method of the present invention;
[0042] Figure 2 It is a schematic structural diagram of the PMU optimal configuration reinforcement learning model of the present invention;
[0043] Figure 3 It is a schematic structural diagram of the neural network of the present invention;
[0044] Figure 4 It is a schematic structural diagram of the search tree of the present invention;
[0045] Figure 5Schematic diagram of the configuration scheme after the PMU optimal configuration method of the power system of the present invention is used on the IEEE-14 standard system;
[0046] Figure 6 Schematic diagram of the configuration scheme after the PMU optimal configuration method of the power system of the present invention is used on the IEEE-30 standard system;
[0047] Figure 7 Variation trend graph of the number of protected buses after the PMU optimal configuration method of the power system of the present invention is used on each standard system. Detailed implementation manners
[0048] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0051] See Figure 1 and 2 In an embodiment of the present invention, a PMU optimal configuration method for a power system is provided, including the following steps.
[0052] S1: Obtain the initial state of the PMU configuration of the power system, and construct a search tree with the initial state of the PMU configuration as the root node and the PMU configuration state as the node.
[0053] Preferably, the specific method for obtaining the initial state of the PMU configuration of the power system is as follows: select the bus with the least number of adjacent buses in the power system to obtain the vulnerable bus; place the PMU on the adjacent bus of the vulnerable bus to obtain the initial state of the PMU configuration of the power system.
[0054] Among them, the state vector s is defined to represent the observability of the power system. The i-th element s i = 1 of the vector s indicates that the state of bus i is observed by the PMU, and s i = 0 indicates that it is not observed. Therefore, the end state of the PMU configuration can be s = [1, 1,..., 1] T .
[0055] Define the action as selecting a bus to place the PMU. Set the binary PMU configuration vector P. The i-th element P i = 1 of the vector P indicates that the PMU is placed on bus i; on the contrary, P i = 0 indicates that the PMU is not placed on bus i. The action can be regarded as setting one element in the vector P to 1. Accordingly, the state can be calculated according to the following formula:
[0056] B = C × P
[0057]
[0058] Among them, the vector B represents the number of PMUs observing each bus. For example, the i-th element B i = 1 of the vector B indicates that bus i is observed by 1 PMU. The adjacency matrix C is calculated as follows:
[0059]
[0060] Among them, i, j = 1, 2,..., n, n is the number of buses, and N i is the set of buses adjacent to bus i.
[0061] Specifically, first assume that the cost paid by the attacker to tamper with each meter measurement value is the same. The injection power and power flow of all buses in the power system are measured by meters, and the attacker will only modify the state variable of one bus each time an attack is carried out. In the data integrity attack, in order to bypass the bad data detection of the system, the attacker needs to tamper with the injection power of the buses adjacent to the attacked bus and the power flow between the two. According to the assumption, in order to minimize the attack cost, the attacker will select the bus with the least number of adjacent buses. Therefore, for a general power system, the bus with the least number of adjacent buses, generally 1, should be selected as the vulnerable bus.
[0062] The PMU can protect the bus at its installation location and the buses adjacent to that bus. To protect the vulnerable bus, the PMU can be placed on the adjacent bus of the vulnerable bus.
[0063] See Figure 3 As shown in Figure 3 , the initial state of the PMU configuration will be used as the root node of the search tree. Since different states will be generated for each node due to actions, multiple child nodes will be generated. If the state of a node is the end state, that is, all nodes are protected, this node is considered a leaf node. Continuously explore and enrich the search tree. If all child nodes of a node are leaf nodes or fully expanded nodes, then this node is also a fully expanded node.
[0064] S2: Use the initial state of the PMU configuration as the current state of the PMU configuration.
[0065] S3: Obtain the node corresponding to the current state of the PMU configuration in the search tree to get the current node; according to the current state of the PMU configuration, through a preset neural network, select a child node from all the un-fully-expanded child nodes of the current node in the search tree as the current action.
[0066] Preferably, for effective exploration, the specific method of selecting a child node from all the un-fully-expanded child nodes of the current node in the search tree as the current action according to the current state of the PMU configuration through a preset neural network is: according to the current state of the PMU configuration, through a preset neural network, obtain the predicted values of each child node among all the un-fully-expanded child nodes of the current node in the search tree; select the child node with the maximum predicted value from all the un-fully-expanded child nodes of the current node in the search tree as the current action.
[0067] Specifically, when selecting the child node with the maximum predicted value from all the un-fully-expanded child nodes of the current node in the search tree as the current action, the ε-greedy strategy is adopted for selection.
[0068] Specifically, the neural network includes a target neural network and a main neural network. The target neural network is used to evaluate the predicted value of each child node, that is, the function action value function q(s, a; θ - ), s is the current state of the PMU configuration, θ - is the parameter of the target neural network. The main neural network is used to update the policy in real time. Its structure is the same as that of the target neural network, but the parameter θ is different. See Figure 4 , which shows the structures of the target neural network and the main neural network. After the current state of the PMU configuration is input into the target neural network, it first passes through a fully connected hidden layer:
[0069] w = f(M1 × s + p1)
[0070] where w is the output of the hidden layer, M1 is the weight of the first fully connected matrix, p1 is the first bias vector, and the function f(·) is the activation function. In this embodiment, the activation function adopts "RELU".
[0071] The hidden layer also outputs through a fully connected layer:
[0072] o = M2 × w + p2
[0073] where M2 is the weight of the second fully connected matrix and p2 is the second bias vector. The output o of this layer is the predicted value of the child node, that is, the value of each bus-configured PMU, and the parameters θ of the target neural network - include M1, M2, p1, and p2.
[0074] S4: According to the current state and current action of the PMU configuration, through a preset PMU optimal configuration reinforcement learning model, obtain the PMU configuration update state and the current action reward, and combine the current state of the PMU configuration, the current action, the current action reward, and the PMU configuration update state to obtain a training sample.
[0075] Among them, the reward function of the PMU optimal configuration reinforcement learning model is:
[0076] r t+1 = (||s t+1 - s t ||0 - 1) × 10
[0077] where r t+1 is the current action reward, s t+1 is the PMU configuration update state, and s t is the current state of the PMU configuration.
[0078] S5: Update the neural network parameters according to the training sample; use the PMU configuration update state as the current state of the PMU configuration, and repeat S3 - S4 until the PMU configuration update state reaches the preset PMU configuration end state, and record the action sequence from the initial state of the PMU configuration to the end state of the PMU configuration.
[0079] Among them, the specific method for updating the neural network parameters according to the training sample is:
[0080] According to the kth training sample (s k , a k , r k , s k '), obtain the target value y k as:
[0081]
[0082] where s k is the current state of the PMU configuration, a k is the current action, r k is the reward obtained by the current action, and s k'Configure the update status for the PMU, q(s k ', a'; θ - ) is the target neural network function, and a' is the action that makes the function q(s k ', a'; θ - ) achieve the maximum value, and θ - is the target neural network parameter.
[0083] According to the target value y k , update the main neural network parameters through the following formula:
[0084]
[0085] where θ t+1 is the updated main neural network parameter, θ t is the current main neural network parameter, τ is the update step size, L is the loss function, N is the number of training samples, is the gradient of the loss function L.
[0086] When the number of updates of the main neural network parameters reaches the preset number of times, generally 200 times, assign the target neural network parameters to the main neural network parameters and re-accumulate the number of updates of the main neural network parameters.
[0087] S6: Repeat S2 - S5 until the preset number of repetitions or the action sequence is less than the preset length. Select the action sequence with the minimum length from all the action sequences to obtain the PMU optimization configuration plan. According to the initial state of the PMU configuration and the PMU optimization configuration plan, perform the PMU optimization configuration of the power system.
[0088] In summary, for the power system PMU optimization configuration method of the present invention, first, considering the vulnerable buses in the power system that are prone to data integrity attacks, the PMU is used to give priority protection; second, in order to achieve the complete observability of the power system, the reinforcement learning-guided search tree method is applied to find the best PMU optimization configuration plan to achieve the goal. Considering the observability and security of the power system comprehensively, a reasonable PMU configuration plan can be finally obtained, which can guide the PMU configuration of the power system when the PMU is insufficient, protect more buses with limited PMUs, fully adapt to the engineering practice, and has better engineering realizability.
[0089] The following details a specific implementation step of the power system PMU optimization configuration method of the present invention, including a pre-configuration stage and a main-configuration stage.
[0090] I. Pre-configuration stage.
[0091] Step 11: Identify the power grid bus with the minimum attack cost under a data integrity attack. To minimize the attack cost, the attacker will select the bus with the fewest adjacent buses. In this embodiment, the bus with the fewest adjacent buses is the bus with 1 adjacent bus, i.e., the vulnerable bus.
[0092] Step 12: Place the PMU to protect the bus identified in Step 11. The PMU can protect the bus at its installation location and the buses adjacent to that bus. To protect a bus with 1 adjacent bus, the PMU can be placed on the adjacent bus of that bus.
[0093] II. Main configuration stage.
[0094] Step 21: Calculate the adjacency matrix of the power grid topology and initialize the parameters of the target network and the main network in the deep reinforcement learning method.
[0095] Step 22: Take the configuration state after the end of the pre-configuration stage as the initial state of the PMU configuration.
[0096] Step 23: Starting from the node corresponding to the current PMU configuration state in the search tree, use the ε-greedy strategy to select an uncompletely expanded child node as the next action.
[0097] Step 24: Execute the next action, obtain the new PMU configuration state and the reward value of the next action, and combine the current PMU configuration state, the next action, the reward value of the next action, and the new PMU configuration state into a training sample, and store it in the experience pool.
[0098] Step 25: Update the parameters of the main neural network by randomly extracting several training samples from the experience pool. Whenever the parameters of the main neural network are updated to a certain extent, in this embodiment, with the limit of 200 updates of the main neural network parameters, assign the parameters of the main neural network to the parameters of the target neural network.
[0099] Step 26: When the current PMU configuration state does not reach the set end state, in this embodiment, the set end state is [1, 1,..., 1] T , go back to Step 23.
[0100] Step 27: If the set end state is reached, record the action sequence from the initial state of the PMU configuration to the set end state.
[0101] Step 28: Repeat the above steps several times, which can be set according to the actual power system. Among the obtained several action sequences, select the action sequence with the minimum length to obtain the PMU optimal configuration scheme, and perform the PMU optimal configuration of the power system according to the initial state of the PMU configuration and the PMU optimal configuration scheme.
[0102] In another embodiment of the present invention, the power grid topologies of multiple IEEE standard systems are used to verify the above-mentioned PMU optimal configuration method for power systems, and the obtained configuration scheme is shown in Table 1.
[0103] Table 1 Configuration Scheme Table
[0104]
[0105]
[0106] See Figure 5 and 6 which show the configuration scheme in Table 1 on the IEEE-14 and IEEE-30 standard systems. The buses with shading in the figure are the buses where PMUs need to be placed, and the numbers marked beside the buses represent the sequence of PMU configuration. See Figure 7 which shows the change graph of the proportion of the total number of protected buses obtained according to the configuration scheme in Table 1 with the PMU configuration process. It can be seen that each curve is roughly convex upward, proving that the configuration actions with a higher configuration sequence can enable the PMU to observe more buses.
[0107] Therefore, the PMU optimal configuration scheme provided by the PMU optimal configuration method for power systems of the present invention can enable limited PMUs to monitor more buses in the case of insufficient PMUs. Thus, it can be seen that the PMU optimal configuration method for power systems of the present invention can provide a relatively reasonable scheme, verifying the effectiveness of the present invention.
[0108] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0109] In another embodiment of the present invention, a PMU optimal configuration system for power systems is provided, which can be used to implement the above-mentioned PMU optimal configuration method for power systems. Specifically, the PMU optimal configuration system for power systems includes an acquisition module, an action determination module, a reinforcement learning module, a training module, and an optimal configuration module.
[0110] Among them, the acquisition module is used to acquire the initial PMU configuration state of the power system, and construct a search tree with the initial PMU configuration state as the root node and the PMU configuration state as the nodes; the action determination module is used to use the initial PMU configuration state as the current PMU configuration state, obtain the node corresponding to the current PMU configuration state in the search tree, and get the current node; according to the current PMU configuration state, through a preset neural network, select a child node from all the unexpanded child nodes of the current node in the search tree as the current action; the reinforcement learning module is used to obtain the updated PMU configuration state and the current action reward through a preset PMU optimal configuration reinforcement learning model according to the current PMU configuration state and the current action, and combine the current PMU configuration state, the current action, the current action reward and the updated PMU configuration state to obtain a training sample; the training module is used to update the neural network parameters according to the training sample, use the updated PMU configuration state as the current PMU configuration state, and repeatedly trigger the action determination module and the reinforcement learning module until the updated PMU configuration state reaches the preset PMU configuration end state, and record the action sequence from the initial PMU configuration state to the PMU configuration end state; the optimal configuration module is used to repeatedly trigger the action determination module, the reinforcement learning module and the first repetition module until the preset number of repetitions or the action sequence is less than the preset length, select the action sequence with the minimum length from all the action sequences, obtain the PMU optimal configuration scheme, and perform PMU optimal configuration of the power system according to the initial PMU configuration state and the PMU optimal configuration scheme.
[0111] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the PMU optimal configuration method of the power system.
[0112] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one magnetic disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for optimizing the configuration of the power system PMU in the above embodiments.
[0113] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0114] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in Figure 1 one flow or multiple flows and / or blocksFigure 1 The functions specified in one or more boxes.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 or more boxes.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for optimizing the configuration of PMUs in a power system, characterized in that, It includes the following steps: S1: Select the bus with the least number of adjacent buses in the power system to obtain the vulnerable bus; Place the PMU on the adjacent bus of the vulnerable bus to obtain the initial state of the PMU configuration of the power system; Construct a search tree with the initial state of the PMU configuration as the root node and the PMU configuration state as the nodes; S2: Use the initial state of the PMU configuration as the current state of the PMU configuration; S3: Obtain the corresponding node of the current state of the PMU configuration in the search tree to get the current node; According to the current state of the PMU configuration, through a preset neural network, obtain the predicted value of each child node among all the unexpanded child nodes of the current node in the search tree; Select the child node with the largest predicted value as the current action from all the unexpanded child nodes of the current node in the search tree; S4: According to the current state of the PMU configuration and the current action, through a preset PMU optimal configuration reinforcement learning model, obtain the updated state of the PMU configuration and the current action reward, and combine the current state of the PMU configuration, the current action, the current action reward, and the updated state of the PMU configuration to obtain a training sample; S5: Update the neural network parameters according to the training sample; Use the updated state of the PMU configuration as the current state of the PMU configuration, repeat S3 - S4 until the updated state of the PMU configuration reaches the preset end state of the PMU configuration, and record the action sequence from the initial state of the PMU configuration to the end state of the PMU configuration; S6: Repeat S2 - S5 until the preset number of repetitions or the action sequence is less than the preset length, select the action sequence with the minimum length from all the action sequences to obtain the PMU optimal configuration scheme, and perform the PMU optimal configuration of the power system according to the initial state of the PMU configuration and the PMU optimal configuration scheme; The neural network includes a target neural network and a main neural network. The specific method for updating the neural network parameters according to the training sample is: According to the k-th training sample (s k , a k , r k , s k '), the target value y k is: Among them, s k is the current state of the PMU configuration, a k is the current action, r k is the reward obtained for the current action, s k ' is the updated state of the PMU configuration, q(s k ', a'; θ - ) is the target neural network function, a' is the action that makes the function q(s k ', a'; θ - ) achieve the maximum value, and θ - are the target neural network parameters; According to the target value y k , update the main neural network parameters by the following formula: where θ t+1 is the updated main neural network parameter, θ t is the current main neural network parameter, τ is the update step size, L is the loss function, N is the number of training samples, is the gradient of the loss function L; When the number of updates of the main neural network parameters reaches the preset number of times, assign the target neural network parameters to the main neural network parameters and re - accumulate the number of updates of the main neural network parameters.
2. The method for optimizing the configuration of PMU in the power system according to claim 1, wherein When selecting the child node with the largest predicted value as the current action from all the unexpanded child nodes of the current node in the search tree, the ε - greedy strategy is used for selection.
3. The power system PMU optimal configuration method according to claim 1, wherein The reward function of the PMU optimal configuration reinforcement learning model is: r t+1 = (||s t+1 - s t ||0 - 1) × 10 Among them, r t+1 is the current action reward, s t+1 is the PMU configuration update status, s t is the current status of the PMU configuration.
4. The power system PMU optimal configuration method according to claim 1, characterized in that The preset number of times is 200 times.
5. A PMU optimal configuration system for a power system, characterized in that, It includes: An acquisition module, used to select the bus with the least number of adjacent buses in the power system to obtain the vulnerable bus; Place the PMU on the adjacent bus of the vulnerable bus to obtain the initial state of the PMU configuration of the power system; Construct a search tree with the initial state of the PMU configuration as the root node and the PMU configuration state as the nodes; An action determination module, configured to use the initial state of the PMU configuration as the current state of the PMU configuration, obtain the node corresponding to the current state of the PMU configuration in the search tree, and obtain the current node; according to the current state of the PMU configuration, through a preset neural network, obtain the predicted values of each child node among all the unexpanded child nodes of the current node in the search tree; select the child node with the largest predicted value as the current action from all the unexpanded child nodes of the current node in the search tree; A reinforcement learning module, configured to obtain the updated state of the PMU configuration and the current action reward according to the current state of the PMU configuration and the current action through a preset PMU optimal configuration reinforcement learning model, and combine the current state of the PMU configuration, the current action, the current action reward, and the updated state of the PMU configuration to obtain a training sample; A training module, configured to update the neural network parameters according to the training sample, use the updated state of the PMU configuration as the current state of the PMU configuration, repeatedly trigger the action determination module and the reinforcement learning module until the updated state of the PMU configuration reaches a preset end state of the PMU configuration, and record the action sequence from the initial state of the PMU configuration to the end state of the PMU configuration; An optimal configuration module, configured to repeatedly trigger the action determination module, the reinforcement learning module, and the first repetition module until a preset number of repetitions or the action sequence is less than a preset length, select the action sequence with the minimum length from all the action sequences to obtain a PMU optimal configuration plan, and perform PMU optimal configuration of the power system according to the initial state of the PMU configuration and the PMU optimal configuration plan; The neural network includes a target neural network and a main neural network. The specific method for updating the neural network parameters according to the training sample is as follows: According to the k-th training sample (s k , a k , r k , s k '), the target value y k is as follows: where s k is the current state of the PMU configuration, a k is the current action, r k is the reward obtained for the current action, s k ' is the updated state of the PMU configuration, q(s k ', a'; θ - ) is the target neural network function, a' is the action that makes the function q(s k ', a'; θ - ) achieve the maximum value, and θ - are the target neural network parameters; According to the target value y k , update the main neural network parameters by the following formula: where, θ t+1 is the updated main neural network parameter, θ t is the current main neural network parameter, τ is the update step size, L is the loss function, N is the number of training samples, is the gradient of the loss function L; When the number of times of updating the main neural network parameters reaches a preset number of times, assign the target neural network parameters to the main neural network parameters and re-accumulate the number of times of updating the main neural network parameters.
6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the power system PMU optimal configuration method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the power system PMU optimal configuration method according to any one of claims 1 to 4.