Power distribution network single-phase earth fault line selection decision-making method and system based on deep reinforcement learning, and related equipment
Through the deep reinforcement learning model combined with the dynamic graph convolutional neural network, the problem of low decision-making reliability in single-phase grounding fault line selection and extraction of distribution network is solved, and efficient, accurate and rapid recovery of line selection and extraction is achieved, improving the safety and reliability of distribution network.
Patent Information
- Application Number
- CN202510754653.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing single-phase grounding fault line selection and pulling method of distribution network fails to effectively combine line operation conditions, load types and line selection and pulling order tables, resulting in low decision-making reliability and deep learning methods fail to fully consider expert experience and technical analysis.
The deep reinforcement learning model is adopted, and by constructing the distribution network feeder switch adjacency matrix and feature matrix, combining dynamic graph convolutional neural network and deep reinforcement learning, fault line selection and pulling strategies are generated to achieve efficient fusion of expert experience and technical analysis.
It improves the accuracy and rapid recovery capability of single-phase grounding fault line selection in the distribution network, and improves the safety, reliability and scientific scheduling and operation decisions of the distribution network.
Smart Images

Figure CN120296542A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of power system research, and specifically relates to a method, system and related equipment for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning. Background Technique
[0002] The distribution network is an important part of the power system, which is used to transmit electric energy from the power plant to the end users and provide power supply for various electrical equipment. It is the last link in the power system and is responsible for converting the electric energy transmitted by the high-voltage transmission line into low-voltage electric energy suitable for household, industrial and commercial use. However, during long-term operation, the distribution network faces various potential fault risks, such as grounding faults, bus faults and distribution line faults. These faults may cause power outages, equipment damage and even fire accidents. Therefore, it is crucial to analyze and solve distribution network faults in a timely and effective manner to ensure the safe and stable operation of the power grid.
[0003] The problem of selecting and pulling a single-phase grounding fault line in a distribution network is one of the important research topics in the construction and development of the distribution network. From the traditional method of judging the fault line based on manual experience and expert knowledge base to the strategy of selecting the fault line by using technologies such as multi-source information fusion and deep learning, the method of selecting and pulling a single-phase grounding fault line in a distribution network is gradually changing from manual experience judgment to intelligent analysis and selection. The selection and pulling of a single-phase grounding fault line in a distribution network not only depends on a scientific, reasonable and reliable fault line selection strategy, but also has a close relationship with the different operating conditions of the line. At present, the method of selecting and pulling a single-phase grounding fault line in a distribution network mainly uses the line selection mechanism of the fault line to achieve, realizes the probability statistical analysis of the fault line through the comprehensive judgment of various electrical quantity information of the distribution network line, and provides auxiliary decision-making for the power grid dispatching and operation personnel according to the fault probability of each line. When selecting and pulling a single-phase grounding fault line in a distribution network, the power grid dispatching and operation personnel not only refer to the fault analysis results provided by the line selection mechanism, but also comprehensively consider the line operating conditions, the types of loads supplied and the line selection and pulling sequence table, etc. to make the final decision.
[0004] Common line selection methods mainly include fault electrical quantity information analysis, multi-information fusion line selection, and line selection based on deep learning methods. Fault electrical quantity information analysis mainly analyzes and selects fault lines based on fault steady-state information or fault transient information. Multi-information fusion line selection usually uses the change characteristics of various electrical quantity information after a fault occurs to effectively identify fault lines, and at the same time uses information intelligent fusion methods such as D-S evidence theory, genetic algorithm, and fuzzy control theory for comprehensive decision-making and judgment. The line selection mechanism based on deep learning methods uses artificial neural networks, deep belief networks, etc. to replace information intelligent fusion methods, learns and identifies the characteristics of various electrical quantity information of single-phase grounding faults in the distribution network, and realizes the analysis and judgment of fault lines. The line selection mechanism for single-phase grounding faults in the distribution network only selects fault lines based on the change characteristics of fault electrical quantity information, and does not involve the comprehensive and complex situations such as line operating conditions, load types supplied, and line selection and pulling sequence tables in the disposal of single-phase grounding faults in the distribution network, and cannot truly combine artificial experience and technical analysis.
[0005] The analysis of single-phase grounding fault line selection in the distribution network based on deep learning has strong characterization and learning performance, avoiding the difficulty of analyzing the change of fault electrical quantities in the distribution network or the difficulty of line selection caused by different grounding operating modes of the distribution network by information intelligent fusion methods such as D-S evidence theory and genetic algorithm, resulting in low reliability and affecting auxiliary decision-making. However, the biggest feature of deep learning is to achieve accurate feature extraction and recognition, and the decision-making for the disposal of single-phase grounding faults in the distribution network needs to be realized by the machine network through simulating the interaction between the dispatcher and the environment. Therefore, reinforcement learning becomes the primary choice. Combining deep learning and reinforcement learning, a deep reinforcement learning model is used to model the control decision-making for the selection and pulling of single-phase grounding fault lines in the distribution network, realizing the efficient integration of expert experience and technical analysis. Summary of the Invention
[0006] To solve one of the above technical defects, the present application provides a method, system and related equipment for decision-making on the selection and pulling of single-phase grounding fault lines in the distribution network based on deep reinforcement learning.
[0007] According to the first aspect of the present application, a method for decision-making on the selection and pulling of single-phase grounding fault lines in the distribution network is provided, which is used to select the line where a single-phase grounding fault occurs in the distribution network, and includes: Obtain the topological connection relationship of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network, and construct an adjacency matrix of the feeder switches in the distribution network G ; According to the collected characteristic data of the electrical quantity information of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network, construct a characteristic matrix of the feeder switches in the distribution network ; where the characteristic data of the electrical quantity information includes: sudden change of active power, sudden change of reactive power, sudden change of current, and zero-sequence current. Using a dynamic graph convolutional neural network to process the adjacency matrix of the feeder switches in the distribution network G and the feature matrix of the feeder switches in the distribution network to obtain the state information feature quantity; Based on the fault line selection and pulling order, generate a fault line selection and pulling strategy; According to the state information feature quantity and the fault line selection and pulling strategy, based on deep reinforcement learning, construct a decision-making model for selecting and pulling the single-phase grounding fault line in the distribution network, which is used to realize the analysis of selecting and pulling the single-phase grounding fault line in the distribution network, and determine the single-phase grounding fault line selection in the distribution network according to the analysis result.
[0008] Preferably, based on the collected feature data of the electrical quantity information of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network, construct the feature matrix of the feeder switches in the distribution network , specifically including: Collect the feature data of the electrical quantity information of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network, and construct the feeder switch matrix of the distribution network at the corresponding moment; the feeder switch matrix of the distribution network at time t' is: ; In the formula, is the feature data of the electrical quantity information of the nth feeder switch at time t', n is the number of feeder switches in the distribution network, is the sudden change in active power of the nth feeder switch at time t', is the sudden change in reactive power of the nth feeder switch at time t', is the sudden change in current of the nth feeder switch at time t', is the zero-sequence current of the nth feeder switch at time t'; Perform normalization calculation on the feeder switch matrix of the distribution network at the corresponding moment to obtain the feature matrix of the feeder switches in the distribution network; the feature matrix of the feeder switches in the distribution network at time t' is: ; In the formula, is the normalized feature data of the electrical quantity information of the nth feeder switch at time t', is the normalized sudden change in active power of the nth feeder switch at time t', is the normalized sudden change in reactive power of the nth feeder switch at time t', is the normalized sudden change in current of the nth feeder switch at time t', is the normalized zero-sequence current of the nth feeder switch at time t'; Construct the feature matrix of the feeder switches in the distribution network : , where is the distribution network feeder switch feature matrix at time N, .
[0009] Preferably, the dynamic graph convolutional neural network includes an input layer, two hidden layers and an output layer; the two hidden layers are a dynamic graph convolutional network layer and a graph convolutional network layer respectively; The dynamic graph convolutional neural network is used to process the distribution network feeder switch adjacency matrix G and the distribution network feeder switch feature matrix to obtain state information feature quantities, specifically including: Input the distribution network feeder switch feature matrix and the distribution network feeder switch adjacency matrix G into the input layer of the dynamic graph convolutional neural network; among them, the number of neurons in the input layer is 4; Input the distribution network feeder switch feature matrix and the distribution network feeder switch adjacency matrix G in the input layer into the dynamic graph convolutional network layer of the hidden layer. First, extract the dynamic relationships of each feeder switch at each time in the distribution network feeder switch feature matrix and the distribution network feeder switch adjacency matrix G through the self-attention mechanism, and then use the dynamic graph convolutional network to calculate the extracted dynamic relationships to obtain the feeder switch feature information quantity, and extract the feeder switch feature information quantity; Input the extracted feeder switch feature information quantity into the graph convolutional network layer of the hidden layer, use the graph convolutional network to calculate the feeder switch feature information quantity to obtain the state information feature quantity, and extract the state information feature quantity; Output the extracted state information feature quantity through the output layer.
[0010] Preferably, the calculation formula of the self-attention mechanism is: ; where M is the weight matrix after the self-attention mechanism calculation, is the normalization function, is the transposed matrix of the distribution network feeder switch feature matrix , D is the degree matrix of the distribution network feeder switch adjacency matrix B ; The calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is: ; where H is the feeder switch feature information quantity, is the activation function, is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product; The calculation formula of the graph convolutional network in the dynamic graph convolutional network layer is: ; In the formula: I is the state information feature quantity, is the weight matrix of the graph convolutional network layer.
[0011] Preferably, generating a fault line selection and switching strategy based on the fault line selection and switching order specifically includes: Generating a fault line selection and switching table based on the fault line selection and switching order: Among them, the fault line selection and switching order is: the switch of the no-load line (branch line), the double-power feed-in switch, the reactive power compensation device and the station service transformer; the selection and switching order of the double-power feed-in switch is: the double-power feed-in of the general user substation, the double-power feed-in of the general public substation, the double-power feed-in of the high-risk user substation and the double-power feed-in of the key public substation; Obtaining the fault phase voltage of the single-phase grounding fault of the distribution network; According to the selection and switching order in the fault line selection and switching table, perform selection and switching on each switch at the same time, and judge whether the fault phase voltage is restored after selecting and switching the corresponding switch; If it is restored, obtain the single-phase grounding fault line of the distribution network and end the selection and switching; Otherwise, continue to perform the next moment's selection and switching on each switch according to this selection and switching order until the single-phase grounding fault line of the distribution network is obtained and the selection and switching ends.
[0012] Preferably, constructing a single-phase grounding fault line selection and switching decision model for the distribution network based on the state information feature quantity and the fault line selection and switching strategy specifically includes: Constructing a deep reinforcement learning network and an experience pool; among them, the deep reinforcement learning network includes a training Q network and a target Q network; Initializing the training Q network and the target Q network, and assigning values to the parameters in the training Q network and the target Q network; According to the extracted state information feature quantity, constructing an initial state space variable , and performing an observation update of the initial state space variable once per cycle; where the cycle is the selection and switching control cycle, , is the state information feature quantity at the t-th moment of each cycle; According to the opening and closing states of each feeder switch, obtaining the corresponding action , and based on each obtained action , form the corresponding action set according to the faulty line selection and switching strategy A ; among them, , , is the action at the t-th moment of each cycle, is the opening / closing state of the i-th feeder switch at the t-th moment of each cycle, and n is the number of feeder switches in the distribution network; Construct the reward function r; Input the initial state space variables of each cycle into the training Q network, and iteratively execute the following steps: According to the strategy, iteratively select an action A from the corresponding action set , generate the action value corresponding to the initial state, and input the selected action into the faulty line selection and switching strategy to execute the selection and switching, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the immediate reward value of this action ; among them, is the neural network weight of the training Q network; Based on the SCADA system, according to the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, obtain the new state information characteristic quantity, and update and calculate the state space variables at the next moment of the corresponding cycle ; Construct the experience data set , and store it in the experience pool; Randomly extract d groups of experience data sets from the experience pool, and input them into the training Q network and the target Q network respectively, and calculate the value of the training Q network and the target value of the target Q network; among them, d < d m , d m is the data capacity of the experience pool; According to the loss function, update the neural network weights of the training Q network with the target value of the target Q network, so that the value of the training Q network is closer to the target value of the target Q network; Judge whether the current iteration number is equal to the maximum iteration number; If it is equal, obtain the trained single-phase grounding fault line selection and switching strategy model of the distribution network; If it is not equal, continue to execute the next iteration.
[0013] Preferably, the selection of an action from the action set A according to the strategy specifically includes: With Select the optimal action corresponding to the maximum Q' value from the action set A with a probability of, in order to Select a random action from the action set A with a probability of; The calculation formula of the policy is: ; In the formula, is the policy function, is the random selection probability and ε < 1, argmaxQ'(a, s) is the maximum Q' value corresponding to the optimal action a, and |A| is the number of actions in the action set.
[0014] Preferably, the calculation formula of the reward function is: ; In the formula, is the fault phase voltage penalty term, is the substation load loss penalty term, is the reactive power imbalance penalty term; Among them, The expression of is: ; In the formula, is the threshold voltage, is the fault phase voltage during single-phase grounding fault of the distribution network, and d is the penalty term; The expression of is: ; In the formula, is the load loss of the general user substation, is the load loss of the general public substation, is the load loss of the high-risk user substation, is the load loss of the key public substation, m 1 ~ m 4 are the general user substation, general public substation, high-risk user substation, and key public substation respectively, k 1 ~ k 4 are the load weight parameters of the corresponding substations respectively, and k 4 >k 3 >k 2 >k 1 ,, l1 ~ l 4 are respectively the sum of the numbers of corresponding substations or corresponding users; The expression of ; In the formula, 、 are respectively the upper limit value and the lower limit value of the reactive power under the normal operation state of the distribution network system; is the reactive power of the distribution network system at time t, and c is the penalty value.
[0015] The calculation formula of the target Q network described above is: ; In the formula, is the target value of the target Q network, h is the current iteration number, T is the maximum iteration number, γ is the discount factor, arg max Q '( a h , s h ) is the maximum Q' value corresponding to the optimal action selected under the current iteration; The update formula of the neural network weights of the training Q network described above is: ; In the formula, is the loss function, is the neural network weights of the training Q network updated at the h -th iteration; is the value of the training Q network at the h -th iteration.
[0016] According to the second aspect of the present application, there is provided a distribution network single-phase grounding fault line selection and pulling decision-making system based on deep reinforcement learning, including a module for implementing the distribution network single-phase grounding fault line selection and pulling decision-making method as described above.
[0017] According to the third aspect of the present application, there is provided a device, including: a memory; a processor; and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the distribution network single-phase grounding fault line selection and pulling decision-making method as described above.
[0018] According to the fourth aspect of the present application, there is provided a computer-readable storage medium, characterized in that a computer program is stored thereon; the computer program is executed by a processor to implement the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning as described above.
[0019] In the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning provided in the present application, by constructing a characteristic matrix of feeder switches in the distribution network , the deficiencies of the traditional matrix are improved, and the constructed adjacency matrix of feeder switches in the distribution network G is taken as the object, which fully adapts to the establishment of the single-phase grounding fault line selection and pulling decision model in the distribution network. At the same time, four variables, namely, the sudden change in active power, the sudden change in reactive power, the sudden change in current, and the zero-sequence current of each feeder switch, are selected as the characteristics of the feeder switch, providing reliable and accurate information support for the selection and pulling of single-phase grounding fault lines in the distribution network; the present application uses a dynamic graph convolutional neural network to dynamically calculate and extract the operation topology and electrical quantity information characteristic data in the adjacency matrix G of the distribution network feeder switches and the characteristic matrix of the distribution network feeder switches, realizing the real-time calculation of the operation topology and electrical quantity information characteristic data during the single-phase grounding selection and pulling process in the distribution network, so as to adapt to the changes in the distribution network topology and power flow information caused by the dynamic selection and pulling of single-phase grounding fault lines in the distribution network, and improving the adaptability and accuracy of the single-phase grounding fault line selection and pulling decision model in the distribution network; at the same time, the present application adopts a complete and practical fault line selection and pulling strategy, which plays a scientific, reasonable, and effective role in the construction of the model, improving the practicability and universality of the single-phase grounding fault line selection and pulling decision model in the distribution network and the accuracy of this method; in addition, according to the state information characteristic quantities identified by the dynamic graph neural network and the fault line selection and pulling strategy, a control model for selecting and pulling a single-phase grounding fault line in the distribution network is built based on the deep reinforcement learning model, which not only realizes the comprehensive analysis and judgment of the selection and pulling of single-phase grounding fault lines in the distribution network, but also realizes the efficient integration of expert experience and technical analysis, improving the rapid recovery ability under single-phase grounding faults in the distribution network, and providing an efficient guarantee for the safe, reliable, scientific, and accurate dispatching operation decision of the distribution network.
[0020] Other features and advantages of the present application will be described in the following description, and some of them will become obvious from the description or be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the content pointed out in the written description and the drawings. Description of the Drawings
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 Schematic flow diagram of a single-phase grounding fault line selection and switching decision method for a distribution network based on deep reinforcement learning provided by an embodiment of the present application; Figure 2 is Figure 1 Schematic flow diagram of the process of processing the distribution network feeder switch adjacency matrix and the distribution network feeder switch feature matrix using a dynamic graph convolutional neural network in the embodiment; Figure 3 is Figure 1 Schematic flow diagram of the process of executing the fault line selection and switching strategy in the embodiment; Figure 4 is Figure 1 Schematic flow diagram of the process of obtaining a trained single-phase grounding fault line selection and switching strategy model for the distribution network in the embodiment; Figure 5 Schematic functional structure diagram of a single-phase grounding fault line selection and switching decision system for a distribution network based on deep reinforcement learning provided by an embodiment of the present application; Figure 6 is Figure 5 Schematic functional structure diagram of the second matrix construction module in the embodiment; Figure 7 is Figure 5 Schematic functional structure diagram of the state information feature quantity acquisition module in the embodiment; Figure 8 is Figure 5 Schematic functional structure diagram of the fault line selection and switching strategy generation module in the embodiment; Figure 9 is Figure 5 Schematic functional structure diagram of the model construction module in the embodiment; In the figure: 10 is a topological connection relationship acquisition module, 20 is a first matrix construction module, 30 is a data acquisition module, 40 is a second matrix construction module, 50 is a state information feature quantity acquisition module, 60 is a faulty line selection and switching strategy generation module, 70 is a model construction module, 401 is a distribution network feeder switch matrix construction unit, 402 is a first calculation unit, 403 is a distribution network feeder switch feature matrix construction unit, 501 is a first input unit, 502 is a first output unit, 503 is a second input unit, 504 is a self-attention mechanism extraction unit, 505 is a second calculation unit, 506 is a first extraction unit, 507 is a third input unit, 508 is a third calculation unit, 509 is a second extraction unit, 510 is a fourth input unit, 511 is a second output unit, 601 is a faulty line selection and switching table generation unit, 602 is a faulty phase voltage acquisition unit, 603 is a selection and switching unit, 604 is a first judgment unit, 605 is a faulty line acquisition unit, 606 is a first execution unit, 701 is a first construction unit, 702 is an initialization and assignment unit, 703 is a second construction unit, 704 is an initial state space variable update unit, 705 is an action acquisition unit, 706 is a third construction unit, 707 is a fourth construction unit, 708 is a fifth input unit, 709 is an action selection unit, 710 is a data generation unit, 711 is a data update unit, 712 is a reward value acquisition unit, 713 is a data acquisition unit, 714 is a state update unit, 715 is a fifth construction unit, 716 is an experience data set storage unit, 717 is an experience data set extraction unit, 718 is a sixth input unit, 719 is a fourth calculation unit, 720 is a neural network weight update unit, 721 is a second judgment unit, 722 is a model acquisition unit, 723 is a third execution unit. Specific implementation manners
[0022] In order to make the technical solutions and advantages in the embodiments of the present application clearer and more understandable, the following further describes the exemplary embodiments of the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0023] Regarding some problems existing in the prior art: In a first aspect, an embodiment of the present application provides a method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning. This method can be executed by a system for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning, or by components configured inside the system for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning, such as chips, chip systems, etc. It can also be implemented by a logic module or software with some or all of the functions of the system for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning. The present application does not limit this.
[0024] Exemplarily, as Figure 1 shown, the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning is used to select the line where a single-phase grounding fault occurs in the distribution network, and includes: Obtain the topological connection relationship of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network, and construct an adjacency matrix of the feeder switches in the distribution network G ; where the adjacency matrix of the feeder switches in the distribution network G is used to reflect the structure of the distribution network, G is an n×n matrix, and n is the number of feeder switches in the distribution network; in the adjacency matrix of the feeder switches in the distribution network G , if there is a connection between feeder switch i and feeder switch j, then the jth element in the ith row of the adjacency matrix of the feeder switches in the distribution network G is 1, otherwise it is 0; According to the characteristic data of the electrical quantity information of each feeder switch at each moment during the operation of the single-phase grounding fault in the distribution network collected, construct a feature matrix of the feeder switches in the distribution network ; where the characteristic data of the electrical quantity information includes: sudden change in active power, sudden change in reactive power, sudden change in current, and zero-sequence current; Use a dynamic graph convolutional neural network to process the adjacency matrix of the feeder switches in the distribution network G and the feature matrix of the feeder switches in the distribution network to obtain the state information feature quantity; Generate a fault line selection and pulling strategy based on the order of fault line selection and pulling; Based on the state information feature quantity and the fault line selection and pulling strategy, construct a decision-making model for selecting and pulling a single-phase grounding fault line in the distribution network based on deep reinforcement learning, which is used to realize the analysis of selecting and pulling a single-phase grounding fault line in the distribution network, and determine the single-phase grounding fault line selection in the distribution network according to the analysis result.
[0025] Based on the above solution, in the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning provided in the present application, by constructing a feature matrix of the feeder switches in the distribution network , the deficiencies of the traditional matrix are improved, and with the constructed adjacency matrix of the feeder switches in the distribution networkG Taking it as the object, it fully adapts to the establishment of the single-phase grounding fault line selection and tripping decision-making model of the distribution network. At the same time, four variables, namely the sudden change of active power, the sudden change of reactive power, the sudden change of current, and the zero-sequence current of each feeder switch, are selected as the characteristics of the feeder switch, providing reliable and accurate information support for the selection and tripping of single-phase grounding fault lines in the distribution network. This application uses a dynamic graph convolutional neural network for the adjacent matrix of distribution network feeder switches G and the characteristic matrix of distribution network feeder switches to perform dynamic calculations and feature extraction on the operation topology and electrical quantity information feature data, realizing the real-time calculation of the operation topology and electrical quantity information feature data during the single-phase grounding selection and tripping process of the distribution network, so as to adapt to the changes in the distribution network topology and power flow information caused by the dynamic selection and tripping of single-phase grounding fault lines in the distribution network, and improving the adaptability and accuracy of the single-phase grounding fault line selection and tripping decision-making model of the distribution network. At the same time, this application adopts a complete and practical fault line selection and tripping strategy, which plays a scientific, reasonable, and effective role in the construction of the model, improving the practicability and universality of the single-phase grounding fault line selection and tripping decision-making model of the distribution network, and improving the accuracy of this method. In addition, according to the state information feature quantities identified by the dynamic graph neural network and the fault line selection and tripping strategy, this application conducts control modeling for the selection and tripping of single-phase grounding fault lines in the distribution network based on a deep reinforcement learning model, not only realizing the comprehensive analysis and judgment of the selection and tripping of single-phase grounding fault lines in the distribution network, but also realizing the efficient integration of expert experience and technical analysis, improving the rapid recovery ability under single-phase grounding faults in the distribution network, and providing an efficient guarantee for the safe, reliable, scientific, and accurate dispatching operation decision-making of the distribution network.
[0026] In some possible implementation manners of the first aspect, constructing the characteristic matrix of distribution network feeder switches according to the electrical quantity information feature data of each feeder switch at each moment during the operation of the single-phase grounding fault of the distribution network , specifically includes: Collect the electrical quantity information feature data of each feeder switch at each moment during the operation of the single-phase grounding fault of the distribution network, and construct the distribution network feeder switch matrix at the corresponding moment; the distribution network feeder switch matrix at time t' is: ; In the formula, is the electrical quantity information feature data of the nth feeder switch at time t', n is the number of distribution network feeder switches, is the sudden change of active power of the nth feeder switch at time t', is the sudden change of reactive power of the nth feeder switch at time t', is the sudden change of current of the nth feeder switch at time t', is the zero-sequence current of the nth feeder switch at time t'; Normalize the distribution network feeder switch matrix at the corresponding moment according to the formula to obtain the distribution network feeder switch feature matrix at the corresponding moment; the distribution network feeder switch feature matrix at time t' is: ; In the formula, is the electrical quantity information feature data of the nth feeder switch normalized at time t', is the sudden change in active power of the nth feeder switch normalized at time t', is the sudden change in reactive power of the nth feeder switch normalized at time t', is the sudden change in current of the nth feeder switch normalized at time t', is the zero-sequence current of the nth feeder switch normalized at time t'.
[0027] Based on the above scheme, each data in the distribution network feeder switch feature matrix after normalization is within the range of [-1, 1], which can clearly reflect the change degree of the electrical quantity information of the feeder switch and eliminate the adverse effects caused by singular data.
[0028] Specifically, the calculation formula for the sudden change in active power is: ; In the formula, is the sudden change in active power of the ith feeder switch at time t', is the active power of the ith feeder switch at time t', is the active power of the ith feeder switch corresponding to the normal operation of the distribution network without failure at time t'; this variable has positive and negative values. A positive value indicates that the feeder switch has power flowing out and is a power source side switch, while a negative value indicates that the feeder switch has power flowing in and is a load side switch.
[0029] The calculation formula for the sudden change in reactive power is: ; In the formula, is the sudden change in reactive power of the ith feeder switch at time t', is the reactive power of the ith feeder switch at time t', is the reactive power of the ith feeder switch corresponding to the normal operation of the distribution network without failure at time t'; this variable has positive and negative values. A positive value indicates that the feeder switch has power flowing out and is a power source side switch, while a negative value indicates that the feeder switch has power flowing in and is a load side switch.
[0030] The calculation formula for the sudden change in current is: ; In the formula, is the sudden change in current of the ith feeder switch at time t', is the current of the ith feeder switch at time t', ' The current of the i-th feeder switch corresponding to the normal operation of the distribution network without a fault at time t'; this variable has positive and negative values. A positive value indicates that the current of the feeder switch flows out and it is a power-side switch; a negative value indicates that the current of the feeder switch flows in and it is a load-side switch.
[0031] The calculation formula for the zero-sequence current is: , ; In the formula, is the zero-sequence current; is the system angular frequency, , is the system frequency; is the total capacitance to the ground of each component in the distribution network, is the zero-sequence voltage, and are the voltages of the three phases A', B', and C' after a fault in the distribution network system , and The sum. In the distribution network, an insulation monitoring device is usually configured to monitor the zero-sequence voltage. If conditions permit, a zero-sequence protection can be configured to obtain the zero-sequence current, so as to achieve the calculation and extraction of the zero-sequence current. Different from the power and current mutation amounts, the zero-sequence current is a vector, including the zero-sequence current amplitude and phase angle, so as to effectively identify the electrical fault characteristics in the operation mode of the distribution network grounded through an arc suppression coil.
[0032] Optionally, the normalization calculation of the distribution network feeder switch matrix at the corresponding moment is performed to obtain the distribution network feeder switch feature matrix at the corresponding moment, specifically: According to the formula Normalization calculation is performed on each element in the distribution network feeder switch matrix at the corresponding moment to obtain the corresponding normalized element; the element is the active power mutation amount or reactive power mutation amount or current mutation amount or zero-sequence current of each feeder switch. In the formula, is the k th element in the distribution network feeder switch matrix at the corresponding moment, , are the maximum and minimum values of this element respectively, is the k th element after normalization processing.
[0033] In some possible implementation manners of the first aspect, as Figure 2 shown, the dynamic graph convolutional neural network includes an input layer, two hidden layers, and an output layer; the two hidden layers are a dynamic graph convolutional network layer and a graph convolutional network layer respectively; The use of the dynamic graph convolutional neural network for the distribution network feeder switch adjacency matrix GAnd the characteristic matrix of the distribution network feeder switch is processed to obtain the characteristic quantities of the state information, specifically including: The characteristic matrix of the distribution network feeder switch and the adjacency matrix of the distribution network feeder switch G are input into the input layer of the dynamic graph convolutional neural network; among them, the number of neurons in the input layer is 4; The characteristic matrix of the distribution network feeder switch in the input layer and the adjacency matrix of the distribution network feeder switch G are input into the dynamic graph convolutional network layer of the hidden layer. First, the dynamic relationship of each feeder switch at each moment in the characteristic matrix of the distribution network feeder switch and the adjacency matrix of the distribution network feeder switch G is extracted through the self-attention mechanism, and then the dynamic graph convolutional network is used to calculate the extracted dynamic relationship to obtain the characteristic information quantity of the feeder switch, and the characteristic information quantity of the feeder switch is extracted; among them, the described dynamic relationship is the dynamic change of the characteristic matrix of the distribution network feeder switch and the adjacency matrix of the distribution network feeder switch G caused by the change of the power grid topology at each moment; The extracted characteristic information quantity of the feeder switch is input into the graph convolutional network layer of the hidden layer, and the graph convolutional network is used to calculate the characteristic information quantity of the feeder switch to obtain the characteristic quantity of the state information, and the characteristic quantity of the state information is extracted; The extracted characteristic quantity of the state information is output through the output layer.
[0034] Optionally, the calculation formula of the self-attention mechanism is: ; In the formula, M is the weight matrix after the calculation of the self-attention mechanism, used to represent the dynamic weight relationship of the distribution network feeder switch, is the normalization function, is the transpose matrix of the characteristic matrix of the distribution network feeder switch , D is the degree matrix of the adjacency matrix of the distribution network feeder switch G .
[0035] Based on the above solution, the present application can perform dynamic update calculation according to the position and size of each input variable through the self-attention mechanism, and then obtain the spatial correlation of the state of the distribution network feeder switch. In addition, the self-attention mechanism can perform parallel calculation, effectively improving the calculation efficiency.
[0036] Specifically, the degree matrix D is a diagonal matrix, and the elements on its diagonal are the degrees of each vertex in the graph. Among them, the degree matrixD The calculation formula is: ; In the formula, is the diagonal element, is the degree of the i-th feeder switch, which represents the number of edges connected to this vertex, that is, the sum of all non-zero elements (element 1) in the i-th row, j' are all elements in the i-th row.
[0037] Optionally, the calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is: ; In the formula, H is the feature information amount of the feeder switch, is the activation function, is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product; Optionally, the calculation formula of the graph convolutional network in the graph convolutional network layer is: ; In the formula: I is the state information feature quantity, is the weight matrix of the graph convolutional network layer.
[0038] Specifically, the activation function is the Sigmoid function, and the calculation formula of the Sigmoid function is: .
[0039] In some possible implementation manners of the first aspect, as Figure 3 shown, generating a fault line selection and pulling strategy based on the fault line selection and pulling order specifically includes: Generating a fault line selection and pulling table based on the fault line selection and pulling order: Among them, the fault line selection and pulling order is: the switch of the unloaded line (branch line), the double-power feed-in switch, the reactive power compensation device and the station service transformer; the selection and pulling order of the double-power feed-in switch is: the double-power feed-in of the general user substation, the double-power feed-in of the general public substation, the double-power feed-in of the high-risk user substation and the double-power feed-in of the key public substation; Obtaining the fault phase voltage at time t of the single-phase grounding fault of the distribution network; According to the selection and pulling order in the fault line selection and pulling table, select and pull each switch at time t + 1, and judge whether the fault phase voltage recovers after selecting and pulling the corresponding switch; select and pull each switch according to the selection and pulling order at the same time, avoiding random selection and pulling at the same time, not distinguishing the line type and load loss, etc., effectively improving the selection and pulling accuracy; If it is restored, the single-phase grounded fault line of the distribution network is obtained, and the selection and tripping are ended; Otherwise, continue the selection and tripping for the next moment for each switch according to this selection and tripping sequence until the single-phase grounded fault line of the distribution network is obtained, and the selection and tripping are ended.
[0040] Optionally, after the selection and tripping of the dual-power incoming line switch, if the single-phase grounding fault has not disappeared, then perform the selection and tripping on the opposite-side switch, that is, the switch at the node of the dual-power feeder line.
[0041] Specifically, after the selection and tripping of the reactive power compensation device and the station transformer, there is no need to judge whether the faulty phase voltage is restored.
[0042] Specifically, the active power, reactive power, and current values of the no-load line (branch line) switch in the normal operation state are all 0, that is , , ; The active power, reactive power, and current values of the dual-power incoming line switch in the normal operation state are all the corresponding power and current of the load. Since the power grid load is generally inductive, therefore, the active power, reactive power, and current values are all negative, indicating the direction of power flow in, that is , , ; The active power, reactive power, and current values of the switch at the node of the dual-power feeder line in the normal operation state are all the corresponding power and current of the load. At this time, the active power, reactive power, and current values are all positive, indicating the direction of power flow out, that is , , ; The active power, reactive power, and current values of the reactive power compensation device and the station transformer in the normal operation state are all the corresponding power and current of the reactive power compensation device and the station transformer. Since the reactive power compensation device and the station transformer are all capacitive elements, the active power of the reactive power compensation device and the station transformer is all 0, and they inject reactive power and reactive current into the distribution network system. Therefore, the reactive power and current values of the reactive power compensation device and the station transformer are all negative, indicating the direction of power flow in, that is , , .
[0043] More specifically, the reactive power compensation device is a capacitor or a static var generator (SVG).
[0044] In this application, the general user substation is the substation of a dedicated line user not included in the high-risk user list of the Energy Bureau. The general public substation is the substation shared by those not connected to critical distribution network loads (urban network users, important users, and high-risk users). The high-risk user substation is the substation of a dedicated line user included in the high-risk user list of the Energy Bureau. The critical public substation is the substation shared by those connected to critical distribution network loads (urban network users, important users, and high-risk users).
[0045] In some possible implementation manners of the first aspect, as Figure 4 shown, the single-phase grounding fault line selection and tripping decision model for the distribution network constructed based on deep reinforcement learning according to the state information feature quantity and the fault line selection and tripping strategy specifically includes: Construct a deep reinforcement learning network (Deep Q Networks, DQN) and an experience pool; among them, the DQN includes a training Q network and a target Q network, and the two network structures are exactly the same; Initialize the training Q network and the target Q network, and assign values to the parameters in the training Q network and the target Q network; Construct an initial state space variable according to the extracted state information feature quantity , and perform an observation update of the initial state space variable once per period; where the period is the distribution network selection and tripping control period, , is the state information feature quantity at the t-th moment of each period; Obtain the corresponding action according to the opening and closing states of each feeder switch in each period, and form a corresponding action set based on each obtained action A , is the action at the t-th moment of each period, and the opening and closing states of each feeder switch are different within each period; where, , , is the opening and closing state of the i-th feeder switch at the t-th moment of each period, and the opening and closing state is represented by 0 and 1, 0 means the switch does not act, 1 means the action is to disconnect, and n is the number of feeder switches in the distribution network; Construct the reward function r. Among them, when a single-phase grounding fault occurs in the distribution network, it should be considered whether the disconnection result leads to the disappearance of the grounding fault. The electrical quantity that directly reflects the single-phase grounding fault in the distribution network is the phase voltage. When a single-phase grounding fault occurs in the distribution network, the voltage of the fault phase decreases and is within the stable range when the voltage recovers. When selecting and disconnecting the single-phase grounding fault line in the distribution network, the load loss caused by the selection and disconnection of the double-power feeders should be considered, and corresponding penalties should be imposed according to the type of double-power substations. At the same time, it is also necessary to consider the reactive power imbalance of the system caused by the withdrawal of reactive power compensation devices such as capacitors and SVG and the station service transformers due to the grounding selection and disconnection, and impose penalties on the reactive power imbalance caused by the selection and disconnection of capacitors, SVG, and station service transformers. The initial state space variables of each cycle are input into the training Q-network, and the following steps are iteratively executed: According to the policy, select an action A from the corresponding action set , generate the action value corresponding to the initial state, and input the selected action into the fault line selection and disconnection policy to perform the selection and disconnection, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the immediate reward value of this action according to the reward function r; among them, ; is the neural network weight of the training Q-network; Based on the SCADA system, according to the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, obtain the new state information characteristic quantity, and update and calculate the state space variables at the next moment of the corresponding cycle; Construct an experience data set , and store it in the experience pool; Randomly extract d groups of experience data sets from the experience pool, and input them into the training Q-network and the target Q-network respectively, and calculate the value of the training Q-network and the target value of the target Q-network; among them, d < d m , d m is the data capacity of the experience pool; According to the loss function, update the neural network weights of the training Q-network with the target value of the target Q-network, so that the value of the training Q-network is closer to the target value of the target Q-network; Judge whether the current iteration number is equal to the maximum iteration number; If it is equal, obtain the trained single-phase grounding fault line selection and disconnection policy model of the distribution network; If it is not equal, continue to execute the next iteration.
[0046] The SCADA (Supervisory Control And Data Acquisition) system is a data acquisition and monitoring control system. This system can upload information such as voltage, current, and power of various components in the distribution network, such as switches, main transformers, and busbars, to the dispatching master station in real time. Dispatching operators can intuitively view the wiring topology, operation information at each moment, etc.
[0047] When a single-phase grounding fault occurs in the distribution network, it is usually allowed to operate in the power grid for 2 hours. During this period, short-term selective switching control can be performed on the power grid. The selective switching control period is divided into multiple periods, each period is 10s, and in each period, the fault line selective switching strategy provided by this application is used to select and switch the single-phase grounding fault line in the distribution network, and the objects of selection and switching are each feeder switch, including branch switches, so as to determine the single-phase grounding fault line in the distribution network. In addition, the opening and closing states of each feeder switch, the topological connection relationship of each feeder switch, and the characteristic data of electrical quantities are different in each period.
[0048] Based on the above scheme, this application breaks the correlation between data through the experience replay mechanism, improves the stability of the deep reinforcement learning network training, and enhances the reliability and accuracy of the single-phase grounding fault line selection and switching strategy model for the distribution network.
[0049] Optionally, assigning values to the parameters in the training Q-network and the target Q-network includes: assigning the number of layers of the training Q-network, the number of layers of the target Q-network, the maximum number of iterations T, the discount factor γ, and the data capacity d of the experience pool m 。
[0050] Specifically, the DQN includes an input layer, a hidden layer, and an output layer; among them, the input layer is the state space variable , and the number of layers is set to the number of elements in the state space variable , that is, the state information feature quantity extracted by the dynamic graph neural network calculation; the number of layers of the hidden layer is set according to specific training sample data; the output layer is the action value of each action in this state , is the action of the i-th feeder switch in the output layer, where , and n is the number of feeder switches in the distribution network. In specific implementation, the parameters of the input layer, hidden layer, and output layer of the training Q-network and the target Q-network are set according to the above parameters.
[0051] Optionally, according to strategy, select the corresponding action A from the action set , specifically including: With probability, select the optimal action corresponding to the maximum Q' value from the action set A, with Select a random action from the action set with a probability of A ; The calculation formula of the policy is: ; In the formula, is the policy function, ε is the random selection probability, and ε is a very small positive number, ε < 1, argmax Q' ( a , s ) is the maximum Q' value corresponding to the optimal action a , and |A| is the number of actions in the action set.
[0052] Optionally, the calculation formula of the reward function is: ; In the formula, is the fault phase voltage penalty term, is the substation load loss penalty term, is the reactive power imbalance penalty term; when making intelligent decisions, this reward function can control the loss cost to the minimum; Among them, The expression of ; In the formula, is the threshold voltage, is the fault phase voltage during single-phase ground fault in the distribution network, d is the penalty term; during single-phase ground fault in the distribution network, the fault phase voltage decreases and the non-fault phase voltage increases. When the fault phase voltage is greater than the threshold voltage , the voltage is normal and no penalty term is added; when the fault phase voltage is lower than the threshold voltage , the ground fault has not disappeared yet, and the penalty term d is added; The expression of ; In the formula, is the load loss of the general user substation, is the load loss of the general public substation, is the load loss of the high-risk user substation, is the load loss of the key public substation, m 1 ~ m 4 are the general user substation, general public substation, high-risk user substation and key public substation respectively, k 1 ~k 4 are the load weight parameters of the corresponding substations, and k 4 >k 3 >k 2 >k 1 , l 1 ~ l 4 are the sums of the numbers of the corresponding substations or corresponding users respectively. The more important the load loss is, the greater the penalty is; The expression of ; In the formula, , are the upper limit value and the lower limit value of the reactive power under the normal operation state of the distribution network system respectively; is the reactive power of the distribution network system at time t, and c is the penalty value.
[0053] Optionally, the calculation formula of the target Q network is: ; In the formula, is the target value of the target Q network, h is the current iteration number, T is the maximum iteration number, is the maximum Q' value corresponding to the optimal action selected under the current iteration, is the discount factor to avoid the reward value being infinite and limit the reward value.
[0054] In specific implementation, when h reaches the set maximum iteration number T , the immediate reward value obtained by training the Q network is the final reward value; when it does not reach the set maximum iteration number T , the long-term cumulative reward value is the final reward value.
[0055] Based on the above scheme, when constructing the reward function in this application, the order of selecting and pulling lines, load types, power-off loads, and selection and pull accuracy are comprehensively considered, improving the accuracy of comprehensive analysis and judgment for line selection and pulling of single-phase grounding faults in the distribution network.
[0056] Optionally, the update formula of the neural network weights of the trained Q network is: ; In the formula, is the loss function, is theh The neural network weights of the updated training Q-network for the iteration h value of the training Q-network for the
[0057] In a second aspect, an embodiment of the present application provides a single-phase grounding fault line selection and pulling decision-making system for a distribution network based on deep reinforcement learning. The system includes a module for implementing the foregoing single-phase grounding fault line selection and pulling decision-making method for a distribution network based on deep reinforcement learning.
[0058] Exemplarily, as Figure 5 shown, the single-phase grounding fault line selection and pulling decision-making system for a distribution network based on deep reinforcement learning includes: Topological connection relationship acquisition module 10: used to acquire the topological connection relationship of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network; First matrix construction module 20: used to construct a distribution network feeder switch adjacency matrix according to the topological connection relationship of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network G ; Data acquisition module 30: used to collect the electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network; Second matrix construction module 40: used to construct a distribution network feeder switch feature matrix according to the electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network ; State information feature quantity acquisition module 50: used to process the distribution network feeder switch adjacency matrix G and the distribution network feeder switch feature matrix by using a dynamic graph convolutional neural network to obtain state information feature quantities; Fault line selection and pulling strategy generation module 60: used to generate a fault line selection and pulling strategy based on the fault line selection and pulling order; Model construction module 70: used to construct a single-phase grounding fault line selection and pulling decision-making model for the distribution network based on deep reinforcement learning according to the state information feature quantities and the fault line selection and pulling strategy.
[0059] In some possible implementation manners of the second aspect, as Figure 6 shown, the second matrix construction module 40 includes: Distribution network feeder switch matrix construction unit 401 corresponding to the moment: used to construct a distribution network feeder switch matrix corresponding to the moment according to the electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network; First calculation unit 402: used to perform normalization calculation on the distribution network feeder switch matrix corresponding to the moment; The corresponding - moment distribution - network feeder - switch feature - matrix construction unit 403: It is used to construct the corresponding - moment distribution - network feeder - switch feature matrix according to the distribution - network feeder - switch matrix at the corresponding moment after normalization calculation; The distribution - network feeder - switch feature - matrix construction unit 404: It is used to construct the distribution - network feeder - switch feature matrix according to the distribution - network feeder - switch feature matrices at each moment .
[0060] Optionally, as Figure 7 shown, the state - information feature - quantity acquisition module 50 includes: The first input unit 501: It is used to input the distribution - network feeder - switch feature matrix and the distribution - network feeder - switch adjacency matrix G into the input layer of the dynamic graph convolutional neural network; The first output unit 502: It is used to output the distribution - network feeder - switch feature matrix of the input layer and the distribution - network feeder - switch adjacency matrix G ; The second input unit 503: It is used to input the distribution - network feeder - switch feature matrix of the input layer and the distribution - network feeder - switch adjacency matrix G into the dynamic graph convolutional network layer of the hidden layer; The self - attention mechanism extraction unit 504: It is used to extract the dynamic relationships of each feeder switch at each moment in the distribution - network feeder - switch feature matrix and the distribution - network feeder - switch adjacency matrix G by using the self - attention mechanism; The second calculation unit 505: It is used to calculate the extracted dynamic relationships by using the dynamic graph convolutional network to obtain the feeder - switch feature information quantity; The first extraction unit 506: It is used to extract the feeder - switch feature information quantity; The third input unit 507: It is used to input the extracted feeder - switch feature information quantity into the graph convolutional network layer of the hidden layer; The third calculation unit 508: It is used to calculate the feeder - switch feature information quantity by using the graph convolutional network to obtain the state - information feature quantity; The second extraction unit 509: It is used to extract the state - information feature quantity; The fourth input unit 510: It is used to input the extracted state - information feature quantity into the output layer; The second output unit 511: It is used to output the state - information feature quantity of the output layer.
[0061] Optionally, as Figure 8 shown, the faulty - line selection - and - switching - off strategy generation module 60 includes: Faulty line selection and switching list generation unit 601: used to generate a faulty line selection and switching list based on the faulty line selection and switching sequence; Faulty phase voltage acquisition unit 602: used to acquire the faulty phase voltage of a single-phase grounding fault in the distribution network; Selection and switching unit 603: used to perform selection and switching on each switch at the same time according to the selection and switching sequence in the faulty line selection and switching list; First judgment unit 604: used to judge whether the faulty phase voltage recovers after selecting and switching the corresponding switch; Faulty line acquisition unit 605: used to obtain the faulty line of the single-phase grounding fault in the distribution network and end the selection and switching when the faulty phase voltage recovers; First execution unit 606: used to continue the selection and switching at the next moment according to the selection and switching sequence when the faulty phase voltage does not recover.
[0062] Optionally, as Figure 9 shown, the model construction module 70 includes: First construction unit 701: used to construct a deep reinforcement learning network and an experience pool; Initialization and assignment unit 702: used to initialize the training Q-network and the target Q-network and assign values to the parameters in the training Q-network and the target Q-network; Second construction unit 703: used to construct an initial state space variable according to the extracted state information feature quantity ; Initial state space variable update unit 704: used to perform an observation update on the initial state space variable once per cycle ; Action acquisition unit 705: used to obtain the corresponding action according to the opening and closing states of each feeder switch per cycle ; Third construction unit 706: used to form an action set according to each obtained action according to the faulty line selection and switching strategy A ; Fourth construction unit 707: used to construct a reward function r; Fifth input unit 708: used to input the initial state space variable into the training Q-network; Action selection unit 709: used to select an action from the action set according to the A policy; ; Data generation unit 710: used to generate the action value corresponding to the initial state ; Data update unit 711: used to update the selected action Execute the selection and pulling in the input fault line selection and pulling strategy to obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch; Reward value acquisition unit 712: used to calculate the immediate reward value of this action according to the reward function r ; ; Data acquisition unit 713: used to obtain new state information characteristic quantities based on the SCADA system according to the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch; State update unit 714: used to update and calculate the state space variables at the next moment according to the new state information characteristic quantities ; Fifth construction unit 715: used to construct an experience data set ; Experience data set storage unit 716: used to store the experience data set into the experience pool; Experience data set extraction unit 717: used to randomly extract d groups of experience data sets from the experience pool; Sixth input unit 718: used to input the randomly extracted d groups of experience data sets into the training Q network and the target Q network respectively; Fourth calculation unit 719: used to calculate the value of the training Q network and the target value of the target Q network; Neural network weight update unit 720: used to update the neural network weights of the training Q network according to the loss function using the target value of the target Q network; Second judgment unit 721: used to judge whether the current iteration number is equal to the maximum iteration number; Model acquisition unit 722: used to obtain the trained distribution network single-phase grounding fault line selection and pulling strategy model when the current iteration number is equal to the maximum iteration number; Third execution unit 723: used to continue to execute the next iteration when the current iteration number is not equal to the maximum iteration number.
[0063] Thirdly, in the embodiments of the present application, a device is provided. The device can be any device capable of implementing the distribution network single-phase grounding fault line selection and pulling decision method based on deep reinforcement learning. The device can be various terminal devices, such as: desktop computers, laptops, tablets, handheld devices, etc., and can be specifically implemented by software and / or hardware.
[0064] Exemplarily, the device includes: A memory; A processor; and A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the above-mentioned method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning.
[0065] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which may be: ROM, RAM, a magnetic disk, an optical disc, etc.
[0066] Exemplarily, a computer program is stored on the computer-readable storage medium; the computer program is executed by a processor to implement the above-mentioned method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning.
[0067] In summary, the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning in the present application provides a scientific and efficient auxiliary decision for the disposal of single-phase grounding faults in the distribution network, promotes the efficiency of selecting and pulling a single-phase grounding line in the distribution network, and at the same time improves the restoration ability of the distribution network.
[0068] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, for example, C language, VHDL language, Verilog language, object-oriented programming language Java, and interpreted scripting language JavaScript, etc.
[0069] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in a block or blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in a block or blocks.
[0072] Furthermore, the terms "first", "second" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of those features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0073] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0074] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A single-phase grounding fault line selection and switching decision method for a distribution network based on deep reinforcement learning, which is used to select the line where a single-phase grounding fault occurs in the distribution network, and is characterized in that: Including: Obtain the topological connection relationship of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, and construct the adjacent matrix of the feeder switches in the distribution network G ; Construct a characteristic matrix of the feeder switches in the distribution network according to the characteristic data of the electrical quantity information of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network collected. Among them, the characteristic data of the electrical quantity information includes: sudden change in active power, sudden change in reactive power, sudden change in current, and zero-sequence current. Using a dynamic graph convolutional neural network to process the adjacency matrix of the distribution network feeder switches G and the feature matrix of the distribution network feeder switches to obtain the state information feature quantity; Generate a fault line selection and disconnection strategy based on the fault line selection and disconnection sequence. Based on the state information feature quantity and the fault line selection and disconnection strategy, construct a decision-making model for selecting and disconnecting single-phase grounded fault lines in a distribution network based on deep reinforcement learning, which is used to analyze the selection and disconnection of single-phase grounded fault lines in the distribution network, and determine the fault line selection for single-phase grounded faults in the distribution network according to the analysis result.
2. The single-phase grounding fault line selection and switching decision-making method for a distribution network based on deep reinforcement learning according to claim 1, characterized in that: Constructing a characteristic matrix of the feeder switches in the distribution network based on the characteristic data of the electrical quantity information of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network collected , specifically including: According to the characteristic data of the electrical quantity information of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network collected, a distribution network feeder switch matrix at the corresponding moment is constructed; the distribution network feeder switch matrix at time t' is as follows: ; Wherein, is the characteristic data of the electrical quantity information of the nth feeder switch at time t', and n is the number of feeder switches in the distribution network; is the sudden change in active power of the nth feeder switch at time t'; is the sudden change in reactive power of the nth feeder switch at time t'; is the sudden change in current of the nth feeder switch at time t'; is the zero-sequence current of the nth feeder switch at time t'. Perform normalization calculation on the distribution network feeder switch matrix at the corresponding moment to obtain the distribution network feeder switch feature matrix at the corresponding moment; the distribution network feeder switch feature matrix at time t' It is as follows: ; Wherein, is the characteristic data of the electrical quantity information of the nth feeder switch normalized at time t'; is the sudden change in active power of the nth feeder switch normalized at time t'; is the sudden change in reactive power of the nth feeder switch normalized at time t'; is the sudden change in current of the nth feeder switch normalized at time t'; is the zero-sequence current of the nth feeder switch normalized at time t'; Construct the characteristic matrix of the distribution network feeder switch : , where is the characteristic matrix of the distribution network feeder switch at time N, .
3. The single-phase grounding fault line selection and switching decision-making method for a distribution network based on deep reinforcement learning according to claim 1, characterized in that: The dynamic graph convolutional neural network includes an input layer, two hidden layers and an output layer; the two hidden layers are a dynamic graph convolutional network layer and a graph convolutional network layer respectively; The dynamic graph convolutional neural network is used to process the adjacent matrix of the distribution network feeder switch G and the feature matrix of the distribution network feeder switch to obtain the state information feature quantity, specifically including: Input the characteristic matrix of the distribution network feeder switch and the adjacency matrix of the distribution network feeder switch G into the input layer of the dynamic graph convolutional neural network; where the number of neurons in the input layer is 4; The distribution network feeder switch feature matrix of the input layer and the distribution network feeder switch adjacency matrix G are input into the dynamic graph convolutional network layer of the hidden layer. First, the dynamic relationships of each feeder switch at each moment in the distribution network feeder switch feature matrix and the distribution network feeder switch adjacency matrix G are extracted through the self-attention mechanism, and then the dynamic graph convolutional network is used to calculate the extracted dynamic relationships to obtain the feature information quantity of the feeder switch, and the feature information quantity of the feeder switch is extracted; Input the extracted feeder switch feature information quantity into the graph convolutional network layer of the hidden layer, use the graph convolutional network to calculate the feeder switch feature information quantity, obtain the state information feature quantity, and extract the state information feature quantity. Output the extracted state information feature quantity through the output layer.
4. The method for selecting and pulling decision of single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 3, characterized in that: The calculation formula of the self-attention mechanism is: ; In the formula, M is the weight matrix after the calculation of the self-attention mechanism, is the normalization function, is the characteristic matrix of the distribution network feeder switch is the transpose matrix, D is the adjacency matrix of the distribution network feeder switch G is the degree matrix; The calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is: ; In the formula, H is the characteristic information quantity of the feeder switch, is the activation function, is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product; The calculation formula of the graph convolutional network in the graph convolutional network layer is: ; In the formula: I is the characteristic quantity of state information, is the weight matrix of the graph convolutional network layer.
5. The single-phase grounding fault line selection and switching decision-making method for a distribution network based on deep reinforcement learning according to claim 1, characterized in that: The generating of the fault line selection and disconnection strategy based on the fault line selection and disconnection sequence specifically includes: Generate a fault line selection and disconnection table based on the fault line selection and disconnection sequence: among them, the fault line selection and disconnection sequence is: the switch of the no-load line (branch line), the double-power-inlet switch, the reactive power compensation device and the station service transformer; the selection and disconnection sequence of the double-power-inlet switch is: the double-power-inlet of the general user substation, the double-power-inlet of the general public substation, the double-power-inlet of the high-risk user substation and the double-power-inlet of the key public substation; Obtain the fault phase voltage of the single-phase grounded fault in the distribution network. According to the selection and disconnection sequence in the fault line selection and disconnection table, select and disconnect each switch at the same time, and judge whether the fault phase voltage is restored after selecting and disconnecting the corresponding switch. If it is restored, obtain the single-phase grounded fault line in the distribution network and end the selection and disconnection. Otherwise, continue to select and disconnect each switch at the next moment according to this selection and disconnection sequence until the single-phase grounded fault line in the distribution network is obtained and the selection and disconnection is ended.
6. The single-phase grounding fault line selection and switching decision method for a distribution network based on deep reinforcement learning according to claim 1, characterized in that: The constructing of the decision-making model for selecting and disconnecting single-phase grounded fault lines in the distribution network based on deep reinforcement learning according to the state information feature quantity and the fault line selection and disconnection strategy specifically includes: Construct a deep reinforcement learning network and an experience pool; among them, the deep reinforcement learning network includes a training Q network and a target Q network; Initialize the training Q network and the target Q network, and assign values to the parameters in the training Q network and the target Q network. Construct an initial state space variable according to the extracted state information feature quantity , and perform an observation update of the initial state space variable once per cycle ; wherein, the cycle is a selected pull control cycle , is the state information feature quantity at time t of each cycle Based on the opening and closing states of each feeder switch in each period, the corresponding actions are obtained , and based on each obtained action , an action set is formed according to the fault line selection and tripping strategy A ; where , , is the action at time t in each period, is the opening and closing state of the i-th feeder switch at time t in each period, and n is the number of feeder switches in the distribution network; Construct a reward function r. The initial state space variables of each cycle are input into the trained Q-network, and the following steps are iteratively executed: According to the policy, select an action A from the corresponding action set , generate the action value corresponding to the initial state, and input the selected action into the fault line selection and switching strategy to perform selection and switching, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the immediate reward value of this action according to the reward function r; where is the neural network weight of the trained Q-network; Based on the SCADA system, according to the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, new state information characteristic quantities are obtained, and the state space variables at the next moment of the corresponding cycle are updated and calculated. ; Construct an empirical data set and store it in the empirical pool; Randomly extract d groups of empirical data sets from the experience pool, and input them into the training Q-network and the target Q-network respectively to calculate the value of the training Q-network and the target value of the target Q-network; where d < d m , d m is the data capacity of the experience pool; According to the loss function, update the neural network weights of the training Q network with the target value of the target Q network, so that the value of the training Q network is closer to the target value of the target Q network. Judge whether the current iteration number is equal to the maximum iteration number. If it is equal, obtain the trained decision-making model for selecting and disconnecting single-phase grounded fault lines in the distribution network. If it is not equal, continue to execute the next iteration.
7. The method for selecting and pulling decision of single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 6, characterized in that: The described according to the policy, select an action from the action set A and specifically includes: select an action Select the optimal action corresponding to the maximum Q' value from the action set A with probability , and select a random action from the action set with probability ; A The calculation formula of the strategy is as follows: ; wherein, is the policy function, is the random selection probability and , is the maximum Q' value corresponding to the optimal action a, is the number of actions in the action set.
8. The single-phase grounding fault line selection and switching decision method for a distribution network based on deep reinforcement learning according to claim 6, wherein: The calculation formula of the reward function is: ; In the formula, is the fault phase voltage penalty term, is the substation load loss penalty term, is the reactive power imbalance penalty term; Among them, The expression of is: ; wherein, is the threshold voltage, is the fault phase voltage during single-phase grounding fault of the distribution network, and d is the penalty term; The expression is: ; Wherein, is the load loss of the general user substation, is the load loss of the general public substation, is the load loss of the high-risk user substation, is the load loss of the key public substation, are the general user substation, the general public substation, the high-risk user substation and the key public substation respectively, are the load weight parameters of the corresponding substations, and , are the sum of the numbers of the corresponding substations or the corresponding users respectively; The expression is: ; In the formula, , are respectively the upper limit value and the lower limit value of the reactive power under the normal operation state of the distribution network system; is the reactive power of the distribution network system at time t, and c is the penalty value; The calculation formula of the target Q network is: ; wherein, is the target value of the target Q network, h is the current iteration number, T is the maximum iteration number, γ is the discount factor, is the maximum Q' value corresponding to the optimal action selected in the current iteration; The update formula of the neural network weights of the training Q network is: ; In the formula, is the loss function, is the neural network weights of the updated training Q-network at the h -th iteration; is the value of the training Q-network at the h -th iteration.
9. A single-phase grounding fault line selection and switching decision-making system for a distribution network based on deep reinforcement learning, characterized in that, Including a module for implementing the decision-making method for selecting and disconnecting single-phase grounded fault lines in the distribution network based on deep reinforcement learning according to any one of claims 1-8.
10. A device, characterized in that, Including: A memory; Processor; and computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, A computer program is stored thereon; the computer program is executed by a processor to implement the method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to any one of claims 1-8.
Citation Information
Patent Citations
Small current grounding fault line selection method and device based on clustering and deep learning
CN116227538A
Pedestrian action prediction method based on dynamic graph convolution attention model
CN118522070A
Graph structure feature-based routing optimization method and system
WO2024037136A1