Power distribution network fault positioning method based on Markov decision process
By applying Markov decision-making process and approximate dynamic planning in the fault positioning of distribution networks, the problems of low fault positioning accuracy and large calculation volume in the existing technology are solved, and the rapid and accurate positioning of fault sections of distribution networks are achieved.
Patent Information
- Application Number
- CN202510304888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as low accuracy, large calculation amount and poor fault tolerance in the fault positioning of distribution networks. Especially when distributed power supply is connected to the grid and information distortion, it is difficult to achieve fast and accurate fault positioning.
Using a fault location method based on Markov decision-making process (MDP), the Markov decision-making process model is constructed and converted into the Bellman equation, and combined with approximate dynamic programming (ADP) solutions, the accurate and fast positioning of the fault segment is achieved.
It improves the accuracy and efficiency of fault positioning of distribution networks, reduces the amount of calculation, enhances fault tolerance performance, and can achieve accurate fault positioning under the grid connection of distributed power supplies and information distortion.
Smart Images

Figure CN120214480A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of distribution network fault location and new energy energy saving, and particularly relates to a distribution network fault location method based on Markov decision process. Background Art
[0002] As a guarantee basis for improving the safety and reliability of the distribution network, a fast and accurate fault location method is essential. With the continuous rapid development of the power network, quickly and accurately locating faults after a fault can not only improve the power supply recovery rate, but also be crucial for improving economic benefits and the quality of residents' lives. In the new development situation, the distribution network presents a new architecture with large-scale grid connection of distributed energy sources (DG) such as wind and light, resulting in a more complex topological network. Moreover, the access of a large number of distributed power sources has significantly changed the power flow distribution and direction of the distribution network, making fault location more difficult and posing a huge challenge to the stability and reliability of the operation of the power system in practical applications. Traditional distribution network fault section location methods mainly rely on matrix algorithms and intelligent optimization algorithms, such as Grey Wolf Optimization Algorithm (GWO), Particle Swarm Optimization Algorithm (PSO), Bat Algorithm (BA), Sine Cosine Algorithm (SCA), etc.
[0003] However, the matrix algorithm generates a matrix that can describe the network characteristics and information of the distribution network based on graph theory knowledge and the distribution network structure, and locates the fault section by solving the matrix. However, the location accuracy depends on the distribution network structure. Once the structure changes, the matrix needs to be reconstructed, with a large amount of calculation and poor fault tolerance; the intelligent optimization algorithm searches for the global optimal solution in the solution space based on the fault hypothesis theory, that is, converts the fault problem into a mathematical problem to determine the fault section. However, this type of method is usually prone to falling into local optimal solutions, resulting in low fault location accuracy. Moreover, as the topological scale of the distribution network expands, the dimension of its solution space increases, easily causing the curse of dimensionality. Both of these methods perform fault location based on the data provided by feeder terminal units (FTUs) or fault passage indicators (FPIs). However, due to the possible information distortion of the data uploaded by FTUs or FPIs under the condition of a large number of DGs connected to the grid, it has uncertainty and randomness. And when the distribution network topological structure changes, the location accuracy of traditional fault section location methods often cannot meet the actual application requirements. Therefore, how to accurately and quickly locate the fault section of the distribution network, especially the fault section location under the condition of distributed power grid connection and information distortion, has become a current research hotspot. Summary of the Invention
[0004] In order to solve the challenges in the distribution network fault section location, this paper proposes a distribution network fault location method based on Markov decision process to locate the fault section of the distribution network.
[0005] To implement the above technical solution, the specific steps are as follows:
[0006] S1. Determine the basic elements required to establish a Markov decision process (MDP) model according to the fault section location objective function during the grid connection of distributed generation (DG) and the information distortion of the feeder terminal unit (FTU). Construct a Markov decision process based on the fault section location objective function and convert it into a solvable Bellman equation form;
[0007] S1.1. Develop the coding form of the fault current of node j in the distribution network;
[0008] The development process is as follows: When distributed generation (DG) is connected to the distribution network, the power flow distribution and direction in the grid will change. When there is no DG in the feeder section of the grid system, the positive direction of the current is from the system power source to the load; otherwise, the positive direction of the current is from the system power source to the DG. In this way, the fault current coding of each node (switch) I in the distribution network j is divided into three cases, and the form is as follows:
[0009]
[0010] S1.2. Construct a switching function based on the coding form I of the fault current of the distribution network node j and the operating state of the feeder section L of node j j to obtain the expected fault current state sequence;
[0011] The expression of the switching function is as follows:
[0012]
[0013] In the formula, L d and L u represent the operating states of the downstream d and upstream u feeder sections of node j, respectively; n1 and n2 are the numbers of the upstream and downstream feeder sections at node j; D u and D d represent the switch coefficients of the DGs upstream and downstream of node j, respectively, and the expressions are as follows:
[0014]
[0015]
[0016] In the formula, M1 and M2 represent the total numbers of DGs connected upstream and downstream of node j, respectively; m represents the number of DGs connected; and represent the operating states of the sections H1 and H2 from node j to the upstream power source M1 and the downstream power source M2, respectively;
[0017] S1.3. Use the actually uploaded information of the Feeder Terminal Unit (FTU) and the switching function to construct the fault location model F it (Objective function);
[0018] The fault location model is represented by the difference between the actually uploaded information of the FTU and the switching function. When the difference between the two is the smallest, the optimal solution is obtained, which means that the actual fault current sequence is most similar to the expected function value and the error is the smallest. However, due to the influence of the distribution network structure and power flow direction on the actually uploaded information of the FTU, problems such as information distortion (false information reporting and information omission) will occur. For example, when a fault occurs in the distribution network, the actual fault current code of a certain node should be 1, but the information uploaded by the FTU is -1 or 0, or the current code information of this node is missed. Therefore, in order to improve the accuracy and fault tolerance performance of the fault location model,
[0019] When considering the existence of information distortion, the expression for constructing the fault location model is as follows:
[0020]
[0021] In the formula, I j represents the actual fault current coding situation of node j, that is, the actual operating state; I j * (L) represents the expected fault current coding situation of node j, that is, the expected operating state; K represents the total number of nodes in the distribution network with DG; α is the weight coefficient, which is taken as 0.5 according to the minimum fault theory to prevent misjudgment; N represents the total number of feeder sections in the distribution network with DG; i represents the number of feeder sections in the distribution network with DG; is the sum of the states of all feeder sections; and respectively represent the total sum of information omission e j and false information reporting situation g j ; η e and η g respectively represent the coefficients of information omission and false information reporting situations, η e = η g = 1;
[0022] S1.4. Construct a Markov Decision Process (MDP) according to the fault location model. First, the basic elements of the model need to be defined, including: state variables, decision variables, influencing factors, and state transition functions;
[0023] The definitions are as follows:
[0024] State variable S t: The current state of the distribution network includes the operating state of each feeder section L at the current time period t, defined as N is the total number of feeder sections in the distribution network with DG, that is, the number of MDP states;
[0025] Decision variable A t : It includes all possible coding action sequences for coding each feeder section, defined as
[0026] Influence factor e t : It is described by the deviation between the actual information uploaded by the FTU without considering information distortion at time t and the switch function structure, defined as e t ={|I - I * | t};
[0027] State transition function P: It reflects the probability that the state variable S t transfers from the current state to the next state under the action of the decision variable A t and the influence factor e t , defined as:
[0028]
[0029] S1.5. According to the basic elements and the fault section location objective function, establish the Bellman equation, the steps are as follows:
[0030] S1.5.1. Represent the fault section location objective function under the action of the decision variable and the influence factor;
[0031] The objective function has an optimal solution, that is, the absolute error between the actual information I j uploaded by the FTU considering information distortion and the expected switch function is the smallest. Described by the probability idea, it is the probability of the event Y that the absolute error between the two approaches 0 is the largest;
[0032] Under the action of the decision variable A t and the influence factor e t in the basic elements, represent the objective function as:
[0033]
[0034] S1.5.2. Initialize the state of the feeder section in the basic elements as S1, set the MDP model to make decisions in T time periods, and maximize the probability of the event Y under the decision sequence {A1, A2,..., A T}, the expression is as follows:
[0035]
[0036] S1.5.3. Derive the expression obtained in S1.5.2 according to the definitions of conditional probability and Markov property as follows:
[0037]
[0038] Take the logarithm of both sides of the Fit t expression to convert it into the form of Bellman equation with a recursive form, and we can get:
[0039]
[0040] Replace ln Fit t with Z t (S t ), replace lnP(S t+1 |S t ) with R t (S t ,A t ), replace ln Fit t+1 with Z t+1 (S t+1 ), and the obtained Bellman equation is:
[0041]
[0042] In the formula, Z t (S t ) is the objective value function of state S t ; Z t+1 (S t+1 ) is the objective value function of state S t+1 ; R t (S t ,A t ) is the immediate probability reward.
[0043] S2. Solve the Bellman equation using approximate dynamic programming based on a value table, and store the value function corresponding to the decision variable sequence in the value function table;
[0044] S2.1. Obtain the value functions for all time periods and states using a backward traversal calculation method;
[0045] For the state S T at time period T, the expression is as follows:
[0046]
[0047] In the formula, num(S T ) represents the number of occurrences of state S T at time period T; represents the number of occurrences of state S T when event Y occurs;
[0048] The remaining time period S t The expression is as follows:
[0049]
[0050] In the formula, num(S t →S t+1 ) represents the number of times from state S t to state S t+1 ;
[0051] S2.2. Obtain the value functions of all states in all time periods through S2.1, and store each value function and its corresponding decision variable sequence in the value function table.
[0052] S3. Obtain the optimal decision from the value function table through policy iteration. The optimal decision is the location of the faulty section;
[0053] The steps are as follows:
[0054] S3.1. According to the value function table obtained in S2, adopt the policy iteration method to search for the maximum Z t (S t ) in the value function table and record its corresponding decision A t as the optimal decision; then calculate the state S t+1 of the next time period at time t until the states of all T time periods are traversed;
[0055] The calculation expression is as follows:
[0056]
[0057] S3.2. Set the number of iterations to k. Each iteration executes S2 to S3.1 in sequence until Z t (S t ) converges and the iteration is completed;
[0058] S3.3. Output the decision corresponding to when Z t (S t ) converges as the optimal decision, that is, the corresponding section with section code 1 or -1 is the faulty section. The beneficial effects of the present invention:
[0059] The beneficial effects of the present invention:
[0060] (1) Based on the current situation of the distribution network development (large scale, complex DG grid-connected topological structure), considering the large-scale grid connection of DG and the distortion of the operating status information uploaded by FTU, a distribution fault section location model is established to improve the location accuracy. Aiming at the problems of low efficiency and low accuracy in distribution network fault location (compared with traditional matrix algorithms and optimization algorithms), according to the advantages of solving the optimal solution in the Markov decision process and analyzing the Markov properties of the constructed fault location model, the fault location model (objective function) is converted into an MDP model, thus transforming the minimization problem of the mathematical model into the maximization problem of the probability model. And compared with the matrix-based location method, its computational complexity is greatly reduced, and the model can be flexibly changed according to the scale and structure of the distribution network.
[0061] (2) Utilizing the advantages of ADP in solving the MDP model (using the idea of policy iteration to improve the convergence speed), the established fault location MDP model is solved by the ADP method based on the value function table to obtain the optimal coding decisions for all sections. And compared with the optimization algorithm-based location method, its convergence speed and accuracy are higher, realizing the accurate and rapid location of the distribution network fault section.
[0062] (3) By establishing a distribution network fault location model considering FTU information distortion and different DG penetration rates, the requirements for the accuracy of fault location that are more in line with actual applications are realized, and the reliability and stability of the distribution network operation are improved.
[0063] (4) MDP is flexible and effective in solving the problems of decision-making results depending on uncertain factors and random factors during the search process. It can guide the agent to develop in the direction of maximizing benefits to achieve the expected objective function value. Based on the strong advantages of MDP in modeling decision-making problems under the influence of uncertain factors, and combined with the distribution network topology model, MDP modeling is carried out for the fault section location problem. At the same time, the approximate dynamic programming is used to solve the model to obtain the optimal coding decision that satisfies the global optimality. This method overcomes the defects of traditional methods such as long solution time and curse of dimensionality to obtain the global optimal convergence value and improve the fault location accuracy. Description of the Drawings
[0064] Figure 1 is the step flow chart of the present invention;
[0065] Figure 2 is the calculation flow chart of the present invention;
[0066] Figure 3 is the Friedman test result chart of the embodiment of the present invention and other algorithms;
[0067] Figure 4 is the convergence result chart of the embodiment of the present invention and other algorithms. Detailed Embodiment
[0068] The present invention will be further described in detail below in conjunction with specific embodiments.
[0069] Taking the 33-node distribution network model as an example, the present invention uses the state information at T = 10 moments uploaded by the FTU (in the example, the dynamic change of the FTU information is not considered, and the FTU information at 10 moments is set to be the same, which is to ensure the accuracy of fault location for a specific fault scenario) to verify the accurate location of the fault section.
[0070] As Figure 1 shown, a distribution network fault location method based on the Markov decision process includes the following steps:
[0071] S1. Determine the basic elements required to establish the Markov decision process MDP model according to the fault section location objective function when the distributed power source DG is connected to the grid and the information of the feeder terminal unit FTU is distorted. Construct the Markov decision process according to the fault section location objective function and convert it into a solvable Bellman equation form;
[0072] S1.1. Develop the coding form of the fault current of node j in the distribution network;
[0073] The development process is as follows: When the distributed generation (DG) is connected to the distribution network, the power flow distribution and direction of the power grid will change. When there is no DG in the feeder section of the power grid system, the direction from the system power source to the load is the positive direction of the current; otherwise, the direction from the system power source to the DG is the positive direction of the current; thus, the fault current coding of each node (switch) I j (j = 1, 2,..., 33) in the distribution network is divided into three cases, and the form is as follows:
[0074]
[0075] S1.2. Construct a switching function, that is, the expected fault current state sequence, according to the coding form I j of the fault current of the distribution network node and the operating state of the feeder section L j of node j;
[0076] The expression of the switching function is as follows:
[0077]
[0078] In the formula, L d and L u respectively represent the operating states of the downstream d and upstream u feeder sections of node j; n1 and n2 are the numbers of the upstream and downstream feeder sections at node j respectively; D u and D dThey respectively represent the opening coefficients of the DGs upstream and downstream of node j, and the expressions are as follows:
[0079]
[0080] In the formula, M1 and M2 respectively represent the total number of DGs connected upstream and downstream of node j; m represents the number of DGs connected; and They respectively represent the operating states of sections H1 and H2 from node j to upstream power source M1 and downstream power source M2;
[0081] S1.3. Use the actually uploaded information of the Feeder Terminal Unit (FTU) and the switching function to construct the fault location model F it (Objective function);
[0082] The fault location model is represented by the difference between the actually uploaded information of the FTU and the switching function. When the difference between the two is the smallest, the optimal solution is obtained, which means that the actual fault current sequence is most similar to the expected function value and the error is the smallest. However, due to the influence of the distribution network structure and power flow direction on the actually uploaded information of the FTU, problems such as information distortion (false alarms and missed reports) will occur. For example, when a fault occurs in the distribution network, the actual fault current code of a certain node should be 1, but the information uploaded by the FTU is -1 or 0, or the current code information of this node is missed. Therefore, in order to improve the accuracy and fault tolerance performance of the fault location model,
[0083] When considering the existence of information distortion, the expression for constructing the fault location model is as follows:
[0084]
[0085] In the formula, I j represents the actual fault current coding situation of node j, that is, the actual operating state; represents the expected fault current coding situation of node j, that is, the expected operating state; K represents the total number of nodes in the distribution network with DGs; α is the weight coefficient, and 0.5 is taken according to the minimum fault theory to prevent misjudgment; N represents the total number of feeder sections in the distribution network with DGs; i represents the number of feeder sections in the distribution network with DGs; is the sum of the states of all feeder sections; and respectively represent the total sum of the missed reports e j and false alarms g j of all sections; η e and η g respectively represent the coefficients of the missed reports and false alarms. η e = η g = 1;
[0086] S1.4. As shown Figure 2 in the figure, according to the fault location model, a Markov Decision Process (MDP) is constructed. First, the basic elements of the model need to be defined, including: state variables, decision variables, influencing factors, and state transition functions;
[0087] The definitions are as follows:
[0088] State variable S t : The current state of the distribution network includes the operating state of each feeder section L at the current time period t, defined as N = 1, 2, 3,..., 32 is the total number of feeder sections in the distribution network with DG, that is, the number of MDP states;
[0089] Decision variable A t : It includes all possible coded action sequences for coding each feeder section, defined as
[0090] Influencing factor e t : It is described by the deviation between the actual information uploaded by the FTU without considering information distortion at time t and the switch function structure, defined as e t ={|I - I * | t};
[0091] State transition function P: It reflects the probability that the state variable S t transfers from the current state to the next state under the action of the decision variable A t and the influencing factor e t , and is defined as:
[0092]
[0093] S1.5. According to the basic elements and the fault section location objective function, establish the Bellman equation. The steps are as follows:
[0094] S1.5.1. Represent the fault section location objective function under the action of the decision variable and the influencing factor;
[0095] The objective function has an optimal solution, that is, under the condition of considering information distortion, the absolute error between the actual information I j uploaded by the FTU and the expected switch function I j * (L) is the smallest. Described by probability theory, it is the probability of the event Y that the absolute error between the two approaches 0 is the largest;
[0096] Among the basic elements, the decision variable A t and the influencing factor e tUnder the action, the objective function is expressed as:
[0097]
[0098] S1.5.2. Initialize the feeder section status of the basic elements to S1, set the MDP model to make decisions in T time periods, and maximize the probability of event Y under the decision sequence {A1, A2, …, A T} as follows:
[0099]
[0100] S1.5.3. Derive the expression obtained in S1.5.2 according to the definition of conditional probability and the Markov property as follows:
[0101]
[0102] Take the logarithm of both sides of the Fit t expression to convert it into the form of the Bellman equation with a recursive form, and we can get:
[0103]
[0104] Replace ln Fit t with Z t (S t ), replace lnP(S t+1 |S t ) with R t (S t , A t ), and replace ln Fit t+1 with Z t+1 (S t+1 ). The obtained Bellman equation is:
[0105]
[0106] In the formula, Z t (S t ) is the objective value function of state S t ; Z t+1 (S t+1 ) is the objective value function of state S t+1 ; R t (S t , A t ) is the immediate probability reward.
[0107] S2. Use approximate dynamic programming based on the value table to solve the Bellman equation, and store the value function corresponding to the decision variable sequence in the value function table;
[0108] S2.1. Obtain the value functions of all time periods and states by using the reverse traversal calculation method;
[0109] For the state S at time period T = 10 10 , the expression is as follows:
[0110]
[0111] In the formula, num(S 10 ) represents the number of times the state S appears at time period T; 10 ; represents the number of times the state S appears when event Y occurs; represents the number of times the state S appears when event Y occurs; 10 ;
[0112] For the remaining time periods, the expression of S t is as follows:
[0113]
[0114] In the formula, num(S t → S t+1 ) represents the number of times from state S t to state S t+1 ;
[0115] S2.2. Obtain the value functions of all states in all time periods through S2.1, and store each value function and its corresponding decision variable sequence in the value function table.
[0116] S3. Obtain the optimal decision from the value function table through policy iteration. The optimal decision is the location of the faulty section;
[0117] The steps are as follows:
[0118] S3.1. According to the value function table obtained in S2, search for the maximum Z t (S t ) in the value function table by traversing, and record its corresponding decision A t as the optimal decision; then calculate the state S t+1 at the next time period of time period t until the states of 10 time periods are all traversed;
[0119] The calculation expression is as follows:
[0120]
[0121] S3.2. Set the number of iterations to k. Each iteration executes S2 to S3.1 in sequence until Z t (S t ) converges and the iteration is completed; in this embodiment, k = 100;
[0122] S3.3. Output Z t (S t ) The decision corresponding to the convergence is the optimal decision, that is, the corresponding section with section coding of 1 or -1 is the faulty section.
[0123] To verify the effectiveness of the present invention, Friedman test and convergence effect test were carried out with the Grey Wolf Optimization Algorithm (GWO), Particle Swarm Optimization Algorithm (PSO), Bat Algorithm (BA), and Sine Cosine Algorithm (SCA). The experimental settings were k = 100 and T = 10.
[0124] The lower the average value of the Friedman test, the better the overall performance of the algorithm; from Figure 3 it can be seen that the present invention has the lowest average value among all results, indicating that the present invention has the best performance.
[0125] The convergence effect is as Figure 4 shown. Compared with the Grey Wolf Optimization Algorithm (GWO), Particle Swarm Optimization Algorithm (PSO), Bat Algorithm (BA), and Sine Cosine Algorithm (SCA), the present invention has better convergence results.
[0126] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A distribution network fault location method based on Markov decision process, characterized in that: The following steps are involved: S1. According to the fault section location objective function when the distributed generation DG is connected to the grid and the feeder terminal unit FTU information is distorted, the basic elements required to establish the Markov decision process MDP model are determined, and the Markov decision process is constructed according to the fault section location objective function, and converted into a solvable Bellman equation form; S2. Use approximate dynamic programming based on value tables to solve the Bellman equation, and store the value functions corresponding to the decision variable sequence in the value function table; S3. Obtain the optimal decision from the value function table through strategy iteration calculation, and the optimal decision is the fault section location.
2. A distribution network fault location method based on Markov decision process according to claim 1, characterized in that: According to the fault section location objective function when distributed generation DG is connected to the grid and feeder terminal unit FTU information is distorted, the basic elements required to establish the Markov decision process MDP model are determined, and the Markov decision process is constructed according to the fault section location objective function, and the steps to convert it into a solvable Bellman equation form are as follows: S1.
1. Formulate the coding form of the fault current of the distribution network node j; Fault current coding is divided into 3 cases, the form is as follows: S1.2, according to the coding form of the fault current of the distribution network node I j Feeder section L with node j j The running state builds a switch function, the expression is as follows: Where, L d and L u represents the operating status of the downstream d and upstream u feeder sections of node j respectively; n1 and n2 are the numbers of the upstream and downstream feeder sections at node j respectively; D u and D d They represent the switching coefficients of the upstream and downstream DGs of node j, respectively, and are expressed as follows: Where, M1 and M2 represent the total number of DG accesses upstream and downstream of node j, respectively; m represents the number of DG accesses; and The operating states of the sections H1 and H2 from the node j to the upstream power source M1 and the downstream power source M2, respectively; S1.
3. Construct the fault section location objective function F using the actual uploaded information of the feeder terminal unit and the switch function it , the expression is as follows: In the formula, I j Indicates the actual fault current encoding of node j; I j * (L) represents the expected fault current encoding of node j; K represents the total number of nodes; α is the weight coefficient; N represents the total number of feeder sections in the distribution network containing DG; i represents the number of feeder sections in the distribution network containing DG; is the sum of all feeder section states; and Indicates that all segments have information omissions e j and false positives j The sum of η e and η g The coefficients representing information omission and false alarm respectively; S1.
4. Define the basic elements of the Markov decision model according to the fault section location objective function, including: state variables, decision variables, influencing factors and state transfer function; S1.
5. Establish the Bellman equation based on the basic elements and the fault section location objective function.
3. A distribution network fault location method based on Markov decision process according to claim 2, characterized in that: The steps of establishing the Bellman equation according to the basic elements and the fault section location objective function are as follows: S1.5.
1. Use event Y to indicate that the fault section location objective function has an optimal value, and represent the fault section location objective function under the influence of decision variables and influencing factors; S1.5.
2. Initialize the feeder section status in the basic elements, set the MDP model to make decisions in T time periods, and maximize the probability of event Y under the decision sequence; S1.5.
3. According to the definition of conditional probability and Markov property, the expression obtained in S1.5.2 is derived, and the obtained Bellman equation is: In the formula, Z t (S t ) indicates state S t The objective value function of Z t+1 (S t+1 ) indicates state S t+1 The objective value function of R t (S t ,A t ) represents the immediate probability reward.
4. The method for locating distribution network faults based on a Markov decision process according to claim 1, characterized in that: The steps of using approximate dynamic programming based on a value table to solve the Bellman equation and storing the value function corresponding to the decision variable sequence in the value function table are as follows: S2.
1. Use the reverse traversal calculation method to obtain the value function of all time periods and states; For the state S in period T T , the expression is as follows: In the formula, num(S T ) represents the state S at time period T T Number of occurrences; Indicates that when event Y occurs, state S T Number of occurrences; The rest of the time S t The expression is as follows: In the formula, num(S t →S t+1 ) indicates that from state S t To state S t+1 The number of times; S2.
2. Obtain the value functions of all states in all time periods through S2.1, and store each value function and its corresponding decision variable sequence in the value function table.
5. A distribution network fault location method based on Markov decision process according to claim 1, characterized in that: The optimal decision is obtained from the value function table through strategy iteration calculation. The optimal decision is the location of the fault section. The steps are as follows: S3.
1. Based on the value function table obtained in S2, use the strategy iteration method to search for the largest Z in the value function table. t (S t ) and its corresponding decision A t , recorded as the optimal decision; then the state S of the next period of period t is obtained by calculation t+1 , until all the states of T time periods are traversed; The expression to be calculated is as follows: S3.2, set the number of iterations to k, and execute S2 to S3.1 in sequence for each iteration until Z t (S t ) converges and the iteration is completed; S3.3, Output Z t (S t ) The corresponding decision when convergence It is the optimal decision, that is, the corresponding segment with segment code 1 or -1 is the fault segment.