Distribution Network Defect Management Method and Device Based on Deep Deterministic Policy Gradient

By establishing a power grid environment database and utilizing a deep deterministic strategy gradient method, power grid defects are automatically identified and classified, and defect elimination plans are optimized. This solves the problems of low efficiency and high uncertainty in existing power grid defect management, and achieves stable and safe operation of the power grid.

CN119151721BActive Publication Date: 2025-11-11ZHANJIANG POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411599373.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-11-11
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing power grid defect management methods rely on manual inspection, which is inefficient, results in untimely and inaccurate defect handling, and makes it difficult to handle the rapid processing and analysis of large amounts of defect data. This leads to an impact on power grid stability and security, and the unreasonable allocation of construction resources makes it difficult to effectively carry out defect elimination operations.

Method used

A distribution network defect management method based on deep deterministic strategy gradient is adopted. By establishing a power grid environment database, the agent interacts with the database to train the DDPG model, and the noise is optimized using a quasi-random fractal search algorithm to construct a QRFS-DDPG model. The defect elimination plan is determined to achieve automatic identification, classification and intelligent scheduling of power grid defects.

Benefits of technology

It has improved the intelligence level of power grid defect management, enhanced the accuracy and response speed of defect handling, optimized defect elimination plans, ensured efficient allocation of resources, improved the operating efficiency and safety of the power grid, reduced power outage time, and improved the overall reliability of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151721B_ABST
    Figure CN119151721B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of power technology and provides a method and apparatus for distribution network defect management based on deep deterministic strategy gradient. The method includes: establishing a power grid environment database, wherein the database includes a defect record database, a weather status database, a line data database, and a human resource database; training a DDPG model through interactive learning between an agent and the power grid environment database; optimizing the noise in the DDPG model using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model; and inputting the acquired power grid environment information into the QRFS-DDPG model to determine the defect elimination plan for the distribution network. This disclosure enables the automatic identification, classification, and intelligent scheduling of power grid defects using a deep deterministic strategy gradient algorithm optimized by a quasi-random fractal search algorithm, improving the efficiency and quality of defect handling, ensuring the stable and safe operation of the power grid, and enhancing the operational efficiency and security of the power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of power technology, specifically to a distribution network defect management method and apparatus based on deep deterministic strategy gradient. Background Technology

[0002] In the stable operation of the power system, power grid defect management plays a crucial role. However, current defect management methods face many challenges, mainly reflected in the following aspects: (1) Existing power grid defect management largely relies on manual inspection and recording, which is not only inefficient, but also limited by the number of personnel and their subjective experience level, making it difficult to guarantee the timeliness and accuracy of defect handling. (2) If new defects cannot be included in the schedule in a timely manner, it will affect the operating status of equipment and may even cause power safety accidents, threatening the stability and safety of the power grid. (3) The shortage of construction forces and the unreasonable schedule for defect elimination will make it difficult to carry out defect elimination effectively, especially when defect elimination work needs to be carried out in conjunction with power outage plans, the insufficiency of resource allocation will be more prominent. (4) The imperfection of manual inspection and data collection leads to data omissions, and traditional methods are difficult to handle the rapid processing and analysis of a large amount of defect data, which greatly increases the complexity and uncertainty of power grid defect management. (5) The impact of severe weather conditions on the construction progress also greatly increases the complexity and uncertainty of power grid defect management. Summary of the Invention

[0003] The purpose of this disclosure is to provide a distribution network defect management method and apparatus based on deep deterministic strategy gradient, so as to at least solve the technical problems of the high complexity and uncertainty of power grid defect management in the prior art, which affect the stability and safe operation of the power grid.

[0004] To solve the above-mentioned technical problems, the embodiments of this disclosure adopt the following technical solutions:

[0005] In a first aspect, embodiments of this disclosure provide a distribution network defect management method based on an improved deep deterministic strategy, including:

[0006] Establish a power grid environment database, which includes a defect record database, a weather status database, a line data database, and a human resources database.

[0007] The DDPG model is trained through interaction between the agent and the power grid environment database.

[0008] The noise in the DDPG model was optimized using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model;

[0009] The acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan for the distribution network.

[0010] In some embodiments, training a DDPG model through interaction between an agent and the power grid environment database includes:

[0011] Construct a strategy network and a value network, wherein the strategy network outputs dispatch order decisions based on the current power grid status, and the value network evaluates the expected returns of the dispatch order decisions;

[0012] The policy network and value network are trained using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy.

[0013] In some embodiments, constructing a policy network and a value network includes:

[0014] Determine the policy function and value function , wherein the strategy function , indicating that under a given state s, according to the parameters Given a defined strategy distribution, output action a, the value function Predict the action a to be performed in state s and follow the policy. The expected return is given by the state s, which represents the power grid environment in which the agent is located, and the action a, which represents the defect work order dispatched by the agent. The defect work order includes the assignment to a preset defect team and the expected defect elimination time. Indicates policy network parameters, Indicates the parameters of the value network;

[0015] Determine the objective function The objective function is defined as starting from the current state s and following the policy. The sum of the discounts received is expressed as:

[0016] ,

[0017] In the formula, The discount factor for the time step. The immediate reward obtained at time t;

[0018] The policy network and value network are trained using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy, including:

[0019] The policy network parameters are updated using policy gradients to optimize the objective function and learn the optimal policy.

[0020] The value network parameters are updated by minimizing the prediction error of the Q-value to learn the value of the state-action pair and obtain the optimal value.

[0021] In some embodiments, updating the policy network parameters via policy gradients includes:

[0022] The policy network parameters are optimized using gradient ascent, as follows:

[0023] ,

[0024] in, It is the action output by the policy network based on the current state s.

[0025] In some embodiments, updating the value network parameters by minimizing the prediction error of the Q-value includes:

[0026] The loss function for updating the value network is determined as follows:

[0027] ,

[0028] In the formula, y is the target Q value, and the formula for calculating the target Q value is:

[0029] ,

[0030] In the formula, It's an instant reward. Q' is the discount factor, and Q' is the output of the target value network.

[0031] In some embodiments, the method further includes:

[0032] Obtain distribution network experience data, wherein the distribution network experience data includes at least one of the following: defect status, work order execution status, and the difference between the actual defect elimination time and the planned elimination time;

[0033] The strategy estimate and value estimate are updated based on the aforementioned distribution network experience data;

[0034] The DDPG model is trained using the updated policy estimate and value estimate.

[0035] In some embodiments, the acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan for the distribution network, including:

[0036] The importance of eliminating distribution network defects in the power grid environment information is scored using the QRFS-DDPG model, and the scoring results are obtained.

[0037] Based on the scoring results, the handling of the distribution network defects is scheduled and work orders are dispatched.

[0038] In some embodiments, the noise in the DDPG model is optimized using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model, including:

[0039] The optimal noise level under a preset power grid environment is searched iteratively by gradually narrowing the search space;

[0040] Adding the optimal noise to the action output by the policy network is represented as:

[0041] ,

[0042] in, The action selected at time step t. It is the action output by the Actor network. It is the noise term extracted from the Ornstein-Uhlenbeck process at each step.

[0043] In some embodiments, before determining the policy function and the value function, the method further includes:

[0044] The agent initializes the policy network parameters, value network parameters, and corresponding target network parameters.

[0045] Secondly, embodiments of this disclosure provide a distribution network defect management device based on an improved deep deterministic strategy, comprising:

[0046] The database creation module is configured to create a power grid environment database, which includes a defect record database, a weather status database, a line database, and a human resources database.

[0047] The model training module is configured to train the DDPG model through interaction between the agent and the power grid environment database.

[0048] The model optimization module is configured to use a quasi-random fractal search algorithm to optimize the noise in the DDPG model, thereby obtaining a QRFS-DDPG model;

[0049] The defect management module is configured to input the acquired power grid environment information into the QRFS-DDPG model to determine the defect elimination plan for the distribution network.

[0050] Thirdly, embodiments of this disclosure provide an electronic device, including at least a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above-described distribution network defect management method based on an improved deep deterministic strategy when executing the computer program in the memory.

[0051] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described distribution network defect management method based on an improved deep deterministic strategy.

[0052] This disclosure provides a distribution network defect management method and apparatus based on deep deterministic strategy gradient. It establishes a power grid environment database, including a defect record database, a weather status database, a line data database, and a human resource database. A Direct Distribution Generation Program (DDPG) model is trained through interactive learning between an agent and the database. Noise in the DDPG model is optimized using a quasi-random fractal search algorithm to obtain a Quasi-Randomized Fractal Search (QRFS-DDPG) model. The acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan. This method utilizes reinforcement learning algorithms to achieve automatic identification, classification, and intelligent scheduling of power grid defects, reducing reliance on human experience, improving the intelligence level of power grid defect management, increasing the accuracy and response speed of defect handling, and improving the efficiency and quality of defect handling, thus ensuring the stable and safe operation of the power grid. Simultaneously, it can optimize the defect elimination plan, scientifically and rationally arranging defect elimination plans to ensure efficient allocation and use of resources, improving the operational efficiency and safety of the power grid. For example, it can better coordinate human resources, optimize the defect elimination plan, reduce power outage time, and improve the overall reliability of power grid operation. Furthermore, since the DDPG algorithm has the advantage of stability and adaptability to high-dimensional continuous action space, and distribution network defect management decisions are easily affected by the environment, this embodiment improves the deep deterministic policy gradient method in reinforcement learning to obtain a distribution network defect management method based on the optimized DDPG network. This method can enhance the adaptability and flexibility of power grid defect management. Through the adaptability of the algorithm, the defect management system can adapt to various complex environments and emergencies, ensuring that the power grid can still maintain stable operation when facing uncertainty. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A flowchart of a distribution network defect management method based on an improved deep deterministic strategy according to an embodiment of this disclosure;

[0055] Figure 2 This describes the interaction mode between the intelligent agent and the environment in embodiments of this disclosure.

[0056] Figure 3 This is the QRFS algorithm framework of this disclosure embodiment;

[0057] Figure 4 This is a schematic diagram of the structure of a distribution network defect management device based on an improved deep deterministic strategy according to an embodiment of this disclosure. Detailed Implementation

[0058] Various embodiments and features of this disclosure are described herein with reference to the accompanying drawings.

[0059] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this disclosure will be apparent to those skilled in the art.

[0060] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present disclosure and, together with the general description of the disclosure given above and the detailed description of the embodiments given below, serve to explain the principles of the disclosure.

[0061] These and other features of this disclosure will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.

[0062] It should also be understood that although this disclosure has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this disclosure, which have the features described in the claims and are therefore all within the scope of protection defined herein.

[0063] The above and other aspects, features and advantages of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.

[0064] Specific embodiments of this disclosure are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this disclosure, which may be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure this disclosure. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely to serve as the basis and representative basis for the claims to teach those skilled in the art to use this disclosure in a variety of substantially any suitable detailed structures.

[0065] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in still another embodiment,” all of which may refer to one or more of the same or different embodiments according to this disclosure.

[0066] First, the terms that may be used in the following text will be explained. Table 1 shows the proper nouns that may be used and their definitions.

[0067] Table 1. Definitions of potentially relevant proper nouns

[0068]

[0069] Example 1

[0070] Figure 1 A flowchart of a distribution network defect management method based on an improved deep deterministic strategy according to an embodiment of this disclosure is shown, such as... Figure 1 As shown in the figure, an embodiment of this disclosure provides a distribution network defect management method based on an improved deep deterministic strategy, comprising:

[0071] S101: Establish a power grid environment database.

[0072] The power grid data includes a defect record database, a weather status database, a line data database, and a human resources database.

[0073] The defect record database includes the following statuses: 1. Defect category; 2. Defect severity; 3. Defect urgency. The weather status database includes: 1. Defect elimination date; 2. Daily high and low temperatures; 3. Maximum wind speed; 4. Humidity. The line data database includes: 1. Line service life; 2. Line trip records. The human resources database includes: 1. Roster of available personnel for each shift; 2. Number of operators at different skill levels.

[0074] The four databases mentioned above together form an environment in which intelligent agents can learn and interact.

[0075] S102: Train the DDPG model through interaction between the agent and the power grid environment database.

[0076] Deep Deterministic Policy Gradient (DDPG) is an advanced deep reinforcement learning algorithm specifically designed to handle problems in continuous action spaces. By combining the policy-value (Actor-Critic) framework with deep learning techniques, it utilizes deterministic policy gradients for optimization, effectively learning the optimal action to take given a state. The core idea of ​​the DDPG algorithm is to solve reinforcement learning problems in continuous action spaces by combining deterministic policy gradient methods and deep learning techniques.

[0077] like Figure 2 As shown, in this step, after establishing the power grid environment database, the agent interacts and learns with the environment to train the DDPG model and obtain a strategy model for managing distribution network defects.

[0078] Step S102 specifically includes:

[0079] S1021: Construct a strategy network (Actor network) and a value network (Critic network), wherein the strategy network outputs dispatch order decisions based on the current power grid status, and the value network evaluates the expected return of the dispatch order decisions;

[0080] S1022: The policy network and value network are trained using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy.

[0081] At each time step, the agent interacts with the power grid environment by executing actions (dispatching defect work orders) and collects the resulting state transitions and reward information to derive a distribution network defect management strategy. This process is iterative; in each iteration, the agent adjusts its strategy based on the latest state transitions and reward information to obtain the optimal strategy, i.e., the best defect management model. The Actor network and Critic network collaborate: the Actor network generates actions, and the Critic network provides feedback, jointly driving the strategy towards higher rewards.

[0082] S103: The noise in the DDPG model is optimized using the quasi-random fractal search algorithm to obtain the QRFS-DDPG model.

[0083] The DDPG model contains some noise, which affects its accuracy. Therefore, in power grid defect management, optimizing noise parameters can improve the adaptability and flexibility of the agent in complex environments, thereby achieving a better defect handling strategy.

[0084] S104: Input the acquired power grid environment information into the QRFS-DDPG model to determine the defect elimination plan for the distribution network.

[0085] After constructing the QRFS-DDPG model, the acquired power grid environment information can be input into the QRFS-DDPG model. The QRFS-DDPG model processes the power grid environment information, identifies power grid defects, classifies the power grid defects, and sets the best treatment plan for the power grid defects based on the identified power grid defects or the classification results. It also intelligently schedules the treatment of each defect to the optimal defect elimination plan.

[0086] Step S104 specifically includes:

[0087] S1041: Using the QRFS-DDPG model, score the importance of eliminating distribution network defects in the power grid environment information to obtain the score results;

[0088] S1042: Schedule and dispatch work orders for the handling of the distribution network defects based on the scoring results.

[0089] When the power grid environment information contains multiple defects, the QRFS-DDPG model can score the importance of eliminating each defect, and schedule the handling of distribution network defects and dispatch work orders according to the importance from high to low, so as to obtain the optimal defect elimination plan.

[0090] The distribution network defect management method based on an improved deep deterministic strategy provided in this disclosure establishes a power grid environment database, including a defect record database, a weather status database, a line data database, and a human resource database. A DDPG model is trained through interactive learning between an agent and the power grid environment database. Noise in the DDPG model is optimized using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model. The acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan. This method utilizes reinforcement learning algorithms to achieve automatic identification, classification, and intelligent scheduling of power grid defects, reducing reliance on human experience, improving the intelligence level of power grid defect management, increasing the accuracy and response speed of defect handling, and improving the efficiency and quality of defect handling, thus ensuring the stable and safe operation of the power grid. Simultaneously, it can optimize the defect elimination plan, scientifically and rationally arranging the elimination of power grid defects to ensure efficient allocation and use of resources, improving the operational efficiency and safety of the power grid. For example, it can better coordinate human resources, optimize the elimination plan, reduce power outage time, and improve the overall reliability of power grid operation.

[0091] Furthermore, since the DDPG algorithm has the advantage of stability and adaptability to high-dimensional continuous action space, and distribution network defect management decisions are easily affected by the environment, this embodiment improves the deep deterministic policy gradient method in reinforcement learning to obtain a distribution network defect management method based on the optimized DDPG network. This method can enhance the adaptability and flexibility of power grid defect management. Through the adaptability of the algorithm, the defect management system can adapt to various complex environments and emergencies, ensuring that the power grid can still maintain stable operation when facing uncertainty.

[0092] In some embodiments, step S1021, constructing the policy network and the value network, includes:

[0093] S201: Determine the strategy function and value function .

[0094] Wherein, the strategy function , indicating that under a given state s, according to the parameters Given a defined strategy distribution, output action a, the value function Predict the action a to be performed in state s and follow the policy. The expected return is given by the state s, which represents the power grid environment in which the agent is located, and the action a, which represents the defect work order dispatched by the agent. The defect work order includes the assignment to a preset defect team and the expected defect elimination time. Indicates policy network parameters, Indicates the parameters of the value network;

[0095] S202: Determine the objective function .

[0096] The objective function clarifies the learning objective of the agent, namely, maximizing the expected reward. Specifically, the objective function is defined as starting from the current state s and following the policy... The sum of the discounts received is expressed as:

[0097] (1)

[0098] In equation (1), The discount factor for the time step. The instantaneous reward obtained at time t.

[0099] Before constructing the policy network and value network, it is necessary to initialize the policy network parameters, value network parameters, and corresponding target network parameters through the agent.

[0100] Furthermore, in step S1022, the policy network and value network are trained using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy, including:

[0101] S301: Update the policy network parameters through policy gradients to optimize the objective function and learn the optimal policy;

[0102] S302: Update the value network parameters by minimizing the prediction error of the Q value to learn the value of the state-action pair and obtain the optimal value.

[0103] Policy gradient descent uses the policy gradient theorem to derive how to update policy parameters to optimize the objective function. For deterministic policies, according to the policy gradient theorem, gradient ascent can be used to optimize the policy network parameters. :

[0104] (2)

[0105] in, These are the parameters of the Actor network. Critic network parameters, Actor network parameters Update the policy using policy gradients to learn the optimal policy. This refers to the action output by the Actor network based on the current state s. Critic network parameters. The value of the state-action pair is learned by updating the prediction error by minimizing the Q-value.

[0106] To stabilize the learning process, a target network is used. and These are slow-updated copies of the main network parameters. The Actor target network is used to provide the policy for the next state. The Critic target network is used to evaluate the policy for the next state.

[0107] The loss function of the Critic network is the mean squared error between the predicted Q-value and the target Q-value. The loss function measures the difference between the Critic network's current prediction of the action's value and the actual reward obtained (estimated by the target network). It follows the formula below:

[0108] (3)

[0109] Where y is the target Q value.

[0110] Calculating the target Q-value is a crucial intermediate step required for updating the Critic network. The calculation method is as follows:

[0111] (4)

[0112] Formula (4) uses the Bellman equation, where, It's an instant reward. Q' is the discount factor, and Q' is the output of the target Critic network.

[0113] After the action is executed, both the Actor network and the Critic network need to update their parameters. Specifically, gradient ascent and policy gradient are used to update the Critic network and Actor network respectively, following the formulas below:

[0114] (5)

[0115] (6)

[0116] In addition, the target network parameters also need to be soft-updated to ensure that the target network smoothly follows the main network, thereby maintaining the stability of the learning process. The update follows the following formula:

[0117] (7)

[0118] in, It is a positive number close to 0, used to control the speed at which the target network parameters are updated.

[0119] In some embodiments, the method further includes:

[0120] S301: Obtain distribution network experience data, wherein the distribution network experience data includes at least one of the following: defect status, work order execution status, and the difference between the actual defect elimination time and the planned elimination time;

[0121] S302: Update the strategy estimate and value estimate based on the aforementioned distribution network experience data;

[0122] S303: Train the DDPG model using the updated policy estimate and value estimate.

[0123] In this embodiment, an experience playback mechanism is introduced, whereby the agent plays back the collected experience data. Store to playback buffer In this process, these experiences are used for subsequent batch learning and policy evaluation, updating the policy and value estimates, and reducing the correlation between samples through empirical data. Among these, the reward... It is based on the difference between the actual time to eliminate the defect and the planned time to eliminate it. If the actual time to eliminate the defect is earlier than or equal to the planned time, a positive reward is given; if it is later than the planned time, a negative reward is given. This simple binary reward function can be adjusted as needed to reflect the reward intensity under different circumstances.

[0124] The agent uses reward or penalty signals (based on a comparison between the actual and planned elimination time of the defect). In this embodiment, this immediate, performance-based reward mechanism incentivizes the agent to optimize work order generation, thereby reducing defect processing time, improving grid operation efficiency, and enhancing the ability to intelligently, rationally, and efficiently schedule distribution network defects.

[0125] In some embodiments, step S103 involves optimizing the noise in the DDPG model using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model, including:

[0126] S1031: Iteratively search for the optimal noise under a preset power grid environment by gradually narrowing the search space;

[0127] S1032: Add the optimal noise to the action output by the policy network.

[0128] In the DDPG algorithm, noise is added to the action space via an Ornstein-Uhlenbeck process to enhance the agent's exploratory behavior during training. The Ornstein-Uhlenbeck process is a time-dependent stochastic process that generates continuous, biased noise. However, noise itself also has some drawbacks: (1) The noise level is difficult to determine: if the noise is too high, it may cause the agent to take ineffective or harmful actions; if the noise is too low, it may not be enough to encourage sufficient exploration. (2) Noise may affect learning stability: in some cases, adding noise may make the learning process unstable, especially in complex or high-risk environments. (3) Noise parameter tuning may be complex: finding a suitable noise level may require a lot of tuning and experimentation.

[0129] To address the aforementioned issues, this disclosure introduces a quasi-random fractal search (QRFS) algorithm to optimize noise in the DDPG. The QRFS algorithm can automatically adjust noise parameters to achieve better policy performance. Through the iterative search process of the algorithm, the optimal exploration noise level under specific conditions can be found, thereby guiding the agent to conduct effective exploration, accelerating the learning process, and improving the intelligence level of power grid defect management.

[0130] QRFS is a novel metaheuristic algorithm that utilizes fractal geometry, low-dispersion sequences, and intelligent search space partitioning techniques to solve complex global optimization problems. Based on the generation of fractals, the algorithm seeks the global minimum of the objective function by carefully selecting low-dispersion sequences to initialize the population. This process strategically narrows the fractal search space, guiding each population towards the most promising regions. The key operator of the algorithm is the initialization of different fractal populations. Compared to random initialization, initializing these populations with low-dispersion sequences comprehensively enhances the algorithm's ability to explore the search space. This favorable initialization allows the algorithm to identify promising regions early in the iteration process. At the end of the iteration, the focus shifts to refining the found optimal regions. This refinement is facilitated by the dynamic properties of the population controlled by the sigmoid function. The sigmoid function coordinates the gradual reduction of the population. Therefore, the QRFS algorithm focuses on exploring optimal solutions in the initial stage and continuously improves these solutions until the end of the iteration process. The flowchart of QRFS is as follows. Figure 2 As shown.

[0131] QRFS incorporates population management methods where population initialization is coordinated by low-discrepancy sequences, referred to as quasi-random or quasi-random number sequences. Unlike purely random sequences, these deterministic sequences provide a more uniform spatial distribution, facilitating better coverage of the search space or data. In QRFS, the population for each fractal includes:

[0132] (8)

[0133] Where npop represents the population size decreasing according to the sigmoid function, i.e., the population scale. npop simulates the gradually decreasing population size throughout the iteration process; l b and u b The lower and upper bounds are dynamically changed during the iteration process to adapt to the ever-changing problem space; the dimension of the problem is captured by dims; sequence represents one of four randomly selected low-discretion values.

[0134] The QRFS algorithm uses the sigmida function contained in equation (9) below to manage the overall reduction of each fractal. This equation coordinates the reduction of the population size from its initial value to its final value, and the parameters therein can adjust the rate of reduction.

[0135] (9)

[0136] In equation (9), npop represents the population size. and These represent the initial and final population sizes, respectively. I represents the total number of iterations, and I represents the current iteration number.

[0137] Initialization is one of the key operations in the QRFS algorithm and is crucial in its exploration and development. As the search space is gradually narrowed and the number of algorithms is reduced during the iterative process, the initial phase becomes a major driving force for these fundamental aspects. The initialization of the algorithm begins by creating a named population (nfractals), where each population represents a fractal and is selected through a random sequence of low dissimilarity. Then, the fitness of each solution across all fractal populations is determined, including identifying the best and worst solutions, followed by computation of the best solution. And the worst-case solution The Euclidean distance (ED) between the two sides is calculated, and then the boundary of the fractal is redefined by adding and subtracting the aforementioned Euclidean distance from the optimal solution. The formula for calculating the Euclidean distance is:

[0138] (10)

[0139] The above process transforms the optimal solution into the centroid, and the two new solutions obtained by adding or subtracting the Euclidean distance become the updated upper and lower bounds of the fractal. Redefining the boundary of the fractal using the Euclidean distance is to ensure that the new constraints do not exceed the search space defined by the function being evaluated.

[0140] The iterative process of the QRFS algorithm involves gradually narrowing the search space by reducing the upper and lower bounds of each fractal and simultaneously reducing the overall size of the algorithm. Initially, the population of nfractals is established according to their respective new constraints, i.e., determined by the initialization process defined in Equation (8) above.

[0141] After overall initialization, the fitness of the solutions within each fractal population is evaluated to determine the optimal and worst solutions. The optimal and worst solutions are crucial for determining the centroid of the fractal.

[0142] In each iteration, the centroid with the highest best fit is designated as the overall best centroid, and a new centroid for each fractal is calculated using formula (11).

[0143] (11)

[0144] In formula (11), This indicates the location of the optimal solution for a given fractal. Let j represent the optimal solution among all fractals, where j ranges from 1 to nfractals.

[0145] Then, the distance between the worst solution of each fractal and the new centroid is calculated using formula (12):

[0146] (12)

[0147] In formula (12), This represents the location of the worst-case solution for a given fractal. This indicates the location of the new centroid of the fractal. Subsequently, from the new centroid ( Add or subtract this distance (ED) from the original text. Ncj This ensures that the obtained solution becomes the new upper and lower bounds of the fractal. Understandably, the above solution remains within the search space defined by the evaluated function. Repeat this process until the iterative loop is complete.

[0148] Using a quasi-random fractal search algorithm, the search space is gradually narrowed iteratively until the optimal noise under the preset power grid environment is found. In the implementation of the DDPG algorithm, the optimized optimal noise is added to the action output by the Actor network. The formula can be expressed as:

[0149] (13)

[0150] in, The action selected at time step t. The action output by the Actor network This is the noise term extracted from the Ornstein-Uhlenbeck process at each step. This noise term is the noise optimized using the aforementioned quasi-random fractal search algorithm.

[0151] The Ornstein-Uhlenbeck process is a time-dependent stochastic process, and its update rule can be expressed as:

[0152] (14)

[0153] in, It is the attenuation rate. It is the mean of the noise. It is the standard deviation of noise. It is a random sample drawn from the standard normal distribution. It is the time step.

[0154] In power grid defect management, optimizing noise parameters through the aforementioned quasi-random fractal search algorithm can improve the adaptability and flexibility of the agent in complex environments, thereby enabling a better defect handling strategy.

[0155] It should be noted that the swarm intelligence optimization (population management) algorithm used to optimize noise parameters can include various algorithms, such as particle swarm optimization and ant colony optimization. Other optimization algorithms can also optimize noise parameters. Furthermore, embodiments of this disclosure can employ multi-agent reinforcement learning algorithms to simulate the collaborative work of multiple agents (such as different defect elimination teams) in a power grid. Each agent can learn the optimal defect elimination strategy based on its own task allocation, resource status, and environmental feedback, thereby obtaining the best defect elimination arrangement for different defect elimination teams and improving the applicability of the defect management scheme.

[0156] In summary, the distribution network defect management method based on an improved deep deterministic strategy provided in this disclosure collects power grid environmental information, including defect record databases, weather status databases, line data databases, and human resource databases. It uses a deep deterministic strategy gradient algorithm optimized by a quasi-random fractal search algorithm to learn information such as defect data, weather status, lines, and human resources. Based on the difference between the actual defect elimination time and the planned elimination time as a reward or penalty signal, it can formulate a more comprehensive and reasonable defect elimination strategy, effectively improve the automatic management capability of distribution network defects, and ensure the stable operation of the power grid and the reliability of power supply.

[0157] Example 2

[0158] Figure 4 A schematic diagram of the structure of a distribution network defect management device based on an improved deep deterministic strategy according to an embodiment of this disclosure is shown, as follows: Figure 4As shown, this disclosure provides a distribution network defect management device based on an improved deep deterministic strategy, comprising:

[0159] Model building module 10 is configured as a database building module, specifically for building a power grid environment database, which includes a defect record database, a weather status database, a line database, and a human resources database.

[0160] The model training module 20 is configured to train the DDPG model through interaction between the agent and the power grid environment database.

[0161] The model optimization module 30 is configured to use a quasi-random fractal search algorithm to optimize the noise in the DDPG model to obtain a QRFS-DDPG model;

[0162] The defect management module 40 is configured to input the acquired power grid environment information into the QRFS-DDPG model to determine the defect elimination plan for the distribution network.

[0163] In some embodiments, the model training module 20 includes:

[0164] The construction unit is configured to construct a strategy network and a value network, wherein the strategy network outputs dispatch order decisions based on the current power grid status, and the value network evaluates the expected return of the dispatch order decisions.

[0165] The training unit is configured to train the policy network and the value network using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy.

[0166] In some embodiments, the building unit is further configured as follows:

[0167] Determine the policy function and value function , wherein the strategy function , indicating that under a given state s, according to the parameters Given a defined strategy distribution, output action a, the value function Predict the action a to be performed in state s and follow the policy. The expected return is given by the state s, which represents the power grid environment in which the agent is located, and the action a, which represents the defect work order dispatched by the agent. The defect work order includes the assignment to a preset defect team and the expected defect elimination time. Indicates policy network parameters, Indicates the parameters of the value network;

[0168] Determine the objective function The objective function is defined as starting from the current state s and following the policy. The sum of the discounts received is expressed as:

[0169] ,

[0170] In the formula, The discount factor for the time step. The immediate reward obtained at time t;

[0171] The training unit is further configured as follows:

[0172] The policy network parameters are updated using policy gradients to optimize the objective function and learn the optimal policy.

[0173] The value network parameters are updated by minimizing the prediction error of the Q-value to learn the value of the state-action pair and obtain the optimal value.

[0174] In some embodiments, the training unit is further configured as follows:

[0175] The policy network parameters are optimized using gradient ascent, as follows:

[0176] ,

[0177] in, It is the action output by the policy network based on the current state s.

[0178] In some embodiments, the training unit is further configured as follows:

[0179] The loss function for updating the value network is determined as follows:

[0180] ,

[0181] In the formula, y is the target Q value, and the formula for calculating the target Q value is:

[0182] ,

[0183] In the formula, It's an instant reward. Q' is the discount factor, and Q' is the output of the target value network.

[0184] In some embodiments, the model training module 20 is further configured to:

[0185] Obtain distribution network experience data, wherein the distribution network experience data includes at least one of the following: defect status, work order execution status, and the difference between the actual defect elimination time and the planned elimination time;

[0186] The strategy estimate and value estimate are updated based on the aforementioned distribution network experience data;

[0187] The DDPG model is trained using the updated policy estimate and value estimate.

[0188] In some embodiments, the defect management module 40 is further configured as follows:

[0189] The importance of eliminating distribution network defects in the power grid environment information is scored using the QRFS-DDPG model, and the scoring results are obtained.

[0190] Based on the scoring results, the handling of the distribution network defects is scheduled and work orders are dispatched.

[0191] In some embodiments, the model optimization module 30 is further configured as follows:

[0192] The optimal noise level under a preset power grid environment is searched iteratively by gradually narrowing the search space;

[0193] Adding the optimal noise to the action output by the policy network is represented as:

[0194] ,

[0195] in, The action selected at time step t. It is the action output by the Actor network. It is the noise term extracted from the Ornstein-Uhlenbeck process at each step.

[0196] In some embodiments, the construction unit is further configured to: initialize policy network parameters and value network parameters, as well as corresponding target network parameters, through an agent before determining the policy function and value function.

[0197] The distribution network defect management device based on the improved deep deterministic strategy provided in this disclosure corresponds to the distribution network defect management method based on the improved deep deterministic strategy in the above embodiments. Any option in the embodiments of the distribution network defect management method based on the improved deep deterministic strategy is also applicable to the embodiments of the distribution network defect management device based on the improved deep deterministic strategy, and will not be repeated here.

[0198] Example 3

[0199] This disclosure also provides an electronic device, including at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described distribution network defect management method based on an improved deep deterministic strategy when executing the computer program in the memory.

[0200] In some embodiments, the processor executing a computer program may be a processing device that includes one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that runs other instruction sets, or a processor that runs a combination of instruction sets. The processor may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), system-on-a-chip (SoCs), etc.

[0201] The memory may be a read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash drives or other forms of flash memory, cache, registers, static memory, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape cassette or other magnetic storage devices, or any other possible non-transitory medium used to store information or instructions that can be accessed by computer equipment.

[0202] The electronic devices disclosed herein may include, but are not limited to, fixed terminal devices such as servers, desktop computers, and digital TVs, as well as mobile terminal devices such as in-vehicle devices (e.g., head-up displays), handheld devices (e.g., mobile phones, tablets, etc.), and wearable devices (e.g., smartwatches, smart bracelets, etc.).

[0203] Example 4

[0204] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described distribution network defect management method based on an improved deep deterministic strategy.

[0205] The computer-readable storage medium of this disclosure can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device; for example, it can be the memory described above.

[0206] The computer programs of embodiments of this disclosure can be organized into one or more computer-executable components or modules. Various aspects of this disclosure can be implemented with any number and combination of such components or modules. For example, aspects of this disclosure are not limited to the specific computer-executable instructions or particular components or modules shown in the drawings and described herein. Other embodiments may include different computer-executable instructions or components having more or fewer functions than those shown and described herein.

[0207] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A distribution network defect management method based on an improved deep deterministic strategy, characterized in that, include: Establish a power grid environment database, which includes a defect record database, a weather status database, a line data database, and a human resources database. The DDPG model is trained through interaction between the agent and the power grid environment database. The noise in the DDPG model was optimized using a quasi-random fractal search algorithm to obtain a QRFS-DDPG model; The acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan for the distribution network. The step of training the DDPG model through interaction between the agent and the power grid environment database includes: Construct a strategy network and a value network, wherein the strategy network outputs dispatch order decisions based on the current power grid status, and the value network evaluates the expected returns of the dispatch order decisions; The policy network and value network are trained by deterministic policy gradient algorithm and deep learning to obtain distribution network defect management policy. The agent interacts with the power grid environment by performing actions at each time step and collects the resulting state transition and reward information to obtain the distribution network defect management policy. The actions include dispatching defect work orders. The step of optimizing the noise in the DDPG model using a quasi-random fractal search algorithm to obtain the QRFS-DDPG model includes: The optimal noise level under a preset power grid environment is searched iteratively by gradually narrowing the search space; The optimal noise is added to the action output by the policy network.

2. The method according to claim 1, characterized in that, Constructing policy networks and value networks includes: Determine the policy function and value function , wherein the strategy function , indicating that under a given state s, according to the parameters Given a defined strategy distribution, output action a, the value function Predict the action a to be performed in state s and follow the policy. The expected return is given by the state s, which represents the power grid environment in which the agent is located, and the action a, which represents the defect work order dispatched by the agent. The defect work order includes the assignment to a preset defect team and the expected defect elimination time. Indicates policy network parameters, Indicates the parameters of the value network; Determine the objective function The objective function is defined as starting from the current state s and following the policy. The sum of the discounts received is expressed as: , In the formula, The discount factor for the time step. The immediate reward obtained at time t; The policy network and value network are trained using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy, including: The policy network parameters are updated using policy gradients to optimize the objective function and learn the optimal policy. The value network parameters are updated by minimizing the prediction error of the value function to learn the value of state-action pairs and obtain the optimal value.

3. The method according to claim 2, characterized in that, Updating the policy network parameters via policy gradients includes: The policy network parameters are optimized using gradient ascent, as follows: , in, It is the action output by the policy network based on the current state s.

4. The method according to claim 2, characterized in that, Updating the value network parameters by minimizing the prediction error of the value function includes: The loss function for updating the value network is determined as follows: , In the formula, y is the target Q value, and the formula for calculating the target Q value is: , In the formula, It's an instant reward. Q' is the discount factor, and Q' is the output of the target value network.

5. The method according to claim 1, characterized in that, The method further includes: Obtain distribution network experience data, wherein the distribution network experience data includes at least one of the following: defect status, work order execution status, and the difference between the actual defect elimination time and the planned elimination time; The strategy estimate and value estimate are updated based on the aforementioned distribution network experience data; The DDPG model is trained using the updated policy estimate and value estimate.

6. The method according to claim 1, characterized in that, The acquired power grid environment information is input into the QRFS-DDPG model to determine the defect elimination plan for the distribution network, including: The importance of eliminating distribution network defects in the power grid environment information is scored using the QRFS-DDPG model, and the scoring results are obtained. Based on the scoring results, the handling of the distribution network defects is scheduled and work orders are dispatched.

7. The method according to claim 1, characterized in that, The action of adding the optimal noise to the output of the policy network is expressed as: , in, The action selected at time step t. It is the action output by the Actor network. It is the noise term extracted from the Ornstein-Uhlenbeck process at each step.

8. The method according to claim 2, characterized in that, Before determining the policy function and the value function, the method further includes: The agent initializes the policy network parameters, value network parameters, and corresponding target network parameters.

9. A distribution network defect management device based on an improved deep deterministic strategy, characterized in that, include: The database creation module is configured to create a power grid environment database, which includes a defect record database, a weather status database, a line database, and a human resources database. The model training module is configured to train the DDPG model through interaction between the agent and the power grid environment database. The model optimization module is configured to use a quasi-random fractal search algorithm to optimize the noise in the DDPG model, thereby obtaining a QRFS-DDPG model; The defect management module is configured to input the acquired power grid environment information into the QRFS-DDPG model to determine the defect elimination plan for the distribution network. The model training module includes: The construction unit is configured to construct a strategy network and a value network, wherein the strategy network outputs dispatch order decisions based on the current power grid status, and the value network evaluates the expected return of the dispatch order decisions. The training unit is configured to train the policy network and the value network using a deterministic policy gradient algorithm and deep learning to obtain a distribution network defect management policy. The agent interacts with the power grid environment by performing actions at each time step and collects the resulting state transition and reward information to obtain the distribution network defect management policy. The actions include dispatching defect work orders. The model optimization module is further configured as follows: The optimal noise level under a preset power grid environment is searched iteratively by gradually narrowing the search space; The optimal noise is added to the action output by the policy network.

Citation Information

Patent Citations

  • Power distribution network fault intelligent repair method and device based on deep reinforcement learning

    CN111401769A