An improved DQN method for determining the equipotential path for live working personnel

By optimizing the live working path through an improved DQN algorithm and combining simulation modeling and dynamic strategies, the problem of high simulation test cost and inability to select the optimal path in the existing technology is solved, and efficient and safe equipotential path selection is achieved.

CN119312999BActive Publication Date: 2025-09-16GUO JIA DIAN WANG YOU XIAN GONG SI XI NAN FEN BU +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411352858.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-09-16
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

When evaluating the equipotential path for live workers to enter, existing technologies involve extremely costly simulation tests that are unable to screen out the optimal path and are unable to effectively evaluate safety and selectivity.

Method used

An improved deep Q-network (DQN) algorithm is adopted, combined with a dynamic greedy strategy and information importance distinction. Through simulation modeling, the surface field strength and discharge hazard rate are calculated, and the reward function and action rules are designed to optimize the entry into the equipotential path.

Benefits of technology

It improves the safety and comfort of live working personnel entering the equipotential path, reduces the cost of physical testing, and realizes fast and efficient optimal path selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312999B_ABST
    Figure CN119312999B_ABST
Patent Text Reader

Abstract

The present invention provides an improved DQN method for determining the equipotential path for live workers. It simulates and models live work performed using the hanging basket method, calculating the surface field strength and discharge hazard rate of workers at various spatial locations that they may pass through when entering the equipotential phase. Based on the actual conditions of live work, the DQN reward function is redesigned, adopting an improved dynamic greedy strategy and distinguishing the importance of samples, ultimately obtaining the optimal path for live workers to enter the equipotential phase. The present invention uses reinforcement learning to optimize the equipotential path for live work, improving the safety and comfort of workers entering the equipotential phase. The reward function, action rules, and time-varying dynamic greedy strategy of the DQN algorithm are designed to achieve autonomous learning for entering the equipotential path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention provides an improved DQN method for determining when live-work personnel enter an equipotential path, and belongs to the technical field of high-voltage power transmission. Background Art

[0002] With the rapid development of society and the economy, large-scale, long-distance power transmission projects have been built and put into operation one after another. In this process, to ensure the stable operation of these critical transmission projects, the corresponding maintenance technologies are constantly improving and refining. In this context, live working on DC transmission lines, as a core and critical link, plays an increasingly important role. Live working on DC transmission lines refers to the technology that inspects and maintains DC transmission lines while the power system is in operation, without interrupting power transmission. Live working on DC transmission lines can promptly detect and address line faults, reducing power outages caused by line faults, and improving grid reliability and line efficiency. This technology ensures stable operation of the power system while also bringing significant economic benefits.

[0003] Currently, the safety assessment of access paths to equipotentials primarily relies on simulation tests on real-world towers, observing whether discharges occur to determine the safety of the paths. While simulation tests offer a straightforward and efficient way to evaluate specific access paths, they are often challenging and involve significant risks. Especially in the case of suspended-basket access, with numerous possible paths available, conducting large-scale tests to verify the safety of each would undoubtedly be labor-intensive and time-consuming. Furthermore, this assessment method only confirms path safety but fails to identify the optimal entry path. If safety assessment parameters for the access area, including the discharge hazard rate and the field strength at the operator's body surface, could be calculated first, and then optimized using intelligent algorithms, the expensive and cumbersome physical testing required could be avoided while enabling the rapid and efficient identification of the optimal entry path. Summary of the Invention

[0004] To address the shortcomings and deficiencies of the aforementioned existing technologies, this paper proposes an improved Deep Q Network (DQN) method for determining when live workers enter equipotential paths. By incorporating strategies such as dynamic greed and information importance differentiation, the improved Deep Q Network (DQN) significantly improves convergence speed, robustness, and generalization, while reducing overfitting. Combined with effective exploration strategies and the ability to adapt to dynamic environments, DQN demonstrates higher accuracy and performance in the field of reinforcement learning, making it a valuable tool for solving complex decision-making problems.

[0005] The specific technical solutions are:

[0006] This paper proposes an improved DQN method for determining the path for live working personnel to enter the equipotential. First, a simulation model is built for live working using the hanging basket method, and the surface field strength and discharge hazard rate of the workers at various spatial positions that the workers may pass through when entering the equipotential are calculated. Based on the actual situation of live working, the deep Q network (DQN) is improved, the reward function of the DQN is redesigned, an improved dynamic greedy strategy is adopted, and the importance of samples is distinguished. Finally, the optimal path for live working personnel to enter the equipotential is obtained.

[0007] The specific steps include:

[0008] Step 1: Simulation modeling;

[0009] Finite element software was used to establish a simulation model, and the range of motion of the hanging basket method was meticulously divided into N rows and M columns with equal spacing, totaling N×M calculation nodes. The discharge hazard rate of each node and the electric field intensity on the human body surface were calculated respectively.

[0010] (1) Calculation of surface field strength

[0011] In the finite element software, the electrostatic physics field and the charge transfer physics field are coupled to establish a finite element simulation model for live operation of UHVDC transmission lines. The corresponding positive and negative potentials are applied to the two pole conductors respectively, and the surface field strength value is calculated by formula (1):

[0012]

[0013] Where m is the conductor surface roughness coefficient; δ is the relative air density; and r is the radius of the split conductor. The operator model was established with reference to GB 10000-88.

[0014] The surface field strength of the human body is calculated by the finite element simulation model. The surface field strength value of the human body at each coordinate is calculated separately, and finally the surface field strength values ​​of N×M individuals are counted. (2) Calculate the combined gap and discharge hazard rate of each point

[0015] Combined clearance refers to the sum of the minimum distances between the equipotential and ground potential when the operator is at an intermediate potential during the equipotential process. A coordinate system is established using the area enclosed by the towers and conductors. Combined clearance is then calculated based on the operator's position within this area. Finally, a calculation script is used to generate N x M combined clearance values.

[0016] The probability of gap insulation damage during live working is called discharge hazard rate, which is R o The calculation method is as follows:

[0017]

[0018] Where P0(U) is the probability density function of the switching overvoltage amplitude; P d (U) The probability distribution function of air gap breakdown when the overvoltage amplitude is U. The calculation method is shown in formulas (3)-(6):

[0019]

[0020] U 50% =k g 1080ln(0.46D+1) (6)

[0021] Where U a is the average value of the switching overvoltage; σ0 is the standard deviation of the switching overvoltage; σ d is the standard deviation of the air gap discharge voltage; U 50% 50% discharge voltage of air gap; U 0.13% is the maximum operating overvoltage; [σ0] is the relative standard deviation of overvoltage; k g is the clearance coefficient; D is the combined clearance.

[0022] Step 2: Data preprocessing;

[0023] The maximum value normalization is used to map the surface field strength and the discharge hazard rate to the same level, and a fusion method of the surface field strength and the gap discharge hazard rate based on the dynamic weighting of the distance between the operator and the conductor is adopted. The expression is shown in Equation (7).

[0024]

[0025] Where d i 、E i 、R i They represent the distance from the current position to the end point, the surface field strength of the operator at the current position, and the discharge hazard rate at the current position, respectively. max Indicates the distance between the starting point and the end point, E max 、R max They represent the maximum surface field strength and maximum discharge hazard rate within the range of movement of the operator, F i Indicates the fusion value of the body surface field strength and discharge hazard rate at the current location.

[0026] in is the weight of the surface field strength, is the weight of the discharge hazard rate. Initially, the surface field strength is prioritized to mitigate the discomfort caused by the strong electric field to the operator. As the operator approaches the conductor, the weight of the discharge hazard rate increases, while the weight of the surface field strength decreases, minimizing the risk to the operator when entering the equipotential zone.

[0027] Step 3: Live working path planning based on improved DQN algorithm;

[0028] (1) Design of reward function based on the real situation of live working;

[0029] The rewards are set as follows:

[0030]

[0031] The endpoint reward, exploration reward, and inspiration reward are set to positive values, while the out-of-bounds penalty is a large negative value. The specific reward settings are as follows:

[0032]

[0033]

[0034] (2) Improved greedy strategy;

[0035] The exploration factor % determines whether to perform random actions or use the network to perform actions. The dynamic exploration factor decreases slowly in the early stage of learning, encouraging the algorithm to perform more random actions and explore more, thereby accumulating more comprehensive learning samples. In the later stage of learning, the exploration factor decreases quickly, which encourages the agent to select actions to perform based on the previous learning results. The specific expression is shown in formula (11):

[0036]

[0037] Where t represents the current number of iterations, t max Indicates the maximum number of iterations when setting % to 0.

[0038] (3) Differentiation of the importance of information in the replay experience pool;

[0039] The reinforcement learning algorithm needs to extract data from the replay experience pool for learning, and divide the replay experience pool into important information and normal information; the important information stores information when crossing the boundary or reaching the end point. When sampling, it samples from important information and normal information according to a certain ratio to accelerate the convergence speed of DQN and the success rate of path planning.

[0040] (4) Design of action rules;

[0041] Based on the hanging basket method to enter the equipotential, considering the practicality of the hanging basket method, the agent's action space is designed to be three, with its action directions being downward, rightward, and rightward and downward;

[0042] (5) Design of Q network;

[0043] The Q network is a multi-layer perceptron with 2 input layer neurons and the current coordinates of the agent as input. The output layer has 3 neurons, corresponding to the value of the action space. The hidden layer has 128 neurons and 2 hidden layers. The activation function is the ReLU function.

[0044] (6) Model training.

[0045] Compared with the prior art, the present invention has the following technical effects:

[0046] (1) Use reinforcement learning to optimize the equipotential path for live working to improve the safety and comfort of workers entering the equipotential area;

[0047] (2) Aiming at the actual situation of live-line operation, the reward function and action rules of the DQN algorithm were designed to achieve autonomous learning to enter the equipotential path;

[0048] (3) A time-varying dynamic strategy is designed. In the early stage, the exploration factor is large, and the agent focuses on exploration in the early stage, thereby accumulating more comprehensive learning information. In the later stage, the exploration factor decreases rapidly, and the agent uses the learned experience to independently select action strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the change of the exploration factor of the present invention;

[0050] Figure 2 is a schematic diagram of the movement direction of the intelligent body of the present invention;

[0051] Figure 3 It is a schematic diagram of the Q network structure of the present invention;

[0052] Figure 4 This is a flow chart of the improved DQN algorithm of the present invention;

[0053] Figure 5 Schematic diagram of the live working model of this embodiment;

[0054] Figure 6 This is a schematic diagram of the motion range of the conventional hanging basket method of this embodiment;

[0055] Figure 7 Schematic diagram of the grid division of the motion range of this embodiment;

[0056] Figure 8 This is a schematic diagram of the field intensity distribution on the body surface within the range of motion of this embodiment;

[0057] Figure 9 This is a schematic diagram of the discharge risk rate distribution within the movable range of this embodiment;

[0058] Figure 10This is a schematic diagram of the distribution of the fusion value of the surface field strength and discharge risk rate within the movable range of this embodiment;

[0059] Figure 11 This is a schematic diagram of the reward changes of the improved DQN algorithm in this embodiment;

[0060] Figure 12 This is a schematic diagram of the reward changes of the DQN algorithm in this embodiment;

[0061] Figure 13 2 is a schematic diagram for comparing entry paths in this embodiment. DETAILED DESCRIPTION

[0062] The specific method steps of the present invention are:

[0063] Step 1: Simulation modeling;

[0064] The hanging basket method is widely used as a common safe entry method. Using finite element software, a simulation model was established, meticulously dividing the basket's range of motion into N rows and M columns at equal intervals, totaling N × M computational nodes. The discharge hazard rate and the electric field strength on the human body surface were calculated for each node. This detailed division provides a more detailed and rigorous foundation for equipotential path planning.

[0065] (1) Calculation of surface field strength

[0066] In the finite element software, the electrostatic physics field and the charge transfer physics field are coupled to establish a finite element simulation model for live operation of UHVDC transmission lines. The corresponding positive and negative potentials are applied to the two pole conductors respectively, and the surface field strength value is calculated by formula (1):

[0067]

[0068] Where m is the conductor surface roughness coefficient; δ is the relative air density; and r is the radius of the split conductor. The operator model was established with reference to GB 10000-88.

[0069] The finite element software simulation model is used to calculate the surface field strength. The coordinates of each point on the human body are adjusted by programming, and the surface field strength value at each coordinate is calculated. Finally, the surface field strength values ​​of N×M individuals are calculated. (2) Calculate the combined gap and discharge hazard rate of each point

[0070] Combined clearance refers to the sum of the minimum distances between the equipotential and ground potential when the operator is at an intermediate potential during the equipotential process. A coordinate system is established using the area enclosed by the towers and conductors. The combined clearance is then calculated using a program based on the human position within this area. Finally, a calculation script is used to generate N × M combined clearance values.

[0071] The probability of gap insulation damage during live working is called the discharge hazard rate. GB / T 18037-2008 stipulates that the gap discharge hazard rate R≤10 -5 It is considered safe, and the discharge hazard rate is calculated as follows:

[0072]

[0073] Where P0(U) is the probability density function of the switching overvoltage amplitude; P d (U) The probability distribution function of air gap breakdown when the overvoltage amplitude is U. The calculation method is shown in formulas (3)-(6):

[0074]

[0075] U 50% =k g 1080ln(0.46D+1) (6)

[0076] Where U a is the average value of the switching overvoltage; σ0 is the standard deviation of the switching overvoltage; σ d is the standard deviation of the air gap discharge voltage; U 50% 50% discharge voltage of air gap; U 0.13% is the maximum operating overvoltage; [σ0] is the relative standard deviation of overvoltage; k g is the clearance coefficient; D is the combined clearance.

[0077] Step 2: Data preprocessing;

[0078] In live working, both the gap discharge hazard rate and the worker's body surface field strength are important factors to consider when planning the worker's path to the equipotential. To this end, we first use maximum normalization to map the body surface field strength and the discharge hazard rate to the same magnitude. We then design a method to fuse the body surface field strength and the gap discharge hazard rate based on the dynamic weighting of the worker's distance from the conductor. The expression is shown in Equation (7).

[0079]

[0080] Where d i 、E i 、R i They represent the distance from the current position to the end point, the surface field strength of the operator at the current position, and the discharge hazard rate at the current position, respectively. max Indicates the distance between the starting point and the end point, E max 、R max They represent the maximum surface field strength and maximum discharge hazard rate within the range of movement of the operator, F i Indicates the fusion value of the body surface field strength and discharge hazard rate at the current location.

[0081] in is the weight of the surface field strength, is the weight of the discharge hazard rate. Initially, the surface field strength is prioritized to mitigate the discomfort caused by the strong electric field to the operator. As the operator approaches the conductor, the weight of the discharge hazard rate increases, while the weight of the surface field strength decreases, minimizing the risk to the operator when entering the equipotential zone.

[0082] Step 3: Live working path planning based on improved DQN algorithm;

[0083] (1) Reward function design based on the actual situation of live working

[0084] In reinforcement learning, the design of the reward function is crucial, and its setting is related to the effectiveness of the algorithm's path planning. In path planning based on the DQN algorithm, traditional rewards are generally set as distance rewards to obtain the shortest distance path between the starting point and the end point. However, in live working entering the equipotential path, path distance should not be the primary factor. Therefore, in this invention, distance rewards are abandoned to prevent the intelligent agent from tending to the shortest path between the starting point and the end point. The reward settings of this invention are as follows:

[0085]

[0086] The endpoint reward, exploration reward, and inspiration reward are set to positive values, while the out-of-bounds penalty is a large negative value. The specific reward settings are as follows:

[0087]

[0088] The endpoint reward is the reciprocal of the average fusion value of the path from the start point to the end point. Multiplying it by 10 indicates that the endpoint reward is more important than other rewards, and the agent is more likely to complete the entire path than other actions. The exploration reward is the reciprocal of the fusion value. A lower fusion value indicates greater safety for live-line workers, and thus a larger reward is given. The inspiration reward is the reciprocal of the fusion value of the current position minus the reciprocal of the fusion value of the previous position. This further encourages movement towards positions with lower fusion values. If movement does not reach a lower position, no penalty is given, and a reward of 0 is given instead, preventing the algorithm from favoring the shortest path.

[0089] (2) Improved Greedy Strategy

[0090] In the process of algorithm learning, the initial learning samples come from the random exploration process of the algorithm, and the exploration factor ε determines whether to perform random actions or use the network to perform actions. Therefore, the design of the exploration factor ε is crucial. The present invention designs a dynamic exploration factor. In the early stage of learning, the exploration factor decreases slowly, encouraging the algorithm to perform more random actions and explore more, thereby accumulating more comprehensive learning samples. In the later stage of learning, the exploration factor decreases quickly, which encourages the intelligent agent to select actions to perform based on the previous learning effect. The specific expression is shown in formula (11). The change of the exploration factor with the number of iterations is as follows: Figure 1 shown.

[0091]

[0092] Where t represents the current number of iterations, t max Indicates the maximum number of iterations when setting ε to 0.

[0093] (3) Differentiation of the importance of information in the replay experience pool

[0094] The reinforcement learning algorithm needs to extract data from the replay experience pool for learning, and divide the replay experience pool into important information and normal information; important information stores information when crossing the boundary or reaching the end point. When sampling, it will sample from important information and normal information according to a certain ratio, thereby accelerating the convergence speed of DQN and the success rate of path planning.

[0095] (4) Design of action rules

[0096] The present invention is based on the hanging basket method to enter the equipotential. Considering the practicability of the hanging basket method, the action space of the agent is designed to be 3, and its action directions are downward, rightward and right-downward. Figure 2 shown.

[0097] (5) Design of Q network

[0098] In the present invention, the Q network is a multi-layer perceptron, the number of neurons in its input layer is 2, the input is the current coordinate of the agent, the number of neurons in the output layer is 3, corresponding to the value of the action space, the number of neurons in the hidden layer is 128, the number of hidden layers is 2, the activation function is the ReLU function, and its structure is as follows: Figure 3 shown.

[0099] (6) Model training

[0100] The improved DQN algorithm process is as follows Figure 4 Its training parameters are shown in Table 1

[0101] Table 1 DQN algorithm training parameters

[0102]

[0103]

[0104] The specific contents of this embodiment are:

[0105] First, according to step 1, a finite element simulation model of live operation of a ±1100kV DC transmission line is established, such as Figure 5 As shown. ±1100 kV potentials were applied to the two polar conductors, and the surface field strength was calculated using Equation (1), where m is 0.49, δ is 1, and r is the radius of the split conductor. The calculated result is 17.46 kV / cm. The surface field strength at each point within the range of motion was simulated. The operator model was established with reference to GB 10000-88, and the dimensions of each component are shown in Table 2.

[0106] Table 2 Worker dimensions

[0107]

[0108] The model and the operator's range of motion are as follows Figure 6 As shown in the figure, the range of motion of the hanging basket method is divided into 20×20 continuous grids, with a total of 400 computing nodes. Figure 7 shown.

[0109] First, calculate the combined gap table, and then calculate the discharge risk rate based on this table. According to formulas (3)-(6), σ d Take 5%, U 0.13% The design standard of Jiquan Line is 1.58pu, [σ0] is 12%, and k g Take 1.129. Place the operator at Figure 7 The discharge hazard rate of each point within the movable range was simulated and calculated.

[0110] The distribution of body surface field strength and discharge risk rate within the range of motion is obtained as follows Figure 8 and Figure 9 shown.

[0111] According to formula (7), the surface field strength and discharge hazard rate were preprocessed to obtain the fusion value distribution of the surface field strength and discharge hazard rate within the movable range as follows: Figure 10 shown.

[0112] according to Figure 7 The improved DQN algorithm flow chart shown in the figure enables the agent to learn autonomously. Considering that the algorithm has a certain degree of randomness, the improved DQN algorithm and the DQN algorithm were run 5 times in total. The average fusion value of each path was used to evaluate the quality of each path. The performance comparison of the algorithms is shown in Table 3. The reward changes corresponding to the optimal fusion value are shown in Table 3. Figure 11 and Figure 12As shown, the optimal path pair is Figure 13 shown.

[0113] Table 3 Algorithm performance comparison

[0114]

[0115] In order to evaluate the effectiveness of the autonomous learning path, the learned optimal path was compared with the conventional entry paths (45° entry, horizontal entry, and vertical entry). The average fusion values ​​of each path are shown in Table 4.

[0116] Table 4 Comparison of security of each path

[0117]

[0118] It can be seen from Table 4 that the optimal path proposed by the present invention is far better than the conventional entry path, and can greatly improve the safety and comfort of live working.

Claims

1. An improved DQN method for determining whether live working personnel enter an equipotential path, characterized in that: The following processes are included: Firstly, a simulation model is built for live working with a hanging basket method, and the surface field strength and discharge hazard rate of the workers at various spatial locations that they may pass through when entering the equipotential state are calculated. In view of the actual situation of live-line working, the DQN reward function was redesigned, an improved dynamic greedy strategy was adopted, and the importance of samples was differentiated. Ultimately, the optimal path for live-line workers to enter the equipotential zone was obtained. The following steps are involved: Step 1: Simulation modeling; Finite element software was used to establish a simulation model, and the range of motion of the hanging basket method was meticulously divided into N rows and M columns with equal spacing, totaling N×M calculation nodes. The discharge hazard rate of each node and the electric field intensity on the human body surface were calculated respectively. (1) Calculation of surface field strength In the finite element software, the electrostatic physics field and the charge transfer physics field are coupled to establish a finite element simulation model for live operation of UHVDC transmission lines. The corresponding positive and negative potentials are applied to the two pole conductors respectively, and the surface field strength value is calculated by formula (1): Where m is the conductor surface roughness coefficient; δ is the relative density of air; r is the radius of the split conductor; the operator model is established with reference to GB 10000-88; The surface field strength of the body is calculated by combining the finite element simulation model, and the surface field strength value at each coordinate is calculated respectively, and finally the surface field strength values ​​of N×M individuals are counted; (2) Calculate the combined gap and discharge hazard rate at each point The combined clearance refers to the sum of the minimum distances between the equipotential and the ground potential when the operator is at an intermediate potential during the process of entering the equipotential. A coordinate system is established using the regional structure enclosed by the towers and conductors. The combined clearance is then calculated based on the position of the human body in this area. Finally, a calculation script is used to calculate N×M combined clearance values. The probability of gap insulation damage during live working is called discharge hazard rate, which is R o The calculation method is as follows: Where P0(U) is the probability density function of the switching overvoltage amplitude; P d (U) Probability distribution function of air gap breakdown when the overvoltage amplitude is U; Step 2: Data preprocessing; The maximum value normalization is used to map the surface field strength and discharge hazard rate to the same level, and a method of fusing the surface field strength and gap discharge hazard rate based on the dynamic weighting of the distance between the operator and the conductor is adopted. The data preprocessing expression is shown in formula (7); Where d i 、E i 、R i They represent the distance from the current position to the end point, the surface field strength of the operator at the current position, and the discharge hazard rate at the current position, respectively. max Indicates the distance between the starting point and the end point, E max 、R max They represent the maximum surface field strength and maximum discharge risk rate within the range of movement of the operator, F i Indicates the fusion value of the surface field strength and discharge hazard rate at the current location; in is the weight of the surface field strength, is the weight of the discharge hazard rate; in the initial stage of entering the range, the surface field strength is given priority consideration to reduce the discomfort caused by the strong electric field to the operator. As the operator gradually approaches the conductor, the weight of the discharge hazard rate gradually increases, and the weight of the surface field strength gradually decreases, to reduce the danger to the operator when entering the equipotential range; Step 3: Live working path planning based on improved DQN algorithm; (1) Design of reward function based on the real situation of live working; The rewards are set as follows: The endpoint reward, exploration reward, and inspiration reward are set to positive values, and the out-of-bounds penalty is a large negative value; The specific reward settings are as follows: (2) Improved greedy strategy; The exploration factor ε determines whether to perform random actions or use the network to perform actions. The dynamic exploration factor decreases slowly in the early stage of learning, encouraging the algorithm to perform more random actions and explore more, thereby accumulating more comprehensive learning samples. In the later stage of learning, the exploration factor decreases quickly, which encourages the agent to select actions based on the previous learning results. The specific expression is shown in formula (11): Where t represents the current number of iterations, t max Indicates the maximum number of iterations when ε is set to 0; (3) Differentiation of the importance of information in the replay experience pool; The reinforcement learning algorithm needs to extract data from the replay experience pool for learning, and divide the replay experience pool into important information and normal information. The important information stores information when crossing the boundary or reaching the end point. When sampling, the sampling is carried out according to a certain ratio from the important information and normal information, which accelerates the convergence speed of DQN and the success rate of path planning. (4) Design of action rules; Based on the hanging basket method to enter the equipotential, considering the practicality of the hanging basket method, the agent's action space is designed to be three, with its action directions being downward, rightward, and rightward and downward; (5) Design of Q network; The Q network is a multi-layer perceptron with 2 input layer neurons and the current coordinates of the agent as input. The output layer has 3 neurons, corresponding to the value of the action space. The hidden layer has 128 neurons and 2 hidden layers. The activation function is the ReLU function. (6) Model training.

2. The improved DQN method for determining whether a live working worker enters an equipotential path according to claim 1 is characterized in that: In step 1, P d The calculation method of (U) is shown in formulas (3)-(6): U 50% =k g 1080ln(0.46D+1) (6) Where U a is the average value of the switching overvoltage; σ0 is the standard deviation of the switching overvoltage; σ d is the standard deviation of the air gap discharge voltage; U 50% 50% discharge voltage of air gap; U 0.13% is the maximum operating overvoltage; [σ0] is the relative standard deviation of overvoltage; k g is the clearance coefficient; D is the combined clearance.

Citation Information

Patent Citations

  • Extra-high voltage direct current line live-line work in-out electric field path optimization method

    CN113971356A

  • Unmanned ship path planning algorithm based on improved DQN

    CN117055549A