Unmanned aerial vehicle inspection path planning method, device and equipment and storage medium

By combining the zero-controlled displacement deviation and zero-controlled speed deviation framework and the near-end strategy optimization algorithm, the drone patrol path planning model is trained, and the problem of inefficiency in drone patrol path planning is solved, active attention to key points of interest and optimization of fuel consumption is achieved, and the complex power grid environment is adapted.

CN120540337APending Publication Date: 2025-08-26SHAOGUAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510628250.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When focusing on the path from the takeoff point to the target end point, the existing drone patrol path planning method ignores the active attention of key points of interest, resulting in inefficient patrol inspections and failure to effectively consider multiple constraints such as fuel consumption.

Method used

Combining the zero-controlled displacement deviation and zero-controlled speed deviation framework and the near-end strategy optimization algorithm, the inspection path planning model is trained to obtain, actively pay attention to and cover the target interest points corresponding to the inspection task, and optimize key parameters through the near-end strategy optimization algorithm to realize path planning.

Benefits of technology

It improves patrol efficiency, avoids repeated flights, saves drone fuel consumption, and has good robustness and practicality, adapts to complex dynamic power grid environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540337A_ABST
    Figure CN120540337A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle inspection path planning method and device, equipment and a storage medium, and relates to the technical field of electric power inspection. The method comprises the following steps: acquiring related information of an unmanned aerial vehicle inspection task, wherein the related information comprises first position information and first speed information of a starting point of the inspection task, second position information and second speed information of a terminal point and third position information of a target interest point corresponding to the inspection task; the first position information, the first speed information, the second position information, the second speed information and the third position information are input into a routing inspection path planning model, a target routing inspection path output by the routing inspection path planning model is obtained, and the routing inspection path planning model is based on a zero-control displacement deviation and zero-control speed deviation framework. And carrying out optimization training through a near-end strategy optimization algorithm. The target interest point corresponding to the inspection task can be actively concerned and covered, so that the inspection efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power inspection technology, and in particular to a method, device, equipment and storage medium for planning a drone inspection path. Background Art

[0002] With the rapid development of drone technology, drones are gradually being used to perform inspection tasks, such as inspecting power transmission lines. Among them, drone inspection path planning is crucial for the application of drones.

[0003] Currently, when planning drone inspection routes, the focus is on planning the path from the takeoff point to the target destination. However, in some scenarios, when drones perform inspections along these planned routes, they experience low inspection efficiency. Summary of the Invention

[0004] The present application provides a drone inspection path planning method, device, equipment and storage medium to solve the problem of low inspection efficiency when drones conduct inspections according to drone inspection paths planned by current methods in some scenarios.

[0005] In a first aspect, the present application provides a method for planning a drone inspection path, comprising:

[0006] Obtain relevant information of the drone inspection task, including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task;

[0007] The first position information, first speed information, second position information, second speed information and third position information are input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework and is obtained through optimization training of the proximal strategy optimization algorithm.

[0008] Optionally, the inspection path planning model is trained by the following method: obtaining training samples, the training samples include relevant information samples of the UAV inspection task samples, the relevant information samples include the first position information sample and the first speed information sample of the starting point sample of the inspection task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the point of interest sample corresponding to the inspection task sample; inputting the training samples into the inspection path planning model for iterative training, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and optimizes the key parameters in the zero-control displacement deviation and zero-control speed deviation framework through the proximal strategy optimization algorithm. When the reward function of the proximal strategy optimization algorithm converges, the iteration is stopped to obtain the trained inspection path planning model.

[0009] Optionally, the key parameters include a first weighting coefficient and a second weighting coefficient, the first weighting coefficient being the weighting coefficient corresponding to the first position error between the second position information sample and the total flight time corresponding to the inspection task sample when no control is applied, the second weighting coefficient being the weighting coefficient corresponding to the second position error between the drone and the point of interest sample. The key parameters in the zero-control displacement deviation and zero-control velocity deviation framework are optimized by the proximal strategy optimization algorithm, including: based on the training sample, obtaining the state space corresponding to the current state of the drone, the state space including the current position vector and velocity vector of the drone; based on the state space, under the condition of satisfying the dynamic constraints and the task time limit corresponding to the inspection task sample, with the goal of maximizing the coverage of the point of interest sample and minimizing the fuel consumption of the drone, the proximal strategy optimization algorithm based on the actor-critic architecture is used to solve the action space to obtain the optimized key parameters, and the action space is composed of the first weighting coefficient and the second weighting coefficient.

[0010] Optionally, the first weighting coefficient and the second weighting coefficient satisfy the following first formula:

[0011]

[0012] Where a represents the command acceleration in the zero-control displacement deviation and zero-control velocity deviation framework; t go represents the remaining flight time; θ1 represents the first weighting coefficient; ZEM represents the first position error; θ2 represents the second weighting coefficient; ZEM int represents the second position error; ZEV represents the error between the second velocity information sample and the final velocity reached by the drone when the total flight time is reached without applying control.

[0013] Optionally, the reward function satisfies the following second formula:

[0014] R t =α1rpoi +α2r time -α3r penalty Second formula

[0015] Among them, R t represents the reward function; r poi Indicates the positive reward given for successfully inspecting the point of interest; α1 represents r poi The corresponding first weight; r time represents the reward for completing the task on time; α2 represents r time The corresponding second weight; r penalty It represents the penalty for deviation from the command trajectory in the zero-control displacement deviation and zero-control velocity deviation framework or when the acceleration change of the UAV exceeds the set acceleration change threshold; α3 represents r penalty The corresponding third weight.

[0016] Optionally, the target points of interest are obtained in the following manner: obtaining multiple points of interest corresponding to the inspection task; using a multi-factor weighted model to obtain a priority score corresponding to each of the multiple points of interest, where the multiple factors include at least the historical failure probability of the point of interest, the operating environment, and the distance from the transmission line tower; and selecting a preset number of points of interest from high to low according to the priority score as the target points of interest.

[0017] Optionally, after obtaining the target inspection path output by the inspection path planning model, the UAV inspection path planning method further includes: controlling the UAV to perform the inspection task according to the target inspection path.

[0018] Optionally, after controlling the drone to perform the inspection task, the drone inspection path planning method further includes: receiving inspection data corresponding to the inspection task sent by the drone; and generating an inspection report based on the inspection data.

[0019] In a second aspect, the present application provides a drone inspection path planning device, comprising:

[0020] An acquisition module is used to obtain relevant information of the UAV inspection task, including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task;

[0021] The processing module is used to input the first position information, the first speed information, the second position information, the second speed information and the third position information into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and is obtained by optimization training through the proximal strategy optimization algorithm.

[0022] Optionally, the UAV inspection path planning device also includes a training module, which is used to train and obtain an inspection path planning model by the following method: obtaining training samples, the training samples include relevant information samples of the UAV inspection task samples, the relevant information samples include the first position information sample and the first speed information sample of the starting point sample of the inspection task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the point of interest sample corresponding to the inspection task sample; inputting the training samples into the inspection path planning model for iterative training, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and optimizes the key parameters in the zero-control displacement deviation and zero-control speed deviation framework through the proximal strategy optimization algorithm. When the reward function of the proximal strategy optimization algorithm converges, the iteration is stopped to obtain the trained inspection path planning model.

[0023] Optionally, the key parameters include a first weighting coefficient and a second weighting coefficient, the first weighting coefficient being the weighting coefficient corresponding to the first position error between the second position information sample and the total flight time corresponding to the inspection task sample when no control is applied, and the second weighting coefficient being the weighting coefficient corresponding to the second position error between the drone and the point of interest sample. When the training module is used to optimize the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework through the proximal strategy optimization algorithm, it is specifically used to: based on the training sample, obtain the state space corresponding to the current state of the drone, the state space including the current position vector and velocity vector of the drone; based on the state space, under the condition of satisfying the dynamic constraints and the task time limit corresponding to the inspection task sample, with the goal of maximizing the coverage of the point of interest sample and minimizing the fuel consumption of the drone, the proximal strategy optimization algorithm based on the actor-critic architecture is used to solve the action space to obtain the optimized key parameters, and the action space is composed of the first weighting coefficient and the second weighting coefficient.

[0024] Optionally, the first weighting coefficient and the second weighting coefficient satisfy the following first formula:

[0025]

[0026] Where a represents the command acceleration in the zero-control displacement deviation and zero-control velocity deviation framework; t go represents the remaining flight time; θ1 represents the first weighting coefficient; ZEM represents the first position error; θ2 represents the second weighting coefficient; ZEM int represents the second position error; ZEV represents the error between the second velocity information sample and the final velocity reached by the drone when the total flight time is reached without applying control.

[0027] Optionally, the reward function satisfies the following second formula:

[0028] R t =α1r poi +α2r time -α3r penalty Second formula

[0029] Among them, R t represents the reward function; r poi Indicates the positive reward given for successfully inspecting the point of interest; α1 represents r poi The corresponding first weight; r time represents the reward for completing the task on time; α2 represents r time The corresponding second weight; r penalty It represents the penalty for deviation from the command trajectory in the zero-control displacement deviation and zero-control velocity deviation framework or when the acceleration change of the UAV exceeds the set acceleration change threshold; α3 represents r penalty The corresponding third weight.

[0030] Optionally, the acquisition module is also used to obtain target points of interest in the following manner: obtaining multiple points of interest corresponding to the inspection task; using a multi-factor weighted model to obtain a priority score corresponding to each point of interest in the multiple points of interest, where the multiple factors include at least the historical failure probability of the point of interest, the operating environment, and the distance from the transmission line tower; and selecting a preset number of points of interest from high to low according to the priority score as target points of interest.

[0031] Optionally, the processing module is further used to: after obtaining the target inspection path output by the inspection path planning model, control the UAV to perform the inspection task according to the target inspection path.

[0032] Optionally, the processing module is further used to: after controlling the drone to perform the inspection task, receive inspection data corresponding to the inspection task sent by the drone; and generate an inspection report based on the inspection data.

[0033] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0034] Memory stores computer-executable instructions;

[0035] The processor executes the computer-executable instructions stored in the memory to implement the drone inspection path planning method as described in the first aspect of this application.

[0036] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer program instructions. When the computer program instructions are executed, the drone inspection path planning method described in the first aspect of the present application is implemented.

[0037] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed, implements the drone inspection path planning method as described in the first aspect of the present application.

[0038] The present application provides a method, device, equipment and storage medium for planning an unmanned aerial vehicle inspection path. The method obtains relevant information of the unmanned aerial vehicle inspection task, including the first position information and first speed information of the starting point of the inspection task, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task; the first position information, first speed information, second position information, second speed information and third position information are input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model; wherein the inspection path planning model is based on a zero-control displacement deviation and zero-control speed deviation framework, and is obtained by optimization training through a proximal strategy optimization algorithm. The target inspection path output by the inspection path planning model can actively pay attention to and cover the target points of interest corresponding to the inspection task, thereby avoiding repeated flights, effectively improving inspection efficiency, and saving fuel consumption of the unmanned aerial vehicle. It has good robustness and practicality, and can better adapt to the complex and dynamic inspection environment in the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0040] Figure 1 A flowchart of a method for planning a drone inspection path according to an embodiment of the present application;

[0041] Figure 2 A flowchart of a method for planning a drone inspection path according to another embodiment of the present application;

[0042] Figure 3 A schematic diagram of a top view of a target inspection path provided in one embodiment of the present application;

[0043] Figure 4 A schematic diagram of a front view of a target inspection path provided in one embodiment of the present application;

[0044] Figure 5 A flowchart of a method for training an inspection path planning model provided in one embodiment of the present application;

[0045] Figure 6 A schematic diagram of a cumulative reward curve during the training process provided in one embodiment of the present application;

[0046] Figure 7 A schematic diagram of the structure of a UAV inspection path planning device provided in one embodiment of the present application;

[0047] Figure 8 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application.

[0048] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0049] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0051] With the rapid development of drone technology, drones are increasingly being used to perform inspection tasks, such as power transmission line inspections, due to their ease of control and high maneuverability. Among these tasks, drone inspection path planning is crucial for drone applications.

[0052] Currently, drone inspection route planning primarily focuses on the path from the takeoff point to the target destination. However, in some scenarios, when drones conduct inspections along these planned routes, they often neglect to proactively monitor key points of interest (POIs). These POIs may indicate areas with potential equipment failures or require special attention, such as critical nodes on high-voltage lines and aging equipment. Potential failures in these areas can pose significant risks to the safety and stability of the power system, and therefore require specific attention during inspections. However, current drone inspection route planning methods generally lack the ability to proactively identify and cover these POIs, resulting in inefficient inspections and requiring multiple flights to repeatedly cover them, wasting significant time and energy. Furthermore, most current drone inspection route planning methods ignore multiple practical constraints, such as fuel consumption, when selecting routes. Improving the efficiency and reliability of drone inspections requires comprehensive consideration of multiple constraints during route planning and effectively covering all POIs.

[0053] In related technologies, for example, the zero displacement deviation (ZEM) / zero velocity deviation (ZEV) algorithm, as an energy guidance algorithm, can achieve path optimization under multiple constraints by controlling the flight attitude and thrust of the drone, close to the open-loop optimal solution, and has lower fuel consumption. Therefore, the application of this algorithm in the drone inspection process helps to improve flight efficiency and extend endurance. By dividing the flight time into several parts and applying the ZEM / ZEV algorithm in each part, the control effect can be made closer to the optimal solution while maintaining good closed-loop feedback control robustness. However, the existing ZEM / ZEV algorithm still has some limitations when applied, that is, although it can optimize the inspection path selection and improve the inspection accuracy, it lacks flexibility and real-time adaptability.

[0054] Based on the above problems, the present application provides a UAV inspection path planning method, device, equipment and storage medium, which trains and obtains an inspection path planning model by combining the zero-controlled displacement deviation and zero-controlled speed deviation framework with the proximal strategy optimization algorithm, wherein a point of interest bias vector and key parameters to be optimized by the proximal strategy optimization algorithm are constructed in the zero-controlled displacement deviation and zero-controlled speed deviation framework; the target inspection path corresponding to the UAV inspection task is obtained through the inspection path planning model, which can actively pay attention to and cover the target points of interest corresponding to the inspection task, thereby avoiding repeated flights, effectively improving inspection efficiency, and saving UAV fuel consumption. It has good robustness and practicality, and can better adapt to the complex and dynamic inspection environment in the power grid.

[0055] It should be noted that the drone inspection path planning method provided in the embodiment of the present application can be applied in a server, which can be an independent server, or a service cluster, etc.

[0056] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0057] Figure 1 This is a flow chart of a method for planning a drone inspection path according to an embodiment of the present application. Figure 1 As shown, the drone inspection path planning method of the embodiment of the present application includes:

[0058] S101. Obtain relevant information of the drone inspection task, including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task.

[0059] In the embodiment of the present application, the inspection task of the drone includes not only the existing main inspection task (i.e., the arrival task), but also the point of interest inspection task for the target point of interest corresponding to the inspection task. Therefore, in this step, the relevant information of the drone inspection task is obtained, and the relevant information includes the relevant information from the starting point to the end point of the inspection task, i.e., the first position information and first speed information of the starting point, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task.

[0060] Optionally, the target points of interest are obtained in the following manner: obtaining multiple points of interest corresponding to the inspection task; using a multi-factor weighted model to obtain a priority score corresponding to each of the multiple points of interest, where the multiple factors include at least the historical failure probability of the point of interest, the operating environment, and the distance from the transmission line tower; and selecting a preset number of points of interest from high to low according to the priority score as the target points of interest.

[0061] For example, during the inspection process, drones can use the multimodal sensing devices (such as optical cameras, infrared thermal imagers, etc.) on board to achieve multi-scale and multi-angle observation of transmission lines and their ancillary structures (such as the important points of interest in Table 1), and use the deep convolutional neural network (CNN) model to perform real-time recognition of the image data obtained by the drone. Based on the prior experience of historical inspection tasks, multiple points of interest corresponding to the inspection tasks can be obtained, and each point of interest can be scored and evaluated according to its importance. Table 1 shows the important points of interest and their defect types for transmission line inspections.

[0062] Table 1

[0063] Serial number Important points of interest during inspection Defect Type 1 Power supply tower Deformation, position shift, corrosion 2 Insulated terminals Cracks, ruptures, deformations, burns 3 Metal utensils Corrosion, deformation, cracks 4 Lightning protection device Discharge gap change 5 transmission lines Damage, corrosion, aging, etc.

[0064] When evaluating each POI based on its importance, a multi-factor weighted model can be used to determine the priority score for each of the multiple POIs. The multi-factor model includes at least the POI's historical failure probability, operating environment, and distance from the transmission line tower. After obtaining the priority score for each POI, a POI ranking table can be created based on the scoring results. This allows a preset number of POIs to be selected as target POIs, ranked from highest to lowest priority.

[0065] S102. Input the first position information, the first speed information, the second position information, the second speed information and the third position information into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework and is obtained through optimization training of the proximal strategy optimization algorithm.

[0066] It can be understood that the zero displacement error (ZEM) and zero velocity error (ZEV) frameworks, based on traditional terminal control, add constraints on points of interest that appear dynamically in the inspection path, and establish a dynamic guidance mechanism for the UAV from its current position to the target point of interest, thereby achieving a highly dynamic, low-energy path generation method that takes points of interest into account. It is suitable for irregular and unstructured inspection tasks in transmission lines, and can make the overall performance of the control effect close to the open-loop optimal solution while maintaining the robustness of closed-loop feedback control. In this embodiment, the zero-displacement deviation and zero-velocity deviation framework is combined with the Proximal Policy Optimization (PPO) algorithm. The zero-displacement deviation and zero-velocity deviation framework provides an initial energy optimal path as a priori control basis. The PPO algorithm uses this as a guide and, through real-time interaction with the dynamic environment, learns and adjusts the control parameters in the guidance law of the zero-displacement deviation and zero-velocity deviation framework. This balances the inspection path between global optimal control and local environmental adaptability, achieves active attention to structural points of interest and fine-tuning control of the path, and thus achieves the comprehensive optimization of priority coverage of points of interest, minimized fuel consumption, and flight safety.

[0067] In this step, the inspection path planning model is obtained by optimizing and training the model through the proximal strategy optimization algorithm based on the zero-control displacement deviation and zero-control speed deviation framework. For the specific training method to obtain the inspection path planning model, please refer to the subsequent embodiments and will not be repeated here. The first position information and first speed information of the starting point of the inspection task, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task can be input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, so that the UAV can be controlled to perform the inspection task according to the target inspection path to obtain the inspection data corresponding to the inspection task.

[0068] The UAV inspection path planning method provided in the embodiment of the present application obtains relevant information of the UAV inspection task, the relevant information including the first position information and first speed information of the starting point of the inspection task, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task; the first position information, first speed information, second position information, second speed information and third position information are input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model; wherein, the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and is obtained by optimization training through the proximal strategy optimization algorithm. The target inspection path output by the inspection path planning model can actively pay attention to and cover the target point of interest corresponding to the inspection task, thereby avoiding repeated flights, effectively improving the inspection efficiency, and saving the fuel consumption of the UAV. It has good robustness and practicality, and can better adapt to the complex and dynamic inspection environment in the power grid.

[0069] Figure 2 This is a flow chart of a method for planning a drone inspection path provided by another embodiment of the present application. Based on the above embodiment, this embodiment of the present application further illustrates the method for planning a drone inspection path. Figure 2 As shown, the drone inspection path planning method of the embodiment of the present application may include:

[0070] S201. Obtain relevant information of the drone inspection task, including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task.

[0071] The detailed description of this step can be found in Figure 1 The relevant description of S101 in the illustrated embodiment will not be repeated here.

[0072] S202. Input the first position information, the first speed information, the second position information, the second speed information and the third position information into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and is obtained by optimization training through the proximal strategy optimization algorithm.

[0073] The detailed description of this step can be found in Figure 1 The relevant description of S102 in the embodiment shown. For example, Figure 3 A schematic diagram of a top view of a target inspection path provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the units of X coordinate, Y coordinate and Z coordinate are all meters. Figure 3The target inspection path generated by the drone in the test scenario is shown, including the starting point 301, end point 302 and target point of interest 303 of the inspection task; as can be seen from the top view, the target inspection path is not a simple linear connection, but produces a certain curvature deviation when approaching the target point of interest 303 area, which is a reflection of the point of interest attention mechanism in the PPO algorithm. Figure 4 A schematic diagram of a front view of a target inspection path provided in an embodiment of the present application, such as Figure 4 As shown, based on Figure 3 , showing the Figure 3 The top view corresponds to the front view of the target inspection path, and the front view shows the high point in the air of the UAV during the flight, thereby achieving effective observation of the target point of interest 303.

[0074] S203: Control the UAV to perform the inspection task according to the target inspection path.

[0075] For example, after obtaining the target inspection path corresponding to the drone inspection mission, the drone can be controlled to fly along the target inspection path to perform the inspection mission. During the flight, the drone can use onboard sensors to perform operations such as image acquisition, infrared temperature measurement, and structural identification of target points of interest and key parts of the transmission line.

[0076] S204: Receive inspection data corresponding to the inspection task sent by the drone.

[0077] For example, inspection data obtained by a drone during an inspection mission can be compressed in real time and transmitted via a wireless link to an electronic device executing an embodiment of the present method, such as a server on a ground control system. Once a data connection is established between the ground control system and the drone, the ground control system can receive and analyze the inspection data in real time, displaying the drone's location information, attitude parameters, and sensor status according to the communication protocol.

[0078] S205: Generate an inspection report based on the inspection data.

[0079] For example, after the drone completes the inspection mission, the electronic device executing the embodiment of the present method can automatically generate an inspection report based on the inspection data, and mark abnormal situations and issue warning prompts.

[0080] The UAV inspection path planning method provided in the embodiment of the present application obtains relevant information of the UAV inspection task, the relevant information including the first position information and first speed information of the starting point of the inspection task, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task; the first position information, first speed information, second position information, second speed information and third position information are input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model; wherein, the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and is obtained by optimization training through the proximal strategy optimization algorithm; the target inspection path output by the inspection path planning model can actively pay attention to and cover the target point of interest corresponding to the inspection task, thereby avoiding repeated flights, effectively improving the inspection efficiency, and saving the fuel consumption of the UAV; and then according to the target inspection path, the UAV is controlled to perform the inspection task, and the inspection data corresponding to the inspection task can be obtained more accurately and comprehensively, and an inspection report is generated based on the inspection data, which helps to improve the inspection quality and reliability of the power system.

[0081] Based on the above embodiments, Figure 5 This is a flow chart of a training method for an inspection path planning model provided in one embodiment of the present application. Figure 5 As shown, the training method of the inspection path planning model of the embodiment of the present application may include:

[0082] S501. Obtain training samples, where the training samples include relevant information samples of the drone inspection task samples, where the relevant information samples include the first position information sample and the first speed information sample of the starting point sample of the inspection task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the point of interest sample corresponding to the inspection task sample.

[0083] In this step, training samples can be constructed based on existing drone inspection tasks. The training samples include relevant information samples of the drone inspection task samples. The relevant information samples include the first position information sample and the first speed information sample of the starting point sample of the inspection task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the point of interest sample corresponding to the inspection task sample.

[0084] S502. Input the training samples into the inspection path planning model for iterative training, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and optimizes the key parameters in the zero-control displacement deviation and zero-control speed deviation framework through the proximal policy optimization algorithm. When the reward function of the proximal policy optimization algorithm converges, the iteration is stopped to obtain the trained inspection path planning model.

[0085] For example, for transmission line inspection, the system dynamics equation in the inspection path planning process of the drone is shown in the following formula 1:

[0086]

[0087] Among them, r represents the position of the UAV, represents the derivative of r; v represents the speed of the drone, v represents the derivative of v; T represents the thrust generated by the drone's engine on the inspection drone; a represents the acceleration generated by the drone's engine on the inspection drone; g represents the acceleration of gravity; m represents the mass of the drone. Depending on the optimization objectives, ZEM and ZEV guidance include the optimal fuel consumption problem with fixed time and the optimal time problem under non-fixed time. The embodiment of the present application considers the optimal control problem under fixed time. Assume that the initial and final moments of a given inspection task are t0 and t f , and the initial position r0 and speed v0 of the aircraft's lander; at the same time, in order to prevent the lander from colliding with the ground during the path planning process, causing bouncing, overturning and other problems, the UAV is guaranteed to arrive at the tower landing position at the given target position r f When , the velocity requirement at the final arrival time is zero. Therefore, the endpoint constraint can be expressed as the following formula 2:

[0088]

[0089] The performance index (corresponding to fuel consumption) is expressed in the form of the square integral of acceleration as shown in the following formula 3:

[0090]

[0091] Construct a Hamiltonian function, input J into the Hamiltonian function and calculate the partial derivative with respect to J. The command acceleration shown in the following formula 4 can be obtained:

[0092]

[0093] Among them, t go =t f -t represents the remaining flight time, t represents the flight time; ZEM represents the target position r f Compared with the final position r(t f ) between the position error; ZEV represents the target speed v f The final speed v(t f ) between them.

[0094] The position and speed without control instructions can be expressed as the following formula 5:

[0095]

[0096] Here, τ represents a constant.

[0097] Based on flight dynamics modeling, the ZEM and ZEV state vectors can be calculated in real time to evaluate the distance and speed difference from the target point when no control input is required in the current UAV flight state, and use this as the basis for path adjustment. This algorithm has good trajectory smoothness and target point convergence capabilities in aircraft guidance, but it only uses the initial point and the end point as guidance objects and lacks the ability to process points of interest during flight. To solve this problem, the embodiment of the present application constructs a point of interest bias vector and key parameters in the zero-control displacement deviation and zero-control velocity deviation framework. By introducing the point of interest position error and key parameters as dynamic correction terms into the guidance formula, the UAV can not only converge at the end point, but also respond to the position of the point of interest in the path.

[0098] Optionally, the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework include a first weighting coefficient and a second weighting coefficient. The first weighting coefficient is the weighting coefficient corresponding to the first position error between the second position information sample and the total flight time corresponding to the inspection task sample when no control is applied. The second weighting coefficient is the weighting coefficient corresponding to the second position error between the drone and the point of interest sample. Based on the above formula 4, the first weighting coefficient and the second weighting coefficient satisfy the following formula 6 (i.e., the first formula):

[0099]

[0100] Where a is the command acceleration; t go is the remaining flight time; θ1 represents the first weighting coefficient; ZEM represents the first position error; θ2 represents the second weighting coefficient; ZEM int represents the second position error; ZEV represents the error between the second velocity information sample and the final velocity reached by the drone at the end of the total flight time without control. Formula 6 above can be understood as a calculation expression for the command acceleration based on the improved formula 4 above.

[0101] Based on the above formula 6, optimizing the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework through the proximal policy optimization algorithm can include: obtaining the state space corresponding to the current state of the UAV based on the training samples, where the state space includes the current position vector and velocity vector of the UAV; based on the state space, under the conditions of satisfying the dynamic constraints and the task time limit corresponding to the inspection task samples, with the goal of maximizing the coverage of the point of interest samples and minimizing the fuel consumption of the UAV, using the proximal policy optimization algorithm based on the actor-critic architecture to solve the action space to obtain the optimized key parameters, where the action space is composed of the first weighting coefficient and the second weighting coefficient.

[0102] It can be understood that the embodiment of the present application adopts a two-layer control structure that combines the zero-control displacement deviation and zero-control speed deviation framework with the proximal strategy optimization algorithm. The zero-control displacement deviation and zero-control speed deviation framework provides the initial energy optimal path as the prior control basis, and the proximal strategy optimization algorithm uses this as a guide to correct the control strategy through real-time interaction with the dynamic environment, so that the inspection path achieves a balance between global optimal control and local environmental adaptability, and realizes active attention to structural points of interest and path fine-tuning control. Specifically, the proximal strategy optimization algorithm learns and adjusts the key parameters in the zero-control displacement deviation and zero-control speed deviation framework to achieve a comprehensive optimization of priority coverage of points of interest, minimization of fuel consumption and flight safety. By interacting with the environment to obtain feedback information, the flight strategy of the drone is optimized in continuous trial and learning, so that it maximizes coverage of key points of interest and minimizes the fuel consumption of the drone while meeting the dynamic constraints and mission time limit, so as to achieve a comprehensive optimal effect.

[0103] For example, when a UAV is performing an inspection task, it perceives its current state from the environment, and its state space S is represented as: s t =[p t ,v t ];in, Represents the current position vector of the drone, Represents the current velocity vector of the UAV. The action space A corresponds to the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework, expressed as: a t =[θ1,θ2]; This key parameter can be understood as the adjustment parameter of the target point of interest, thereby incorporating the information of the target point of interest into the calculation formula of the ZEM and ZEV guidance rate.

[0104] Optionally, the reward function of the proximal policy optimization algorithm satisfies the following formula 7 (i.e., the second formula):

[0105] R t =α1r poi +α2r time -α3r penalty Formula 7

[0106] Among them, R t represents the reward function; r poi Indicates the positive reward given for successfully inspecting the point of interest; α1 represents r poi The corresponding first weight; r time represents the reward for completing the task on time; α2 represents r time The corresponding second weight; r penalty It represents the penalty for deviation from the command trajectory in the zero-control displacement deviation and zero-control velocity deviation framework or when the acceleration change of the UAV exceeds the set acceleration change threshold; α3 represents r penalty The corresponding third weight.

[0107] Through the above reward function, multi-objective optimization such as priority coverage of target points of interest and shortest path can be achieved.

[0108] The intelligence of the proximal policy optimization algorithm depends on the reasonable design of the state space, action space, and reward function, and changes dynamically with flight time and mission progress to adapt to decision-making needs in complex environments. For example, a proximal policy optimization algorithm based on the actor-critic architecture can be used, and the corresponding training structure is as follows:

[0109] Actor network: input is state space s t , the output is the action space a t , the clip function (a function used to limit the range of a value) is used to constrain the PPO strategy update to avoid violent shocks of the PPO strategy; Critic network: estimates the value function V(s t ) is used to guide the optimization of the Actor direction; the strategy update process uses mini-batch sampling and updates the main network weights after multiple rounds of iterations; the circuit breaker mechanism is used during training to dynamically adjust the exploration rate when the curve corresponding to the reward function fluctuates or decreases for a long time. The optimization goal of the PPO strategy is:

[0110] L(θ)=maxΕ t [min(r t (θ)A t ,clip(r t (θ),1-ε,1+ε,)A t )]

[0111] Among them, the target L(θ) is based on Adjust the policy parameter θ to make good actions appear more and bad actions appear less; r t (θ) represents the ratio of the current PPO strategy to the old PPO strategy; represents the advantage estimation function; ε represents the clipping coefficient, which ensures that the policy update is not too large to prevent training instability.

[0112] The action space is solved by a proximal policy optimization algorithm based on the actor-critic architecture, and the key parameters in the optimized zero-control displacement deviation and zero-control velocity deviation framework can be obtained.

[0113] To verify the feasibility and effectiveness of the PPO algorithm-based path strategy in actual power transmission inspection tasks, the present embodiment further conducted visualization of the reinforcement learning process and strategy convergence testing. After training, a deterministic flight strategy ready for actual deployment was generated, namely the trained inspection path planning model. This inspection path planning model no longer relies on sampling, but instead performs forward reasoning based on the optimal strategy network obtained by the PPO algorithm, ensuring controllable behavior and stable paths during deployment.

[0114] In addition, during the training process, as the training iterations proceed, the reward function shows a trend of gradual convergence. In the initial stage, the UAV cannot effectively focus on the point of interest samples, resulting in a slow increase in the reward; but with the continuous optimization of the PPO strategy network, the inspection path planning model gradually learns to cover key points of interest (i.e., point of interest samples) while meeting path safety, and finally the reward function converges within a stable area. Figure 6 A schematic diagram of a cumulative reward curve during the training process provided in one embodiment of the present application is shown as follows: Figure 6 As shown in the figure, as training iterations progress, the reward function's reward value and average reward value show a gradual convergence trend. After approximately 8,000 training iterations, the inspection path planning model has essentially reached convergence. The trained inspection path planning model can be used for actual UAV flight control, manifested by adjusting the output acceleration under the same initial state. The trained inspection path planning model was able to generate inspection paths with high coverage and low energy consumption in multiple test tasks, and successfully observed target points of interest along the inspection paths.

[0115] The training method of the inspection path planning model provided in the embodiment of the present application is to obtain training samples, the training samples including relevant information samples of the unmanned aerial vehicle inspection task samples, the relevant information samples including the first position information sample and the first speed information sample of the starting point sample of the inspection task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the interest point sample corresponding to the inspection task sample; input the training samples into the inspection path planning model for iterative training, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, optimizes the key parameters in the zero-control displacement deviation and zero-control speed deviation framework through the proximal policy optimization algorithm, stops the iteration when the reward function of the proximal policy optimization algorithm converges, and obtains the trained inspection path planning model. By combining the zero-control displacement deviation and zero-control speed deviation framework with the proximal policy optimization algorithm, the inspection path can achieve a balance between global optimal control and local environmental adaptability, realize active attention to structural interest points and path fine-tuning control, so that when the trained inspection path planning model is used for reasoning, the target inspection path corresponding to the unmanned aerial vehicle inspection task can be obtained more accurately.

[0116] Optionally, the trained inspection path planning model can be embedded into the drone's main control chip to enable online policy reasoning. During real-world flight, the drone dynamically generates control commands based on its current state and the output of the inspection path planning model, enabling fine-tuning of the path and responding to points of interest.

[0117] In summary, the inspection path planning model obtained by optimizing and training the proximal strategy optimization algorithm based on the zero-controlled displacement deviation and zero-controlled speed deviation framework in the embodiment of the present application can realize dynamic optimization of path selection. It can not only respond to environmental changes in real time, but also maximize the inspection coverage of the drone while ensuring safety and efficiency, thereby improving the inspection quality and reliability of the power system. It has high robustness, high intelligence and high practicality, can adapt to the complex and dynamic inspection environment in the power grid, has significant advantages in improving inspection efficiency and reducing operational risks, and has broad engineering application prospects.

[0118] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0119] Figure 7 This is a schematic diagram of the structure of the UAV inspection path planning device provided in one embodiment of the present application. Figure 7 As shown, the drone inspection path planning device 700 of the embodiment of the present application includes: an acquisition module 701 and a processing module 702. Among them:

[0120] The acquisition module 701 is used to obtain relevant information of the drone inspection task, including the first position information and first speed information of the starting point of the inspection task, the second position information and second speed information of the end point, and the third position information of the target point of interest corresponding to the inspection task.

[0121] Processing module 702 is used to input the first position information, first speed information, second position information, second speed information and third position information into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and is obtained through optimization training of the proximal strategy optimization algorithm.

[0122] In some embodiments, the UAV inspection path planning device 700 may further include a training module ( Figure 7 (not shown in the figure), which is used to train and obtain a patrol path planning model by the following method: obtaining training samples, the training samples include relevant information samples of the UAV patrol task samples, the relevant information samples include the first position information sample and the first speed information sample of the starting point sample of the patrol task sample, the second position information sample and the second speed information sample of the end point sample, and the third position information sample of the point of interest sample corresponding to the patrol task sample; inputting the training samples into the patrol path planning model for iterative training, wherein the patrol path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework, and optimizes the key parameters in the zero-control displacement deviation and zero-control speed deviation framework through the proximal strategy optimization algorithm. When the reward function of the proximal strategy optimization algorithm converges, the iteration is stopped to obtain the trained patrol path planning model.

[0123] Optionally, the key parameters include a first weighting coefficient and a second weighting coefficient, the first weighting coefficient being the weighting coefficient corresponding to the first position error between the second position information sample and the total flight time corresponding to the inspection task sample when no control is applied, and the second weighting coefficient being the weighting coefficient corresponding to the second position error between the drone and the point of interest sample. When the training module is used to optimize the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework through the proximal strategy optimization algorithm, it can be specifically used to: based on the training sample, obtain the state space corresponding to the current state of the drone, the state space including the current position vector and velocity vector of the drone; based on the state space, under the condition of satisfying the dynamic constraints and the task time limit corresponding to the inspection task sample, with the goal of maximizing the coverage of the point of interest sample and minimizing the fuel consumption of the drone, the proximal strategy optimization algorithm based on the actor-critic architecture is used to solve the action space to obtain the optimized key parameters, and the action space is composed of the first weighting coefficient and the second weighting coefficient.

[0124] Optionally, the first weighting coefficient and the second weighting coefficient satisfy the following first formula:

[0125]

[0126] Where a represents the command acceleration in the zero-control displacement deviation and zero-control velocity deviation framework; t go represents the remaining flight time; θ1 represents the first weighting coefficient; ZEM represents the first position error; θ2 represents the second weighting coefficient; ZEM int represents the second position error; ZEV represents the error between the second velocity information sample and the final velocity reached by the drone when the total flight time is reached without applying control.

[0127] Optionally, the reward function satisfies the following second formula:

[0128] R t =α1r poi +α2r time -α3r penalty Second formula

[0129] Among them, R t represents the reward function; r poi Indicates the positive reward given for successfully inspecting the point of interest; α1 represents r poi The corresponding first weight; r time represents the reward for completing the task on time; α2 represents r time The corresponding second weight; r penalty It represents the penalty for deviation from the command trajectory in the zero-control displacement deviation and zero-control velocity deviation framework or when the acceleration change of the UAV exceeds the set acceleration change threshold; α3 represents r penalty The corresponding third weight.

[0130] Optionally, the acquisition module 701 can also be used to obtain target points of interest in the following manner: obtain multiple points of interest corresponding to the inspection task; use a multi-factor weighted model to obtain the priority score corresponding to each point of interest in the multiple points of interest, where the multiple factors include at least the historical failure probability of the point of interest, the operating environment, and the distance from the transmission line tower; and select a preset number of points of interest from high to low according to the priority score as the target points of interest.

[0131] Optionally, the processing module 702 may also be configured to: after obtaining the target inspection path output by the inspection path planning model, control the UAV to perform the inspection task according to the target inspection path.

[0132] Optionally, the processing module 702 may also be used to: after controlling the drone to perform the inspection task, receive inspection data corresponding to the inspection task sent by the drone; and generate an inspection report based on the inspection data.

[0133] The device of the embodiment of the present application can be used to execute the technical solution of any of the above-mentioned method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.

[0134] Figure 8 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 8 As shown, the electronic device 800 may include: at least one processor 801 and a memory 802.

[0135] The memory 802 is used to store programs. Specifically, the programs may include program codes, and the program codes include computer-executable instructions.

[0136] The memory 802 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0137] The processor 801 is used to execute the computer-executable instructions stored in the memory 802 to implement the drone inspection path planning method described in the aforementioned method embodiment. Among them, the processor 801 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. Specifically, when implementing the drone inspection path planning method described in the aforementioned method embodiment, the electronic device can be, for example, an electronic device with processing capabilities such as a server.

[0138] Optionally, the electronic device 800 may further include a communication interface 803. In a specific implementation, if the communication interface 803, the memory 802, and the processor 801 are implemented independently, the communication interface 803, the memory 802, and the processor 801 may be interconnected via a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc., but this does not mean that there is only one bus or only one type of bus.

[0139] Optionally, in a specific implementation, if the communication interface 803, the memory 802 and the processor 801 are integrated on a chip, the communication interface 803, the memory 802 and the processor 801 can complete communication through an internal interface.

[0140] The present application also provides a computer-readable storage medium, which stores computer program instructions. When a processor executes the computer program instructions, the above-mentioned drone inspection path planning method is implemented.

[0141] The present application also provides a computer program product, including a computer program, which, when executed, implements the above-mentioned drone inspection path planning method.

[0142] The computer-readable storage medium may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0143] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may be located in an application-specific integrated circuit. Of course, the processor and the readable storage medium may also exist as discrete components in the drone inspection path planning device.

[0144] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for planning a UAV inspection path, characterized in that: include: Obtain relevant information of the drone inspection task, the relevant information including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task; The first position information, the first speed information, the second position information, the second speed information and the third position information are input into the inspection path planning model to obtain the target inspection path output by the inspection path planning model, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control speed deviation framework and is obtained through optimization training using a proximal strategy optimization algorithm.

2. The UAV inspection path planning method according to claim 1, characterized in that: The inspection path planning model is trained by the following method: Obtaining training samples, wherein the training samples include relevant information samples of the UAV inspection task sample, the relevant information samples including a first position information sample and a first speed information sample of a starting point sample of the inspection task sample, a second position information sample and a second speed information sample of an end point sample, and a third position information sample of a point of interest sample corresponding to the inspection task sample; The training samples are input into the inspection path planning model for iterative training, wherein the inspection path planning model is based on the zero-control displacement deviation and zero-control velocity deviation framework, and the key parameters in the zero-control displacement deviation and zero-control velocity deviation framework are optimized by the proximal strategy optimization algorithm. When the reward function of the proximal strategy optimization algorithm converges, the iteration is stopped to obtain a trained inspection path planning model.

3. The UAV inspection path planning method according to claim 2, characterized in that: The key parameters include a first weighting coefficient and a second weighting coefficient. The first weighting coefficient is a weighting coefficient corresponding to a first position error between the second position information sample and the final position reached by the drone when reaching the total flight time corresponding to the inspection task sample without applying control. The second weighting coefficient is a weighting coefficient corresponding to a second position error between the drone and the point of interest sample. The key parameters in the zero-control displacement deviation and zero-control velocity deviation framework are optimized by the proximal strategy optimization algorithm, including: Based on the training samples, a state space corresponding to a current state of the drone is obtained, where the state space includes a current position vector and a current velocity vector of the drone; Based on the state space, under the condition of satisfying the dynamic constraints and the task time limit corresponding to the inspection task sample, with the goal of maximizing the coverage of the point of interest samples and minimizing the fuel consumption of the drone, a proximal policy optimization algorithm based on the actor-critic architecture is adopted to solve the action space to obtain optimized key parameters, where the action space is composed of the first weighting coefficient and the second weighting coefficient.

4. The UAV inspection path planning method according to claim 3 is characterized in that: The first weighting coefficient and the second weighting coefficient satisfy the following first formula: Wherein, a represents the command acceleration in the zero-control displacement deviation and zero-control velocity deviation framework; t go represents the remaining flight time; θ1 represents the first weighting coefficient; ZEM represents the first position error; θ2 represents the second weighting coefficient; ZEM int represents the second position error; ZEV represents the error between the second velocity information sample and the final velocity reached by the drone when the total flight time is reached without applying control.

5. The UAV inspection path planning method according to claim 3, characterized in that: The reward function satisfies the following second formula: R t =α1r poi +α2r time -α3r penalty Second formula Among them, R t represents the reward function; r poi Indicates the positive reward given for successfully inspecting the point of interest; α1 represents r poi The corresponding first weight; r time represents the reward for completing the task on time; α2 represents r time The corresponding second weight; r penalty It represents the penalty given when the command trajectory deviates from the zero-control displacement deviation and zero-control velocity deviation framework or the acceleration change of the UAV exceeds the set acceleration change threshold; α3 represents r penalty The corresponding third weight.

6. The method for planning a UAV inspection path according to any one of claims 1 to 5, characterized in that: The target point of interest is obtained by: Acquire multiple points of interest corresponding to the inspection task; Obtaining a priority score corresponding to each of the plurality of points of interest using a multi-factor weighted model, where the multi-factors include at least a historical failure probability of the point of interest, an operating environment, and a distance from a transmission line tower; A preset number of points of interest are selected from high to low according to the priority scores as the target points of interest.

7. The method for planning a UAV inspection path according to any one of claims 1 to 5, characterized in that: After obtaining the target inspection path output by the inspection path planning model, the method further includes: According to the target inspection path, the drone is controlled to perform the inspection task.

8. The UAV inspection path planning method according to claim 7, characterized in that: After controlling the UAV to perform the inspection task, the method further includes: Receiving inspection data corresponding to the inspection task sent by the drone; Generate an inspection report based on the inspection data.

9. A drone inspection path planning device, characterized in that: include: An acquisition module is used to obtain relevant information of the UAV inspection task, the relevant information including first position information and first speed information of the starting point of the inspection task, second position information and second speed information of the end point, and third position information of the target point of interest corresponding to the inspection task; A processing module is used to input the first position information, the first speed information, the second position information, the second speed information and the third position information into a patrol path planning model to obtain a target patrol path output by the patrol path planning model, wherein the patrol path planning model is based on a zero-control displacement deviation and zero-control speed deviation framework and is obtained through optimization training using a proximal strategy optimization algorithm.

10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the UAV inspection path planning method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the UAV inspection path planning method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed, the UAV inspection path planning method according to any one of claims 1 to 8 is implemented.