Path generation method and device of unmanned aerial vehicle, unmanned aerial vehicle and storage medium

By combining the flower pollination algorithm and reinforcement learning algorithm, the problem of low path planning and risk avoidance accuracy of drones in power inspection is solved, and more efficient and safe path planning and obstacle avoidance capabilities are achieved.

CN120085665APending Publication Date: 2025-06-03SHAOGUAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510239501.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Drones face complex terrain and dynamic obstacles during power inspection, resulting in poor accuracy in path planning and emergency risk avoidance.

Method used

A drone path generation method is adopted, combining flower pollination algorithm and reinforcement learning algorithm to optimize path planning. The specific steps include obtaining the path generation request, generating the initial path based on the pollination algorithm, and iteratively optimizing the path and generating the target path through reinforcement learning algorithms, especially dual-deep Q networks and empirical memory/multiplexing strategies.

Benefits of technology

It improves the accuracy of path planning and emergency risk avoidance of drones in complex environments, improves the real-time and efficiency of obstacle avoidance and path planning, enhances the autonomous obstacle avoidance capabilities of drones, and ensures mission safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085665A_ABST
    Figure CN120085665A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a path generation method and device of an unmanned aerial vehicle, the unmanned aerial vehicle and a storage medium. Relates to the technical field of unmanned aerial vehicles. The method comprises the following steps: acquiring a path generation request; wherein the path generation request comprises a starting point and an ending point. And based on the path generation request, performing path optimization processing on a path between the starting point and the ending point according to a preset flower pollination algorithm to generate an initial path. And obtaining path sample data according to a preset reinforcement learning algorithm, and performing path iteration processing on the path sample data and the initial path to generate a target path. The method is used for achieving the effect of improving the accuracy of path planning and emergency risk avoiding of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unmanned aerial vehicles, and in particular, to a method and apparatus for generating a path of an unmanned aerial vehicle, an unmanned aerial vehicle, and a storage medium. Background Art

[0002] Currently, with the development of remote sensing and deep learning technologies, unmanned aerial vehicles (UAVs) provide great opportunities for promoting power line inspection. Using UAVs for power line inspection greatly improves the efficiency of power system operation and maintenance, and ensures the safe and stable operation of the power system. With the expansion of the industrial scale and the in-depth research of technologies, the methods of UAV power line detection have gradually matured, and there are corresponding UAV equipment and detection methods for different detection tasks. UAVs play an important role in power line inspection with their flexible flight mode and high-efficiency data acquisition ability. UAV obstacle avoidance and local path planning are the fundamental guarantees for the flight safety and task completion of UAVs, and require extremely high reliability and practicality.

[0003] However, the actual power scene environment is complex. Power lines often pass through complex terrain environments such as mountains and forests, and the line tower structures are complex. At the same time, the obstacles in the power scene include not only fixed obstacles such as trees and buildings, but also dynamic obstacles such as cables and birds. This makes UAVs lack reliable obstacle perception ability and obstacle handling ability in the face of the complex and changeable environment at low altitude, resulting in poor accuracy in path planning and emergency avoidance for inspection flight robots. Summary of the Invention

[0004] Embodiments of this application provide a method and apparatus for generating a path of an unmanned aerial vehicle, an unmanned aerial vehicle, and a storage medium, so as to achieve the effect of improving the accuracy of path planning and emergency avoidance of the unmanned aerial vehicle.

[0005] In a first aspect, an embodiment of this application provides a method for generating a path of an unmanned aerial vehicle, including:

[0006] Obtaining a path generation request; wherein, the path generation request includes a starting point and an ending point;

[0007] Based on the path generation request, optimizing the path between the starting point and the ending point according to a preset flower pollination algorithm to generate an initial path;

[0008] According to a preset reinforcement learning algorithm, obtaining path sample data, and performing path iteration processing on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the unmanned aerial vehicle between the starting point and the ending point.

[0009] In a possible implementation, the path generation request is based on the preset flower pollination algorithm to optimize the path between the starting point and the ending point to generate an initial path, including:

[0010] Based on the path generation request, according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm, initialize the preset Q-table to obtain the target Q-value of each initial Q-value; wherein, the preset Q-table includes the initial Q-values of multiple state-action pair data;

[0011] According to the target Q-value of each initial Q-value, combine and generate the optimized initial path between the starting point and the ending point.

[0012] In a possible implementation, the path generation request is based on the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to initialize the preset Q-table to obtain the target Q-value of each initial Q-value, including:

[0013] Based on the path generation request, according to the global pollination algorithm, the local pollination algorithm and the control probability set by the user in the preset flower pollination algorithm, initialize the preset Q-table to obtain the target Q-value of each initial Q-value.

[0014] In a possible implementation, the preset reinforcement learning algorithm is used to obtain path sample data, and path iteration processing is performed on the path sample data and the initial path to generate a target path, including:

[0015] According to the experience memory strategy in the preset reinforcement learning algorithm, store the quadruple data generated when the drone executes actions in the environment into a preset experience replay pool; wherein, the quadruple data is the recycled path sample data; the quadruple data includes immediate rewards;

[0016] For the quadruple data in the experience replay pool, according to the experience reuse strategy in the preset reinforcement learning algorithm, set a high priority for the quadruple data where the immediate reward is higher than the preset reward threshold;

[0017] According to the preset reinforcement learning algorithm, perform parameter initialization processing on the main parameters of the main network and the target parameters of the target network in the preset double deep Q-network to obtain the main parameters of the main network and the target parameters of the target network; wherein, the main parameters are equal to the target parameters;

[0018] Repeat the following steps until the preset termination condition is reached:

[0019] Randomly obtain a preset number of high-priority path sample data from a preset experience replay pool; wherein, the path sample data includes a plurality of state-action pair data;

[0020] Select an action in any state according to the main network; and determine the target Q value of the action in the any state according to the target network;

[0021] Update the main parameters of the main network according to the target Q value; at a preset step interval, update the target parameters to the updated main parameters;

[0022] Wherein, the target Q value obtained when the preset termination condition is reached is used to generate a target path.

[0023] In a possible implementation manner, the obtaining path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path includes:

[0024] Based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current start point and end point is similar to the historical flight environment or the historical task target, then call and reuse the quadruple data corresponding to the flight environment in the preset experience replay pool, or call and reuse the quadruple data corresponding to the task target to generate a target path; wherein, the preset experience replay pool includes quadruple data generated and recycled when the unmanned aerial vehicle executes actions in the environment.

[0025] In a possible implementation manner, the path generation request further includes at least one transition point, and the path between the start point and the end point includes a first path between the start point and the transition point and a second path between the transition point and the end point;

[0026] The performing path optimization processing on the path between the start point and the end point according to a preset flower pollination algorithm based on the path generation request to generate an initial path includes:

[0027] Based on the path generation request, perform path optimization processing on the first path according to a preset flower pollination algorithm to generate a first initial path;

[0028] Perform path optimization processing on the second path according to a preset flower pollination algorithm to generate a second initial path;

[0029] Concatenate the first initial path and the second initial path to generate an initial path.

[0030] In a possible implementation, after obtaining path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path, the method further includes:

[0031] Performing smoothing processing on the target path to obtain a smoothed target path.

[0032] In a second aspect, an embodiment of the present application provides a path generation device for a drone, including:

[0033] An acquisition module, configured to acquire a path generation request; wherein, the path generation request includes a starting point and an ending point;

[0034] A first generation module, configured to perform path optimization processing on the path between the starting point and the ending point based on the path generation request according to a preset flower pollination algorithm, and generate an initial path;

[0035] A second generation module, configured to obtain path sample data according to a preset reinforcement learning algorithm, perform path iteration processing on the path sample data and the initial path, and generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the drone between the starting point and the ending point.

[0036] In a possible implementation, the first generation module includes:

[0037] An initialization module, configured to perform initialization processing on a preset Q table based on the path generation request according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm, to obtain a target Q value for each initial Q value; wherein, the preset Q table includes initial Q values of multiple state-action pair data;

[0038] A generation module, configured to generate an optimized initial path between the starting point and the ending point according to the target Q value of each initial Q value.

[0039] In a possible implementation, the initialization module is specifically configured to:

[0040] Perform initialization processing on a preset Q table based on the path generation request according to the global pollination algorithm, the local pollination algorithm, and the control probability set by the user in the preset flower pollination algorithm, to obtain a target Q value for each initial Q value.

[0041] In a possible implementation, the second generation module is specifically configured to:

[0042] According to the experience memory strategy in the preset reinforcement learning algorithm, store the quadruple data generated when the drone executes actions in the environment into a preset experience replay pool; wherein, the quadruple data is the recovered path sample data; the quadruple data includes an immediate reward.

[0043] For the quadruple data in the experience replay pool, according to the experience reuse strategy in the preset reinforcement learning algorithm, set a high priority for the quadruple data where the immediate reward is higher than the preset reward threshold.

[0044] According to the preset reinforcement learning algorithm, perform parameter initialization processing on the main parameters of the main network and the target parameters of the target network in the preset double deep Q network to obtain the main parameters of the main network and the target parameters of the target network; wherein, the main parameters are equal to the target parameters.

[0045] Repeat the following steps until a preset termination condition is reached:

[0046] Randomly obtain a preset number of high-priority path sample data in the preset experience replay pool; wherein, the path sample data includes a plurality of state-action pair data.

[0047] Select an action in any state according to the main network; and determine the target Q value of the action in the any state according to the target network.

[0048] Update the main parameters of the main network according to the target Q value; at a preset step interval, update the target parameters to the updated main parameters.

[0049] Wherein, the target Q value obtained when the preset termination condition is reached is used to generate a target path.

[0050] In a possible implementation manner, the second generation module is specifically configured to:

[0051] Based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current start point and end point is similar to the historical flight environment or historical task target, then call and reuse the quadruple data corresponding to the flight environment in the preset experience replay pool, or call and reuse the quadruple data corresponding to the task target to generate a target path; wherein, the preset experience replay pool includes quadruple data generated and recovered when the drone executes actions in the environment.

[0052] In a possible implementation manner, the path generation request further includes at least one transition point, and the path between the start point and the end point includes a first path between the start point and the transition point and a second path between the transition point and the end point.

[0053] The first generation module is specifically configured to:

[0054] Based on the path generation request, for the first path, perform path optimization processing on the first path according to the preset flower pollination algorithm to generate a first initial path;

[0055] For the second path, perform path optimization processing on the second path according to the preset flower pollination algorithm to generate a second initial path;

[0056] Concatenate the first initial path and the second initial path to generate an initial path.

[0057] In a possible implementation manner, the device is further specifically configured to:

[0058] After obtaining path sample data according to the preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path to generate a target path, perform smoothing processing on the target path to obtain a smoothed target path.

[0059] In a third aspect, an embodiment of the present application provides a drone, including: a memory, a processor;

[0060] The memory stores computer execution instructions;

[0061] The processor executes the computer execution instructions stored in the memory, so that the processor executes the various possible implementation manners in the first aspect as above.

[0062] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the various possible implementation manners in the first aspect as above.

[0063] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the various possible implementation manners in the first aspect as above.

[0064] The path generation method, device, drone, and storage medium provided by the embodiments of the present application obtain a path generation request; wherein, the path generation request includes a starting point and an ending point. Based on the path generation request, the path between the starting point and the ending point is optimized according to the preset flower pollination algorithm to generate an initial path. According to the preset reinforcement learning algorithm, path sample data is obtained, and path iteration processing is performed on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the drone between the starting point and the ending point. In this solution, for the environmental changes and dynamic obstacles encountered by the drone during power inspection, by introducing the global search and local search characteristics of the flower pollination algorithm (FPA), the initialization of the reinforcement learning (Q-learning) is optimized to generate an optimal initial path, so that the subsequent reinforcement learning algorithm starts learning from a better initial path, thereby accelerating the convergence speed, enabling the drone and other intelligent agents to quickly generate effective obstacle avoidance paths in complex environments, improving the real-time performance and efficiency of obstacle avoidance and path planning, thus realizing efficient local path obstacle avoidance and path planning, improving the efficiency of early learning and the accuracy of path planning, effectively enhancing the autonomous obstacle avoidance ability of the drone, enabling it to identify and avoid obstacles in real time during flight, ensuring mission safety, and achieving the effect of improving the accuracy of the drone's path planning and emergency avoidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0066] Figure 1 Schematic flowchart of a path generation method for a drone provided by an embodiment of the present application Figure 1 ;

[0067] Figure 2 Schematic flowchart of another path generation method for a drone provided by an embodiment of the present application Figure 2 ;

[0068] Figure 3 Schematic flowchart of a path generation method for a drone provided by an embodiment of the present application Figure 3 ;

[0069] Figure 4 Schematic diagram of the scenario of a path generation method for a drone provided by an embodiment of the present application;

[0070] Figure 5 Schematic flowchart of a path generation method for a drone provided by an embodiment of the present application Figure 4 ;

[0071] Figure 6 It is a schematic structural diagram of a path generation device for a drone provided by an embodiment of the present application;

[0072] Figure 7 It is a schematic structural diagram of another path generation device for a drone provided by an embodiment of the present application;

[0073] Figure 8 It is a schematic structural diagram of a drone provided by an embodiment of the present application.

[0074] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0075] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0076] Currently, with the development of remote sensing and deep learning technologies, drones have provided great opportunities for promoting power line inspections. Using drones for power line inspections has greatly improved the efficiency of power system operation and maintenance, and ensured the safe and stable operation of the power system. With the expansion of the industrial scale and the in-depth research of technologies, the methods of drone power line detection have gradually matured, and there are corresponding drone equipment and detection methods for different detection tasks. Drones play an important role in power line inspections with their flexible flight methods and high-efficiency data acquisition capabilities. Drone obstacle avoidance and local path planning are the fundamental guarantees for drone flight safety and task completion, and require extremely high reliability and practicality.

[0077] However, the actual power scenario has a complex environment. Power lines often pass through complex terrain environments such as mountains and forests, and the line tower structures are complex. At the same time, the obstacles in the power scenario include not only fixed obstacles such as trees and buildings, but also dynamic obstacles such as cables and flying birds. Traditional path planning and obstacle avoidance algorithms often have difficulty quickly coping with the characteristics of these changing obstacles. Existing sensors (such as lidar, vision sensors, etc.) are prone to being affected by environmental conditions (such as light, weather, etc.) when identifying small and distant power equipment and surrounding obstacles, resulting in inaccurate and untimely data. Therefore, in the existing technology, for UAVs in terms of autonomous obstacle avoidance and path planning, facing the complex and changeable environment at low altitude, at present, UAVs lack reliable obstacle perception ability and obstacle handling ability, resulting in poor accuracy of path planning and emergency avoidance for inspection flying robots.

[0078] In one example, during the UAV inspection process, there may be an urgent obstacle avoidance requirement near power equipment. Traditional path planning algorithms, such as the A* algorithm or RRT (Rapidly-exploring Random Tree) algorithm based on graph search, have a high computational complexity and are difficult to complete efficient local path planning in a short time. Many algorithms often consume a large amount of time in local path optimization to avoid complex obstacles, affecting the overall inspection efficiency. At present, when UAVs perform local path planning in a complex environment, they often do not fully consider the trade-off relationship between battery power consumption and the length of the obstacle avoidance path, which may lead to the inability to complete the inspection task when the battery runs out. For large-scale power line inspection tasks, path planning needs to be both accurate and fast, and it is difficult to achieve a balance between the two.

[0079] In one example, in the power inspection task, the inspection area is usually very large, and both global path planning is required to complete the overall task and local path planning is required to address detailed obstacle avoidance needs. However, many existing algorithms lack a mechanism to organically combine global path planning and local path obstacle avoidance, resulting in poor coordination between global path planning and local path planning in complex scenarios. Under the influence of various obstacles and environmental changes, the robustness of local path planning is poor, easily leading to local optimality or failure and being unable to complete the inspection task smoothly.

[0080] To sum up, although many scholars have conducted research on UAV power inspection obstacle avoidance and local planning, in actual power inspection applications, further research and optimization are still needed to better cope with various challenges in large-scale and complex power scenarios.

[0081] Combined with the above scenarios, in the existing technology, there is a technical problem that the accuracy of path planning and emergency avoidance for UAVs is poor.

[0082] The path generation method for the drone provided by this application solves the technical problem of poor accuracy in path planning and emergency avoidance for drones.

[0083] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0084] Figure 1 Flow schematic of a path generation method for a drone provided by an embodiment of this application Figure 1 , as Figure 1 shown, this method includes:

[0085] S101. Obtain a path generation request; wherein, the path generation request includes a starting point and an ending point.

[0086] Exemplarily, the execution subject of this embodiment can be an intelligent agent such as a drone, an electronic device, or a terminal device, or a path generation device or equipment of a drone, or other devices or equipment that can execute this embodiment, and there is no limitation thereto. In this embodiment, the execution subject is introduced as a drone.

[0087] First, the drone is an inspection flight robot. The drone needs to obtain a path generation request, which includes a starting point and an ending point. Specifically, geographical spatial data can be obtained first, such as terrain elevation data, building outlines, road networks, etc. Then, the geographical spatial data is converted into a grid form. For example, terrain grids can be generated using altitude data. Since a three-dimensional grid map organizes and represents data by dividing the geographical space into grid cells, each grid cell (or "grid unit") represents a specific spatial range with a fixed size and shape. Based on the grid-form geographical spatial data, the starting point and the ending point in the path generation request are obtained.

[0088] S102. Based on the path generation request, perform path optimization processing on the path between the starting point and the ending point according to the preset flower pollination algorithm to generate an initial path.

[0089] Exemplarily, since all initial values in traditional reinforcement learning are set to zero in complex environments, the path selection in the initial stage is random and the convergence speed is slow. The Flower Pollination Algorithm (FPA) optimizes the initial values to guide the drone to start exploration from a more advantageous initial path, greatly improving the convergence speed of the algorithm, reducing the random exploration of the drone in complex scenarios, and shortening the planning time. The flower pollination algorithm mimics the process of pollen transfer in nature and uses the global and local propagation characteristics of pollen to achieve objective optimization. Its basic mechanisms include: Global pollination: Through biotic (biological vector transfer) and cross-pollination, pollen is spread over long distances, with the characteristics of global exploration. Local pollination: Through abiotic (such as wind) and self-pollination, pollen is spread in local areas, with the characteristics of local optimization.

[0090] The switching between global and local pollination is determined by a control probability p ∈ [0, 1], which is used to control the switching of the algorithm between different modes. The mathematical model of the flower pollination algorithm is as follows:

[0091] Global pollination formula: In global pollination, the position update of pollen follows the Levy flight random distribution, which can promote large-scale search and prevent the algorithm from falling into local optima. The update formula is:

[0092]

[0093] Where: represents the position of the i-th pollen in the (t + 1)-th iteration; represents the position of the pollen at the t-th iteration; γ is the step size scaling factor, usually taking a small value; g * represents the currently found global optimal solution; L(λ) is the Levy flight distribution, which is used to generate a random flight path with a large step size to increase the exploration range.

[0094] Local pollination formula: Local pollination only occurs between nearby flowers to meet the needs of local search. Its update formula is:

[0095]

[0096] Where: and represent two randomly selected pollens from the same flower species; ∈ is a random number, following a uniform distribution on [0, 1], which controls the amplitude of local variation.

[0097] The flower pollination algorithm is used in Q - learning to initialize the Q - table, and the Q - value of each state - action pair data is initialized to the best possible value. The process is as follows:

[0098] Initialize the pollen positions: Each state - action pair data is regarded as a pollen, and the position of each pollen represents the Q - value of the corresponding state - action pair data.

[0099] Global and local update of Q - values: Through the global and local pollination processes, appropriate Q - values are generated during the initialization process to avoid random behavior in initial learning. According to the iteration of the flower pollination algorithm, path optimization is performed on the path between the starting point and the ending point. High - potential - value Q - values are assigned to the corresponding state - action pair data, and then multiple high - potential - value Q - values are obtained. Based on these multiple high - potential - value Q - values, an initial path is generated, enabling the drone to move towards the target area faster.

[0100] Specifically, the initialization steps of the Q - table are as follows:

[0101] Q(s, a)=Q init +γL(λ)(g * -Q init )

[0102] where Q(s, a) represents the initial Q - value of state s and action a; Q init is the initial path determined according to the global or local pollination rule.

[0103] S103. According to the preset reinforcement learning algorithm, obtain path sample data, and perform path iteration processing on the path sample data and the initial path to generate a target path; where the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the drone between the starting point and the ending point.

[0104] Exemplarily, the Double Deep Q - Network (DDQN) is an improved Deep Q - Network (DQN) algorithm used to solve the bias problem of traditional DQN in value estimation. The DQN algorithm tends to overestimate the Q - value of actions during decision - making. Especially in complex and dynamic environments, this overestimation may lead to decision - making biases and cause the algorithm to fall into a sub - optimal solution. DDQN effectively reduces this bias by introducing a double - network structure, which cooperates in action selection and Q - value estimation, improving the stability and convergence speed of the algorithm.

[0105] Furthermore, DDQN contains two neural networks, namely the main network (Current Network) which is used to select actions in the current state, and the target network (Target Network) which is used to calculate the target Q value to update the parameters of the main network. These two networks have the same structure but independent parameters. The parameters of the target network are copied from the main network after a certain number of steps, so that the target Q value remains relatively stable during the training process, preventing the training from being unstable or divergent.

[0106] Furthermore, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy. Among them, the experience memory strategy is used to recycle the generated quadruple data, i.e., the path sample data and experience, to a preset experience replay pool after an agent such as a drone executes an action in the environment. The parameters in the quadruple data (s, a, r, s′) include state, action, reward, and the next state in sequence; the experience reuse strategy is used to assign priorities to each quadruple data in the experience replay pool and can preferentially call relevant experiences when it detects that the current environment is similar to the past without having to explore again. Therefore, through the experience sharing mechanism of the experience reuse strategy, agents such as drones can reuse the already formed locally optimal paths, avoid repeated exploration, and improve the task execution efficiency.

[0107] In this step, according to the double deep Q network, a preset number of high-priority path sample data are randomly selected from the preset experience replay pool, and path iteration processing is performed on the path sample data and the initial path according to the preset reinforcement learning algorithm to generate a target path, which is the actual flight path of the drone between the starting point and the ending point. The high-priority path sample data can be randomly selected from multiple path sample data greater than the preset priority level threshold, or the priorities can be sorted and a preset number of high-priority path sample data at the front positions can be selected. There is no limitation on this.

[0108] The path generation method for the drone provided by the embodiment of the present application obtains a path generation request; wherein, the path generation request includes a starting point and an ending point. Based on the path generation request, the path between the starting point and the ending point is optimized by the preset flower pollination algorithm to generate an initial path. According to the preset reinforcement learning algorithm, path sample data is obtained, and path iteration processing is performed on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the drone between the starting point and the ending point. In this solution, for the environmental changes and dynamic obstacles encountered by the drone during power inspection, by introducing the global search and local search characteristics of the flower pollination algorithm (FPA), the initialization of the reinforcement learning (Q-learning) is optimized to generate an optimal initial path, so that the subsequent reinforcement learning algorithm starts learning from a better initial path, thereby accelerating the convergence speed, enabling intelligent agents such as drones to quickly generate effective obstacle avoidance paths in complex environments, improving the real-time performance and efficiency of obstacle avoidance and path planning, thus realizing efficient local path obstacle avoidance and path planning, improving the efficiency of early learning and the accuracy of path planning, effectively enhancing the autonomous obstacle avoidance ability of the drone, enabling it to identify and avoid obstacles in real time during flight, ensuring mission safety, and achieving the effect of improving the accuracy of path planning and emergency avoidance of the drone.

[0109] Figure 2 Schematic flow of a method provided by the present application Figure 2 As Figure 2 shown, on the basis of the Figure 1 embodiment, the path generation method for the drone is described in detail. The method includes:

[0110] S201. Obtain a path generation request; wherein, the path generation request includes a starting point and an ending point.

[0111] Exemplarily, this step can refer to Figure 1 step 101 in

[0112] and will not be elaborated herein.

[0113] In one example, S202 includes: based on the path generation request, the preset Q-table is initialized according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to obtain the target Q-value of each initial Q-value; wherein, the preset Q-table includes the initial Q-values of multiple state-action pair data; according to the target Q-value of each initial Q-value, the optimized initial path between the starting point and the ending point is combined and generated.

[0114] In one example, "generating a request based on a path, and initializing a preset Q-table according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to obtain the target Q-value of each initial Q-value" includes: generating a request based on a path, and initializing a preset Q-table according to the global pollination algorithm, the local pollination algorithm, and the control probability set by the user in the preset flower pollination algorithm to obtain the target Q-value of each initial Q-value.

[0115] In one example, S202 includes: the path generation request further includes at least one transition point, and the path between the starting point and the ending point includes a first path between the starting point and the transition point and a second path between the transition point and the ending point; based on the path generation request, for the first path, performing path optimization processing on the first path according to the preset flower pollination algorithm to generate a first initial path; for the second path, performing path optimization processing on the second path according to the preset flower pollination algorithm to generate a second initial path; splicing the first initial path and the second initial path to generate an initial path.

[0116] Exemplarily, the flower pollination algorithm mimics the process of pollen transfer in nature and uses the global and local propagation characteristics of pollen to achieve target optimization. Its basic mechanisms include: global pollination: through biological (biological vector transmission) and cross-pollination, pollen is spread over a long distance, with the characteristics of global exploration. Local pollination: through abiotic (such as wind) and self-pollination methods, pollen is spread in a local area, with the characteristics of local optimization.

[0117] The switching between global and local pollination is determined by a control probability p ∈ [0, 1], which is used to control the switching of the algorithm between different modes. The mathematical model of the flower pollination algorithm is as follows:

[0118] Global pollination formula: In global pollination, the position update of pollen follows the Levy flight random distribution, which can promote large-scale search and prevent the algorithm from falling into local optima. The update formula is:

[0119]

[0120] Where: represents the position of the i-th pollen in the (t + 1)-th iteration; represents the position of the pollen at the t-th iteration; γ is the step size scaling factor, usually taking a relatively small value; g * represents the currently found global optimal solution; L(λ) is the Levy flight distribution, which is used to generate a random flight path with a large step size to increase the exploration range.

[0121] Local pollination formula: Local pollination is only carried out between nearby flowers to meet the requirements of local search. Its update formula is as follows:

[0122]

[0123] Where: and represent two randomly selected pollens from the same flower species; ∈ is a random number, following a uniform distribution on [0, 1], which controls the amplitude of local variation.

[0124] The flower pollination algorithm is used to initialize the Q - table in Q - learning, initializing the Q - value of each state - action pair data to the possible best value. This process is as follows:

[0125] Initialize the pollen positions: Consider each state - action pair data as a pollen, and the position of each pollen represents the Q - value of the corresponding state - action pair data.

[0126] Global and local update of Q - values: Through the global and local pollination processes, appropriate Q - values are generated during the initialization process to avoid random behavior in initial learning. According to the iteration of the flower pollination algorithm, path optimization is performed on the path between the starting point and the ending point. High - potential - value Q - values are assigned to the corresponding state - action pair data, and then multiple high - potential - value Q - values are obtained. Based on these multiple high - potential - value Q - values, an optimized initial path between the starting point and the ending point is combined, enabling the UAV to move towards the target area faster.

[0127] Specifically, the initialization steps of the Q - table are as follows:

[0128] Q(s, a)=Q init +γL(λ)(g * -Q init ) where Q(s, a) represents the initial Q - value of state s and action a; Q init is the initial path determined according to the global or local pollination rule.

[0129] Optionally, the control probability p controls the weights of global pollination and local pollination. Global pollination is the global optimization of the path to avoid falling into local optimal solutions, and local pollination ensures the detailed optimization of local paths. This parameter is an empirical value, and by modifying this parameter for testing, the optimality of the path is guaranteed. When p is large (close to 1), global pollination dominates, and the algorithm tends to perform large - step - size searches to find the global optimal solution in a large search space. When p is small (close to 0), local pollination dominates, and the algorithm tends to perform small - step - size optimizations in the neighborhood of the current solution to improve the accuracy of the solution.

[0130] Optionally, the path generation request further includes at least one transition point. Taking one transition point as an example, the path between the starting point and the ending point includes a first path between the starting point and the transition point and a second path between the transition point and the ending point. Based on the path generation request, for the first path, the first path is processed for path optimization according to the preset flower pollination algorithm to generate a first initial path. For the second path, the second path is processed for path optimization according to the preset flower pollination algorithm to generate a second initial path. Finally, the first initial path and the second initial path are spliced to generate an initial path.

[0131] Therefore, this application utilizes the global and local search characteristics of the flower pollination algorithm to optimize the initialization of the Q table, enhancing the convergence speed and path optimization effect of the algorithm in the early stage. Through the random search method of Levy flight, the exploration ability of the algorithm in a complex environment with dense obstacles is improved, avoiding falling into local optimal solutions.

[0132] S203. According to the preset reinforcement learning algorithm, obtain path sample data, and perform path iteration processing on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the unmanned aerial vehicle between the starting point and the ending point.

[0133] In one example, S203 includes: according to the experience memory strategy in the preset reinforcement learning algorithm, storing the quadruple data generated when the unmanned aerial vehicle performs actions in the environment into a preset experience replay pool; wherein, the quadruple data is the recycled path sample data; the quadruple data includes an immediate reward; for the quadruple data in the experience replay pool, according to the experience reuse strategy in the preset reinforcement learning algorithm, set a high priority for the quadruple data where the immediate reward is higher than the preset reward threshold; according to the preset reinforcement learning algorithm, perform parameter initialization processing on the main parameters of the main network and the target parameters of the target network in the preset double deep Q network to obtain the main parameters of the main network and the target parameters of the target network; wherein, the main parameters are equal to the target parameters; repeat the following steps until a preset termination condition is reached:

[0134] Randomly obtain a preset number of high-priority path sample data in the preset experience replay pool; wherein, the path sample data includes multiple state-action pair data; select an action in any state according to the main network; and determine the target Q value of the action in any state according to the target network; update the main parameters of the main network according to the target Q value; at preset step intervals, update the target parameters to the updated main parameters; wherein, the target Q value obtained when the preset termination condition is reached is used to generate the target path.

[0135] In one example, S203 includes: based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current starting point and ending point is similar to the historical flight environment or historical task objective, then call and reuse the quadruple data corresponding to the flight environment in the preset experience replay pool, or call and reuse the quadruple data corresponding to the task objective to generate a target path; wherein, the preset experience replay pool includes the quadruple data generated and recycled when the drone executes actions in the environment.

[0136] Exemplarily, Figure 3 is a schematic flow chart of a method for generating a path of a drone provided by an embodiment of the present application Figure 3 , as Figure 3 shown, in DDQN, action selection and target Q-value calculation are performed separately:

[0137] Action Selection: Use the main network (Current Network) to select an action a' in the current state s':

[0138]

[0139] where θ is the parameter of the main network. This step is only used to select the next action, rather than directly calculating the Q-value.

[0140] Target Q-value calculation: Input the selected action a' into the target network (Target Network) to calculate the target Q-value. The target Q-value is defined as:

[0141] y = r + γQ(s', a'; θ - )

[0142] where: r is the immediate reward; γ is the discount factor; a' is the action selected in state s'; a is the action; Q(s', a'; θ - ) is the Q-value estimate; y is the target Q-value; θ - is the parameter of the target network, used to calculate the Q-value of the action selected in state s'.

[0143] This separate calculation method effectively avoids the problem of using the same network to select and evaluate the action Q-value simultaneously in traditional DQN, thereby reducing the overestimation bias.

[0144] The goal of DDQN is to minimize the error between the predicted Q-value and the target Q-value. The loss function is usually defined as the mean square error:

[0145] L(θ) = E[(y - Q(s, a; θ)) 2

[0146] ​where: y is the target Q value; Q(s,a;θ) is the Q value estimation of the main network for action a in state s.

[0147] By minimizing this loss function, the parameters θ of the main network can be continuously updated to converge to the true Q value.

[0148] In DDQN, the parameters θ of the target network - are not updated at each iteration. Instead, the parameters θ of the main network are copied to the target network every certain number of steps (for example, every 500 steps, which is not limited here). The advantage of this design is that the Q value of the target network remains stable for a period of time, reducing the update fluctuations and enhancing the stability of the algorithm.

[0149] In this step, the complete DDQN algorithm process is as follows:

[0150] 1. Initialize the main parameters θ and target parameters θ of the main network and target network in the double deep Q network - , and let θ = θ - .

[0151] 2. Repeat the following steps until the preset termination condition is reached:

[0152] (1) Randomly draw a preset number of a batch of high-priority path sample data (s,a,r,s′) from the preset experience replay pool; the path sample data includes multiple state-action pair data;

[0153] (2) Action selection: Use the main network to select action a′ in state s′:

[0154] a′ = arg max a Q(s′,a;θ)

[0155] (3) Target Q value calculation: Use the target network to calculate the target Q value of this action a′: y = r + γQ(s′,a′;θ - ); The parameter meanings are as described above and will not be elaborated here.

[0156] (4) Update the main network parameters θ: Minimize the loss function L(θ) through backpropagation to update the main network parameters θ; thus, the parameter update is achieved according to minimizing the loss function.

[0157] (5) Every preset number of step intervals, copy the value of θ to the parameters θ of the target network - , that is, update the target parameters to the updated main parameters.

[0158] Finally, the target Q-value obtained when the preset termination condition is reached can generate the target path. Among them, the preset termination condition is any one or more of the following: the training reaches the maximum number of iterations, the training reward converges, the Q-value change converges, the success rate of the agent reaches the target, and the training time exceeds the limit.

[0159] Optionally, this step introduces the experience memory strategy and the experience reuse strategy. In reinforcement learning, the experience memory (Experience Replay) and experience reuse (Experience Reuse) strategies are used to improve the learning efficiency and generalization ability of the agent. The experience memory strategy stores past experience samples to avoid bias problems caused by data correlation, while the experience reuse strategy enables the agent to reuse valuable experiences, which is particularly suitable for path planning problems in dynamic environments.

[0160] The experience memory strategy is a key technology in reinforcement learning to optimize learning by storing historical data. The experience replay pool is used to store the quadruple data (s, a, r, s') of the states, actions, rewards, and next states encountered by agents such as drones at multiple past moments. During training, agents such as drones randomly draw samples from the experience replay pool for training to reduce the temporal correlation of the data.

[0161] The process of the experience memory strategy:

[0162] (1) Storing experience: After an agent such as a drone executes an action in the environment, the generated quadruple data (s, a, r, s') is stored in the experience replay pool. The experience replay pool has a maximum capacity. When this capacity is exceeded, the old experience will be eliminated to ensure that the data in the experience replay pool is the latest and representative.

[0163] (2) Random sampling: In each training, a small batch of data is randomly drawn from the experience replay pool. This random sampling breaks the temporal correlation of the data, makes the learning process more stable, and improves the generalization ability of the model.

[0164] (3) Batch update: The network parameters are updated through gradient descent or other optimization algorithms of the batch samples. Due to the independence and diversity of random sampling, the agent can learn a wider range of state and action relationships and avoid biases towards a certain type of action.

[0165] Experience reuse further utilizes the existing experience memory in specific scenarios, which is particularly important in path planning and dynamic environment adaptation. Experience reuse enables the agent to reuse the existing experience to re-plan the path and avoid obstacles when the environment changes, improving the learning efficiency and adaptability. The experience reuse strategy includes:

[0166] (1) Prioritized Experience Replay: Assign priorities to each sample in the experience replay pool and preferentially use those experiences that are more helpful for training. For example, for samples with higher reward values or larger errors, higher priorities are assigned, thereby increasing the usage frequency of these samples. Prioritized Experience Replay is particularly effective in path planning and obstacle avoidance tasks, helping the agent focus on critical paths or obstacle avoidance points.

[0167] (2) Experience Sharing and Reuse: In path planning scenarios, when agents such as drones detect that the environment is similar to the past (e.g., similar obstacle distributions), relevant experiences can be preferentially invoked without having to re-explore. Through the experience sharing mechanism, agents such as drones can reuse the locally optimal paths that have already been formed, avoiding repeated exploration and improving task execution efficiency.

[0168] Combined with the DDQN algorithm, the complete process of combining experience memory and experience reuse in this step is as follows:

[0169] (1) Execute actions and store experiences: Agents such as drones execute actions in the environment, and the obtained experiences (state, action, reward, next state) are stored in the experience replay pool. The experiences are quadruple data.

[0170] (2) Priority sorting: Store the experiences according to their priorities, and experiences with higher priorities will be used more frequently in subsequent sampling.

[0171] (3) Batch sampling to update the network: At each training step, randomly sample a batch of experiences with higher priorities, calculate the target Q value through the target network of DDQN, and update the parameters of the main network.

[0172] (4) Experience reuse: Based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current starting point and ending point is similar to the historical flight environment or historical task objective, call the experiences with specific priorities corresponding to the flight environment for reuse, or call the experiences corresponding to the task objective for reuse to generate the target path to achieve dynamic environment adaptation.

[0173] Therefore, this application adopts a Double Deep Q-Network (DDQN) structure. Through the division of labor between the main network and the target network, the problem of overestimating Q-values is reduced, and the accuracy and stability of path planning are improved. Combined with the introduced experience memory and priority experience replay mechanism, historical experiences are stored in the experience replay pool. When the environment changes dynamically, past key experiences can be reused, and key experiences helpful for obstacle avoidance are preferentially reused to quickly respond to newly emerging obstacles and re-plan the path, improving learning efficiency and dynamic adaptability. Priority experience replay further strengthens the ability to reuse key obstacle avoidance experiences, enabling the intelligent agent to continuously generate safe and stable paths in complex dynamic environments. Moreover, this application integrates dynamic obstacle avoidance and safety distance detection functions in path planning. By real-time monitoring the distance between intelligent agents such as drones and obstacles, it ensures that the intelligent agent always maintains a safe distance. Once the distance is lower than the set threshold, intelligent agents such as drones will automatically adjust the path or decelerate, thus effectively avoiding the risk of collision and enhancing the safety in high-density obstacle areas. By designing reward and punishment strategies, the intelligent agent is encouraged to move towards the target area while avoiding approaching obstacles, improving the overall safety of path planning.

[0174] S204. Smooth the target path to obtain the smoothed target path.

[0175] Exemplarily, optimizing path smoothness and safety is the key to ensuring the smooth and stable completion of tasks by drones. Path smoothness focuses on the feasibility and smoothness of the path to reduce unnecessary turns or sharp turns, save energy, and improve task execution efficiency; while safety optimization ensures that the intelligent agent maintains a sufficient safe distance from obstacles in a complex environment to prevent collisions.

[0176] The optimization of path smoothness aims to eliminate sharp turning angles in the path and make the path more smooth and natural. For drones or mobile robots, a smooth path can reduce the number and angle of turns, thereby reducing energy consumption and improving navigation efficiency. Commonly used smoothing techniques include Bezier curves, which are not limited herein. The Bezier curve is a commonly used path smoothing method. Its advantages are simple calculation, good continuity and smoothness, and it is very suitable for smoothing the path.

[0177] Second-order Bezier curve: A smooth curve is generated by three control points and is defined as follows:

[0178] B(t) = (1 - t) 2 P 0 + 2(1 - t)tP 1 + t 2 P 2 , t ∈ [0, 1]

[0179] Among them, P 0 、P1 , P 2 are control points respectively, and t is a parameter.

[0180] Cubic Bézier curve: A smooth curve is generated by four control points, and its formula is:

[0181] B(t) = (1 - t) 3 P 0 + 3(1 - t) 2 tP 1 + 3(1 - t)t 2 P 2 + t 3 P 3 , t ∈ [0, 1]

[0182] where P 0 , P 1 , P 2 , P 3 are control points. The cubic Bézier curve has higher smoothness and is suitable for fine smoothing of paths.

[0183] In this step, Figure 4 is a schematic diagram of the scenario of a path generation method for a drone provided by an embodiment of the present application. As Figure 4 shown, the path to be smoothed includes b 0 , b 0 1 , b 1 , b 1 1 , b 2 . By interpolating b 0 1 , b 1 1 with a Bézier curve, a smoother target path b 0 , b 0 2 , b 2 is generated, enabling the drone or robot to move along the smooth curve and avoiding unnecessary sharp turns and vibrations.

[0184] Therefore, through the above iteration, a multi-resolution grid with semantic information is generated according to different power equipment, and based on this map, a power equipment inspection route for global optimization is generated. The paths generated by traditional methods are prone to problems such as sharp turns and discontinuous curvatures, resulting in frequent sudden stops and accelerations of the agent on the path, increasing energy consumption and affecting stability. By introducing the smoothing processing of Bézier curves and smoothing spline curves in the present application, the generated paths are more continuous and smooth, reducing sharp turns, improving the execution efficiency and smoothness of the paths. The smooth paths also effectively reduce the energy consumption during the navigation process and extend the battery life of the equipment.

[0185] The path generation method of the unmanned aerial vehicle provided by the embodiment of the present application obtains a path generation request; wherein, the path generation request includes a starting point and an ending point. Based on the path generation request, the path between the starting point and the ending point is optimized according to the preset flower pollination algorithm to generate an initial path. According to the preset reinforcement learning algorithm, path sample data is obtained, and path iteration processing is performed on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the unmanned aerial vehicle between the starting point and the ending point. The target path is smoothed to obtain a smoothed target path. In this solution, for the environmental changes and dynamic obstacles encountered by the unmanned aerial vehicle during power inspection, by introducing the global search and local search characteristics of the flower pollination algorithm (FPA), the initialization of the reinforcement learning (Q-learning) is optimized to generate an optimal initial path, so that the subsequent reinforcement learning algorithm starts learning from a better initial path, thereby accelerating the convergence speed, enabling intelligent agents such as unmanned aerial vehicles to quickly generate effective obstacle avoidance paths in complex environments, improving the real-time performance and efficiency of obstacle avoidance and path planning, thus realizing efficient local path obstacle avoidance and path planning, improving the efficiency of early learning and the accuracy of path planning, effectively enhancing the autonomous obstacle avoidance ability of the unmanned aerial vehicle, enabling it to identify and avoid obstacles in real time during flight, ensuring mission safety, and achieving the effect of improving the accuracy of path planning and emergency avoidance of the unmanned aerial vehicle. And also aiming at the deficiencies of the existing inspection route planning, a dynamically adjustable path planning method is proposed, enabling the unmanned aerial vehicle to flexibly respond to emergencies in complex and changeable environments and ensuring the smooth progress of the inspection mission.

[0186] Figure 5 It is a schematic flow chart of a path generation method of an unmanned aerial vehicle provided by an embodiment of the present application Figure 4 , as Figure 5 shown, start; three-dimensional grid map, waypoints, where the waypoints include a starting point and an ending point, etc.; initial path guidance of the unmanned aerial vehicle based on FDA; constructing a double deep Q-network structure to achieve rapid path screening; introducing an experience memory module and a reuse strategy to record the obstacle avoidance experience of the unmanned aerial vehicle in real time; path smoothing and safety optimization; end.

[0187] Figure 6 It is a schematic structural diagram of a path generation device provided by the present application. As Figure 6 shown, the device 40 provided in this embodiment includes:

[0188] An acquisition module 41, configured to acquire a path generation request; wherein, the path generation request includes a starting point and an ending point;

[0189] The first generation module 42 is configured to generate an initial path by performing path optimization processing on the path between the starting point and the ending point according to a preset flower pollination algorithm based on a path generation request.

[0190] The second generation module 43 is configured to obtain path sample data according to a preset reinforcement learning algorithm, and perform path iteration processing on the path sample data and the initial path to generate a target path; wherein, the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the unmanned aerial vehicle between the starting point and the ending point.

[0191] Figure 7 This is a schematic structural diagram of another path generation device for an unmanned aerial vehicle provided by an embodiment of the present application. On the basis of the embodiment shown in Figure 6 as shown in Figure 7 as shown, the first generation module 42 includes:

[0192] An initialization module 421 is configured to perform initialization processing on a preset Q-table based on a path generation request according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to obtain a target Q-value for each initial Q-value; wherein, the preset Q-table includes initial Q-values of multiple state-action pair data.

[0193] A generation module 422 is configured to generate an optimized initial path between the starting point and the ending point by combining according to the target Q-value of each initial Q-value.

[0194] In a possible implementation manner, the initialization module 421 is specifically configured to:

[0195] Perform initialization processing on a preset Q-table based on a path generation request according to the global pollination algorithm, the local pollination algorithm, and the control probability set by the user in the preset flower pollination algorithm to obtain a target Q-value for each initial Q-value.

[0196] In a possible implementation manner, the second generation module 43 is specifically configured to:

[0197] According to the experience memory strategy in the preset reinforcement learning algorithm, store the quadruple data generated when the unmanned aerial vehicle performs an action in the environment into a preset experience replay pool; wherein, the quadruple data is the recycled path sample data; the quadruple data includes an immediate reward.

[0198] For the quadruple data in the experience replay pool, according to the experience reuse strategy in the preset reinforcement learning algorithm, set a high priority for the quadruple data where the immediate reward higher than the preset reward threshold is located.

[0199] Initialize the main parameters of the main network and the target parameters of the target network in the preset double deep Q-network according to the preset reinforcement learning algorithm to obtain the main parameters of the main network and the target parameters of the target network; wherein, the main parameters are equal to the target parameters.

[0200] Repeat the following steps until the preset termination condition is reached:

[0201] Randomly obtain a preset number of high-priority path sample data in the preset experience replay pool; wherein, the path sample data includes a plurality of state-action pair data.

[0202] Select an action in any state according to the main network; and determine the target Q-value of the action in the any state according to the target network.

[0203] Update the main parameters of the main network according to the target Q-value; update the target parameters to the updated main parameters at preset step intervals.

[0204] Wherein, the target Q-value obtained when the preset termination condition is reached is used to generate a target path.

[0205] In a possible implementation manner, the second generation module 43 is specifically configured to:

[0206] Based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current starting point and ending point is similar to the historical flight environment or the historical task target, then call and reuse the quadruple data corresponding to the flight environment in the preset experience replay pool, or call and reuse the quadruple data corresponding to the task target to generate a target path; wherein, the preset experience replay pool includes quadruple data generated and recycled when the unmanned aerial vehicle executes actions in the environment.

[0207] In a possible implementation manner, the path generation request further includes at least one transition point, and the path between the starting point and the ending point includes a first path between the starting point and the transition point and a second path between the transition point and the ending point.

[0208] The first generation module 42 is specifically configured to:

[0209] Based on the path generation request, for the first path, perform path optimization processing on the first path according to the preset flower pollination algorithm to generate a first initial path.

[0210] For the second path, perform path optimization processing on the second path according to the preset flower pollination algorithm to generate a second initial path.

[0211] Concatenate the first initial path and the second initial path to generate an initial path.

[0212] In a possible implementation, the device is further specifically configured to:

[0213] After obtaining path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path, perform smoothing processing on the target path to obtain a smoothed target path.

[0214] The device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0215] Figure 8 It is a schematic structural diagram of the unmanned aerial vehicle provided in this application. As Figure 8 shown, the unmanned aerial vehicle 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the unmanned aerial vehicle 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus 504.

[0216] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0217] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0218] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0219] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0220] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0221] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the above method.

[0222] The present application also provides a computer-readable storage medium storing computer-executable instructions, and when the processor executes the computer-executable instructions, the above method is implemented.

[0223] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0224] An exemplary readable storage medium is coupled to the processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0225] The division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0226] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0227] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit.

[0228] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0229] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0230] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation schemes of the present invention. The present invention aims to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for generating a path for an unmanned aerial vehicle, characterized in that: include: Obtaining a path generation request; wherein the path generation request includes a start point and an end point; Based on the path generation request, performing path optimization processing on the path between the starting point and the end point according to a preset flower pollination algorithm to generate an initial path; According to a preset reinforcement learning algorithm, path sample data is obtained, and path iteration processing is performed on the path sample data and the initial path to generate a target path; wherein the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the UAV between the starting point and the end point.

2. The method according to claim 1, characterized in that The step of optimizing the path between the starting point and the end point based on the path generation request and generating an initial path according to a preset flower pollination algorithm includes: Based on the path generation request, the preset Q table is initialized according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to obtain the target Q value of each initial Q value; wherein the preset Q table includes the initial Q values ​​of multiple state-action pair data; According to the target Q value of each initial Q value, an optimized initial path between the starting point and the end point is generated in combination.

3. The method according to claim 2, characterized in that The generating request based on the path, initializing the preset Q table according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm to obtain the target Q value of each initial Q value, including: Based on the path generation request, the preset Q table is initialized according to the global pollination algorithm and the local pollination algorithm in the preset flower pollination algorithm and the control probability set by the user to obtain the target Q value of each initial Q value.

4. The method according to claim 1, characterized in that: The step of obtaining path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path includes: According to the experience memory strategy in the preset reinforcement learning algorithm, the quadruple data generated when the drone performs actions in the environment is stored in a preset experience playback pool; wherein the quadruple data is the recovered path sample data; the quadruple data includes an immediate reward; For the quadruple data in the experience replay pool, according to the experience reuse strategy in the preset reinforcement learning algorithm, a high priority is set for the quadruple data containing the immediate rewards higher than the preset reward threshold; According to a preset reinforcement learning algorithm, parameter initialization processing is performed on the main parameters of the main network and the target parameters of the target network in the preset dual-depth Q network to obtain the main parameters of the main network and the target parameters of the target network; wherein the main parameters are equal to the target parameters; Repeat the following steps until the preset termination condition is reached: Randomly obtain a preset number of high-priority path sample data from a preset experience replay pool; wherein the path sample data includes a plurality of state-action pair data; Selecting an action in any state according to the main network; and determining a target Q value of the action in any state according to the target network; According to the target Q value, the main parameters of the main network are updated; according to a preset step interval, the target parameters are updated to the updated main parameters; The target Q value obtained when the preset termination condition is reached is used to generate a target path.

5. The method according to claim 1, characterized in that The step of obtaining path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path includes: Based on the experience reuse strategy in the preset reinforcement learning algorithm, if it is detected that the current task corresponding to the current starting point and the end point is similar to the historical flight environment or the historical task goal, the four-tuple data corresponding to the flight environment is called for reuse in the preset experience replay pool, or the four-tuple data corresponding to the task goal is called for reuse to generate a target path; wherein the preset experience replay pool includes the four-tuple data generated and recovered when the drone performs actions in the environment.

6. The method according to claim 1, characterized in that The path generation request further includes at least one transition point, and the path between the starting point and the end point includes a first path between the starting point and the transition point, and a second path between the transition point and the end point; The step of performing path optimization processing on the path between the starting point and the end point based on the path generation request according to a preset flower pollination algorithm to generate an initial path includes: Based on the path generation request, for the first path, performing path optimization processing on the first path according to a preset flower pollination algorithm to generate a first initial path; For the second path, performing path optimization processing on the second path according to a preset flower pollination algorithm to generate a second initial path; The first initial path and the second initial path are concatenated to generate an initial path.

7. The method according to any one of claims 1 to 6, characterized in that: After acquiring path sample data according to a preset reinforcement learning algorithm, performing path iteration processing on the path sample data and the initial path, and generating a target path, the method further includes: The target path is smoothed to obtain a smoothed target path.

8. A path generation device for an unmanned aerial vehicle, characterized in that: include: An acquisition module, used to acquire a path generation request; wherein the path generation request includes a start point and an end point; A first generation module is used to perform path optimization processing on the path between the starting point and the end point according to a preset flower pollination algorithm based on the path generation request to generate an initial path; The second generation module is used to obtain path sample data according to a preset reinforcement learning algorithm, perform path iteration processing on the path sample data and the initial path, and generate a target path; wherein the preset reinforcement learning algorithm includes an experience memory strategy and / or an experience reuse strategy; the target path is the actual flight path of the UAV between the starting point and the end point.

9. A drone, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Cited By

  • Unmanned aerial vehicle cluster collaborative search method and system based on multi-agent reinforcement learning

    CN121742524A