UAV Swarm Obstacle Avoidance Method Based on Behavior Cloning and Improved DQN Algorithm
By introducing behavioral cloning networks and distance rights into the DQN algorithm, the problem of slow convergence and easy to fall into local optimality in complex environments is solved, and the obstacle avoidance effect with fast convergence and high success rate is achieved.
Patent Information
- Application Number
- CN202211105893.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-09
AI Technical Summary
Traditional UAV cluster obstacle avoidance algorithms are prone to fall into local optimal solutions in complex environments, with slow convergence speed and low task success rate.
The behavioral cloning network is introduced as a subnet of the DQN algorithm, and uses the coordinated update of the distance right and behavioral cloning network to assist the UAV cluster to make obstacle avoidance decisions.
The algorithm convergence speed has been accelerated, the mission success rate has been improved, and the obstacle avoidance ability of the drone cluster in complex environments has been improved.
Smart Images

Figure CN116360479B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of drone technology, and in particular to a drone cluster obstacle avoidance method based on behavior cloning and an improved DQN algorithm. Background Art
[0002] With the rapid development of intelligent unmanned technology, it has become possible to use large-scale drones to form clusters and perform various tasks such as reconnaissance, detection, and strike. As drone clusters have many advantages in terms of mission efficiency, low energy consumption, robustness, fault tolerance, and scalability, they have become the first choice for performing dangerous tasks in the future, showing a trend of gradually replacing manned aircraft. Replacing manned clusters to perform specific tasks has become an inevitable trend. In unmanned systems, drone clusters have many advantages in terms of mission efficiency, low energy consumption, robustness, fault tolerance, and scalability, and therefore have become the preferred representative for performing various tasks.
[0003] The autonomous mission execution of drone swarms is inseparable from the support of technologies such as formation control, task allocation, path planning, collision avoidance and obstacle avoidance. Among them, drone swarm obstacle avoidance planning is to plan the path from the starting point to the target point under certain constraints, so that the specified performance indicators are optimal. Constraints mainly refer to environmental constraints, task constraints, spatial coordination constraints, temporal coordination constraints, and drone constraints. Performance indicators can include path length, path smoothness, path safety, task completion time, etc.
[0004] Traditional path planning algorithm
[0005] A* algorithm. By introducing the heuristic search idea to improve the Dijkstra algorithm, it can find the shortest path faster. The principle is simple and easy to implement.
[0006] Artificial Potential Field (APF) introduces the concept of "potential field" in physics into the UAV cluster mission scenario. The core idea is that obstacles in the scene produce repulsive force on the UAV, and the target point produces gravitational force on the UAV. The UAV moves under the combined force. The algorithm has good real-time performance and smooth planned path, which is suitable for local path planning.
[0007] Sampling method: There is no need to model the entire environment space, and the environment is reconstructed with sampling points, which requires relatively less calculation.
[0008] Intelligent optimization algorithm
[0009] Genetic algorithm. It is an adaptive method based on biological genetic evolution process, which can be used to solve search and optimization problems. Its advantage is that it is not limited by the problem domain and has the ability of fast heuristic search.
[0010] Ant colony algorithm. It is a random search algorithm, and its core idea is to utilize the pheromone of the ant colony to seek the optimal solution of the problem through positive feedback.
[0011] Particle swarm algorithm. It originated from the study of the foraging behavior of bird flocks. The core idea is that each particle in the group shares the extreme value it finds, obtains the optimal value of the entire particle swarm, and then adjusts each particle to finally find the global optimal solution.
[0012] In traditional path planning algorithms, the A* algorithm has a too small search area and too large path turning angles, resulting in an uneven planned path; when the artificial potential field method is applied in UAV swarm tasks, due to the complex scene elements and many points where the resultant force is zero, it is easy to fall into local optima; the sampling algorithm has a large randomness and slow convergence.
[0013] In intelligent optimization algorithms, the disadvantages of genetic algorithms are that they are prone to premature convergence and easy to fall into local optimal solutions; ant colony algorithms have slow self-convergence and are easy to fall into local optimal solutions; while particle swarm algorithms are also easy to fall into local optimal solutions in relatively complex environments. Summary of the Invention
[0014] The embodiments of the present application provide a UAV swarm obstacle avoidance method based on behavior cloning and improved DQN algorithm. By introducing distance weights into the traditional reinforcement learning DQN algorithm and using the neural network of behavior cloning for auxiliary decision-making, the algorithm convergence speed can be greatly accelerated and the task success rate can be improved.
[0015] The embodiments of the present application provide a UAV swarm obstacle avoidance method based on behavior cloning and improved DQN algorithm, including the following steps:
[0016] Pre-configure the UAV obstacle avoidance behavior based on obstacles;
[0017] Train a behavior cloning network based on the configured UAV obstacle avoidance behavior to use the behavior cloning network for behavior cloning guidance;
[0018] Take the behavior cloning network as a sub-network of the DQN network to co-update the DQN network using the updated parameters of the behavior cloning network during the training process;
[0019] Use the trained DQN network for UAV swarm obstacle avoidance.
[0020] Optionally, pre-configuring the UAV obstacle avoidance behavior based on obstacles includes:
[0021] Obtain several frames of image data and convert each frame of image data into a depth scene map based on distance, where the depth scene map is a grayscale image of 0-1, and the closer the distance between the obstacle and the UAV, the darker the color and the closer the grayscale value is to 1, and the grayscale value of non-obstacles is 0;
[0022] Divide the deep scene image into multiple sub-images in a specified direction;
[0023] Configure an obstacle threshold, and mark the pixel points exceeding the obstacle threshold as obstacle pixels to count the proportion of obstacle pixels in each sub-image;
[0024] When the proportion of obstacle pixels in the sub-image exceeds a preset proportion threshold, determine that the sub-image is an obstacle area;
[0025] Execute the action selection of the drone according to the position of the obstacle area in the deep scene image.
[0026] Optionally, dividing into multiple sub-images in the specified direction is performed in the vertical direction;
[0027] Executing the action selection of the drone according to the position of the obstacle area in the deep scene image includes:
[0028] If the middle position of the deep scene image is an obstacle area, deflect from the non-obstacle areas in the left and right sub-images, and preferably deflect to the area with a lower proportion of obstacle pixels.
[0029] Optionally, training the behavior cloning network based on the configured drone obstacle avoidance behavior includes: adding training labels to each frame of image data based on the action selection result to train the behavior cloning network.
[0030] Optionally, taking the behavior cloning network as a sub-network of the DQN network to cooperatively update the DQN network using the updated parameters of the behavior cloning network during training includes:
[0031] Introduce the target convolutional network with distance weights of other drones for the current drone to cooperatively update the DQN network in the following way:
[0032]
[0033] Among them, α, β, and γ all represent attenuation factors, and ε represents the weight coefficient adjustment factor, respectively represent the Q networks of drones i and j and the behavior cloning network of drone i, d i.j represents the distance between drones i and j, N i represents the set of all drones within the communication range of drone i.
[0034] Optionally, taking the behavior cloning network as a sub-network of the DQN network to cooperatively update the DQN network using the updated parameters of the behavior cloning network during training further includes:
[0035] Set the drone reward function of the DQN network to:
[0036] Single-step loss reward R1, obstacle collision reward R2, reaching target point reward R3, target point approaching reward Obstacle approaching reward where d tar represents the distance between the current UAV and the target, and d edg represents the distance between the current UAV and the obstacle.
[0037] The embodiment of the present application also proposes a UAV controller, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the aforementioned UAV cluster obstacle avoidance method are implemented.
[0038] The embodiment of the present application also proposes a UAV, including the aforementioned UAV controller.
[0039] The UAV cluster control method of the embodiment of the present application introduces a distance weight into the traditional reinforcement learning DQN algorithm and uses a neural network of behavior cloning for auxiliary decision-making, which can greatly accelerate the algorithm convergence speed and improve the task success rate.
[0040] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0042] Figure 1 is the basic flowchart of the UAV cluster obstacle avoidance method of the embodiment of the present application;
[0043] Figure 2 is the verification use case of the UAV cluster obstacle avoidance method of the embodiment of the present application;
[0044] Figure 3 is another verification use case of the UAV cluster obstacle avoidance method of the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0046] An embodiment of the present application provides a method for obstacle avoidance of an unmanned aerial vehicle (UAV) cluster based on behavior cloning and an improved DQN algorithm, as Figure 1 shown, including the following steps:
[0047] In step S1011, the UAV obstacle avoidance behavior is pre-configured based on obstacles.
[0048] In some embodiments, pre-configuring the UAV obstacle avoidance behavior based on obstacles includes:
[0049] Obtain several frames of image data, and convert each frame of image data into a depth panorama based on distance, where the depth panorama is a grayscale image of 0-1, and the closer the distance between the obstacle and the UAV, the darker the color and the closer the grayscale value is to 1, and the grayscale value of non-obstacles is 0. In specific applications, a scene can be established in a simulation platform. For example, taking a quadrotor UAV flying in the AIRSIM simulation platform as an example, a simulation scene is established. In the simulation scene, the UAV's built-in camera is used to collect image data, and through the commands in AIRSIM, the color image is converted into a depth panorama based on distance. Since the UAV is constantly flying in the simulation environment, the lower half of the depth panorama is always an infinitely extending ground plane with a grayscale value greater than 0, and only the upper half of the depth panorama can be intercepted for subsequent processing.
[0050] The depth panorama is divided into multiple sub-images in a specified direction. In some embodiments, dividing into multiple sub-images in a specified direction is performed in the vertical direction. For example, the depth panorama can be evenly divided into 3 sub-images horizontally for obstacle avoidance determination.
[0051] Configure an obstacle threshold, and the pixel points exceeding the obstacle threshold are recorded as obstacle pixels. In the case where the proportion of obstacle pixels in the sub-image exceeds a preset proportion threshold, it is determined that the sub-image is an obstacle area; to count the proportion of obstacle pixels in each sub-image, a grayscale threshold can be set to 0.667, and the pixels exceeding this threshold are recorded as obstacle pixels. The proportion of obstacle pixels in the three sub-images is counted respectively. For example, if it exceeds 75% of this image, it is considered that the corresponding field of view area of the sub-image is a "dangerous area", otherwise it is determined as a "safe area".
[0052] According to the position of the obstacle area in the depth panorama, perform action selection of the UAV.
[0053] In this embodiment, taking two-dimensional obstacle avoidance as an example (i.e., horizontal plane obstacle avoidance under the condition of fixed altitude), the states of the UAV are defined as three states (safe, dangerous, collision) (S safe , S danger , S collide ), and the action space of the UAV is defined as three actions (forward, left shift, right shift) (A for , A le , A ri ). In some embodiments, splitting into multiple sub-images in a specified direction is performed in the vertical direction;
[0054] According to the position of the obstacle area in the depth map, the action selection of the UAV includes:
[0055] If the middle position of the depth map is the obstacle area, select a non-obstacle area from the sub-images on the left and right for deflection, and preferentially deflect to the area with a lower proportion of obstacle pixels.
[0056] For example, if the middle sub-image among the three sub-images is a "safe area", the UAV selects the action "forward"; if the middle sub-image is a "dangerous area", then judge the attributes of the two sub-images on its left and right, and deflect to the direction of the "safe area". If both the left and right sub-images are "dangerous areas" or "safe areas", then preferentially deflect to the direction with a lower proportion of obstacle pixels, thus completing the behavior cloning guidance method for selecting the next action according to the image.
[0057] In step S1012, train a behavior cloning network based on the configured UAV obstacle avoidance behavior to use the behavior cloning network for behavior cloning guidance. In some embodiments, training a behavior cloning network based on the configured UAV obstacle avoidance behavior includes: adding training labels to each frame of image data based on the action selection result to train the behavior cloning network. In some specific applications, the network structure of the behavior cloning network may include:
[0058] Input: Grayscale image of 640*540;
[0059] First convolutional layer: kernel_size = 32*32, activation function is relu;
[0060] First max pooling layer: pool_size = 32, sliding step is 8;
[0061] Second convolutional layer: kernel_size = 16*16, activation function is relu;
[0062] Second max pooling layer: pool_size = 8, sliding step is 4;
[0063] Third Convolutional Layer: kernel_size = 2*2, activation function is relu;
[0064] First Fully Connected Layer: length = 88;
[0065] Second Fully Connected Layer: length = 16;
[0066] Third Fully Connected Layer: length = 3;
[0067] The output of the behavior cloning network is a three-dimensional vector for action selection.
[0068] Based on the trained behavior cloning network, it can guide the action selection for the subsequent flight of the UAV, forming a complete UAV swarm obstacle avoidance navigation method based on behavior cloning.
[0069] In step S1013, the behavior cloning network is used as a sub-network of the DQN network to collaboratively update the DQN network by using the updated parameters of the behavior cloning network during the training process.
[0070] In step S102, the trained DQN network is used for UAV swarm obstacle avoidance.
[0071] The UAV swarm control method of the embodiment of the present application introduces distance weights into the traditional reinforcement learning DQN algorithm and uses the neural network of behavior cloning for auxiliary decision-making, which can greatly accelerate the algorithm convergence speed and improve the task success rate.
[0072] In some embodiments, using the behavior cloning network as a sub-network of the DQN network to collaboratively update the DQN network by using the updated parameters of the behavior cloning network during the training process includes:
[0073] Introduce the target convolutional network with distance weights of other UAVs for the current UAV to collaboratively update the DQN network in the following way:
[0074]
[0075] Among them, α, β, γ all represent attenuation factors, respectively represent the Q networks of UAVs i and j and the behavior cloning network of UAV i, i represents the current UAV, j represents a UAV in the set N of other UAVs within the communication range of the current UAV i i one of the UAVs, d i.j represents the distance between UAVs i and j, N iDenote the set of all UAVs within the communication range of UAV \(i\). \(\varepsilon\) represents the weight coefficient adjustment factor, which is used to adjust the proportion of the Q-network value of the current UAV \(i\) to the weighted sum of the Q-network values of other UAVs within the communication range of UAV \(i\) in the DQN algorithm based on distance weights. In this embodiment, \(\varepsilon\) can determine a corresponding optimal proportion according to different task scenarios. Similarly, \(\gamma\) is used to adjust the proportion of the Q-network value obtained by the behavior cloning method to the Q-network value obtained by the entire distance weight method. Through such a design, when updating the Q-network target value of the current UAV \(i\), the Q-network values of all other nearby UAVs are also referred to. The weight of the Q-network values of other UAVs is inversely proportional to their distance from this UAV \(i\).
[0076] In this example, the Q-network structure of the DQN network includes:
[0077] Input: UAV 7-dimensional parameter vector (3-dimensional actions, 3-dimensional velocities in each action direction, UAV state);
[0078] The first fully connected layer: length = 10;
[0079] The second fully connected layer: length = 6;
[0080] The third fully connected layer: length = 3;
[0081] Output: 3D vector for action selection
[0082] In some embodiments, taking the behavior cloning network as a sub-network of the DQN network to collaboratively update the DQN network using the updated parameters of the behavior cloning network during training further includes:
[0083] Set the UAV reward function of the DQN network as:
[0084] Single-step loss reward \(R_1\), obstacle collision reward \(R_2\), reaching the target point reward \(R_3\), target point approaching reward Obstacle approaching reward where \(d\) tar represents the distance between the current UAV and the target, and \(d\) edg represents the distance between the current UAV and the obstacle.
[0085] The applicant set up three quadrotors in the AIRSIM Block environment, fixed their height at 10 meters, and the combined horizontal velocity at 5 meters per second in the stable state, and verified the method of this application. The specific results are as Figure 2 、 Figure 3It shows the influence of the number of samples on the algorithm accuracy. There are a total of 500 experiments. Among them, Tra DQN represents the traditional DQN algorithm, that is, the baseline, while MADQN&BC represents the multi-agent distance-weighted DQN&behavior cloning hybrid algorithm of this application, and MADQN represents the multi-agent distance-weighted DQN algorithm that only uses the second part of this application.
[0086] This application collects image information through a camera, uses expert experience for behavior cloning, and at the same time introduces the distance weights of other drones during the Q-network iteration in the DQN algorithm to reduce the number of iterations and play a role in accelerating the algorithm convergence speed. By introducing the distance-weighted Q-network of other drones within the communication range of the drone during the update process of the Q-network target value in the DQN algorithm, this application can reduce the number of iterations of the Q-network, thereby accelerating the algorithm convergence speed and improving the task success rate. The method of this application can predict the optimal actions for drone obstacle avoidance and provide support for the rapid convergence of subsequent obstacle avoidance algorithms.
[0087] The embodiment of this application also proposes a drone controller, which includes a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, it realizes the steps of the aforementioned drone swarm obstacle avoidance method.
[0088] The embodiment of this application also proposes a drone, which includes the aforementioned drone controller.
[0089] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0090] The serial numbers of the embodiments of the above application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server or network device, etc.) to execute the methods described in various embodiments of the present application.
[0092] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.
Claims
1. A method for obstacle avoidance of an unmanned aerial vehicle (UAV) swarm based on behavior cloning and improved Deep Q-Network (DQN) algorithm, characterized in that, It includes the following steps: Pre-configure the obstacle avoidance behavior of the drone based on obstacles; Train a behavior cloning network based on the configured obstacle avoidance behavior of the drone to use the behavior cloning network for behavior cloning guidance; Use the behavior cloning network as a sub-network of the DQN network to collaboratively update the DQN network using the updated parameters of the behavior cloning network during the training process; Use the trained DQN network for obstacle avoidance of the drone swarm; Using the behavior cloning network as a sub-network of the DQN network to collaboratively update the DQN network using the updated parameters of the behavior cloning network during the training process includes: Introduce a target convolutional network with distance weights of other drones for the current drone to collaboratively update the DQN network in the following way: Among them, both represent the attenuation factor, ε represents the weight coefficient adjustment factor, respectively represent the Q-network of the drone i and j the behavior cloning network of the current drone i , represents the distance between the drones i and j , represents the set of all other drones within the communication range of the current drone i .
2. The method for obstacle avoidance of an unmanned aerial vehicle cluster according to claim 1, wherein Pre-configuring the obstacle avoidance behavior of the drone based on obstacles includes: Obtain several frames of image data and convert each frame of image data into a depth panorama based on distance, where the depth panorama is a grayscale image of 0 - 1, and the closer the distance between the obstacle and the drone, the darker the color and the closer the grayscale value is to 1, and the grayscale value of non-obstacles is 0; Divide the depth panorama into multiple sub-images in a specified direction; Configure an obstacle threshold, and mark the pixel points exceeding the obstacle threshold as obstacle pixels to count the proportion of obstacle pixels in each sub-image; When the proportion of obstacle pixels in the sub-image exceeds a preset proportion threshold, recognize the sub-image as an obstacle area; Execute the action selection of the drone according to the position of the obstacle area in the depth panorama.
3. The method for obstacle avoidance of an unmanned aerial vehicle cluster according to claim 2, wherein, Dividing into multiple sub-images in a specified direction is performed in the vertical direction; Executing the action selection of the drone according to the position of the obstacle area in the depth panorama includes: If the middle position of the depth panorama is an obstacle area, deflect from the non-obstacle areas in the left and right sub-images, and preferably deflect to the area with a lower proportion of obstacle pixels.
4. The method for obstacle avoidance of an unmanned aerial vehicle cluster according to claim 3, characterized in that, Training the behavior cloning network based on the configured obstacle avoidance behavior of the drone includes: adding training labels to each frame of image data based on the action selection result to train the behavior cloning network.
5. The method for obstacle avoidance of an unmanned aerial vehicle cluster according to claim 1, characterized in that, Using the behavior cloning network as a sub-network of the DQN network to collaboratively update the DQN network using the updated parameters of the behavior cloning network during the training process further includes: Set the drone reward function of the DQN network to: Single-step loss reward R1, obstacle collision reward R2, reaching the target point reward R3, target point proximity reward R4 = , obstacle proximity reward R5 = , where represents the distance between the current drone and the target, represents the distance between the current drone and the obstacle.
6. A drone controller, characterized in that, It includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, it implements the steps of the drone swarm obstacle avoidance method described in any one of claims 1 to 5.
7. A drone, characterized in that, It includes a drone controller as described in claim 6.
Citation Information
Patent Citations
Unmanned aerial vehicle obstacle avoidance and path planning device and method
CN112819253A
Advantage estimation method and device, electronic equipment and storage medium
CN113240118A