A UAV autonomous landing method based on reinforcement learning
By decomposing the autonomous landing task of the drone into sub-tasks and combining learning methods of deep reinforcement learning and auxiliary positioning tasks, the problem of insufficient generalization capabilities of the drone in unknown environments is solved. The dynamic partitioning experience replay sampling method is used to improve training efficiency and achieve efficient and accurate landing of the drone.
Patent Information
- Application Number
- CN202210345360.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-31
AI Technical Summary
The existing autonomous landing method of drone based on reinforcement learning is insufficient in generalization capabilities in unknown environments, and the sparseness and delay of rewards in deep reinforcement learning lead to low training efficiency.
The autonomous landing task of the drone is decomposed into two subtasks: horizontal alignment and vertical drop, combined with deep reinforcement learning and auxiliary positioning tasks for learning, and a dynamic partitioning experience playback sampling method is used to improve learning efficiency.
The landing performance of the drone has been significantly improved, allowing the model to be better generalized to unknown environments and accurately identified and landed on different ground textures.
Smart Images

Figure CN114859943B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous landing of unmanned aerial vehicles, and in particular to an autonomous landing method of unmanned aerial vehicles based on reinforcement learning. Background Art
[0002] In recent years, drones have been used in more and more fields, especially for boring, dirty or dangerous tasks such as search and rescue, remote sensing, precision agriculture, surveillance, and cargo delivery. In these applications, drone landing is a crucial link.
[0003] To solve the problem of drone landing, previous methods use multiple onboard sensors or ground-based auxiliary equipment to estimate the drone's attitude and then control the drone. But onboard sensors are usually expensive and power-consuming, and ground-based auxiliary equipment is not always available. Since drones are usually equipped with a monocular camera, researchers have proposed various vision-based drone landing methods, but these methods are easily affected by changes in lighting and drone attitude, or only work with specially designed ground landing landmarks. Inspired by the significant breakthroughs of deep reinforcement learning in the fields of game and robot control, researchers began to use deep reinforcement learning to solve the problem of autonomous drone landing, but the learned landing strategies only perform well in known environments and cannot generalize well to new unknown environments.
[0004] In order to achieve better generalization of the landing strategy, the present invention first decomposes the autonomous landing task of the drone into two subtasks: aligning with the ground mark in the horizontal direction and descending in the vertical direction, and then uses a deep reinforcement learning model with a ground mark auxiliary positioning task to complete each subtask. The model combines the landing task based on deep reinforcement learning with the auxiliary positioning task based on supervised learning to improve the generalization ability of the learned landing strategy. Ground mark positioning and landing strategy share image feature representation, so the auxiliary positioning task can help the reinforcement learning agent learn useful image features, thereby locating the ground mark with high precision and improving the generalization ability of the landing strategy in a new unknown environment. The present invention designs classification auxiliary positioning tasks and return auxiliary positioning tasks to improve image feature learning. In order to solve the sparse and delayed problems of rewards in deep reinforcement learning, the present invention proposes a novel dynamic sampling method, called dynamic partition experience replay sampling, which uses different sampling ratios to sample from different experience partitions. This sampling method can improve learning efficiency while maintaining low computational complexity. Summary of the invention
[0005] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0006] In view of the problems existing in the above-mentioned existing UAV autonomous landing method based on reinforcement learning, the present invention is proposed.
[0007] Therefore, the purpose of the present invention is to provide a UAV autonomous landing method based on reinforcement learning. In the deep Q-network training, a dynamic partitioning experience replay method is adopted to stabilize and speed up the training process. The auxiliary positioning task is combined with an improved sampling strategy to make the trained model better generalized to unknown environments, and finally significantly improve the landing performance of the UAV.
[0008] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for autonomous landing of a drone based on reinforcement learning, comprising the following steps: S1: collecting image information of a drone camera to form raw data, and storing the collected raw data and position information in a sample collection; S2: sampling the sample collection, training the sampled data with a deep Q network with an auxiliary positioning task, and predicting the drone action Q1 value; according to the Q1 value, a greedy strategy is adopted to select the drone action D1 so that the drone itself is horizontally aligned with the ground mark; the reward function of the deep Q network training in the S2 stage is:
[0009]
[0010] Where s is the state of the drone and a is the action performed by the drone.
[0011] S3: The sample collection is sampled using the dynamic partition experience replay sampling method, and the deep Q network with auxiliary positioning task is trained on the sampled data to predict the Q2 value of the drone action; the greedy strategy is used to select the drone action D2, so that the drone itself descends in the vertical direction and adjusts its position in the horizontal direction to keep it aligned with the ground mark; the reward function of the deep Q network training in the S3 stage is:
[0012]
[0013] Where s is the state of the drone and a is the action performed by the drone.
[0014] S4: Drone landing.
[0015] As a preferred solution of the reinforcement learning-based UAV autonomous landing method described in the present invention, wherein: the deep Q network training with auxiliary positioning tasks includes two auxiliary positioning tasks: classification auxiliary positioning task or regression auxiliary positioning task.
[0016] As a preferred solution of the UAV autonomous landing method based on reinforcement learning described in the present invention, wherein: for the classification-assisted positioning task, the sampling data in the S2 stage is processed by the convolution layer and output as a 23×23 dimensional classification vector, and the sampling data in the S3 stage is processed by the convolution layer and output as a 7-dimensional classification vector.
[0017] As a preferred solution of the UAV autonomous landing method based on reinforcement learning described in the present invention, in which: in the regression-assisted positioning task, a neural network is used to regress and predict the relative coordinates (Δx, Δy, Δz) of the UAV and the marker; wherein the spatial coordinates of the marker are (x marker ,y marker ,z marker ), the spatial coordinates of the drone are (x UAV ,y UAV ,z UAV ), the 3D relative coordinates of the drone and the marker can be expressed as:
[0018] (Δx, Δy, Δz) = (x UAV -x marker ,y UAV -y marker ,z UAV -z marker ).
[0019] As a preferred solution of the UAV autonomous landing method based on reinforcement learning described in the present invention, wherein:
[0020] As a preferred solution of the UAV autonomous landing method based on reinforcement learning described in the present invention, the UAV action D1 value includes 5 actions: forward, backward, left, right and descending. When the Q1 value is descending, the UAV enters the S3 stage.
[0021] As a preferred solution of the reinforcement learning-based autonomous landing method of the UAV described in the present invention, wherein: the dynamic partitioned experience playback sampling method of S3 divides the sample collection into neutral, negative and positive partitions, and samples each partition by weighted priority sampling; the priority of each experience sample is proportional to the absolute value of its time difference error; the average absolute time difference error of the three partitions is normalized, and the normalized results are used as the sampling ratio of the partition samples in each sampling batch; each time a batch of experience is collected, the network parameters are updated, and then the updated network is used to recalculate the time difference error of the batch of experience, and the priority of the batch of experience is updated.
[0022] As a preferred solution of the UAV autonomous landing method based on reinforcement learning described in the present invention, the UAV action D2 value includes 6 actions: forward, backward, left, right, descending, and landing. When D2 is landing, the UAV lands.
[0023] As a preferred solution of the reinforcement learning-based UAV autonomous landing method described in the present invention, the classification-assisted positioning task uses a cross entropy loss function to process data, and the regression-assisted positioning task uses a mean square error loss function to process data.
[0024] The beneficial effects of the present invention are as follows: the present invention performs sampling through only one monocular camera, performs deep Q network training on the sampling results, adopts dynamic partitioning experience playback in training to stabilize and speed up the training process, combines auxiliary positioning tasks with improved sampling strategies, and enables the trained model to be better generalized to unknown environments. During training, the landing point can be accurately identified in samples such as bricks, grass, sand, soil, asphalt, etc., ultimately significantly improving the landing performance of the UAV. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:
[0026] Figure 1 It is the overall flow chart of the present invention.
[0027] Figure 2 In the ground landmark classification assisted positioning task of the present invention, the classification method of the horizontal alignment stage between the UAV and the landmark (left) and the classification method of the vertical descent stage of the UAV (right).
[0028] Figure 3 The present invention represents the relative positions of the UAV and the marker in the ground marker regression assisted positioning task. DETAILED DESCRIPTION
[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0031] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0032] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0033] Example 1
[0034] Reference Figure 1 , which is the first embodiment of the present invention, provides a method for autonomous landing of a UAV based on reinforcement learning, the method comprising the following steps:
[0035] S1 (drone flight preparation landing state): at this time, the yaw angle of the drone is random, and the horizontal coordinates at a fixed height are also random. The camera used in the present invention is a common monocular camera, which collects image information of the drone camera to form raw data, and stores the collected raw data and the drone's position information in a sample collection;
[0036] S2 (the stage of horizontal alignment between the drone and the ground landmark): sample the sample collection, train the deep Q network with auxiliary positioning tasks on the sampled data, and predict the Q1 value of the drone action; according to the Q1 value, a greedy strategy is used to select the drone action D1 so that the drone itself is horizontal with the ground landmark. The drone action D1 value includes five actions: forward, backward, left, right and down; the image features extracted by the convolutional neural network in the deep Q network are shared by the auxiliary positioning task network and the policy network. The policy network is used to predict the Q value of each action. The auxiliary positioning network is located after the last convolutional layer, and the auxiliary convolutional layer learns feature extraction; there are two types of auxiliary positioning tasks: classification auxiliary positioning task and regression auxiliary positioning task. The classification auxiliary positioning task uses cross entropy loss, and the regression auxiliary positioning task uses mean square error loss. The reward function for model training in this stage is:
[0037]
[0038] Where s is the state of the drone and a is the action performed by the drone.
[0039] S3 (drone vertical descent stage): When the D1 value is decreasing, the drone enters the S3 stage, samples the sample collection using the dynamic partition experience playback sampling method, trains the sampled data with a deep Q network with auxiliary positioning tasks, and predicts the drone action Q2 value; the greedy strategy is used to select the drone action D2, so that the drone itself descends in the vertical direction and adjusts its position in the horizontal direction to keep it aligned with the ground mark. The drone action D2 value includes 6 actions: forward, backward, left, right, descending, and landing; the auxiliary positioning network is located after the last convolution layer, and the auxiliary convolution layer learns feature extraction; there are two types of auxiliary positioning tasks: classification auxiliary positioning task and regression auxiliary positioning task; the classification auxiliary positioning task uses cross entropy loss, and the regression auxiliary positioning task uses mean square error loss. The reward function for model training in this stage is:
[0040]
[0041] Where s is the state of the drone and a is the action performed by the drone.
[0042] S4: When D2 is landing, the drone turns off the rotor motor and lands on the ground mark.
[0043] Example 2
[0044] Reference Figure 2 The classification auxiliary positioning method of the present invention for the horizontal alignment stage and the vertical descent stage of the drone and the mark is shown in FIG. Figure 3 The relative positions of the UAV and the landmark are represented in the ground landmark regression-assisted positioning task. The classification-assisted positioning task uses the cross entropy loss, and the regression-assisted positioning task uses the mean square error loss. Adding ground landmark positioning as an auxiliary task can increase the convergence speed and is conducive to learning better UAV landing strategies. The learned strategies perform better in both known and unknown environments.
[0045] The classification auxiliary task divides the drone's field of view into multiple areas, each area represents a category, and is classified according to the area where the center of the ground mark is located. In the stage of horizontal alignment between the drone and the ground mark, since the height of the drone remains unchanged, the size of its field of view and the size of the mark in the field of view are fixed, and the field of view width is 23 times the width of the mark. The image captured by the camera can be divided into areas. After the image is processed by the convolution layer, it is input into the auxiliary positioning network, and a 7-dimensional classification vector is output. In the vertical landing stage, since the height of the drone is constantly changing, the size of the mark in the drone's field of view is also constantly changing. The lower the height, the larger the mark in the field of view. According to the characteristics of the change of the mark in the field of view and combined with the output control instructions, the present invention divides the drone's field of view into 7 areas. The image is divided into 7 categories according to the area where the center of the mark is located. This classification auxiliary task can help the convolution layer learn more accurate mark features. At the same time, different classification categories also contain the corresponding spatial position information of the ground mark, so that the extracted features are more conducive to the learning of landing strategies. After the image is processed by the convolution layer, it is input into the auxiliary positioning network, and a 7-dimensional classification vector is output.
[0046] The regression auxiliary task uses a neural network to regress and predict the relative coordinates of the drone and the marker. Assume that the spatial coordinates of the marker are (x marker ,y marker ,z marker ), the spatial coordinates of the drone are (x UAV ,y UAV ,z UAV ), the relative coordinates of the drone and the marker can be expressed as:
[0047] (Δx, Δy, Δz) = (x UAV -x marker ,y UAV -y marker ,z UAV -z marker );
[0048] The ability to correctly predict the representative network has good feature extraction capabilities, so the present invention designs a regression auxiliary task network, and the output is a 2D relative coordinate. The image is processed by the convolution layer and input into the regression network, and first processed by the spatial softmax layer. The feature maps of the k channels in the features extracted by the convolution layer are processed by the softmax function respectively, and then the processed feature maps are used to predict the 2D relative coordinates, and a total of k groups of relative coordinates are output, that is, a 2k-dimensional coordinate vector; finally, linear regression is performed on all coordinates to predict the 2D relative coordinates (Δx, Δy).
[0049] Example 3
[0050] Reference Figure 1, which is the third embodiment of the present invention. This embodiment is different from the second embodiment in that: in training the deep Q network and controlling the drone to land in the vertical direction, the present invention uses a dynamic partition experience playback sampling method; the sample set is divided into neutral, negative and positive partitions, and each partition is sampled by weighted priority sampling; the priority of each experience sample is proportional to the absolute value of its time difference error. The average absolute time difference error of the three partitions is normalized, and the normalized results are used as the sampling ratio of the partition samples in each sampling batch; each time a batch of experience is collected, the network parameters are updated, and then the updated network is used to recalculate the time difference error of the batch of experience, and the priority of the batch of experience is updated.
[0051] The five models were further tested on six different types of background textures, including brick, grass, pavement, sand, snow, and soil. Each type of texture contains 13 instances, and we tested each agent 100 times on each background sample, for a total of 7,800 tests. We make the following two observations. First, adding marker localization as an auxiliary task significantly improves the learning ability of the general learning strategy. Second, formulating the localization of the marker as a regression task further improves the generalization of the invention. In detail, the average successful regression agent has a rate of more than 0.85 for all background textures. We provide a stronger supervision signal for the location of the marker through the auxiliary localization task, thereby better inferring the background texture of the ground landmarks and settings, which means that the model can learn to mark features more quickly and accurately by adding auxiliary tasks.
[0052] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A UAV autonomous landing method based on reinforcement learning, It is characterized in that The following steps are involved: S1: Collect image information from the drone camera to form raw data, and store the collected raw data and drone location information into a sample collection; S2: Sample the sample collection, train the deep Q network with auxiliary positioning task on the sampled data, and predict the Q1 value of the drone action; according to the Q1 value, adopt the greedy strategy to select the drone action D1 so that the drone itself is horizontally aligned with the ground mark; The reward function for deep Q network training in stage S2 is: Among them, s is the state of the drone, and a is the action performed by the drone; S3: The sample collection is sampled using the dynamic partition experience playback sampling method, and the deep Q network with auxiliary positioning task is trained on the sampled data to predict the Q2 value of the drone action; according to the Q2 value, the greedy strategy is used to select the drone action D2, so that the drone itself descends in the vertical direction and adjusts its position in the horizontal direction to keep it aligned with the ground mark; The reward function for deep Q network training in stage S3 is: Where s is the state of the drone, and a is the action performed by the drone; S4: Drone landing.
2. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 1, Features: The deep Q network training with auxiliary positioning task includes two kinds of auxiliary positioning tasks: classification auxiliary positioning task or regression auxiliary positioning task.
3. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 2, Features: For the classification-assisted positioning task, the sampled data is processed by the convolution layer in the S2 stage and output as a 23×23 dimensional classification vector, and the sampled data is processed by the convolution layer in the S3 stage and output as a 7 dimensional classification vector.
4. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 2, Features: In the regression-assisted positioning task, a neural network is used to regress and predict the relative coordinates (Δx, Δy, Δz) of the drone and the landmark; The spatial coordinates of the marker are (x marker ,y marker ,z marker ), the spatial coordinates of the drone are (x UAV ,y UAV ,z UAV ), the 3D relative coordinates of the drone and the marker can be expressed as: (Δx,Δy,Δz)=(x UAV -x marker ,y UAV -y marker ,z UAV -z marker )。 5. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 4, Features: The drone action D1 value includes 5 actions.
6. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 5, Features: The dynamic partitioned experience playback sampling method of S3 divides the sample collection into neutral, negative and positive partitions, and samples each partition by weighted priority sampling; The priority of each experience sample is proportional to the absolute value of its time difference error; the average absolute time difference error of the three partitions is normalized, and the normalized results are used as the sampling ratio of the partition samples in each sampling batch; each time a batch of experience is collected, the network parameters are updated, and then the updated network is used to recalculate the time difference error of the batch of experience, and update the priority of the batch of experience.
7. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 6, Features: The drone action D2 value includes 6 actions.
8. The autonomous landing method for a drone based on reinforcement learning as claimed in claim 7, Features: The classification-assisted positioning task uses a cross entropy loss function to process data, and the regression-assisted positioning task uses a mean square error loss function to process data.
Citation Information
Patent Citations
Unmanned aerial vehicle autonomous landing method based on deep synergetic neural network
CN107273929A
Reinforcement learning small-sized unmanned rotorcraft autonomous landing method based on data fusion increase
CN110231829A