Quadrotor unmanned aerial vehicle autonomous visual navigation obstacle avoidance method based on deep reinforcement learning
By adopting the deep reinforcement learning method of privileged learning and multi-agent exploration strategies in the autonomous visual navigation obstacle avoidance of drones, the problems of low experience collection efficiency and slow model convergence speed in the existing technology are solved, and more efficient and accurate autonomous navigation obstacle avoidance of drones are achieved.
Patent Information
- Application Number
- CN202510101836.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing drone autonomous visual navigation obstacle avoidance algorithm based on deep reinforcement learning requires a large amount of trial and error to collect experience, resulting in inefficient experience collection, which makes the model convergence slower, and in the case of considerable environmental parts, it is difficult to achieve accurate autonomous navigation obstacle avoidance.
An end-to-end quadrotor drone autonomous visual navigation obstacle avoidance method based on deep reinforcement learning is proposed. It adopts privileged learning and multi-agent exploration strategy to directly output flight control instructions through depth images and other self-perception information, process perceived noise and accelerate model convergence.
Effectively handling perceived noise significantly improves the overall performance of the algorithm, improves the accuracy and efficiency of autonomous visual navigation obstacle avoidance of drones, and enhances the migration ability of the model.
Smart Images

Figure CN119937590A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous navigation of unmanned aerial vehicles, and in particular to an autonomous visual navigation and obstacle avoidance method for a quad-rotor unmanned aerial vehicle based on deep reinforcement learning. Background Art
[0002] The autonomous visual navigation and obstacle avoidance technology of quad-rotor drones is a key link in performing complex tasks in unknown environments. It is widely used in urban logistics, pesticide spraying and crop growth monitoring, field post-disaster rescue, battlefield reconnaissance and other fields. According to the process of task execution, autonomous navigation and obstacle avoidance technology can be divided into two types: multi-task cascade and end-to-end. Among them, the multi-task cascade method usually includes task links such as perception, mapping, planning, decision-making and control. Each link is carried out in sequence, and the previous results are used as inputs for subsequent links, which has high interpretability. As a traditional drone navigation algorithm, this method has been studied in depth and extensively. For example, patent CN114355980B proposes a drone autonomous visual navigation and obstacle avoidance method based on ESDF map and DDPG deep reinforcement learning algorithm. The method first obtains the drone posture through sensors and constructs ESDF map in combination with image information. Subsequently, a path search algorithm is used in the map to generate a series of discrete path points, and then the drone posture vector, ESDF map matrix and discrete path point vector are input into the DDPG network in parallel, and the command to control the motor speed is output to achieve autonomous navigation and obstacle avoidance.
[0003] However, in the multi-task cascade method, the noise and error in the previous task will be gradually transferred and accumulated to the subsequent tasks, eventually resulting in a significant decrease in flight performance. In addition, the mapping task requires real-time generation and storage of maps, which consumes a lot of computing and storage resources, making it difficult for drones to efficiently deal with uncertain factors in the environment such as dynamic obstacles and lighting changes. The end-to-end deep learning method avoids the intermediate multiple processing process by directly mapping the perception data to the control command output, effectively reducing the error accumulation problem. In addition, the deep learning method has a stronger advantage in processing multimodal perception data. Through the multi-branch feature extraction network, the system can simultaneously process images from the camera, IMU status data, GPS positioning information, and real-time data from other sensors such as ultrasound. However, the existing end-to-end methods usually do not fully consider some considerable problems in the actual flight environment, such as incomplete or disturbed perception information. In addition, algorithms based on deep reinforcement learning often require a lot of trial and error to collect experience extensively, resulting in low efficiency in experience collection, which in turn makes the model converge slowly. Therefore, there is an urgent need to study an end-to-end autonomous navigation algorithm that can effectively cope with the challenges brought by the observable part of the environment in UAV navigation tasks and achieve safe and efficient autonomous navigation and obstacle avoidance in unknown environments. Summary of the invention
[0004] The purpose of the present invention is to solve the problem that existing algorithms based on deep reinforcement learning often require a large amount of trial and error to collect experience extensively, resulting in low efficiency in experience collection, which in turn makes the model convergence speed slow and the accuracy of autonomous visual navigation and obstacle avoidance of drones is low. A method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning is proposed.
[0005] A method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning. The specific process is as follows:
[0006] Step 1: Build a simulation environment for autonomous visual navigation and obstacle avoidance of a quadrotor drone, select and initialize the dynamic model of the quadrotor drone, set the starting position of the drone and set the target position;
[0007] Step 2: Set up the state space, observation space, and action space in the simulation environment;
[0008] Step 3: Construct a neural network model for the autonomous visual navigation obstacle avoidance decision-making agent;
[0009] Step 4: training the neural network model of the autonomous visual navigation obstacle avoidance decision-making agent to obtain a trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent;
[0010] Step 5: Based on the trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent, the quadrotor drone to be controlled is subject to obstacle avoidance control.
[0011] The beneficial effects of the present invention are:
[0012] (1) This invention is the first autonomous navigation obstacle avoidance decision maker in the field of autonomous navigation of unmanned aerial vehicles that solves the observable problems of the environment based on privileged learning. The decision maker takes depth images and other self-perception information as input and directly outputs the flight control instructions of the unmanned aerial vehicle, thereby completing the autonomous navigation obstacle avoidance task, which reflects its advanced nature.
[0013] (2) The privileged information-based learning framework and multi-agent experience collection strategy adopted by the present invention can effectively process perceptual noise and accelerate the convergence of deep reinforcement learning models, significantly improving the overall performance of the algorithm.
[0014] (3) Compared with non-end-to-end autonomous navigation and obstacle avoidance algorithms, the method of the present invention has obvious advantages in algorithm efficiency, environmental adaptability and tracking success rate, and improves the accuracy of autonomous visual navigation and obstacle avoidance of UAVs.
[0015] (4) In addition, the autonomous navigation and obstacle avoidance algorithm proposed in the present invention is highly consistent with the perception information that the UAV can obtain in actual flight and the input of the underlying controller in terms of the design of the state and action space, thereby enhancing the UAV's ability to migrate from a simulated environment to a real environment and having good mobility. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0017] Figure 2 Schematic diagram of the UAV autonomous visual navigation and obstacle avoidance simulation environment. (a) A top view of the training environment. (b) A top view of the random environment. (c) A schematic diagram of the UAV flying in the environment.
[0018] Figure 3 The effect of adding noise to the depth image perception information, (a): the depth image obtained by the depth camera; (b): the effect of adding salt and pepper noise to the depth image; (c): the effect of adding Gaussian noise to the image in (b); (d): the effect of adding motion blur to the image in (c);
[0019] Figure 4 It is the neural network model DRPL of the autonomous visual navigation obstacle avoidance decision-making agent based on deep reinforcement learning of the present invention. In 1@80×100, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 8@80×100, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 8@40×50, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 16@40×50, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 16@20×25, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 25@20×25, @ is the number of channels in front, and @ is the dimension of the channels in the back; in 25@10×12, @ is the number of channels in front, and @ is the dimension of the channels in the back;
[0020] Figure 5 It is a schematic diagram of updating parameters of the neural network model DRPL of the present invention based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making agent under POMDP modeling;
[0021] Figure 6 The model training curve diagram of the performance comparison experiment of the autonomous visual navigation and obstacle avoidance algorithm is shown in the figure. (a) SR curve of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning and the common TD3 algorithm during the model training process; (b) AER curve of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning and the common TD3 algorithm during the model training process;
[0022] Figure 7The flight trajectory diagrams of the performance comparison experiment of the autonomous visual navigation and obstacle avoidance algorithm are as follows: (a): the flight trajectory of the UAV in the training environment when the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning of the present invention is deployed; (b): the flight trajectory of the UAV in the training environment when the ordinary TD3 algorithm is deployed; (c): the flight trajectory of the UAV in the training environment when the EGO-Planner-v2 is deployed; (d): the flight trajectory of the UAV in the random environment when the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning of the present invention is deployed; (e): the flight trajectory of the UAV in the random environment when the ordinary TD3 algorithm is deployed; (f): the flight trajectory of the UAV in the random environment when the EGO-Planner-v2 is deployed;
[0023] Figure 8 The model training and evaluation curve diagrams of the ablation experiment of the autonomous visual navigation and obstacle avoidance algorithm for privileged information and multi-agent exploration strategy, (a): the SR curves of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning proposed in the present invention and the algorithm of removing privileged information and multi-agent exploration strategy respectively during the model training process; (b): the AER curves of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning proposed in the present invention and the algorithm of removing privileged information and multi-agent exploration strategy respectively during the model training process, SR represents the success rate, AER represents the average round reward;
[0024] Fig. 9 The model training and evaluation curve diagrams of the ablation experiment of the autonomous visual navigation and obstacle avoidance algorithm for state space and action space, (a): SR curves of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning using 4-dimensional action space proposed in the present invention and the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning using 3-dimensional action space proposed in the present invention during the training process; (b): AER curves of the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning using 4-dimensional action space proposed in the present invention and the neural network model DRPL of the autonomous visual navigation and obstacle avoidance decision-making agent based on deep reinforcement learning using 3-dimensional action space proposed in the present invention during the training process. DETAILED DESCRIPTION
[0025] Specific implementation method 1: This implementation method is a method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning. The specific process is as follows:
[0026] In order to solve the above problems, the present invention proposes an end-to-end quadrotor autonomous visual navigation and obstacle avoidance method based on deep reinforcement learning, which is named DPRL (Distributed Privileged Reinforcement Learning). This method designs an agent network structure, a continuous state and action space, and an efficient reward function suitable for the autonomous navigation and obstacle avoidance tasks of quadrotors through a policy-based deep reinforcement learning algorithm. At the same time, based on the privileged learning method, actor networks and critic networks with different inputs are designed to effectively deal with some considerable problems during the flight process. In addition, a multi-agent exploration strategy is adopted to accelerate model convergence, and finally a strategy that can directly generate flight control instructions based on noisy visual perception data and its own state information is obtained.
[0027] The present invention proposes a method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning, comprising the following steps:
[0028] Step 1: Build a simulation environment for autonomous visual navigation and obstacle avoidance of a quadrotor drone, select and initialize the dynamic model of the quadrotor drone, set the starting position of the drone and set the target position;
[0029] Step 2: Set up the state space, observation space, and action space in the simulation environment;
[0030] Step 3: Construct a neural network model for the autonomous visual navigation obstacle avoidance decision-making agent;
[0031] Step 4: training the neural network model of the autonomous visual navigation obstacle avoidance decision-making agent to obtain a trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent;
[0032] Step 5: Based on the trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent, the quadrotor drone to be controlled is subjected to obstacle avoidance control.
[0033] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that in step 1, a quad-rotor drone autonomous visual navigation obstacle avoidance simulation environment is constructed, a quad-rotor drone dynamics model is selected and initialized, the drone's starting position is set, and the target position is set; the specific process is:
[0034] Step 11: Use Unreal Engine 4 to create a training environment containing obstacles (dense obstacles);
[0035] Use the 3D modeling software Unreal Engine 4 to create a training environment containing obstacles (obstacles are dense);
[0036] Use the 3D modeling software Unreal Engine 4 to create a model evaluation environment with randomly set obstacles. The obstacles in the model evaluation environment are the same as those in the training environment.
[0037] In the training environment, 70 cylindrical obstacles with a radius of 2.5m and a height of 15m were placed in a circle with a radius of 60m and the coordinate origin as the center;
[0038] In the model evaluation environment, 60 obstacles identical to those in the training environment were randomly placed in a circle with a radius of 60 m and the coordinate origin as the center;
[0039] Step 12: Integrate the AirSim plug-in into the simulation environment of Unreal Engine 4, import the quad-rotor UAV dynamics model (Multirotor dynamics model) provided by AirSim, and set the dynamics parameters of the quad-rotor UAV dynamics model;
[0040] The parameters of the UAV dynamics model are the duration of the UAV flight action execution, the UAV speed range on the x-axis, the UAV speed range on the y-axis, the UAV speed range on the z-axis, and the UAV yaw angular velocity range;
[0041] The x-axis, y-axis, and z-axis are the coordinate system of the drone body. The direction of the drone head is the x-axis, the vertical direction of the x-axis is the z-axis, and the vertical xz plane is the y-axis;
[0042] Step 13: Set the initial position and target position of the drone for each training round, define the effective flight space range, and set the conditions for successfully completing the navigation and obstacle avoidance flight as well as the failure conditions.
[0043] Reinforcement learning training is divided into rounds. Termination conditions are set for each round. When the termination conditions are met, the training round ends, the simulation environment is reset, and the next training round begins.
[0044] The other steps and parameters are the same as those in the first embodiment.
[0045] Specific implementation method three: This implementation method is different from specific implementation methods one or two in that in step 2, the state space, observation space and action space are set in the simulation environment; the specific process is:
[0046] Step 21: Model the autonomous navigation task of the UAV as a partially observable Markov decision process, expressed as<s,o,a,r,p,γ> ;
[0047] Among them, s is the state, o is the observation, a is the action, r is the reward, p is the state transition probability, and γ is the discount factor;
[0048] Taking into account the partially observable environmental phenomena caused by factors such as high-speed motion of the UAV, sensor noise, and environmental wind disturbance, the present invention models the autonomous navigation task of the UAV as a partially observable Markov decision process (POMDP), which is expressed as<s,o,a,r,p,γ> ;
[0049] In the UAV autonomous navigation task, the state s represents the state information of the UAV in the current environment (the state information includes two parts: visual perception information and the UAV's own state information; the visual perception information refers to the depth image without noise, and the UAV's own state information refers to the UAV's three-axis position, three-axis speed, the angle between the current flight direction and the target direction, and the current yaw angular velocity without noise);
[0050] Observation o represents the estimation of state s obtained by the UAV through the onboard sensors;
[0051] Action a is the flight control command executed by the UAV at each time step;
[0052] The reward r is used to evaluate the value of the drone performing an action in a certain state to transfer to the next state;
[0053] The state transition probability p defines the probability of taking a certain action in the current state to transfer to the next specific state. This probability is determined by the UAV dynamics and is affected by factors such as control input, environmental conditions, and system noise.
[0054] The discount factor γ determines the importance of the reward of the future time step compared to the reward of the current time step. The closer the discount factor is to 1, the more important the future reward is.
[0055] Based on the idea of privileged learning, the present invention adopts the Learning Using Privileged Information (LUPI) framework and inputs the state information as privileged information into the critic network.
[0056] Step 22: Set the state space to obtain the current state from the simulation environment. The state space in the present invention is a continuous state space. The specific process is:
[0057] The status information includes visual perception information and the drone’s own status information;
[0058] The visual perception information is the depth image acquired by the drone’s onboard depth camera;
[0059] The depth image acquired by the drone's onboard depth camera is input into the feature extraction network, and the feature extraction network outputs the visual perception information vector S1;
[0060] The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence;
[0061] Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU in sequence;
[0062] The drone's own status information includes the three-axis distance from the current drone to the target point [d x ,d y ,d z ], the current three-axis speed of the drone [v x ,v y ,v z ], the yaw angle deviation between the current flight direction and the target direction And the current yaw rate of the drone
[0063] The drone’s own state information is the 8-dimensional drone’s own state information vector S2;
[0064] The current three-axis speed of the drone [v x ,v y ,v z ] is the three-axis speed in the UAV body coordinate system;
[0065] The three-axis distance from the current drone to the target point [d x ,d y ,d z ] is the three-axis distance in the world coordinate system;
[0066] The visual perception information vector S1 and the drone's own state information vector S2 are concatenated to obtain the state information vector S = [S1, S2];
[0067] Step 23: setting an action space to control the flight of the quadrotor drone in the simulation environment. The action space in the present invention is a continuous action space;
[0068] The action space is the flight control instruction A of the drone, which contains the three-axis speed of the current drone [v x ,v y ,v z ] and yaw rate These sequential actions directly control the movement of the drone in the simulation environment;
[0069] The current three-axis speed of the drone [vx ,v y ,v z ] is the three-axis speed in the UAV body coordinate system;
[0070] The action space is 4-dimensional;
[0071] Step 24: setting an observation space to obtain currently observed information from the simulation environment. The observation space in the present invention is a continuous observation space.
[0072] The observation information is to add salt and pepper noise, Gaussian noise, and motion blur to the state information to simulate the state that the drone can observe in some observable environments;
[0073] The dimension of the observation space is the same as the dimension of the state space and can be expressed as O∈R 33 In the framework of LUPI, observation information is the input of the actor network;
[0074] The specific process is:
[0075] Salt and pepper noise, Gaussian noise, and motion blur are added to the drone depth image information in sequence to fully simulate the noise and interference that may be encountered during flight in a real environment;
[0076] In the drone's own status information, the three-axis distance [d x ,d y ,d z ]Add Gaussian noise to simulate positioning noise and other sensor noise;
[0077] In the drone's own state information, the current drone's three-axis speed [v x ,v y ,v z ]Add Gaussian noise to simulate positioning noise and other sensor noise;
[0078] In the drone's own status information, the yaw angle deviation between the current flight direction and the target direction Add Gaussian noise to simulate positioning noise and other sensor noise;
[0079] In the drone's own status information, the current drone yaw angular velocity Add Gaussian noise to simulate positioning noise and other sensor noise;
[0080] Add noise to the drone state information in the simulation environment to simulate some observable phenomena of the environment (partial observable phenomena of the environment is a professional term in the field of reinforcement learning, which is used to describe the inability of intelligent agents to obtain complete and accurate state information when interacting in the environment, in order to simulate the complex situation in the real environment);
[0081] The drone status information includes visual perception information and the drone’s own status information;
[0082] Visual perception information refers to the depth image acquired by the drone’s onboard depth camera;
[0083] The drone's own status information refers to the three-axis distance from the current drone to the target point [d x ,d y ,d z ], the current three-axis speed of the drone [v x ,v y ,v z ], the yaw angle deviation between the current flight direction and the target direction And the current yaw rate of the drone
[0084] The noise includes salt and pepper noise, Gaussian noise and motion blur added to the visual perception information, and Gaussian noise added to the drone’s own state information;
[0085] The specific process is:
[0086] Salt and pepper noise, Gaussian noise, and motion blur are added to the drone depth image information in sequence to fully simulate the noise and interference that may be encountered during flight in a real environment;
[0087] The three-axis distance [d x ,d y ,d z ] Gaussian noise is added to the three-axis position of the drone to simulate positioning noise and other sensor noise;
[0088] The current three-axis speed of the drone in the drone's own status information [v x ,v y ,v z ]Add Gaussian noise to simulate positioning noise and other sensor noise;
[0089] The yaw angle deviation between the current flight direction and the target direction in the drone's own status information Add Gaussian noise to simulate positioning noise and other sensor noise;
[0090] The current yaw rate of the drone in the drone's own status information Add Gaussian noise to simulate positioning noise and other sensor noise.
[0091] The other steps and parameters are the same as those in the first or second embodiment.
[0092] Specific implementation method 4: This implementation method is different from any one of the specific implementation methods 1 to 3 in that a neural network model of an autonomous visual navigation obstacle avoidance decision-making agent is constructed in step 3; the specific process is as follows:
[0093] The neural network model of the autonomous visual navigation obstacle avoidance decision-making agent includes a feature extraction network and a decision network in sequence;
[0094] The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence;
[0095] Each of the first convolution block, the second convolution block, and the third convolution block sequentially comprises a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU;
[0096] The decision network includes the first multi-layer perceptron, the second multi-layer perceptron, and the activation function adopts Leaky ReLU;
[0097] The activation function uses Leaky ReLU to prevent the gradient vanishing problem from causing the actor network to continuously output boundary values.
[0098] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.
[0099] Specific implementation method five: This implementation method is different from any one of specific implementation methods one to four in that in step 4, a neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent is trained to obtain a trained neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent; the specific process is:
[0100] In order to overcome the noise and interference of the perceived information, the present invention designs an asymmetric input actor and critic network model architecture. During the model training phase, the actor network will use noisy inputs, while the critic network will use accurate inputs, thereby enhancing the deep reinforcement learning model's ability to perceive the state;
[0101] Define a reinforcement learning reward function, adopt a multi-agent exploration strategy, and train a neural network model based on deep reinforcement learning autonomous visual navigation and obstacle avoidance decision-making agents, so that the model can continuously obtain training samples and optimize network parameters in the interaction with the environment, thereby improving the success rate of autonomous visual navigation and obstacle avoidance;
[0102] The following steps are involved:
[0103] Step 31, define a reward function. The present invention adopts a combination of sparse rewards and discrete rewards based on the current time step and the current round state;
[0104] The reward function includes sparse rewards and continuous rewards;
[0105] Sparse rewards are: if the current time step round ends, the rewards include positive rewards for successfully reaching the target point, and negative rewards for collisions and exceeding the flight range; the specific process is:
[0106] When the drone enters the target reach range, it gets a positive reward of +10; if the drone collides with an obstacle or flies out of the environment boundary, it gets a penalty of -5;
[0107] If the current training round ends because the target point is successfully reached, the agent is given a positive reward; if the current training round ends because the drone flies out of the specified flight range or collides with an obstacle, the agent is given a negative reward. The positive and negative rewards here will not be obtained at the same time because only one termination condition can be triggered.
[0108] If the current time step round has not ended, the reward includes a positive reward for being close to the target, a negative reward for the target distance error, and a negative reward for the collision probability calculated based on the distance to the obstacle surface;
[0109] In addition, the agent will receive continuous rewards including distance difference positive reward, distance error penalty and obstacle approach penalty at each time step. The distance difference positive reward is used to evaluate the agent's proximity to the target at the current time step compared to the previous time step, and the positive reward is used to encourage the drone to approach the target; the distance error penalty is calculated based on the distance from the drone to the straight line connecting the starting point and the target to ensure that the drone travels along the shortest straight path and improve the final navigation accuracy; the obstacle approach penalty is calculated based on the difference between the shortest distance between the drone and the obstacle and the safe distance at each time step, which helps prevent the drone from approaching obstacles.
[0110] The expression of continuous reward is as follows:
[0111]
[0112] in,
[0113] r e Indicates a reward item; d g Indicates the distance from the starting point to the target, d t and d t-1 Respectively represent the distance between the UAV and the target at the current and previous time steps;
[0114] p p represents penalty term 1; z and z g Represents the z-axis coordinates of the current time step and target position of the drone, d l Indicates the distance from the current position of the drone to the straight line connecting the starting point and the target; the function clip is used to limit the data;
[0115] po represents the penalty term 2; d o Indicates the shortest distance from the drone to the obstacle surface; d c is the collision distance, when d o Less than d c It is considered that a collision has occurred; s is the safe distance, when d o Less than d s There is a risk of collision between the UAVs on the mission;
[0116] r represents the continuous reward value; η r Yes e The scaling factor, η p Yes p The scaling factor, η o Yes o The scaling factor of
[0117] Step 32: Set the experience replay pool parameters and deep reinforcement learning neural network parameters; the specific process is as follows:
[0118] The experience replay pool parameter is the size parameter of the experience replay pool;
[0119] The hyperparameters of deep reinforcement learning neural networks include network structure, number of neurons, learning rate, discount factor, learning rate, number of learning start steps, experience replay pool size, batch size, training frequency, standard deviation of action noise, number of training environments, total number of training steps, and maximum number of steps per round;
[0120] Step 33: Set the action selection strategy to complete experience collection; the specific process is as follows:
[0121] First, an ε-greedy strategy is used to select random actions or actions output by the actor network (the action output by the actor network is the observation input to the actor network, and the actor network outputs the action);
[0122] Adopting the ε-greedy strategy, a random action is selected with a small probability, or an action output by the actor network is selected with a high probability, thereby balancing exploration and exploitation;
[0123] The drone then performs the selected random action or the action output by the actor network and interacts with the environment to generate a new state after the transfer, a new observation after the transfer, and a reward value returned by the environment;
[0124] The current state, current observation, selected action, new state after transfer, new observation after transfer, and reward value returned by the environment t ,o t ,a t ,s t+1 ,o t+1 ,r t >Store it in the experience replay pool, thus completing an experience collection;
[0125] s t Indicates the current state, o t represents the current observation, a t Indicates the currently selected action, s t+1 Indicates the new state after the transfer, o t+1 represents the new observation after transfer, r t Represents the reward value returned by the environment;
[0126] Set the action selection strategy to complete experience collection. In order to balance exploration and utilization, the present invention uses the ε-greedy algorithm for action selection. ε is a probability value close to 0. A random number is generated at each time step. If the random number is greater than or equal to ε, the action with the highest value in the current state is selected according to the reinforcement learning algorithm, otherwise an action is randomly selected from the action space. The ε-greedy algorithm will dynamically adjust the reinforcement learning's exploration of new strategies and the utilization of existing strategies as the training process progresses. The action is then executed in a simulation environment, the state at the next moment is collected and the reward value is calculated, and the current state, observation, action, reward, state at the next moment, and observation are stored in the experience replay pool until the experience replay pool is full or the simulation ends.
[0127] Step 34: Set a multi-agent exploration strategy to complete experience collection; the specific process is as follows:
[0128] Use M processes to run M AirSim environments in parallel, place a drone in each environment to perform experience collection in step 33, and thus complete one experience collection;
[0129] M drones input observations into the same decision network, obtain corresponding actions, execute actions and interact with the environment, generate new states after transfer, new observations after transfer, and reward values returned by the environment; and transform the current state, current observations, selected actions, new states after transfer, new observations after transfer, and reward values returned by the environment into a decision network. t ,o t ,a t ,s t+1 ,o t+1 ,r t >Store in the same experience replay pool, thus completing one experience collection;
[0130] Step 35: Train the neural network model of the deep reinforcement learning-based autonomous visual navigation and obstacle avoidance decision-making intelligent agent to obtain a trained neural network model of the deep reinforcement learning-based autonomous visual navigation and obstacle avoidance decision-making intelligent agent.
[0131] The parameter update mechanism of the neural network model DRPL (Twin Delayed Deep Deterministic Policy Gradient, TD3) based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making agent of the present invention is as follows under the task modeling of POMDP: Figure 5 shown.
[0132] Specifically, the TD3 algorithm consists of six networks: the actor network (with parameters θ), the actor target network (with parameters θ - ), two critic networks (with parameters ω1 and ω2), and two critic target networks (with parameters and ).
[0133] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.
[0134] Specific implementation method 6: This implementation method is different from any one of the specific implementation methods 1 to 5 in that in step 35, a neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent is trained to obtain a trained neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent; the specific process is:
[0135] Step 351: Randomly select 128 experiences from the experience replay pool, and 128 experiences form a batch;
[0136] Each experience is a 6-tuple consisting of the current state, current observation, selected action, new state after transfer, new observation after transfer, and reward value returned by the environment;
[0137] The i-th experience is expressed as i ,o i ,a i ,s i+1 ,o i+1 ,r i >, i = 1, 2, ..., 128;
[0138] Among them, s i represents the state in the i-th experience, o i represents the observation in the i-th experience, a i represents the action in the i-th experience, s i+1 represents the new state in the i-th experience after the transfer, o i+1 represents the new observation in the i-th experience after the transfer, r i Represents the reward value returned by the environment;
[0139] The experiences in the experience pool are counted individually. For example, my experience pool can store 50,000 experiences. Each experience is a 6-tuple. When sampling, 128 experiences will be randomly selected from it to form a batch. i means the number in this batch. This batch is used for one network parameter update.
[0140] Step 352
[0141] The observation o in the i-th experience i Input actor network, actor network output action Output actions from the actor network Input critic network 1 and critic network 2 respectively;
[0142] The state s in the i-th experience i Input critic network 1 and critic network 2 respectively, and critic network 1 and critic network 2 output the value estimation of state-action pair respectively and ω1 is the parameter of critic network 1, ω2 is the parameter of critic network 2;
[0143] The new observation o in the i-th experience after the transfer i+1 Input is the actor-target network, and the actor-target network outputs the selected action Output actions of the actor-target network Input critic target network 1 and critic target network 2 respectively;
[0144] The new state s in the i-th experience after the transfer i+1 Critic target network 1 and critic target network 2 are input respectively, and critic target network 1 and critic target network 2 output the value estimation of state-action pair respectively and are the parameters of the critic target network 1; are the parameters of critic target network 2;
[0145] Two critic-target networks (consistent with the input to the critic network);
[0146] Step 353: Based on as well as Calculating Timing Error
[0147] Indicates taking or
[0148] Indicates taking The smaller value;
[0149] r represents the continuous reward value; γ represents the discount factor;
[0150] Step 354, taking the mean square error (MSE) of the timing error to obtain the loss function L(ω) of the critic network;
[0151] Based on the output of the critic network and Construct the loss function L(θ) of the actor network;
[0152] The actor network performs soft updates on the actor-target network;
[0153] Critic network 1 soft updates critic target network 1;
[0154] Critic network 2 soft updates critic target network 2;
[0155] Step 355: The number of steps in each training round (success or failure is 1 round, 1 round is 1 step) does not exceed the maximum number of steps in the round. When the number of steps in this round reaches the maximum number of steps in the round, the drone reaches the target, collides with an obstacle, or exceeds the flight range, the training round ends and enters the next round;
[0156] When the number of training steps exceeds the number of learning steps, the actor network parameters θ and the actor target network parameters θ are updated according to the training frequency (network parameters are updated every few steps). - , Critic network 1 parameter ω1, Critic network 2 parameter ω2, Critic target network 1 parameter and critic target network 2 parameters Make updates;
[0157] When the reward function converges or reaches the total number of training steps, the training is stopped to obtain a trained neural network model of an autonomous visual navigation obstacle avoidance decision-making agent based on deep reinforcement learning.
[0158] The experience playback pool of the present invention adopts a first-in-first-out strategy. t ,o t ,a t ,s t+1 ,o t+1 ,r t >As a set of samples, a batch size of samples is drawn when the model is updated;
[0159] The actor network is updated less frequently than the critic network to ensure that the critic network has sufficient time to fit a more accurate Q function and avoid the actor network using unstable Q-value estimates to optimize network parameters.
[0160] Actor-target network (with parameters θ- ) and two critic target networks (with parameters and ) is based on a soft update approach, where the parameters of the actor and critic networks are gradually updated.
[0161] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.
[0162] Specific implementation example 7: This implementation example is different from any one of specific implementation examples 1 to 6 in that the network results of the actor network, the actor target network, the critic network 1, the critic network 2, the critic target network 1, and the critic target network 2 are the same;
[0163] Any one of the actor network, the actor target network, the critic network 1, the critic network 2, the critic target network 1, and the critic target network 2 sequentially comprises a feature extraction network and a decision network;
[0164] The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence;
[0165] Each of the first convolution block, the second convolution block, and the third convolution block sequentially comprises a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU;
[0166] The decision network includes the first multi-layer perceptron, the second multi-layer perceptron, and the activation function adopts LeakyReLU.
[0167] The activation function uses LeakyReLU to prevent the gradient vanishing problem from causing the actor network to continuously output boundary values.
[0168] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.
[0169] Specific implementation eight: This implementation differs from any one of specific implementations one to seven in that the mean square error (MSE) of the timing error is taken to obtain the loss function L(ω) of the critic network, which is expressed as:
[0170]
[0171] According to the loss function L(ω) of the critic network, the parameters ω1 of critic network 1 and ω2 of critic network 2 are updated using the gradient descent theorem.
[0172] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.
[0173] Specific implementation method 9: This implementation method is different from one of the specific implementation methods 1 to 8 in that the output of the critic network is based on and Construct the loss function L(θ) of the actor network, expressed as:
[0174]
[0175] According to the loss function L(θ), the parameters θ of the actor network are updated using the gradient descent theorem;
[0176] Where N represents the total number of sampled experiences in a batch in the experience replay pool.
[0177] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.
[0178] Specific implementation method ten: This implementation method is different from any one of specific implementation methods one to nine in that in step 5, the obstacle avoidance control of the quadrotor drone to be controlled is performed based on the neural network model of the trained autonomous visual navigation obstacle avoidance decision-making intelligent agent; the specific process is:
[0179] The drone status is obtained, and the drone status is input into the trained neural network model (actor network) of the autonomous visual navigation and obstacle avoidance decision-making agent. The trained neural network model (actor network) of the autonomous visual navigation and obstacle avoidance decision-making agent outputs actions, and obstacle avoidance control is performed on the quadrotor drone to be controlled based on the output actions.
[0180] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.
[0181] The working process of the actor network model is:
[0182] The observation information of the UAV intelligent body is input into the feature extraction network in the actor network model, and the feature extraction network in the actor network model outputs the features of the depth image. The features of the depth image and the noisy state vector of the UAV itself are spliced and input into the decision network in the actor network model, and the decision network in the actor network model outputs the action; thus, a trained actor network model is obtained;
[0183] The working process of the critic network model is:
[0184] The state information of the UAV intelligent body is input into the feature extraction network in the critic network model, and the feature extraction network in the critic network model outputs the features of the depth image. The features of the depth image and the state vector of the UAV itself are concatenated and input into the decision network of the critic network model, and the decision network in the critic network model outputs the value evaluation; thus, a trained critic network model is obtained;
[0185] The specific structure of the network is shown in Table 3.
[0186] In order to more effectively extract key information from images, the present invention uses a convolutional neural network as a feature extraction network for the actor and critic network models in deep reinforcement learning;
[0187] Test and evaluate the performance of a deep reinforcement learning-based autonomous visual navigation and obstacle avoidance strategy for a quadrotor drone, taking into account some observable factors of the environment.
[0188] The following steps are involved:
[0189] 1. Determine the evaluation indicators;
[0190] The evaluation indicators should be closely related to the UAV autonomous navigation and obstacle avoidance mission, which can not only determine whether the mission is successfully completed, but also quantitatively describe the completion effect of the mission in combination with the reward function.
[0191] The present invention uses the following indicators: Average Episode Reward (AER), Average Steps of Successful Episodes (ASSE), and Success Rate (SR);
[0192] 2. Conduct performance comparison experiments on autonomous visual navigation and obstacle avoidance algorithms. In training and random environments, the proposed DPRL algorithm, the common TD3 algorithm, and the current cutting-edge UAV autonomous navigation algorithm EGO-Planner-v2 are compared. During the experiment, noise is added to the perception information of all algorithms. For each algorithm, 30 rounds of tests are conducted in each environment, and the empirical data of each time step is saved to calculate various evaluation indicators.
[0193] 3. Conduct ablation experiments on the autonomous visual navigation obstacle avoidance algorithm. Based on the proposed algorithm, first remove the privileged information and multi-agent experience collection strategies to verify the impact of these two strategies on the overall performance of the algorithm. Then, adjust the state space and action space of the algorithm to evaluate the rationality and effectiveness of the algorithm design proposed by the present invention.
[0194] The technical solution of the present invention is further described below in conjunction with the embodiments, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.
[0195] like Figure 1 As shown, the specific implementation steps of the quadrotor drone autonomous visual navigation obstacle avoidance method based on deep reinforcement learning of the present invention are as follows:
[0196] Step 1: Build a simulation environment for autonomous visual navigation and obstacle avoidance of a quadrotor drone, select and initialize the dynamic model of the quadrotor drone, set the starting position of the drone and set the target position. The specific steps of step 1 are as follows:
[0197] Step 11. Use the 3D modeling software Unreal Engine 4 to design a training environment with dense obstacles and create a model evaluation environment with random obstacle positions, such as Figure 2 As shown in the figure. In the training environment, 70 cylindrical obstacles with a radius of 2.5m and a height of 15m were placed in a circle with a radius of 60m and a center at the coordinate origin. In the model evaluation environment, 60 obstacles identical to those in the training environment were randomly placed in a circle with a radius of 60m and a center at the coordinate origin.
[0198] Step 12: Integrate the AirSim plug-in into the simulation environment of Unreal Engine 4, use the Multirotor dynamics model provided by AirSim, and set the duration of its action execution to 0.1s to ensure smooth control instructions as much as possible and avoid applying too high a computational load to the simulation environment, so as to achieve the optimal simulation frame rate. The parameters of the UAV dynamics model are shown in Table 1.
[0199] Table 1. Parameter settings of UAV dynamics model
[0200]
[0201] Step 13. In each training cycle, the drone takes off from the coordinate origin of the environment at a height of 5m, randomly selects a target point on a circle with the origin as the center and a radius of 65m, and navigates to the target point. When the drone enters the target point arrival range, the training round is judged to be successful, and when the drone exceeds the allowable flight range of the environment or the distance to the obstacle is less than the collision distance, the round is judged to fail, the position and state of the drone are reset, and the next training round is entered. The parameter settings of the simulation environment are shown in Table 2.
[0202] Table 2 Simulation environment parameter settings
[0203]
[0204] Step 2: Model the problem, simulate some observable phenomena in the environment in the simulation environment, design a reasonable state space, observation space and action space in combination with privileged learning, and build a neural network model of a deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making agent. The specific steps of step 2 are as follows:
[0205] Step 21: Model the problem. Considering the partially observable environmental phenomena caused by factors such as high-speed motion of the UAV, sensor noise, and environmental wind disturbance, the present invention models the autonomous navigation task of the UAV as a partially observable Markov decision process (POMDP), which can be expressed as<s,o,a,r,p,γ> , respectively, state, observation, action, reward, state transition probability and discount factor. In the UAV autonomous navigation task, state s represents the accurate state information of the UAV in the current environment; observation o represents the estimation of state s obtained by the UAV through the onboard sensor; action a is the flight control instruction executed by the UAV at each time step; reward r is used to evaluate the value of the UAV performing a certain action in a certain state to transfer to the next state; state transition probability p defines the probability of taking a certain action in the current state to transfer to the next specific state, which is determined by the UAV dynamics and is affected by factors such as control input, environmental conditions and system noise; discount factor γ determines the importance of the reward of the future time step compared to the reward of the current time step. The closer the discount factor is to 1, the more important the future reward is to the current reward.
[0206] Step 22: simulate some observable phenomena in the environment. The present invention uses a variety of noises to simulate some observable phenomena that may be encountered in the real world. Specifically, the present invention adds a mean value of μ to each dimension of the perception information of its own state. s , standard deviation is σ s Gaussian noise is used to simulate the sensor perception error and noise or the estimation error of the visual inertial odometer, and the noise is limited to prevent excessive interference, thereby ensuring the stability of the training. In addition, the present invention adds the following noises to the depth image perception information in sequence:
[0207] 1) Salt and pepper noise: The present invention firstly calculates the probability p sp Randomly assign extreme values (0 or 255) to the pixels of the depth image. Salt and pepper noise simulates sudden sensor perception errors caused by sudden lighting changes, strong light reflections, or signal loss.
[0208] 2) Gaussian noise: After adding salt and pepper noise, the present invention adds a mean value of μ to the depth image. g , standard deviation is σ g Gaussian noise. Gaussian noise introduces random measurement deviations throughout the image, which, combined with salt and pepper noise, softens some extreme values and simulates a more realistic scenario where multiple noises coexist.
[0209] 3) Motion blur: After adding Gaussian noise, the present invention adds a convolution kernel size of k to the depth image. mbMotion blur. Motion blur simulates the blurring effect caused by camera or object motion and is applied after other noise to simulate the effect of high-speed drone motion or camera shake exacerbating sensor noise in real-world scenes.
[0210] The effect of adding noise to the depth image perception information is as follows Figure 3 shown.
[0211] Step 23, set the state space. Based on the idea of privileged learning, the present invention adopts the Learning Using Privileged Information (LUPI) framework to input the state information as privileged information to the critic network. The state information of the present invention is multimodal information including visual perception information and the drone's own state perception information. Specifically, the visual perception information is a depth image with a resolution of 240×320 acquired by the drone's onboard depth camera, which is adjusted to 80×100 for storage and processing. Subsequently, a feature extraction network is used to extract features from the depth image and perform dimensionality reduction, and the visual perception data D∈R 80×100 Compressed into S1∈R 25 As part of the system state. The other part of the system state is the vector S2∈R containing the 8-dimensional state information of the drone itself 8 , which includes the three-axis distance from the drone to the target point at the current time step [d x ,d y ,d z ]∈R 3 , the three-axis speed of the drone [v x ,v y ,v z ]∈R 3 , the yaw angle deviation Δψ∈R between the current flight direction and the target direction 1 And the current yaw angular velocity ψ∈R of the UAV 1 . The two parts of the state vector are concatenated to obtain the complete system state vector S = [S1, S2] ∈ R 33 .
[0212] Step 24, set the observation space. The observation information of the present invention is to add the noise in step 22 on the basis of the state information to simulate the state that the drone can observe in a partially observable environment. The dimension of the observation space is the same as the dimension of the state space, which can be expressed as O∈R 33 In the LUPI framework, observation information is the input of the actor network.
[0213] Step 25: Set the action space. The action space of the present invention is the flight control instruction A∈R of the UAV. 4 , including three-axis speed instructions [vx ,v y ,v z ]∈R 3 and the yaw angular velocity ψ∈R 1 ,These continuous actions directly control the movement of the UAV in the simulation environment.
[0214] Step 26: Based on the principle and implementation framework of privileged learning, a reinforcement learning autonomous navigation and obstacle avoidance model based on neural network is constructed. The present invention combines deep reinforcement learning with privileged learning to construct an actor and critic network model with asymmetric input, whose inputs are the observation information and state information of the UAV agent respectively. The actor and critic network structures are the same, such as Figure 4 As shown in the figure, both of them include a feature extraction network and a decision network. The feature extraction network consists of convolution blocks, each of which contains a convolution layer, a batch normalization operation, and a maximum pooling layer. The activation function uses ReLU. The decision network is a two-layer multi-layer perceptron with 128 neurons in each layer. The activation function uses Leaky ReLU to prevent the vanishing gradient problem from causing the actor network to continuously output boundary values. The specific structure of the network is shown in Table 3.
[0215] Table 3 Deep reinforcement learning network structure
[0216]
[0217] Step 3: Define the reinforcement learning reward function, adopt a multi-agent exploration strategy, and train the autonomous visual navigation and obstacle avoidance decision model based on deep reinforcement learning, so that the model can continuously obtain training samples and optimize network parameters in the interaction with the environment, thereby improving the success rate of autonomous visual navigation and obstacle avoidance. The specific steps of step 3 are as follows:
[0218] Step 31, define the reward function. The present invention adopts a combination of sparse rewards and continuous rewards based on the current time step and the current round state. At the end of each training round, the agent will obtain sparse rewards including positive rewards for reaching the target, collision penalties, and penalties for exceeding the flight range. Specifically, when the drone enters the target arrival distance range, it will obtain a positive reward of +10; if the drone collides with an obstacle or flies out of the environment boundary, it will obtain a penalty of -5. In addition, the agent will obtain continuous rewards including positive rewards for distance differences, distance error penalties, and obstacle approach penalties at each time step. Among them, the positive reward for distance difference is used to evaluate the degree of proximity of the agent to the target at the current time step compared with the previous time step, and the positive reward is used to encourage the drone to approach the target; the distance error penalty is calculated based on the distance from the drone to the straight line connecting the starting point and the target to ensure that the drone travels along the shortest straight path and improve the final navigation accuracy; the obstacle approach penalty is calculated based on the difference between the shortest distance between the drone and the obstacle and the safe distance at each time step, which helps prevent the drone from approaching the obstacle.
[0219] The expression of continuous reward is as follows:
[0220]
[0221] in,
[0222] r e Indicates a reward item; d g Indicates the distance from the starting point to the target, d t and d t-1 Respectively represent the distance between the UAV and the target at the current and previous time steps;
[0223] p p represents penalty term 1; z and z g Represents the z-axis coordinates of the current time step and target position of the drone, d l Indicates the distance from the current position of the drone to the straight line connecting the starting point and the target; the function clip is used to limit the data;
[0224] p o represents the penalty term 2; d o Indicates the shortest distance from the drone to the obstacle surface; d c is the collision distance, when d o Less than d c It is considered that a collision has occurred; s is the safe distance, when d o Less than d s There is a risk of collision between the UAVs on the mission;
[0225] r represents the continuous reward value; η r Yes eThe scaling factor, η p Yes p The scaling factor, η o Yes o The scaling factor of
[0226] Step 32, set the experience replay pool parameters and the hyperparameters of the agent neural network model. The training hyperparameter settings of the present invention are shown in Table 4. The number of steps in each training round does not exceed the maximum number of steps in the round. When the number of steps in this round reaches the maximum number of steps in the round, the drone reaches the target, collides with an obstacle, or exceeds the flight range, the training round ends and enters the next round. When the number of training steps exceeds the number of steps at the beginning of learning, the model is updated according to the training frequency. The experience replay pool of the present invention adopts a first-in, first-out strategy to t ,o t ,a t ,s t+1 ,o t+1 ,r t >As a set of samples, a batch size of samples is drawn when the model is updated.
[0227] Table 4 Training parameter settings
[0228]
[0229] Step 33, set the action selection strategy to complete experience collection. In order to balance exploration and utilization, the present invention uses the ε-greedy algorithm for action selection. ε is a probability value close to 0. A random number is generated at each time step. If the random number is greater than or equal to ε, the action with the highest value in the current state is selected according to the reinforcement learning algorithm, otherwise an action is randomly selected from the action space. The ε-greedy algorithm will dynamically adjust the reinforcement learning's exploration of new strategies and the utilization of existing strategies as the training process progresses. The action is then executed in a simulation environment, the state at the next moment is collected and the reward value is calculated, and the current state, observation, action, reward, state at the next moment, and observation are stored in the experience replay pool until the experience replay pool is full or the simulation ends.
[0230] Step 34, set the multi-agent exploration strategy. The present invention uses multiple processes to run multiple AirSim environments in parallel, and places a drone in each environment. These drones interact with their respective environments independently, collect experience according to step 33, and store it in the same experience replay pool for training a central model. The model provides decision action outputs for each drone to ensure consistent strategy updates in all environments. The multi-agent exploration strategy proposed in the present invention can realize knowledge sharing between drones, improve the efficiency of experience collection, and thus accelerate model convergence.
[0231] Step 35, training the deep reinforcement learning model. The parameter update of the neural network model DRPL based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making agent of the present invention is based on the TwinDelayed Deep Deterministic Policy Gradient (TD3) algorithm. Under the task modeling of POMDP, its parameter update mechanism is as follows: Figure 5 Specifically, the TD3 algorithm consists of six networks: an actor network (with parameters θ), an actor target network (with parameters θ-), two critic networks (with parameters ω1 and ω2), and two critic target networks (with parameters and ). At each time step, Represents the actor network based on the current observation o i Select the action, is the actor-target network based on the observation o at the next time step i+1 The action selected, s i and i+1 is with o i and i+1 The corresponding accurate status of the drone. and are the value estimates of the two critic networks for a given state-action pair, and are the value estimates of the next state-action pair by the two critic target networks, based on which the timing error can be calculated Taking the mean squared error (MSE) of the timing error of a batch of experience, we can get the loss function of the critic network, as shown in Formula 2. According to the loss function, the parameters of the critic network can be updated using the gradient descent theorem. The loss function of the actor network is constructed based on the Q value output by the critic network, as shown in Formula 3, where N is the number of batches of playback data used for updating. According to the loss function, the parameters of the actor network can be updated using the gradient descent theorem. The update frequency of the actor network should be lower than that of the critic network to ensure that the critic network has sufficient time to fit a more accurate Q function and avoid the actor network using unstable Q value estimates to optimize the network parameters. All target networks are gradually updated based on the parameters of the actor and critic networks based on soft updates.
[0232]
[0233] Step 4: Test and evaluate the performance of the autonomous visual navigation and obstacle avoidance strategy of the quadcopter based on deep reinforcement learning, taking into account some observable factors of the environment. The specific steps of step 4 are as follows:
[0234] Step 41, determine the evaluation index. The evaluation index should be closely related to the UAV autonomous navigation and obstacle avoidance task, which can not only determine whether the task is successfully completed, but also quantitatively describe the completion effect of the task in combination with the reward function. The present invention uses the following indicators: Average Episode Reward (Average Episode Reward, AER), Average Steps of Successful Episodes (Average Steps of Successful Episodes, ASSE) and Success Rate (Success Rate, SR). Among them, AER is used to evaluate the comprehensive performance of the algorithm, including the accuracy of navigation, the safety of obstacle avoidance and the efficiency of the algorithm. ASSE is used to evaluate the efficiency of the algorithm. The smaller the ASSE, the more efficient the UAV can complete the task with a shorter step length. SR is used to measure the success rate of the algorithm. The higher the SR, the safer and more practical the algorithm.
[0235] Step 42, conduct a performance comparison experiment of the autonomous visual navigation and obstacle avoidance algorithm. In a training environment and a random environment, the proposed DPRL algorithm, the ordinary TD3 algorithm, and the current cutting-edge UAV autonomous navigation algorithm EGO-Planner-v2 are compared. During the experiment, noisy visual and self-state perception information are provided to DPRL and TD3, and noisy odometer information is provided to EGO-Planner-v2, thereby interfering with its mapping process. The maximum speed of the UAV of the EGO-Planner-v2 algorithm is consistent with the settings of the other two algorithms, and a PD controller is used to output speed yaw angle control instructions to control the flight of the UAV. For each algorithm, the present invention compares the differences in its training process and tests the performance of the trained model in training and random environments. Figure 6 The curves in (a) and (b) are the SR and AER curves of the DPRL and TD3 algorithms trained under 4 different seeds. Figure 6 As shown in (a), DPRL converges significantly faster than TD3, achieving an average success rate of 85% after 220,000 steps. Figure 6 The convergence trend of the AER curve shown in (b) is consistent with the SR curve. The average reward of DPRL increases rapidly in the early and middle stages of training and stabilizes at a higher reward value after 240,000 steps. In contrast, the average reward of TD3 increases slowly throughout the training process and does not converge at the end of training.
[0236] In the model performance test experiment, the present invention conducts 30 rounds of tests in two environments for each algorithm and saves the empirical data of each time step to calculate various evaluation indicators. The flight trajectories of the three algorithms in the two environments are visualized as follows: Figure 7As shown, the blue trajectory represents the trajectory that successfully reaches the target, and the red trajectory represents the trajectory that collides with the obstacle. It can be seen that DPRL maintains a high success rate in both training and random environments, proving its robustness to environmental migration and some observable phenomena in the environment, and has a strong ability to adapt to new environments. In contrast, the TD3 algorithm has the lowest success rate, performs poorly in both environments, and its trajectory is not as smooth as DPRL. For the EGO-Planner-v2 algorithm, since both environments are unfamiliar environments and the obstacles in the training environment are denser, its performance in the training environment is worse than that in the random environment. EGO-Planner-v2 can plan the smoothest trajectory and has the highest arrival accuracy, but its planning efficiency is the lowest and it requires the most steps to reach the target. The comparison of the evaluation indicators of the three algorithms is shown in Table 5.
[0237] Table 5 Comparison results of evaluation indicators of model performance test experiments
[0238]
[0239] Step 43: Conduct an ablation experiment on the autonomous visual navigation obstacle avoidance algorithm. Based on the proposed algorithm, we first remove the privileged information and multi-agent experience collection strategies, respectively, and name them distributed reinforcement learning and privileged reinforcement learning to verify the impact of these two strategies on the overall performance of the algorithm. The SR and AER curves of the three algorithms during the training process are shown in Figure 4. Figure 8 As shown, Figure 8 (a) shows that DPRL has a faster convergence speed and higher final success rate than distributed reinforcement learning and privileged reinforcement learning. Specifically, the success rate of DPRL remains stable after 220,000 steps, while the success rates of distributed reinforcement learning and privileged reinforcement learning converge after 300,000 steps. Figure 8 The trend of the round average reward curve in (b) is consistent with the success rate curve. The average round reward of DPRL stabilizes to more than 30 after 240,000 steps, while distributed reinforcement learning and privileged reinforcement learning need to be trained for more than 300,000 steps to achieve similar rewards, and the average round reward of distributed reinforcement learning does not exceed 30 even after the end of training. In addition, the comparative analysis of the results of distributed reinforcement learning and privileged reinforcement learning shows that the role of privileged learning is greater than that of multi-agent exploration strategy. In the early stage of training, distributed reinforcement learning has a faster convergence speed than privileged reinforcement learning, which proves the significant role of multi-agent exploration strategy in accelerating experience convergence in the early stage; and after 200,000 steps of training, the performance of privileged reinforcement learning surpasses that of distributed reinforcement learning, and finally reaches higher SR and AER, highlighting the effect of privileged learning in dealing with some observable phenomena in the environment.
[0240] The present invention then conducted an ablation experiment on the design of the state and action space. The present invention designed another set of state and action spaces, and modified the composition of the drone's own state perception vector in the state space, changing the original three-axis distance with the target to the distance on the xy plane and the z-axis distance, and changing the original three-axis speed to the xy plane speed and the z-axis speed. Such modifications reduce the total state vector dimension from 33 dimensions to 31 dimensions. At the same time, the action space is modified by replacing the original x and y axis speeds with xy plane speeds, and the x and y axis speeds are decomposed according to the current yaw angle when controlling the drone flight. Such a state and action space setting can ensure that the drone always flies in the direction of the camera optical axis and keeps obstacles in the field of view. However, this design compresses the action space and reduces the flexibility of the drone's flight. The SR and AER curves of DPRL in the two states and action spaces during the model training process are shown in Figure 2. Fig. 9 As shown, it can be seen that the performance of the DPRL in the 4-dimensional action space designed by the present invention is significantly better than the algorithm in the 3-dimensional action space. Specifically, during the training process, the DPRL in the 3-dimensional action space has almost no successful flight rounds in the early stage of training, and the model's navigation and obstacle avoidance ability is learned very slowly, and only 30% of the success rate can be achieved at the end of training. From the AER curve, it can be seen that the average round reward of the DPRL in the 3-dimensional action space has almost no increase with training, and the final reward only reaches 10, which proves that the compression of the action space greatly affects the learning process of the model, and indirectly proves the rationality of the state and action space designed by the present invention.
[0241] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning, characterized in that: The specific process of the method is: Step 1: Build a simulation environment for autonomous visual navigation and obstacle avoidance of a quadrotor drone, select and initialize the dynamic model of the quadrotor drone, set the starting position of the drone and set the target position; Step 2: Set up the state space, observation space, and action space in the simulation environment; Step 3: Construct a neural network model for the autonomous visual navigation obstacle avoidance decision-making agent; Step 4: training the neural network model of the autonomous visual navigation obstacle avoidance decision-making agent to obtain a trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent; Step 5: Based on the trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent, the quadrotor drone to be controlled is subjected to obstacle avoidance control.
2. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 1, characterized in that: In step 1, a quad-rotor UAV autonomous visual navigation obstacle avoidance simulation environment is constructed, a dynamic model of the quad-rotor UAV is selected and initialized, a starting position of the UAV is set, and a target position is set; the specific process is as follows: Step 11: Use Unreal Engine 4 to create a training environment with obstacles. Step 12: Integrate the AirSim plug-in into the simulation environment of Unreal Engine 4, import the quadcopter UAV dynamics model provided by AirSim, and set the dynamics parameters of the quadcopter UAV dynamics model; The parameters of the UAV dynamics model are the duration of the UAV flight action execution, the UAV speed range on the x-axis, the UAV speed range on the y-axis, the UAV speed range on the z-axis, and the UAV yaw angular velocity range; The x-axis, y-axis, and z-axis are the coordinate system of the drone body. The direction of the drone head is the x-axis, the vertical direction of the x-axis is the z-axis, and the vertical xz plane is the y-axis; Step 13: Set the initial position and target position of the drone for each training round, define the effective flight space range, and set the conditions for successfully completing the navigation and obstacle avoidance flight as well as the failure conditions.
3. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 2, characterized in that: In step 2, the state space, observation space and action space are set in the simulation environment; the specific process is: Step 21, model the autonomous navigation task of the UAV as a partially observable Markov decision process, denoted as s,o,a,r,p,γ>; Among them, s is the state, o is the observation, a is the action, r is the reward, p is the state transition probability, and γ is the discount factor; Step 22: Set the state space to be a continuous state space. The specific process is as follows: The status information includes visual perception information and the drone’s own status information; The visual perception information is the depth image acquired by the drone’s onboard depth camera; The depth image acquired by the drone's onboard depth camera is input into the feature extraction network, and the feature extraction network outputs the visual perception information vector S1; The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence; Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU in sequence; The drone's own status information includes the three-axis distance [dx, dy, dz] from the current drone to the target point, the three-axis speed [vx, vy, vz] of the current drone, and the yaw angle deviation between the current flight direction and the target direction. And the current yaw rate of the drone The drone’s own state information is the 8-dimensional drone’s own state information vector S2; The current three-axis speed of the drone [vx, vy, vz] is the three-axis speed in the drone body coordinate system; The three-axis distance [dx, dy, dz] from the current drone to the target point is the three-axis distance in the world coordinate system; The visual perception information vector S1 and the drone's own state information vector S2 are concatenated to obtain the state information vector S = [S1, S2]; Step 23: Set the action space, which is a continuous action space. The action space is the flight control instruction A of the drone, which includes the three-axis speed [vx, vy, vz] and yaw angular velocity of the current drone. The current three-axis speed of the drone [vx, vy, vz] is the three-axis speed in the drone body coordinate system; Step 24: setting an observation space, where the observation space is a continuous observation space; The specific process is: In the drone depth image information, salt and pepper noise, Gaussian noise and motion blur are added in sequence; In the drone's own state information, Gaussian noise is added to the three-axis distance [dx, dy, dz] from the current drone to the target point; In the drone's own state information, add Gaussian noise to the current drone's three-axis speed [vx, vy, vz]; In the drone's own status information, the yaw angle deviation between the current flight direction and the target direction Add Gaussian noise; In the drone's own status information, the current drone yaw angular velocity Add Gaussian noise.
4. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 3 is characterized in that: In step 3, a neural network model of an autonomous visual navigation obstacle avoidance decision-making agent is constructed; the specific process is as follows: The neural network model of the autonomous visual navigation obstacle avoidance decision-making agent includes a feature extraction network and a decision network in sequence; The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence; Each of the first convolution block, the second convolution block, and the third convolution block sequentially comprises a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU; The decision network includes the first multi-layer perceptron, the second multi-layer perceptron, and the activation function adopts Leaky ReLU.
5. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 4, characterized in that: In step 4, a neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent is trained to obtain a trained neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent; the specific process is: Step 31, define the reward function; The reward function includes sparse rewards and continuous rewards; Sparse rewards are: if the current time step round ends, the rewards include positive rewards for successfully reaching the target point, and negative rewards for collisions and exceeding the flight range; the specific process is: When the drone enters the target reach range, it gets a positive reward of +10; if the drone collides with an obstacle or flies out of the environment boundary, it gets a penalty of -5; The expression of continuous reward is as follows: in, re represents the reward item; dg represents the distance from the starting point to the target, dt and dt-1 represent the distance between the drone and the target at the current and previous time steps respectively; pp represents penalty term 1; z and zg represent the z-axis coordinates of the current time step and target position of the drone, respectively; dl represents the distance from the current position of the drone to the straight line connecting the starting point and the target; the function clip is used to limit the data; po represents penalty term 2; do represents the shortest distance from the drone to the obstacle surface; dc is the collision distance; ds is the safety distance; r represents the continuous reward value; ηr is the scaling factor of re, ηp is the scaling factor of pp, and ηo is the scaling factor of po; Step 32: Set the experience replay pool parameters and deep reinforcement learning neural network parameters; the specific process is as follows: The experience replay pool parameter is the size parameter of the experience replay pool; The hyperparameters of deep reinforcement learning neural networks include network structure, number of neurons, learning rate, discount factor, learning rate, number of learning start steps, experience replay pool size, batch size, training frequency, standard deviation of action noise, number of training environments, total number of training steps, and maximum number of steps per round; Step 33: Set the action selection strategy to complete experience collection; the specific process is as follows: First, a ε-greedy strategy is used to select random actions or actions output by the actor network; The drone then performs the selected random action or the action output by the actor network and interacts with the environment to generate a new state after the transfer, a new observation after the transfer, and a reward value returned by the environment; The current state, current observation, selected action, new state after transfer, new observation after transfer, and reward value returned by the environment <st,ot,at,st +1 ,ot +1 ,rt>Store it in the experience replay pool, thus completing an experience collection; st represents the current state, ot represents the current observation, at represents the currently selected action, st +1 Indicates the new state after the transfer, ot +1 represents the new observation after the transfer, and rt represents the reward value returned by the environment; Step 34: Set a multi-agent exploration strategy to complete experience collection; the specific process is as follows: Use M processes to run M AirSim environments in parallel, place a drone in each environment to perform experience collection in step 33, and thus complete one experience collection; M drones input observations into the same decision network, obtain corresponding actions, execute actions and interact with the environment, generate new states after transfer, new observations after transfer, and reward values returned by the environment; the current state, current observation, selected action, new state after transfer, new observations after transfer, and reward values returned by the environment are st,ot,at,st +1 ,ot +1 ,rt> is stored in the same experience replay pool, thus completing one experience collection; Step 35: Train the neural network model of the deep reinforcement learning-based autonomous visual navigation and obstacle avoidance decision-making intelligent agent to obtain a trained neural network model of the deep reinforcement learning-based autonomous visual navigation and obstacle avoidance decision-making intelligent agent.
6. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 5, characterized in that: In step 35, a neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent is trained to obtain a trained neural network model based on deep reinforcement learning autonomous visual navigation obstacle avoidance decision-making intelligent agent; the specific process is: Step 351: Randomly select 128 experiences from the experience replay pool, and 128 experiences form a batch; Each experience is a 6-tuple consisting of the current state, current observation, selected action, new state after transfer, new observation after transfer, and reward value returned by the environment; The i-th experience is expressed as <si,oi,ai,si +1 ,oi +1 ,ri>,i=1,2,…,128; Among them, si represents the state in the i-th experience, oi represents the observation in the i-th experience, ai represents the action in the i-th experience, si +1 represents the new state in the i-th experience after the transfer, oi +1 represents the new observation in the i-th experience after the transfer, and ri represents the reward value returned by the environment; Step 352 The observation oi in the i-th experience is input into the actor network, and the actor network outputs the action Output actions from the actor network Input critic network 1 and critic network 2 respectively; The state si in the i-th experience is input into critic network 1 and critic network 2 respectively, and critic network 1 and critic network 2 output the value estimation of state-action pair respectively and ω1 is the parameter of critic network 1, ω2 is the parameter of critic network 2; The new observation oi in the i-th experience after the transfer +1 Input is the actor-target network, and the actor-target network outputs the selected action Output actions of the actor-target network Input critic target network 1 and critic target network 2 respectively; The new state si in the i-th experience after the transfer +1 Critic target network 1 and critic target network 2 are input respectively, and critic target network 1 and critic target network 2 output the value estimation of state-action pair respectively and are the parameters of the critic target network 1; are the parameters of critic target network 2; Step 353: Based on as well as Calculating Timing Error Indicates taking or Indicates taking The smaller value; r represents the continuous reward value; γ represents the discount factor; Step 354, taking the mean square error of the timing error to obtain the loss function L(ω) of the critic network; Based on the output of the critic network and Construct the loss function L(θ) of the actor network; The actor network performs soft updates on the actor-target network; Critic network 1 soft updates critic target network 1; Critic network 2 soft updates critic target network 2; Step 355: The number of steps in each training round does not exceed the maximum number of steps in the round. When the number of steps in this round reaches the maximum number of steps in the round, the drone reaches the target, collides with an obstacle, or exceeds the flight range, the training round ends and enters the next round. When the number of training steps exceeds the number of learning start steps, the actor network parameters θ and the actor target network parameters θ are adjusted according to the training frequency. - , Critic network 1 parameter ω1, Critic network 2 parameter ω2, Critic target network 1 parameter and critic target network 2 parameters Make updates; When the reward function converges or reaches the total number of training steps, the training is stopped to obtain a trained neural network model of an autonomous visual navigation obstacle avoidance decision-making agent based on deep reinforcement learning.
7. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 6, characterized in that: The network results of the actor network, the actor target network, the critic network 1, the critic network 2, the critic target network 1, and the critic target network 2 are the same; Any one of the actor network, the actor target network, the critic network 1, the critic network 2, the critic target network 1, and the critic target network 2 sequentially comprises a feature extraction network and a decision network; The feature extraction network includes the first convolution block, the second convolution block, the third convolution block, and the global average pooling layer in sequence; Each of the first convolution block, the second convolution block, and the third convolution block sequentially comprises a convolution layer, a batch normalization layer BN, a maximum pooling layer, and an activation function ReLU; The decision network includes the first multi-layer perceptron, the second multi-layer perceptron, and the activation function adopts LeakyReLU.
8. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 7, characterized in that: The mean square error of the timing error is taken to obtain the loss function L(ω) of the critic network, which is expressed as: According to the loss function L(ω) of the critic network, the parameters ω1 of critic network 1 and ω2 of critic network 2 are updated using the gradient descent theorem.
9. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 8, characterized in that: The critic network output and Construct the loss function L(θ) of the actor network, expressed as: According to the loss function L(θ), the parameters θ of the actor network are updated using the gradient descent theorem; Where N represents the total number of sampled experiences in a batch in the experience replay pool.
10. The method for autonomous visual navigation and obstacle avoidance of a quadrotor drone based on deep reinforcement learning according to claim 9, characterized in that: In step 5, the obstacle avoidance control of the quadrotor drone to be controlled is performed based on the trained neural network model of the autonomous visual navigation obstacle avoidance decision-making agent; the specific process is: The drone status is obtained, and the drone status is input into the neural network model of the trained autonomous visual navigation obstacle avoidance decision-making agent. The neural network model of the trained autonomous visual navigation obstacle avoidance decision-making agent outputs actions, and obstacle avoidance control is performed on the quadrotor drone to be controlled based on the output actions.
Citation Information
Patent Citations
Unmanned aerial vehicle autonomous obstacle avoidance navigation method based on memory reinforcement learning
CN115016534A
Unmanned aerial vehicle intelligent navigation method based on deep reinforcement learning
CN115373415A
Rotor unmanned aerial vehicle obstacle avoidance method, device and equipment based on reinforcement learning SAC
CN115494879A
Robot grabbing method and system based on deep reinforcement learning and TSK-FS fuzzy reasoning
CN117260739A
Unmanned aerial vehicle navigation obstacle avoidance control law design method based on reinforcement learning
CN118605559A
Cited By
Unmanned aerial vehicle autonomous navigation system based on hierarchical reinforcement learning strategy
CN120800385A
Trajectory planning method for unmanned aerial vehicle to quickly pass through window
CN120846342A
Visual navigation method and system based on reinforcement learning
CN120890466A
Rapid training type unmanned aerial vehicle reinforcement learning strategy method suitable for microcontroller
CN121256961A
A fast training type unmanned aerial vehicle reinforcement learning strategy method suitable for a microcontroller
CN121256961B