Unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical pulse reinforcement learning

By adopting a layered pulse reinforcement learning framework and pulse neural network in drone navigation, the problems of poor adaptability and high computing resource consumption in the existing technology are solved, and efficient and low-power drone obstacle avoidance navigation is achieved.

CN119935145APending Publication Date: 2025-05-06DALIAN UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510031709.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing UAV obstacle avoidance navigation technology has problems such as poor adaptability, low decision efficiency and high computing resource consumption, making it difficult to achieve efficient and low-power autonomous navigation in complex environments.

Method used

The hierarchical pulse reinforcement learning framework is adopted and combined with the pulse neural network, which is divided into low-level obstacle avoidance decision-making modules and high-level sub-target point decision-making networks. The operation of both is coordinated through the state machine to achieve coordinated decision-making of local obstacle avoidance and global path planning.

Benefits of technology

It improves the autonomous navigation capability and decision-making efficiency of the drone in complex environments, reduces the consumption of computing resources and power consumption, and enhances the adaptability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119935145A_ABST
    Figure CN119935145A_ABST
Patent Text Reader

Abstract

The invention belongs to the cross technical field of artificial intelligence and robot autonomous navigation, and discloses an unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical pulse reinforcement learning. A layered reinforcement learning framework is adopted for obstacle avoidance navigation, a pulse neural network is adopted as a reinforcement learning model, and a layered pulse reinforcement learning model is obtained; the layered pulse reinforcement learning model comprises a low-level obstacle avoidance decision-making module and a high-level sub-target point decision-making network; the obstacle avoidance decision module is used for local obstacle avoidance; the sub-target point decision network is used for planning a global path; and setting a state machine to coordinate the operation of the obstacle avoidance decision-making module and the sub-target point decision-making network. According to the invention, task requirements of local obstacle avoidance and global navigation can be adaptively balanced, so that the unmanned aerial vehicle can make adaptive adjustment more autonomously when executing a task. And while the decision-making efficiency is ensured, the computing resource consumption and power consumption are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross-technical field of artificial intelligence and robot autonomous navigation, and in particular to an obstacle avoidance navigation method for unmanned aerial vehicles based on hierarchical pulse reinforcement learning. Background Art

[0002] Autonomous obstacle avoidance navigation of unmanned aerial vehicles is an important topic in the field of unmanned aerial vehicles. It aims to enable unmanned aerial vehicles to perceive the surrounding environment information and plan a reasonable flight path in an unknown or known environment to avoid collisions with obstacles, thereby achieving safe and efficient autonomous flight. With the continuous development of artificial intelligence, automation technology and unmanned aerial vehicle technology, unmanned aerial vehicles are increasingly used in industry, agriculture, logistics, rescue, monitoring and other fields. Improving the obstacle avoidance and navigation capabilities of unmanned aerial vehicles in complex environments has become one of the key directions of unmanned aerial vehicle technology research. In practical applications, unmanned aerial vehicles need to face dynamic and static obstacles, complex environmental changes and a variety of different types of mission requirements, which requires unmanned aerial vehicles to have a high degree of autonomy, intelligent decision-making ability and adaptability. However, most of the current unmanned aerial vehicle obstacle avoidance and navigation technologies still have problems such as poor adaptability, low decision-making efficiency and high consumption of computing resources.

[0003] As an advanced artificial intelligence algorithm, deep reinforcement learning has good advantages and extensive research value in exploring autonomous learning in unknown environments and obtaining optimal control decisions. However, its reasoning process requires strong computing power, so it is difficult to deploy on drones with computing power and power consumption constraints. As the third generation of artificial neural networks, the Spiking Neural Network (SNN) is closer to the structure of the real brain in terms of underlying mechanism, so it has the characteristics of high dynamics and low power consumption similar to the biological brain. With its rich spatiotemporal neural dynamics, diverse encoding mechanisms and event-driven advantages, the Spiking Neural Network has attracted widespread attention in the academic community in recent years, surpassing the second generation of artificial neural networks (ANN). Therefore, combining deep reinforcement learning with spiking neural networks provides new possibilities for drones to provide efficient and low-power autonomous obstacle avoidance navigation capabilities in complex scenarios. The following will introduce the relevant background technologies in this field in detail.

[0004] (1) Traditional UAV autonomous obstacle avoidance navigation

[0005] Traditional autonomous obstacle avoidance navigation systems for drones generally include a perception module, a positioning module, a planning module, and a control module. The specific process is as follows: the drone obtains environmental information through sensors such as cameras and radars, the perception module understands the scene and extracts features, the positioning module provides the drone's current posture data, the planning module calculates a collision-free trajectory to the target point, and the control module adjusts the drone's motion parameters based on the planning results to achieve trajectory tracking. Although this sequentially executed modular system architecture has achieved mature implementation and application, cumulative errors are prone to occur between modules, which limits the adaptability and robustness of traditional systems in dynamic and complex environments.

[0006] (2) Autonomous navigation and obstacle avoidance of UAVs based on learning

[0007] In recent years, methods based on reinforcement learning and deep learning have been widely used in obstacle avoidance and navigation of unmanned aerial vehicles. Reinforcement learning learns the optimal obstacle avoidance strategy through the interaction between the agent and the environment, enabling unmanned aerial vehicles to make adaptive decisions in complex and uncertain environments. Compared with traditional methods, obstacle avoidance methods based on learning not only avoid the transmission of cumulative errors between modules, but also can process high-dimensional data and multiple modalities, and have stronger environmental perception and self-improvement capabilities. However, obstacle avoidance and navigation methods based on reinforcement learning and deep learning have the disadvantage of huge computing resource requirements. Therefore, how to reduce computing resource consumption while ensuring the real-time and accuracy of obstacle avoidance decisions is a problem worth studying. Summary of the invention

[0008] Aiming at the autonomous navigation of UAVs in complex unstructured scenarios, the present invention proposes a UAV obstacle avoidance navigation method based on hierarchical pulse reinforcement learning. Different from the traditional single reinforcement learning decision network, the present invention adopts a hierarchical reinforcement learning framework, which includes a two-level decision-making mechanism, and can divide the long-term global navigation task into a number of local obstacle avoidance tasks as few as possible, so that the UAV can avoid local obstacles at a low level and plan global path points at a high level according to the different complexities of the current environment. At the same time, the characteristics of low power consumption and efficient timing processing of pulse neural networks are combined to enhance the adaptability, real-time and robustness of intelligent agent decision-making. A method based on deep reinforcement learning, using pulse neural network decision-making and UAVs to achieve efficient obstacle avoidance and sub-target point planning in complex environments.

[0009] The technical solution of the present invention is as follows: a UAV obstacle avoidance navigation method based on hierarchical pulse reinforcement learning, which adopts a hierarchical reinforcement learning framework for obstacle avoidance navigation, wherein the reinforcement learning models all adopt pulse neural networks to obtain a hierarchical pulse reinforcement learning model; the hierarchical pulse reinforcement learning model includes a low-level obstacle avoidance decision module and a high-level sub-target point decision network; the obstacle avoidance decision module is used for local obstacle avoidance; the sub-target point decision network is used for planning a global path; and a state machine is set to coordinate the operation of the obstacle avoidance decision module and the sub-target point decision network.

[0010] The hierarchical pulse reinforcement learning model collects sensor data, target point data and drone status; the sensor data, target point data and drone status are processed by a sub-target decision module to obtain sub-target points; the sub-target points together with the sensor data and drone status are input into an obstacle avoidance decision module to obtain obstacle avoidance instructions, thereby guiding the drone to perform obstacle avoidance actions.

[0011] The output of the obstacle avoidance decision module is a continuous action, in which the Actor network adopts a spiking neural network SNN, and the Critic network adopts an artificial neural network ANN; the output of the sub-target point decision network is a discrete action, in which the Q network adopts a spiking neural network SNN; the neuron model of the spiking neural network selects the LIF model, and its pulse emission process is as follows:

[0012]

[0013] in, is the input pulse flow, w ij is the weight, is the neuron membrane potential, β is the membrane potential attenuation factor, b j is the bias term; when the membrane potential exceeds the threshold V th When , an output pulse is generated:

[0014]

[0015] The obstacle avoidance decision module adopts the TD3 algorithm; the sub-target point decision network adopts the DQN algorithm.

[0016] The reinforcement learning goal of the obstacle avoidance decision module is to find the optimal execution action based on the current environmental information, avoiding obstacles while approaching the sub-target point to maximize the specific reward of the task;

[0017] The input X of the obstacle avoidance decision module includes the relative position between the drone and the sub-target point, the drone speed and the laser data, which is expressed as follows:

[0018] X={W,V,P}

[0019] Where W = {d w,θ w} is the relative distance from the UAV to the sub-target point and the relative direction from the UAV to the sub-target point; V = {v l ,v a} is the linear velocity and angular velocity of the drone; P is the data of the forward laser radar, which contains the laser scanning results of 180° forward at three consecutive moments, P = {P t ,P t-1 ,P t-2}, each P t Contains several laser scanning points,

[0020] For the Actor network, Poisson encoding is used to convert the input into a pulse sequence X s = {W s ,V s ,P s}, the laser data is extracted through the pulse one-dimensional convolution layer to obtain the spatiotemporal features F P =SpikingConv1d(P s ), F P and {W s ,V s} to obtain The action is obtained through the pulse fully connected layer;

[0021] action=SpikingMLP(F fusion )

[0022] For the Critic network, a conventional artificial neural network is used. The network structure is exactly the same as the Actor network. Actions are added during series connection, and Q value 1 and Q value 2 are obtained through full connection. The smaller one is taken as the Q value.

[0023] The three elements of the obstacle avoidance decision module reinforcement learning are as follows:

[0024] The state is defined as {W,V,P}; the action is defined as {v l ,v a};

[0025] The reward is defined as in formula (3), r g represents the positive reward obtained by the drone when it reaches the target; r o represents the negative reward obtained after the drone collides; α is a scaling factor, d o Indicates the distance from the drone to the obstacle, t g Indicates the threshold for the drone to reach the target; t o Indicates the threshold for reaching an obstacle; the reward provides effective positive and negative feedback for the drone during the obstacle avoidance process;

[0026]

[0027] The goal of the sub-target point decision network reinforcement learning is to explore and obtain the best sampling points within the visible range. The sampling points are used as inputs of the obstacle avoidance decision module to guide the drone to the target point.

[0028] The input of the sub-target point decision network includes laser data and the relative position between the drone and the target point, which is expressed as follows:

[0029] X={G,O,M}

[0030] Where G = {d g ,θ g} is the relative distance from the drone to the target and the relative direction from the drone to the target; O is the data of the omnidirectional laser radar, and the laser radar data includes several laser scanning points, M is the local grid map obtained by transforming the lidar data O;

[0031] Use Poisson encoding to convert the input into a pulse sequence X s = {G s ,O s ,M s}; G s After the pulse fully connected layer, we get F G =SpikingMLP(G s );O s After the pulse one-dimensional convolution layer, we get F O =SpikingConv1d(O s );M s After the pulse two-dimensional convolution, we get F M =SpikingConv2d(M s );

[0032]

[0033] The three features are concatenated to get F fusion , input into the pulse fully connected layer, and output the score vector S:

[0034] S = SpikingMLP(F fusion )

[0035] Finally, the action mask M is used to mask the invalid area, and the candidate sub-target point C with the highest score is selected as the sampling point. ⊙ represents the element-by-element multiplication:

[0036] C = argmax(S⊙M).

[0037] The three elements of the sub-goal point decision network reinforcement learning are as follows:

[0038] The state is defined as {G,O}; the action is defined as C = {d c ,θ c}, sub-target point C represents the relative distance and relative direction between the current UAV and the target point;

[0039] The reward is defined as in formula (4), where r p The penalty term representing the distance between the sub-target point and the current position of the UAV, d max is the maximum reachable distance, d p is the distance between the current position of the UAV and the candidate sub-target point C; r g The reward for the candidate sub-target point C to guide the drone to the target position G depends on the Euclidean distance d between C and G. g ; r f Indicates the exploration value of the candidate sub-target point, which depends on the number of boundary points around C. f ; r θ The penalty term representing the deviation between the direction angle of the drone and the sub-target point;

[0040] reward=λ1r p +λ2r g +λ3r f +λ4r θ (4)

[0041]

[0042] r f =Num f (7)

[0043]

[0044] The state machine includes five states, namely:

[0045] init: initialization state, the system starts or restarts after a collision;

[0046] detection: endpoint detection status, used to detect whether the endpoint is in the local field of view;

[0047] targeting: sub-target point decision status, the system decides the next sub-target point based on the surrounding environment;

[0048] Obstacle avoidance: obstacle avoidance state, the system executes obstacle avoidance strategy after detecting an obstacle;

[0049] collision: collision state, the system enters this state when a collision is detected;

[0050] Adjust state switching based on the following trigger events:

[0051] The system starts or reaches the previous target point, switching from init to detection;

[0052] The destination is not in the local field of view, so switch from detection to targeting;

[0053] The end point is in the local field of view, switching from detection to obstacle avoidance;

[0054] Generate a sub-target point and switch from targeting to obstacle avoidance;

[0055] Reach the destination successfully and switch from obstacle avoidance to init;

[0056] Successfully reach the sub-target point and switch from obstacle avoidance to targeting;

[0057] Determine collision with the environment and switch from obstacle avoidance to collision;

[0058] After system recovery or fault handling is complete, the system switches from collision to init.

[0059] Beneficial effects of the present invention:

[0060] (1) Hierarchical decision-making enhances autonomous navigation capabilities

[0061] Common deep learning-based drone navigation methods usually rely on a single-level decision-making mechanism, which is difficult to cope with dynamically changing or unstructured scenes. However, the present invention uses a hierarchical decision-making method. The low-level obstacle avoidance decision module focuses on local obstacle avoidance, and the high-level sub-target point decision network is responsible for global path planning. It can adaptively balance the task requirements of local obstacle avoidance and global navigation, so that drones can make adaptive adjustments more autonomously when performing tasks.

[0062] (2) Synergy between decision-making efficiency and energy consumption optimization

[0063] Current methods based on large models or traditional mapping and navigation have high requirements for computing and storage, and resource-constrained drones find it difficult to maintain stable performance in long-term missions. Unlike these methods, the present invention combines pulse neural networks and hierarchical decision-making frameworks for the first time. The proposed hierarchical pulse reinforcement learning model reduces computing resource consumption and power consumption while ensuring decision-making efficiency. Through the synergy of efficient decision-making and low power consumption, drones can maintain long working time and stable performance in long-term autonomous navigation missions, especially in scenarios with strict power consumption restrictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a framework diagram of the drone obstacle avoidance navigation method based on hierarchical pulse reinforcement learning.

[0065] Figure 2 It is the network structure diagram of the obstacle avoidance decision module.

[0066] Figure 3 It is the network structure diagram of the sub-target point decision module. DETAILED DESCRIPTION

[0067] The present invention will be further described in detail below in conjunction with specific implementation modes, but the present invention is not limited to the specific implementation modes.

[0068] An obstacle avoidance navigation method for unmanned aerial vehicles based on hierarchical impulse reinforcement learning includes a low-level obstacle avoidance decision module and a high-level sub-target point decision network.

[0069] The goal of the obstacle avoidance decision module is to find the best execution action based on the current environmental information, avoiding obstacles while approaching the target point to maximize the specific reward of the task. The three elements of reinforcement learning are as follows:

[0070] The state is defined as {W, V, P, where W = {d w ,θ w} is the relative distance from the UAV to the sub-target point and the relative direction from the UAV to the sub-target point; V = {v l ,v a} is the linear velocity and angular velocity of the drone; P is the data of the forward laser radar, which contains the laser scanning results of 180° forward at three consecutive moments, P = {P t ,P t-1 ,P t-2}, each P t Contains several laser scanning points,

[0071] Action action is defined as {v l ,v a}, the agent is a multi-rotor flying drone, v l and v a represent the horizontal linear velocity and angular velocity respectively.

[0072] The reward is defined as in formula (1), r g represents the positive reward obtained by the drone when it reaches the target; r o represents the reward after a collision. α is a scaling factor, d o Indicates the distance from the drone to the obstacle, t g and t o Represents the threshold of reaching the drone’s target or obstacle, respectively. The reward provides effective positive and negative feedback for the drone in the process of obstacle avoidance.

[0073]

[0074] The obstacle avoidance decision module uses the TD3 algorithm as the reinforcement learning algorithm and uses a spiking neural network as the Actor network. The neuron model uses the LIF (Leaky Integrated and Fire) spiking neuron model. The pulse emission process is as follows:

[0075]

[0076] in, is the input pulse flow, w ij is the weight, is the neuron membrane potential, β is the membrane potential attenuation factor, b j is the bias term. When the membrane potential exceeds the threshold V th When , an output pulse is generated:

[0077]

[0078] The drone obstacle avoidance decision module is used to extract features from sensor data. The input X includes laser data, drone speed and target relative information, which is expressed as follows:

[0079] X={G,V,L}

[0080] Use Poisson encoding to convert the input into a pulse sequence X s = {G s ,V s ,L s}, the laser data is extracted through the spiking convolution layer (Spiking Conv1d) to obtain the spatiotemporal features F L =SpikingConv1d(L s ), F L and {G s ,Vs} to obtain The action is obtained through the spiking fully connected layer (Spiking MLP).

[0081] action=SpikingMLP(F fusion )

[0082] The goal of the sub-target point decision network is to mainly explore and obtain the best sampling points within the visible range. The sampling points serve as the input of the low-level obstacle avoidance decision module to guide the intelligent agent to reach the target point safely and efficiently. The three elements of the sub-target point decision network reinforcement learning are as follows:

[0083] The state is defined as {G,O}, where G = {d g ,θ g} is the relative distance from the drone to the target and the relative direction from the drone to the target; O is the data of the omnidirectional laser radar, which contains several laser scanning points,

[0084] Action action is defined as C = {d c ,θ c}, is the relative distance and relative direction between the sub-target point C and the current robot.

[0085] The reward is defined as in formula (4), where r p The penalty term representing the distance between the sub-target point and the current position of the UAV, d max is the maximum reachable distance, d p is the distance between the current position of the UAV and the candidate sub-target point C; r g The reward for the candidate sub-target point C to guide the drone to the target position G depends on the Euclidean distance d between C and G. g ; r f Indicates the exploration value of the candidate sub-target point, which depends on the number of frontier points around C. f ; r θ A penalty term that represents the deviation between the robot's orientation and the direction angle of the sub-goal point.

[0086] reward=λ1r p +λ2r g +λ3r f +λ4r θ (4)

[0087]

[0088] r f =Num f (7)

[0089]

[0090] The sub-goal point decision network reinforcement learning algorithm uses the DQN algorithm, which also uses the spiking neural network and LIF neuron model. The input X includes {G, L}. G passes through the spiking fully connected layer (SpikingMLP) to obtain F G =SpikingMLP(G), L is passed through a pulse one-dimensional convolution layer (SpikingConv1d) to obtain F L =SpikingConv1d(L), M is the local grid map obtained by transforming the radar data L, and F is obtained by the pulse two-dimensional convolution (SpikingConv2d) M =SpikingConv2d(M).

[0091]

[0092] The three features are concatenated to get F fusion , input to the spiking fully connected layer (SpikingMLP), and output the score vector S:

[0093] S = SpikingMLP(F fusion )

[0094] Finally, the action mask M is used to mask the invalid area, and the candidate sub-target point C with the highest score is selected. ⊙ represents the element-by-element multiplication:

[0095] C = argmax(S⊙M)

[0096] In the autonomous navigation system, the state machine plays a core role in scheduling and coordination. Through the design of the state machine, the system can switch between different states in an orderly manner, thereby coordinating the operation of the obstacle avoidance decision module and the sub-goal decision module to ensure that the system completes the task efficiently and stably. In this system, there are the following states:

[0097] init: Initialization state, the system starts or restarts after a crash.

[0098] detection: endpoint detection status, used to detect whether the endpoint is in the local field of view.

[0099] targeting: Sub-target point decision status. The system decides the next sub-target point based on the surrounding environment.

[0100] Obstacle avoidance: obstacle avoidance state, the system executes obstacle avoidance strategy after detecting an obstacle.

[0101] collision: collision state. The system enters this state when a collision is detected.

[0102] The switching between states is triggered by events and conditions. The specific switching logic is as follows:

[0103]

[0104]

[0105] A UAV obstacle avoidance navigation method based on hierarchical pulse reinforcement learning includes training and testing scene construction and network model training and testing.

[0106] (1) Training and testing scenario construction

[0107] The simulation environment of the present invention is built based on Gazebo of ROS.

[0108] The training is conducted in stages. The drone navigates in four increasingly complex and difficult environments in turn. The size of each training environment is 12m*12m*5m. At the same time, obstacles are randomly added to the scene in each episode to increase the complexity and uncertainty of the environment. In order to enable the drone to learn the ability to avoid obstacles, the starting position and target position of the navigation are randomly sampled in specific areas of the four environments.

[0109] The test scene is a closed environment of 20m*20m*5m in size. Obstacles of various shapes and sizes are set inside the scene, including rectangles, triangles, diamonds and irregular polygons, etc., which can simulate complex situations in the real environment, such as narrow passages, blind spots, stacked obstacles and open spaces.

[0110] (2) Network training

[0111] For the obstacle avoidance decision module, the reinforcement learning algorithm uses TD3, where the delay policy update frequency is 2. The episode is set to 1000, the optimizer uses Adam, and the learning rate is 1e -4 , the learning rate decay is 0.99, the batch size of the model is 256, and the experience replay buffer size is 1e 5 , the action noise is initially set to 0.1 and decays exponentially. In addition, the scaling factor in the reinforcement learning reward is set to 10.

[0112] For the sub-goal point decision network, the reinforcement learning algorithm uses DQN, the episode is set to 400, and the other parameters are the same as TD3. In addition, the weights of the four parts of the reinforcement learning reward are [0.3, 0.3, 0.2, 0.2].

[0113] For the spiking neural network, the time period T is set to 5, the pulse emission threshold is set to 0.5, the membrane potential attenuation coefficient is 0.9, the training method uses STBP, and the encoding method uses Poisson coding.

[0114] A test consists of 200 episodes. The starting point and end point of each episode are randomly generated, and the distance between the two points is more than 15m. Each episode has three results: 1) Success, the agent successfully reaches the target location; 2) Failure, the agent collides with an obstacle; 3) Timeout, the execution steps exceed 1000 times.

Claims

1. A drone obstacle avoidance navigation method based on hierarchical pulse reinforcement learning, characterized in that: A hierarchical reinforcement learning framework is used for obstacle avoidance navigation, wherein the reinforcement learning models all use pulse neural networks to obtain a hierarchical pulse reinforcement learning model; the hierarchical pulse reinforcement learning model includes a low-level obstacle avoidance decision module and a high-level sub-target point decision network; the obstacle avoidance decision module is used for local obstacle avoidance; the sub-target point decision network is used for global path planning; a state machine is set to coordinate the operation of the obstacle avoidance decision module and the sub-target point decision network.

2. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 1 is characterized in that: The hierarchical pulse reinforcement learning model collects sensor data, target point data and drone status; the sensor data, target point data and drone status are processed by a sub-target decision module to obtain sub-target points; the sub-target points together with the sensor data and drone status are input into an obstacle avoidance decision module to obtain obstacle avoidance instructions, thereby guiding the drone to perform obstacle avoidance actions.

3. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 1 is characterized in that: The output of the obstacle avoidance decision module is a continuous action, in which the Actor network adopts a spiking neural network SNN, and the Critic network adopts an artificial neural network ANN; the output of the sub-target point decision network is a discrete action, in which the Q network adopts a spiking neural network SNN; the neuron model of the spiking neural network selects the LIF model, and its pulse emission process is as follows: in, is the input pulse flow, w ij is the weight, is the neuron membrane potential, β is the membrane potential attenuation factor, b j is the bias term; when the membrane potential exceeds the threshold V th When the output pulse is:

4. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical pulse reinforcement learning according to claim 3 is characterized in that: The obstacle avoidance decision module adopts the TD3 algorithm; the sub-target point decision network adopts the DQN algorithm.

5. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 1 or 2, characterized in that: The reinforcement learning goal of the obstacle avoidance decision module is to find the optimal execution action based on the current environmental information, avoiding obstacles while approaching the sub-target point to maximize the specific reward of the task; The input X of the obstacle avoidance decision module includes the relative position between the drone and the sub-target point, the drone speed and the laser data, which is expressed as follows: X={W,V,P} Where W = {d w ,θ w } is the relative distance from the UAV to the sub-target point and the relative direction from the UAV to the sub-target point; V = {v l , v a } is the linear velocity and angular velocity of the drone; P is the data of the forward laser radar, which contains the laser scanning results of 180° forward at three consecutive moments, P = {P t , P t-1 , P t-2 }, each P t Contains several laser scanning points, For the Actor network, Poisson encoding is used to convert the input into a pulse sequence X s = {W s , V s , P s }, the laser data is extracted through the pulse one-dimensional convolution layer to obtain the spatiotemporal features F P =SpikingConv1d(P s ), F P and {W s , V s } to obtain The action is obtained through the pulse fully connected layer; action=SpikingMLP(F fusion ) For the Critic network, a conventional artificial neural network is used. The network structure is exactly the same as the Actor network. Actions are added during series connection, and Q value 1 and Q value 2 are obtained through full connection. The smaller one is taken as the Q value.

6. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 5 is characterized in that: The three elements of the obstacle avoidance decision module reinforcement learning are as follows: The state is defined as {W, V, P}; the action is defined as {v l , v a }; The reward is defined as in formula (3), r g represents the positive reward obtained by the drone when it reaches the target; r o represents the negative reward obtained after the drone collides; α is a scaling factor, d o Indicates the distance from the drone to the obstacle, t g Indicates the threshold for the drone to reach the target; t o Indicates the threshold for reaching an obstacle; the reward provides effective positive and negative feedback for the drone during the obstacle avoidance process; 7. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 1 or 2, characterized in that: The goal of the sub-target point decision network reinforcement learning is to explore and obtain the best sampling points within the visible range. The sampling points are used as inputs of the obstacle avoidance decision module to guide the drone to the target point. The input of the sub-target point decision network includes laser data and the relative position between the drone and the target point, which is expressed as follows: X={G,O,M} Where G = {d g ,θ g } is the relative distance from the drone to the target and the relative direction from the drone to the target; O is the data of the omnidirectional laser radar, and the laser radar data includes several laser scanning points, M is the local grid map obtained by transforming the lidar data O; Use Poisson encoding to convert the input into a pulse sequence X s = {G s , O s , M s }; G s After the pulse fully connected layer, we get F G =SpikingMLP(G s );O s After the pulse one-dimensional convolution layer, we get F O =SpikingConv1d(O s );M s After the pulse two-dimensional convolution, we get F M =SpikingConv2d(M s ); The three features are concatenated to get F fusion , input into the pulse fully connected layer, and output the score vector S: S=SpikingMLP(F fusion ) Finally, the action mask M is used to mask the invalid area, and the candidate sub-target point C with the highest score is selected as the sampling point. ⊙ represents the element-by-element multiplication: C = argmax(S ⊙M).

8. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 7 is characterized in that: The three elements of the sub-goal point decision network reinforcement learning are as follows: The state is defined as {G, O}; the action is defined as C = {d c ,θ c }, sub-target point C represents the relative distance and relative direction between the current UAV and the target point; The reward is defined as in formula (4), where r p The penalty term representing the distance between the sub-target point and the current position of the UAV, d max is the maximum reachable distance, d p is the distance between the current position of the UAV and the candidate sub-target point C; r g The reward for the candidate sub-target point C to guide the drone to the target position G depends on the Euclidean distance d between C and G. g ; r f Indicates the exploration value of the candidate sub-target point, which depends on the number of boundary points around C. r ; r θ The penalty term representing the deviation between the direction angle of the drone and the sub-target point; reward=λ1r p +λ2r g +λ3r f +λ4r θ (4) r f =Number f (7) 9. The unmanned aerial vehicle obstacle avoidance navigation method based on hierarchical impulse reinforcement learning according to claim 1 is characterized in that: The state machine includes five states, namely: init: initialization state, the system starts or restarts after a collision; detection: endpoint detection status, used to detect whether the endpoint is in the local field of view; targeting: sub-target point decision status, the system decides the next sub-target point based on the surrounding environment; Obstacle avoidance: obstacle avoidance state, the system executes obstacle avoidance strategy after detecting an obstacle; collision: collision state, the system enters this state when a collision is detected; Adjust state switching based on the following trigger events: The system starts or reaches the previous target point, switching from init to detection; The destination is not in the local field of view, so switch from detection to targeting; The end point is in the local field of view, switching from detection to obstacle avoidance; Generate a sub-target point and switch from targeting to obstacle avoidance; Reach the destination successfully and switch from obstacle avoidance to init; Successfully reach the sub-target point and switch from obstacle avoidance to targeting; Determine collision with the environment and switch from obstacle avoidance to collision; After system recovery or fault handling is complete, the system switches from collision to init.

Citation Information

Cited By

  • Mobile robot obstacle avoidance motion planning method based on pulse hybrid reinforcement learning

    CN120406474A

  • Obstacle avoidance motion planning method for mobile robots based on pulse hybrid reinforcement learning

    CN120406474B

  • Man-machine hybrid autonomous navigation system in unknown dynamic environment

    CN120467325A

  • Human-robot hybrid autonomous navigation system in unknown dynamic environment

    CN120467325B

  • Layered adaptive safety reinforcement learning system and method applied to patrol robot

    CN121300346A