Unmanned aerial vehicle navigation obstacle avoidance method based on selective imitation enhancement deep reinforcement learning
By constructing a non-expert navigation strategy based on artificial potential fields and selective imitation enhancement learning algorithm, combined with a Q-value-driven dynamic decision-making mechanism, the problem of autonomous navigation and obstacle avoidance of drones under sparse reward conditions is solved, and efficient navigation and obstacle avoidance effects are achieved.
Patent Information
- Application Number
- CN202510391017.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
The existing drone navigation obstacle avoidance technology is difficult to achieve efficient autonomous navigation and obstacle avoidance under sparse reward conditions, and the existing deep reinforcement learning methods rely on intensive reward functions designed by manual design, resulting in high engineering costs and prone to strategy suboptimal problems.
By accessing the privileged state information that cannot be obtained by the drone, an artificial potential field generates non-expert navigation strategy is built, and deep reinforcement learning and imitation learning are integrated to build a selective imitation-enhanced actuator-evaluator learning algorithm, and combined with a Q-value-driven dynamic decision-making mechanism to optimize the navigation strategy of the drone.
It significantly improves the autonomous navigation and obstacle avoidance efficiency of drones under sparse reward conditions, improves flight safety and generalization capabilities, and realizes end-to-end mapping learning from noise-containing sensor input to flight control instructions.
Smart Images

Figure CN120353234A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence and unmanned aerial vehicle (UAV) system technology, and particularly relates to a UAV navigation and obstacle avoidance method based on selective imitation enhanced deep reinforcement learning. Background Art
[0002] The wide application of UAVs in complex scenarios such as agricultural monitoring and disaster relief has put forward higher requirements for their autonomous navigation and obstacle avoidance capabilities. Current UAV navigation and obstacle avoidance technologies have made remarkable progress. Traditional methods based on model predictive control and obstacle Lyapunov functions rely on environmental and dynamic prior knowledge. Although they perform well in structured scenarios, they are difficult to adapt to complex dynamic environments lacking prior information. Deep reinforcement learning autonomously optimizes control strategies through trial-and-error interactions to adapt to dynamic environments. However, existing deep reinforcement learning navigation methods rely on manually designed dense reward functions, which not only have high engineering costs but also are more likely to cause suboptimal policy problems due to conflicting reward objectives. The sparse reward mechanism provides a more general solution by simplifying the reward design process. However, its lack of immediate feedback characteristics leads to slow convergence of agents and difficulty in obtaining optimal policies. For this reason, methods that integrate imitation learning and deep reinforcement learning have emerged, accelerating the learning process by drawing on expert experience. However, the strong dependence of existing technologies on high-quality expert data severely restricts practical applications. How to use non-expert data to assist policy learning in sparse reward scenarios remains a key problem to be solved urgently. Summary of the Invention
[0003] Object of the Invention: The object of the present invention is to provide a UAV navigation and obstacle avoidance method based on selective imitation enhanced deep reinforcement learning. It is used to solve the navigation and obstacle avoidance problems of UAVs under sparse reward conditions, and has superiority in significantly improving the deep reinforcement learning efficiency, flight safety, and generalization ability of UAV autonomous navigation and obstacle avoidance.
[0004] Technical Solution: A UAV navigation and obstacle avoidance method based on selective imitation enhanced deep reinforcement learning of the present invention includes the following steps:
[0005] Step 1: By accessing privileged state information that cannot be obtained by the UAV, construct an artificial potential field to generate a non-expert navigation strategy;
[0006] Step 2: Integrate deep reinforcement learning and imitation learning, construct a selective imitation enhanced actor-critic learning algorithm, and while autonomously learning in a sparse reward environment, intelligently screen and integrate non-expert demonstration actions to optimize the non-expert navigation strategy generated by the artificial potential field;
[0007] Step 3: Adopt a Q-value-driven dynamic decision-making mechanism. By comparing the expected benefits of the instructor's recommended actions and the actions generated by the learner, determine the actions of the UAV to achieve the exploration-exploitation balance and avoid ineffective exploration and local optimal decisions.
[0008] Further, Step 1 specifically includes the following steps:
[0009] Step 1.1: Construct a non-expert instructor module to obtain additional information from the environment. This information is invisible to the agent, and such information is called privileged state information.
[0010] Step 1.2: Construct a non-expert policy generation mechanism based on the artificial potential field. Calculate the target attraction and obstacle repulsion forces through the privileged state information.
[0011] Step 1.3: For the dynamic characteristics of the fixed-wing UAV, convert the resultant potential force into the forward acceleration u t and the heading angle rate ω t , so that the agent can navigate to the preset target point while avoiding collisions.
[0012] Further, Step 1.2 is specifically as follows: Based on the privileged state information, adopt the artificial potential field framework to construct a non-expert navigation policy. The target point generates an attractive potential field to pull the UAV towards the target. The attraction force F att is given by the formula:
[0013]
[0014] where p t is the current position of the UAV, p g is the target position, and k att is a positive gain coefficient. This potential field guides the UAV to move asymptotically towards the target point along the gravitational gradient direction.
[0015] The obstacle realizes collision avoidance through the repulsive potential field mechanism: When the UAV enters the action radius r o of the obstacle p o , that is, the privileged state, the repulsive force F rep is defined as follows:
[0016]
[0017] where η(p t ) = R p / (||p t - p o || - r o ) - 1, and k rep is another positive gain coefficient. The intensity of this force field is negatively correlated with the relative distance between the UAV and the obstacle and decays to zero when the distance exceeds R p .
[0018] Further, step 1.3 is specifically as follows: The navigation decision-making process is realized through the coupling of multiple force fields: The total acting force F is the vector sum of the attractive force and the repulsive forces of each obstacle:
[0019]
[0020] where represents the repulsive force of the i-th obstacle;
[0021] Determine the non-expert actions of the UAV according to the total acting force: For a fixed-wing UAV, its actions of forward acceleration and heading angle rate are calculated by the following formula:
[0022]
[0023] where, ψ t and v t respectively represent the heading and speed of the UAV; respectively represent the accelerations of the UAV in the x-axis and y-axis directions.
[0024] Further, step 2 specifically includes the following steps:
[0025] Step 2.1: Integrate deep reinforcement learning and imitation learning to construct an actor-critic learning algorithm enhanced by selective imitation. By selectively imitating non-expert instructors while learning sparse rewards, fuse the advantages of imitation learning and deep reinforcement learning to achieve policy learning, and combine imitation learning to improve the learning process;
[0026] Step 2.2: Construct an actor network: The actor network outputs a deterministic policy μ(o t ∣θ μ ), which maps the observation o t to the action a t ; The parameters θ μ of the actor network are updated by combining the reinforcement learning objective and the behavior imitation objective, and the implementation includes two stages: First, calculate the mean squared error of the action as the basic imitation loss, and then perform sample-level filtering through a binary mask;
[0027] Step 2.3: Construct a critic network: As the core of value evaluation, the critic network optimizes the Q-value function Q(o t ,a t ∣θ Q ) by minimizing the temporal difference error, providing a value benchmark for policy iteration;
[0028] Step 2.4: The experience filter realizes experience optimization through a double Q-value evaluation mechanism.
[0029] Furthermore, in step 2.1, the specific selective imitation enhanced actor-critic learning algorithm is as follows: First, randomly initialize the parameters of the actor network μ(o t |θ μ ) and the critic network Q(o t , a t |θ Q ), and imitate to generate target networks μ′(o t |θ μ′ ) μ′ and Q′ with the same initial parameters; construct an experience pool with a maximum capacity of N to store state transition data; during the training process of a total of M rounds, at the beginning of each round, the agent obtains the initial observation o t from the environment; then loop to execute the following operations until reaching the target or colliding: The Q-value drives the policy to select the action a t , and after executing this action, obtain the environmental feedback reward r t and the new observation o t+1 ; evaluate the current state transition value through an experience filter, and if it satisfies , then store the state transition in the experience pool. Randomly sample N b samples from the experience pool at each step of training, and update the actor and critic networks by minimizing the actor network loss and the critic network loss respectively. The target network parameters are gradually synchronized using a soft update mechanism every C steps.
[0030] Furthermore, step 2.2 is specifically as follows: Construct the actor network to output a deterministic policy μ(o t |θ μ ), map the observation o t to the action a t , and update the actor network parameters θ μ by combining the reinforcement learning objective and the behavior imitation objective. The total loss function of the actor network is designed as follows:
[0031]
[0032] Among them, is the reinforcement learning loss, is the selective imitation loss, and λ SBC is the weight factor for balancing the two objectives;
[0033] After randomly sampling a small batch of samples from the experience pool , construct the reinforcement learning loss function based on the policy gradient:
[0034]
[0035] Among them, Q(o t , μ(o t|θ μ )) is the action value function, i.e., the Q-value function; minimizing this loss function along the gradient direction drives the optimization of the actuator network parameters towards maximizing the long-term return;
[0036] In the update of the actuator network, a behavior imitation loss is added to encourage the policy to imitate the behavior of the non-expert instructor, making the action recommended by the non-expert instructor, the action generated by the actuator, and the selective imitation loss is defined as:
[0037]
[0038] The leading term of represents and the mean squared error of, and by minimizing this error term, the policy network is forced to approximate the non-expert demonstration trajectory in the given state;
[0039] A masking mechanism is introduced to selectively use the actions of the non-expert instructor based on action quality. The masker evaluates the Q-value to determine whether the action recommended by the instructor is better than the action generated by the actuator. If the Q-value of the instructor's action is higher than that of the actuator's action, it is considered worthy of imitation. Formally, the masking criterion can be expressed as:
[0040]
[0041] By integrating the masking mechanism, the actuator network triggers imitation learning when the value of the guided action is better than the autonomous policy, while ignoring potentially harmful imitation behaviors.
[0042] Furthermore, step 2.3 is specifically as follows: Construct an evaluator network: The evaluator network evaluates the Q-value function Q(o t , a t |θ Q ) of the given state-action pair, and updates the evaluator network by minimizing the temporal difference error:
[0043]
[0044] where D represents the experience pool; o t and r t represent the observation and the reward respectively; and represent the action recommended by the non-expert instructor and the action generated by the actuator respectively; y t represents the target Q-value.
[0045] If the action provided by the instructor produces a better Q-value, it indicates that this action optimizes the future state transition; therefore, a better state transition is used to construct the target Q-value y t , and its calculation method is:
[0046]
[0047] Synchronously consider the expected returns of the next - moment policy action and the instructor's action through the max function. The max operation obtains a better Q - value estimate and accelerates the training process; maintain independent target networks for the actor and the critic, denoted as μ′(o t ∣θ μ′ ) and Q′(o t ,a t ∣θ Q′ ); The target network parameters are updated as follows:
[0048] θ Q′ ←τθ Q +(1 - τ)θ Q′
[0049] θ μ′ ←τθ μ +(1 - τ)θ μ′
[0050] where τ is the soft - update rate.
[0051] Furthermore, step 2.4 is specifically as follows: Construct an experience filter: In the selective imitation enhanced actor - critic learning algorithm, the experience filter screens high - value experiences and stores them in the experience pool. Selective imitation enhanced actor - critic learning uses the experience pool to store the historical experience data generated by the interaction between the agent and the environment. Each piece of experience data is stored in the form of a state - action - reward tuple as By randomly sampling a small batch of transition data from the buffer, stable updates of the actor and critic networks are achieved; this experience replay mechanism improves the learning efficiency and training stability of the algorithm by eliminating the correlation between sequential samples;
[0052] Design an experience filter module for screening and storing state transition data with learning value; during the interaction between the agent and the environment, the system evaluates the value of each state transition by comparing the expected returns of the actions generated by the reinforcement learning policy and the auxiliary guidance policy. The Q - values of the actions corresponding to the two policies are calculated respectively through the critic network. The experience filter decides whether to store the current state transition data in the experience pool according to the Q - value comparison result. When the Q - value of the action generated by the deep reinforcement learning policy is better than the guidance policy, that is , the state transition will be retained, and it is determined that the state transition data has learning value. At this time, the system will retain the data sample, otherwise it will be eliminated through the filter mechanism.
[0053] Furthermore, step 3 specifically includes the following steps:
[0054] Step 3.1: The action decision-making module determines the optimal action selection of the agent by integrating the current learning strategy and the instructor signal, and reuses the evaluator network to construct a Q-value optimization strategy.
[0055] Step 3.2: In the action optimization process, the Q-values of different actions are quantitatively evaluated, and an exploration factor is introduced to achieve decision balance. The evaluator network synchronously calculates the Q-values of the current policy action and the guidance policy action : When the condition is satisfied, the instructor's recommended action is preferentially executed otherwise, the action generated by the strategy is adopted The system superimposes Gaussian noise sampled from a random process on the selected actions to expand the agent's behavior space and prevent the strategy from falling into local optima.
[0056] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages:
[0057] (1) The present invention proposes an innovative end-to-end learning framework for the problem of UAV autonomous navigation under sparse reward conditions. By integrating imitation learning and deep reinforcement learning techniques, this framework realizes end-to-end mapping learning from noisy sensor inputs to flight control commands.
[0058] (2) The present invention proposes a selective imitation enhanced actor-evaluator learning algorithm to learn the optimal navigation strategy. By integrating an experience filter and a Q-value-based action selector to selectively imitate non-expert strategies, both the sample efficiency and the learning performance are significantly improved. Brief Description of the Drawings
[0059] Figure 1 is a selective imitation enhanced deep reinforcement learning method for UAV navigation and obstacle avoidance in a sparse reward scenario;
[0060] Figure 2 is a schematic diagram of the actor and evaluator network architectures in an embodiment of the present invention;
[0061] Figure 3 is a hardware-in-the-loop simulation flight trajectory diagram of a fixed-wing UAV in the present invention. Detailed Embodiments
[0062] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0063] As Figure 1 shown, the present invention is a UAV autonomous navigation and obstacle avoidance control method based on a selective imitation enhanced actor-evaluator learning method, and its specific content includes:
[0064] 1. By accessing the privileged state information that cannot be obtained by the learning agent, construct a non-expert instructor module based on the artificial potential field, and generate action instructions with potential advantages based on heuristic rules.
[0065] 1.1 The instructor module assists the intelligent agent in learning and optimization by constructing a non-expert strategy. This module can access the environmental privileged information that the intelligent agent cannot obtain (including the exact coordinates and action radius of obstacles), and such information is defined as the privileged state, providing prior knowledge support for strategy optimization.
[0066] 1.2 Based on the privileged state information, construct a non-expert navigation strategy using the artificial potential field framework. As a classic method in the field of robot path planning, the core concept of the artificial potential field method is to regard the UAV as a particle moving in the potential field generated by obstacles and targets. By constructing virtual forces, the UAV is guided to move towards the target while staying away from obstacles, thus achieving safe and efficient navigation.
[0067] The calculation process of non-expert actions based on the artificial potential field method: The target point generates an attractive potential field to pull the UAV towards the target, and the attractive force F att is defined by the formula
[0068]
[0069] where p t is the current position of the UAV, p g is the target position, and k att is the positive gain coefficient. This potential field guides the UAV to move asymptotically towards the target point along the gravitational gradient direction.
[0070] The obstacle realizes collision avoidance through the repulsive potential field mechanism: When the UAV enters the action radius r o of the obstacle p o (i.e., the privileged state), the repulsive force F rep is defined as follows:
[0071]
[0072] where η(p t ) = R p / (||p t - p o || - r o ), and k rep is the positive gain coefficient. The intensity of this force field is negatively correlated with the relative distance between the UAV and the obstacle, and decays to zero when the distance exceeds R p .
[0073] 1.3 The navigation decision-making process is achieved through the coupling of multiple force fields: The total acting force F is the vector sum of the attractive force and the repulsive forces of each obstacle:
[0074]
[0075] where represents the repulsive force of the i-th obstacle.
[0076] Determine the non-expert actions of the UAV according to the total force: For a fixed-wing UAV, its actions (i.e., forward acceleration and heading angle rate) are calculated by the following formula:
[0077]
[0078] where, ψ t and v t represent the heading and speed of the UAV, respectively; represent the accelerations of the UAV in the x-axis and y-axis directions, respectively.
[0079] 2. Continuously optimize the navigation strategy according to the observation state and sparse reward signal by integrating imitation learning and deep reinforcement learning through real-time interaction with the environment.
[0080] 2.1 Integrate deep reinforcement learning and imitation learning, utilize the non-expert actions provided by the instructor, and improve the sample efficiency of the learner. Although the instructor based on the artificial potential field can generate preliminary navigation instructions, its inherent local minimum defect and parameter sensitivity (i.e., k att and k rep ), its essence still belongs to the non-expert level. Therefore, construct an actor-critic algorithm enhanced by selective imitation. While autonomously learning in a sparse reward environment, intelligently screen and integrate non-expert demonstration actions, and integrate the advantages of imitation learning and deep reinforcement learning to achieve more effective policy learning.
[0081] Selective imitation enhanced actor-critic learning algorithm: First, randomly initialize the parameters of the actor network μ(o t ∣θ μ ) and the critic network Q(o t ,a t ∣θ Q ), and imitate to generate the target networks μ′(o t ∣θ μ′ ) μ′ and Q′ with the same initial parameters. Construct an experience pool D with a maximum capacity of N to store state transition data. During the training process of a total of M episodes, at the beginning of each episode, the agent obtains the initial observation o t from the environment. Subsequently, loop and execute the following operations until reaching the target or colliding: The Q value drives the policy to select the action a t , after executing this action, obtain the environmental feedback reward r t and the new observation o t+1 . Evaluate the current state transition value through the experience filter. If it satisfies Then the state transition is stored in the experience pool. At each step of training, N samples are randomly drawn from the experience pool b samples, and the actor and critic networks are updated by minimizing the actor network loss and the critic network loss respectively. The target network parameters are gradually synchronized every C steps using a soft update mechanism.
[0082] The Selective Imitation Enhanced Actor-Critic (SIAC) algorithm is an extension of the Deep Deterministic Policy Gradient (DDPG) algorithm that incorporates behavior imitation to improve the learning process (especially in scenarios where non-expert demonstrator online demonstrations are available). The core modules include an actor network, a critic network, an experience filter, and an experience pool.
[0083] In this embodiment, the network architectures of the actor and critic are as Figure 2 shown, consisting of three fully connected (FC) layers, all using the rectified linear unit (ReLU) activation function
[0084] 2.2 Construct the actor network to output a deterministic policy μ(o t |θ μ ), which maps the observation o t to the action a t . The actor network parameters θ μ are updated by combining the reinforcement learning objective and the behavior imitation objective. The total loss function of the actor network is designed as follows:
[0085]
[0086] where is the reinforcement learning loss, is the selective imitation loss, and λ SBC is the weight factor that balances the two objectives.
[0087] After randomly sampling a mini-batch of samples from the experience pool , the reinforcement learning loss function is constructed based on policy gradients:
[0088]
[0089] where Q(o t , μ(o t |θ μ )) is the action value function (i.e., the Q-value function). Minimizing this loss function along the gradient direction drives the actor network parameters to optimize towards maximizing the long-term return.
[0090] In addition, a behavior imitation loss is added to the actor network update to encourage the policy to imitate the behavior of the non-expert demonstrator. Let be the non-expert demonstrator's recommended action, and be the action generated by the actor. The selective imitation loss is defined as:
[0091]
[0092] The leading term of and the mean squared error. By minimizing this error term, the policy network is forced to approximate the non-expert demonstration trajectory in a given state.
[0093] Due to the non-expert nature of the instructor, not all of its actions are worthy of imitation. Therefore, a masking mechanism is introduced. Different from weighting the behavioral imitation loss through reward information, the non-expert instructor's actions are selectively used based on action quality. The masker determines whether the instructor's recommended action is better than the action generated by the executor by evaluating the Q-value. If the Q-value of the instructor's action is higher than that of the executor's action, it is considered worthy of imitation. Formally, the masking criterion can be expressed as:
[0094]
[0095] By integrating the masking mechanism, the executor network can trigger imitation learning when the value of the guiding action is better than the autonomous policy, while ignoring potentially harmful imitation behaviors.
[0096] 2.3 Constructing the evaluator network: The evaluator network evaluates the Q-value function Q(o t , a t ∣θ Q ) of a given state-action pair. The evaluator network is updated by minimizing the temporal difference error:
[0097]
[0098] If the action provided by the instructor produces a better Q-value, it indicates that this action optimizes the future state transition; therefore, a better state transition is used to construct the target value y t , and its calculation method is:
[0099]
[0100] By taking the maximum value function, the expected returns of the policy action at the next moment and the instructor's action are considered simultaneously. The maximum value operation can obtain a better Q-value estimate and accelerate the training process. This ensures that the intelligent agent can learn effective experience from non-expert guidance. Then, independent target networks are maintained for the executor and the evaluator, denoted as μ′(o t ∣θ μ′ ) and Q′(o t , a t ∣θ Q′ ), respectively. Compared with the main network, the target network is updated at a lower rate, aiming to provide a more stable target for learning. The update of the target network parameters is as follows:
[0101] θ Q ′ ← τθ Q +(1 - τ)θ Q ′
[0102] θ μ ′ ← τθ μ +(1 - τ)θ μ ′
[0103] where τ is the soft update rate.
[0104] 2.4 Constructing an Experience Filter: In the selective imitation enhanced actor-critic learning algorithm, the experience filter plays a core role by screening high-value experiences and storing them in the experience pool. Similar to other off-policy deep reinforcement learning algorithms, the selective imitation enhanced actor-critic learning uses an experience pool to store the historical experience data generated by the agent's interaction with the environment. Each piece of experience data is stored in the form of a state-action-reward tuple as This algorithm realizes the stable update of the actor and critic networks by randomly sampling a small batch of transition data from the buffer. This experience replay mechanism can effectively improve the learning efficiency and training stability of the algorithm by eliminating the correlation between sequential samples.
[0105] Since the path guidance provided by the non-expert instructor based on the artificial potential field is not globally optimal, this algorithm specially designs an experience filter module to screen and store the state transition data with learning value. During the interaction between the agent and the environment, the system evaluates the value of each state transition by comparing the expected rewards of the actions generated by the reinforcement learning strategy and the auxiliary guidance strategy, and calculates the Q-values of the corresponding actions of the two strategies through the critic network respectively. The experience filter decides whether to store the current state transition data in the experience pool according to the comparison result of the Q-values. When the Q-value of the action generated by the deep reinforcement learning strategy is better than the guidance strategy, that is
[0106]
[0107] , this state transition will be retained. Then it is determined that this state transition data has learning value. At this time, the system will retain this data sample, otherwise it will be excluded through the filter mechanism.
[0108] It should be emphasized that the experience filter can significantly improve the policy learning efficiency by ensuring that high-quality state transition data is stored in the experience pool. This design is especially suitable for application scenarios where the auxiliary guidance strategy is non-professional and sub-optimal.
[0109] 3. Adopt a Q-value-driven dynamic decision-making mechanism. By comparing the expected rewards of the actions recommended by the instructor and the actions generated by the learner, determine the actions of the UAV, achieve the exploration-exploitation balance, and effectively avoid ineffective exploration and local optimal decisions.
[0110] 3.1 The action decision-making module determines the optimal action selection of the agent by integrating the current learning strategy and the instructor signal. It reuses the evaluator network to construct a Q-value selection strategy without introducing additional computational modules. This design can continuously select high-value actions for execution, thereby accelerating the algorithm convergence speed and improving the strategy performance;
[0111] 3.2 The evaluator network synchronously calculates the Q-values of the current policy action and the guiding policy action : When the condition holds, the instructor's recommended action is preferentially executed Otherwise, the action generated by the policy is adopted To further coordinate exploration and exploitation, the system superimposes Gaussian noise sampled from a random process on the selected action. This design can effectively expand the agent's behavior space and prevent the policy from falling into local optima.
[0112] The specific implementation process of the proposed Q-value selection action decision-making strategy is as follows: First, obtain the current observation state ot, and generate two candidate actions through the actuator network μ(·∣θ μ ) and the non-expert instructor respectively and Subsequently, call the evaluator network Q(·,·∣θ Q ) to calculate the Q-values of the two actions ([[]] and ) in parallel. Make a selection decision by comparing the magnitudes of the two values: When , the instructor's recommended action is preferentially adopted, otherwise the action generated by the policy is selected. Finally, superimpose random noise ε that conforms to the distribution on the selected action to form the final execution action a t with both exploitation and exploration characteristics. This strategy realizes action optimization by reusing the existing network structure, effectively controlling the computational complexity while ensuring the decision-making quality.
[0113] In this embodiment, the flight trajectory of the hardware-in-the-loop simulation system is as Figure 3 shown. Six cylindrical no-fly zones with a radius of 25m (red dashed boundaries) are set in the 800m×400m corridor airspace, and the fixed-wing UAV completes collision-free navigation from the starting point to the target point.
Claims
1. A method for UAV navigation and obstacle avoidance based on selective imitation to enhance deep reinforcement learning, characterized in that, It includes the following steps: Step 1: Construct an artificial potential field to generate a non-expert navigation strategy by accessing privileged state information that cannot be obtained by the drone; Step 2: Integrate deep reinforcement learning and imitation learning to construct an actor-critic learning algorithm enhanced by selective imitation. While autonomously learning in a sparse reward environment, intelligently screen and integrate non-expert demonstration actions to optimize the artificial potential field for generating a non-expert navigation strategy; Step 3: Adopt a Q-value-driven dynamic decision-making mechanism. By comparing the expected rewards of the actions suggested by the instructor and the actions generated by the learner, determine the actions of the drone to achieve the exploration-exploitation balance and avoid ineffective exploration and local optimal decisions.
2. The method for UAV navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 1, wherein Step 1 specifically includes the following steps: Step 1.1: Construct a non-expert instructor module to obtain additional information from the environment. This information is invisible to the agent, and such information is called privileged state information; Step 1.2: Construct a non-expert policy generation mechanism based on the artificial potential field, and calculate the target attraction and obstacle repulsion forces through the privileged state information; Step 1.
3. For the dynamic characteristics of the fixed-wing UAV, convert the potential field force into the forward acceleration u t and the heading angle rate ω t , so that the agent navigates to the preset target point while avoiding collisions.
3. A method for UAV navigation and obstacle avoidance based on selective imitation to enhance deep reinforcement learning according to claim 1, characterized in that, Step 1.2 specifically includes: Based on the privileged state information, an artificial potential field framework is adopted to construct a non-expert navigation strategy. The target point generates an attractive potential field to pull the UAV towards the target, and the attraction force F generated by the target point att is given by the formula: Among them, p t is the current position of the UAV, and p g is the target position. k att is a positive gain coefficient, and this potential field guides the UAV to move asymptotically towards the target point along the gravitational gradient direction; Collision avoidance is achieved through the repulsive potential field mechanism for obstacles: When the drone enters the action radius r o of the obstacle p o within the range, i.e., the privileged state, the repulsive force F rep is defined as follows: where η(p t ) = R p / (||p t - p o || - r o ) - 1, k rep is another positive gain coefficient, and the force field intensity is negatively correlated with the relative distance between the UAV and the obstacle, and decays to zero when the distance exceeds R p .
4. The method for UAV navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 2, wherein Step 1.3 specifically is: The navigation decision-making process is realized through the coupling of multiple force fields: The total force F is the vector sum of the attraction force and the repulsion forces of each obstacle: wherein represents the repulsive force of the i-th obstacle; Determine the non-expert actions of the drone according to the total force: For a fixed-wing drone, its actions of forward acceleration and heading angle rate are calculated by the following formula: Among them, ψ t and v t respectively represent the heading and speed of the UAV; respectively represent the accelerations of the UAV in the x-axis and y-axis directions.
5. A method for UAV navigation and obstacle avoidance based on selective imitation to enhance deep reinforcement learning according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: Integrate deep reinforcement learning and imitation learning to construct an actor-critic learning algorithm enhanced by selective imitation. By selectively imitating the non-expert instructor while learning sparse rewards, integrate the advantages of imitation learning and deep reinforcement learning to achieve policy learning, and combine imitation learning to improve the learning process; Step 2.2, Construct the executor network: The executor network outputs a deterministic policy μ(o t ∣θ μ ), which maps the observation o t to the action a t ; The parameters θ μ of the executor network are updated by combining the reinforcement learning objective and the behavior imitation objective, and the implementation includes two stages: First, calculate the mean squared error of the action as the basic imitation loss, and then perform sample-level filtering through a binary mask; Step 2.3: Construct the evaluator network: As the core of value evaluation, the evaluator network optimizes the Q-value function Q(o t , a t |θ Q ) by minimizing the temporal difference error, providing a value benchmark for policy iteration; Step 2.4: The experience filter realizes experience optimization through a double Q-value evaluation mechanism.
6. The method for unmanned aerial vehicle navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 5, wherein In step 2.1, the specific selective imitation enhanced actor-critic learning algorithm is as follows: First, randomly initialize the parameters of the actor network μ(o t |θ μ ) and the critic network Q(o t , a t |θ Q ), and imitate to generate the target networks μ′(o t |θ μ′ ) μ′ and Q′ with the same initial parameters; construct an experience pool D with a maximum capacity of N to store state transition data; during the training process of a total of M rounds, at the beginning of each round, the agent obtains the initial observation o t from the environment; then loop to execute the following operations until reaching the target or colliding: The Q value drives the policy to select the action a t , and after executing this action, obtain the environmental feedback reward r t and the new observation o t+1 ; evaluate the current state transition value through the experience filter, if it satisfies , then store the state transition in the experience pool, randomly sample N b samples from the experience pool at each step of training, update the actor and critic networks respectively by minimizing the actor network loss and the critic network loss, and gradually synchronize the target network parameters using a soft update mechanism every C steps.
7. A method for UAV navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 5, characterized in that Step 2.2 specifically is: Construct an actor network to output a deterministic policy μ(o t |θ μ ), map the observation o t to the action a t , and update the actor network parameters θ μ by combining the reinforcement learning objective and the behavior imitation objective. The total loss function of the actor network is designed as follows: Among them, is the reinforcement learning loss, is the selective imitation loss, and λ SBC is the weight factor for balancing the two objectives; From the experience pool After randomly sampling a small batch of samples from it, a reinforcement learning loss function is constructed based on policy gradients: Among them, Q(o t , μ(o t | θ μ )) is the action value function, that is, the Q-value function; minimizing this loss function along the gradient direction drives the optimization of the actuator network parameters towards maximizing the long-term return; Add a behavior imitation loss to the actuator network update to encourage the policy to imitate the behavior of a non-expert instructor, and let be the action recommended by the non-expert instructor, be the action generated by the actuator. The selective imitation loss is defined as: The leading term represents and the mean squared error of, by minimizing this error term, forcing the policy network to approximate the non-expert demonstration trajectory in a given state; Introduce a masking mechanism to selectively use the actions of the non-expert instructor based on the action quality. The masker evaluates the Q-value to judge whether the action suggested by the instructor is better than the action generated by the actor. If the Q-value of the instructor's action is higher than that of the actor's action, it is considered worthy of imitation. Formally, the masking criterion can be expressed as: By integrating the masking mechanism, the actor network triggers imitation learning when the value of the guided action is better than the autonomous policy, while ignoring potentially harmful imitation behaviors.
8. The method for UAV navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 5, wherein, Step 2.3 is specifically as follows: Construct an evaluator network: The evaluator network evaluates the Q-value function Q(o t , a t |θ Q ) of a given state-action pair, and updates the evaluator network by minimizing the temporal difference error: Among them, represents the experience pool; o t and r t represent the observation and the reward respectively; and represent the action suggested by the non-expert instructor and the action generated by the actuator respectively; y t represents the target Q value; If the action provided by the instructor yields a better Q-value, it indicates that the action optimizes future state transitions; thus, a better state transition is used to construct the target value y t , and its calculation method is as follows: By using the maximum function to synchronously consider the expected rewards of the next - moment policy action and the instructor's action, the maximum operation obtains a better Q - value estimate and accelerates the training process; maintain independent target networks for the actor and the critic, denoted as μ′(o t ∣θ μ′ ) and Q′(o t ,a t ∣θ Q′ ); the target network parameters are updated as follows: θ Q′ ← τθ Q +(1 - τ)θ Q′ θ μ′ ← τθ μ +(1 - τ)θ μ′ where τ is the soft update rate.
9. The method for UAV navigation and obstacle avoidance based on selective imitation enhanced deep reinforcement learning according to claim 5, wherein, Step 2.4 is specifically as follows: Construct an experience filter: In the selective imitation enhanced actor-critic learning algorithm, the experience filter stores high-value experiences in the experience pool by screening. The selective imitation enhanced actor-critic learning uses the experience pool to store the historical experience data generated by the interaction between the agent and the environment. Each piece of experience data is stored in the form of a state-action-reward tuple as By randomly sampling a small batch of transition data from the buffer, stable updates of the actor and critic networks are achieved; this experience replay mechanism improves the learning efficiency and training stability of the algorithm by eliminating the correlation between sequential samples; Design an experience filter module to screen and store state transition data with learning value; during the interaction between the agent and the environment, the system generates the expected return of actions by comparing the reinforcement learning policy with the auxiliary guidance policy, evaluates the value of each state transition, calculates the Q-values of the actions corresponding to the two policies respectively through the evaluator network, and the experience filter decides whether to store the current state transition data in the experience pool according to the Q-value comparison result. When the Q-value of the action generated by the deep reinforcement learning policy is better than the guidance policy, that is at this time, the state transition will be retained, and it is determined that the state transition data has learning value. At this time, the system will retain the data sample, otherwise it will be excluded through the filter mechanism.
10. A method for drone navigation and obstacle avoidance based on selective imitation to enhance deep reinforcement learning according to claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: The action decision module determines the optimal action selection of the agent by comprehensively considering the current learning strategy and the instructor signal, and reuses the critic network to construct a Q-value selection strategy; Step 3.
2. In the action optimization process, the Q-values of different actions are evaluated quantitatively, and an exploration factor is introduced to achieve decision-making balance. The evaluator network synchronously calculates the Q-values of the current policy actions and the guiding policy actions : When the condition is satisfied, the action recommended by the instructor is preferentially executed otherwise, the action generated by the policy is adopted The system superimposes Gaussian noise sampled from a random process on the selected actions to expand the agent's behavior space and prevent the policy from falling into local optimality.
Citation Information
Cited By
Multi-mode attitude control method for two-wheeled mobile platform
CN121386532A
Unmanned aerial vehicle autonomous obstacle avoidance method and system based on multi-modal fusion and deep reinforcement learning
CN121527664A
Scene generalization end-to-end autonomous navigation method based on imitation-reinforcement learning
CN122083939A