Sine attenuation A3C algorithm applied to autonomous navigation training of unmanned aerial vehicle
By dynamically adjusting the prior strategy through the sinusoidal decay A3C algorithm, the problem of excessive influence of the prior strategy in UAV autonomous navigation training is solved, achieving a more stable and accurate navigation effect.
Patent Information
- Application Number
- CN202510688948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the deep reinforcement learning of UAV autonomous navigation training, the influence of prior strategies in the later stages of training is too high, resulting in mismatch of learning states and affecting the navigation effect.
The sinusoidal decay A3C algorithm is adopted. By designing a sinusoidal decay function to dynamically adjust the decay parameters of the prior strategy, the learning strategy and the prior strategy are combined to generate the behavior strategy, thereby optimizing the A3C algorithm training process.
Effectively reduce the impact of prior strategies, improve the stability and accuracy of UAV autonomous navigation training, and enhance the model's autonomous optimization capabilities.
Smart Images

Figure CN120597199A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep reinforcement learning navigation technology, and specifically to a sinusoidal decay A3C algorithm applied to autonomous navigation training of unmanned aerial vehicles. Background Art
[0002] Traditional navigation methods typically rely on precise environmental models and regularized control strategies, but often struggle to adapt to complex, dynamic, and unknown environments. Deep reinforcement learning, on the other hand, uses interactive learning between an agent and its environment to directly extract features from high-dimensional sensory data (such as geographic space and sensor data) and autonomously optimize the decision space, enabling more precise navigation.
[0003] Nowadays, deep reinforcement learning navigation technology is widely used in robot navigation, autonomous driving, drones and other fields. As the country advocates the development of a "low-altitude economy", the combination of deep reinforcement learning and drone navigation has become a current research hotspot. Its core principle is to use a deep learning model to process sensor data to obtain environmental characteristics, and then generate a navigation strategy through a reinforcement learning algorithm, and finally use the control module to convert the decision results into control instructions to complete the operation. However, deep reinforcement learning requires a large amount of interactive data for model training. In order to increase the convergence speed of the model, many researchers have adopted the more efficient A3C algorithm and incorporated prior strategies to provide early guidance for the model. However, in the later stages of training, it is easy for the prior strategy to have too high an influence and not match the learning state. Therefore, it is necessary to conduct in-depth research on how to attenuate the influence of the prior strategy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a sinusoidal decay A3C algorithm applied to autonomous navigation training of unmanned aerial vehicles.
[0005] In order to solve the above problems, the technical solution of the present invention comprises the following steps:
[0006] S1. Establish an A3C network and set the state space and action space of the drone. The state space includes the drone's flight speed, distance to the target, angle with the target, distance to obstacles, and the drone's first-person perspective angle; the action space includes the drone's flight speed and deflection angle.
[0007] S2. Design a learning strategy function and a priori strategy function, wherein both the learning strategy function and the priori strategy function are approximately represented by Gaussian distribution;
[0008] S3. Fusing the learning strategy and the prior strategy to generate a behavior strategy, wherein the mean and standard deviation of the behavior strategy are jointly determined by the mean and standard deviation of the learning strategy and the prior strategy;
[0009] S4. Design a sparse reward function to give rewards when the distance between the drone and the target is less than a preset threshold;
[0010] S5. Construct a sinusoidal decay function to dynamically adjust the prior strategy decay parameter. The sinusoidal decay function is:
[0011]
[0012] Among them, α0 represents the initial value of the prior strategy decay parameter, t is the current training number, and T is the total training number.
[0013] S6. Switch the attenuation mode according to the reward threshold. When the cumulative reward value is lower than the threshold, a fixed attenuation parameter is used. When the cumulative reward value reaches or exceeds the threshold, a sinusoidal attenuation function is used to adjust the attenuation parameter.
[0014] Furthermore, the Gaussian distribution of the learning strategy function and the prior strategy function in step S2 is expressed as:
[0015]
[0016] Among them, μ l (s) and σ l (s) is the mean and standard deviation of the learning strategy, μ p (s) and σ p (s) is the mean and standard deviation of the prior strategy.
[0017] Furthermore, the sparse reward function in step S4 is defined as:
[0018]
[0019] Among them, (x t ,y t ) is the coordinate of the drone, (x d ,y d ) is the target coordinate, and C is the preset threshold.
[0020] Furthermore, the gradient update formula of the sub-Actor network is:
[0021]
[0022] Among them, R t (n) is the cumulative reward for n steps, E c (s t ) is the evaluation value of the Critic network.
[0023] Furthermore, the adjustment of the a priori strategy attenuation parameter in step S2 is achieved by the following formula:
[0024] σ p(t+1)(s)=σ p(t) (s)(1+αN)
[0025] Where N is the number of interactions between the model and the environment, and α is the attenuation parameter.
[0026] Furthermore, the specific definition of the state space and action space in step S1 is that the state space is the drone flight speed, distance to the target, angle to the target, distance to the obstacle and the drone first-person perspective angle as state values, and the action space is the drone flight speed and deflection angle as action values.
[0027] Furthermore, the mean and standard deviation of the behavior strategy in step S2 are calculated using the following formula:
[0028]
[0029] Furthermore, the A3C network in step S1 includes a global Actor-Critic network and multiple sub-Actor-Critic networks, and the sub-Actor networks and sub-Critic networks are used to generate learning strategies and evaluate state values, respectively.
[0030] Furthermore, the loss function of the sub-critic network is:
[0031] J E =∑ t [R t (n)-E c (s t )] 2 .
[0032] Furthermore, the switching condition of the reward threshold in step S6 is:
[0033]
[0034] Among them, Z is the cumulative reward value, C Z is the preset reward threshold.
[0035] The advantages of the present invention compared with the existing technology are:
[0036] 1. The present invention provides a sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles. The algorithm designs a sinusoidal decay function. First, the initial value of the prior strategy decay parameter is determined. Then, based on the properties of the sinusoidal function, the relationship between the prior strategy decay parameter and the number of training iterations is constructed, so that the algorithm grows nonlinearly.
[0037] 2. The present invention provides a sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles. The A3C algorithm is optimized. During the A3C network training process, the prior policy influence after the sinusoidal function decay is integrated with the policy generated by the learning network, thereby achieving dynamic adjustment of the prior policy influence during the training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The present invention is a schematic diagram of an A3C algorithm framework for a sinusoidal decay A3C algorithm applied to autonomous navigation training of unmanned aerial vehicles and an optimized A3C algorithm.
[0039] Figure 2 The invention discloses a sinusoidal decay A3C algorithm applied to autonomous navigation training of unmanned aerial vehicles, and optimizes the observation map of the unmanned aerial vehicle in a two-dimensional environment of the A3C algorithm. DETAILED DESCRIPTION
[0040] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0042] A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles includes the following steps:
[0043] S1. Establish an A3C network and set the state space and action space of the drone. The state space includes the drone's flight speed, distance to the target, angle with the target, distance to obstacles, and the drone's first-person perspective angle; the action space includes the drone's flight speed and deflection angle.
[0044] S2. Design the learning strategy function and the prior strategy function. Both the learning strategy function and the prior strategy function are approximated by Gaussian distribution.
[0045] S3. The learning strategy and the prior strategy are integrated to generate a behavioral strategy. The mean and standard deviation of the behavioral strategy are determined by the mean and standard deviation of the learning strategy and the prior strategy.
[0046] S4. Design a sparse reward function to give rewards when the distance between the drone and the target is less than a preset threshold;
[0047] S5. Construct a sinusoidal decay function and dynamically adjust the decay parameters of the prior strategy. The sinusoidal decay function is:
[0048]
[0049] Among them, α0 represents the initial value of the prior strategy attenuation parameter, t is the current training number, and T is the total training number.
[0050] S6. Switch the attenuation mode according to the reward threshold. When the cumulative reward value is lower than the threshold, a fixed attenuation parameter is used. When the cumulative reward value reaches or exceeds the threshold, a sinusoidal attenuation function is used to adjust the attenuation parameter.
[0051] Furthermore, the Gaussian distribution of the learned policy function and the prior policy function in step S2 is expressed as:
[0052]
[0053] Among them, μ l (s) and σ l (s) is the mean and standard deviation of the learning strategy, μ p (s) and σ p (s) is the mean and standard deviation of the prior strategy.
[0054] Furthermore, the mean and standard deviation of the behavior strategy in step S2 are calculated using the following formula:
[0055]
[0056] Furthermore, the sparse reward function in step S4 is defined as:
[0057]
[0058] Among them, (x t ,y t ) is the coordinate of the drone, (x d ,y d ) is the target coordinate, and C is the preset threshold.
[0059] Furthermore, the A3C network in step S1 includes a global Actor-Critic network and multiple sub-Actor-Critic networks, where the sub-Actor network and sub-Critic network are used to generate learning strategies and evaluate state values, respectively.
[0060] Furthermore, the gradient update formula of the sub-Actor network is:
[0061]
[0062] Among them, R t (n) is the cumulative reward for n steps, E c (s t ) is the evaluation value of the Critic network.
[0063] Furthermore, the loss function of the sub-critic network is:
[0064] J E =∑ t [R t (n)-E c (s t )] 2 .
[0065] Furthermore, the adjustment of the prior strategy attenuation parameter in step S2 is achieved by the following formula:
[0066] σ p(t+1) (s)=σ p(t) (s)(1+αN)
[0067] Where N is the number of interactions between the model and the environment, and α is the attenuation parameter.
[0068] Furthermore, the switching condition of the reward threshold in step S6 is:
[0069]
[0070] Among them, Z is the cumulative reward value, C Z is the preset reward threshold.
[0071] Example
[0072] refer to Figure 1 and Figure 2 As shown, first set the state space and action space of the drone. In order to simplify the calculation, the design of this patent is for the case where the altitude remains unchanged. Therefore, the drone's flight speed, distance to the target, angle with the target (based on the true north direction), distance to the obstacle, and the drone's first-person perspective angle (based on the true north direction) are used as state values to form the state space; the drone's flight speed and deflection angle are used as action values to form the action space, which can be expressed by the following formula:
[0073] a t =(v t ,θ t )
[0074] Among them, a t represents the action performed by the drone at time t, v t ,θ t They represent the flight speed and deflection angle of the drone respectively.
[0075] Since the actions in the action space are continuously distributed, the learning policy function can be approximated by Gaussian distribution, as shown in the following formula:
[0076]
[0077] Among them, S l (a|s) represents the learning strategy function, a l is the action value under the learning strategy, μ l (s) and σ l (s) represent the mean and standard deviation of the action distribution under the learning strategy. This formula simplifies the learning strategy function.
[0078] Similarly, the prior strategy function Sp can be obtained as shown in the following formula:
[0079]
[0080] Its mean μ p (s) is the change in deflection angle μ p1 (s) and acceleration metric μ p2 (s) is a two-dimensional vector, which is expressed by the following formula:
[0081] μ p (s)=[μ p1 (s),μ p2 (s)]
[0082] μ in the formula p1 (s) and μ p2 (s) is obtained by the following formula:
[0083]
[0084] μ p2 (s)=min(max(v e -v,-0.3),0.3
[0085] For μ p1 (s), θ p Indicates the deflection angle of the drone, which is determined by the first-view angle β of the drone and the angle between the drone and the target. The result of subtraction is:
[0086]
[0087] When θ p >0, the drone rotates clockwise; when θ p <0, the drone rotates counterclockwise; when θ p = 0, the drone does not need to rotate. lIndicates the distance between the drone and the obstacle on the left side of the first-person perspective, d r Indicates the distance between the drone and the obstacle on the right side of the first-person perspective. If d l <d r , indicating that the drone is close to the obstacle on the left and there is a risk of collision. It needs to be moved in θ p Then rotate clockwise If d l =d r , the drone maintains θ p If d l >d r , indicating that the drone is close to the obstacle on the right and there is a risk of collision. It needs to p Then rotate counterclockwise
[0088] For μ p2 (s), v e represents the desired speed of the drone, and v represents the actual speed of the drone. This formula limits the acceleration to the interval [-0.3, 0.3] to prevent the acceleration from being too large or too small, ensuring the smoothness of the movement.
[0089] The behavior strategy ultimately executed by the drone is influenced by both the prior strategy and the learning strategy. The relationship between the three can be expressed as follows:
[0090] S(a|s)∝S l (a|s)·S p (a|s)
[0091] This formula indicates that the behavior strategy is proportional to the product of the prior strategy and the learned strategy. Since both the prior strategy and the learned strategy are approximated by Gaussian distribution, the behavior strategy also follows a Gaussian distribution, and its mean μ(s) and standard deviation σ(s) can be calculated using the following formula:
[0092]
[0093] From the above two formulas, we can see that for μ(s), if σ p (s)>>σ l (s), then Approaching 0, the behavior strategy mean is closer to the learning strategy mean; for σ(s), if σ p (s)>>σ l (s), then σ(s)≈σ l (s), the behavioral strategy variance is dominated by the learning strategy. Therefore, when it is necessary to reduce the influence of the prior strategy, this patent achieves this by increasing the standard deviation of the prior strategy. The formula used is as follows:
[0094] σ p(t+1)(s)=σ p(t) (s)(1+αN)
[0095] Where, σ p(t) (s) represents the standard deviation of the prior strategy at time t, σ p(t+1) (s) represents the standard deviation of the prior strategy at time t+1, N represents the number of times the model interacts with the external environment at time t, and α is the prior strategy attenuation parameter.
[0096] Next, we set up a reward function. This patent transforms the drone navigation problem into a goal-driven model, meaning that navigation is considered successful only if the drone successfully reaches the target location. To this end, this patent designs a sparse reward function, which is expressed as follows:
[0097]
[0098] Where r t represents the reward at time t, (x t ,y t ) represents the coordinates of the drone at time t, (x d ,y d ) represents the coordinates of the target, and C is a preset threshold with an acceptable error. As long as the distance between the drone and the target is less than or equal to C, the drone receives a reward of 1. Otherwise, no reward is received.
[0099] After designing the relevant functions, it is necessary to establish the corresponding A3C network architecture. This patent designs a pair of global Actor-Critic networks and multiple pairs of sub-Actor-Critic networks as the network structure of A3C. The sub-Actor strategy network is constructed by a multi-layer artificial neural network to approximate the learning strategy function. Its network proportion is Status t As the network input, the network is fused with the prior strategy to output the action instruction; for the sub-Critic network, it can use the information of the internal value function to estimate the direction of the gradient update in the sub-Actor network. The proportion of this network is c, and the state s t and action a t As the network input, the evaluation value E is the network output.
[0100] For the gradient update of the sub-Actor network, this patent adopts the more common gradient descent method, and the implementation formula is as follows:
[0101]
[0102] Where S l (a t |s t ) represents the learning strategy function at time t, E c (st ) represents state s t The evaluation value when R t (n)-E c (s t ) is actually the advantage function term in the A3C algorithm, which has the effect of reducing the variance of the policy gradient estimate, thereby improving learning efficiency and stability. t (n) represents the cumulative return of n steps starting at time t, and its calculation formula is as follows:
[0103]
[0104] Where ρ represents the reward discount factor, balancing the importance of current rewards and future rewards; E c (s t+n ) represents the predicted value of the critic network at time t+n. This formula can provide more accurate long-term return estimates and can control variance while reducing estimation bias.
[0105] As for the parameter update of the critic network, this patent uses the minimum mean square error method for calculation, and the formula is as follows:
[0106]
[0107] Where, J E Represents the loss function of the Critic network, namely the mean square error. This formula minimizes the predicted value E c (s t ) and the target value R t (n) gap, improve the accuracy of the Critic network, and provide a reliable advantage function term for the Actor network.
[0108] After each sub-network independently calculates the gradient, the global network parameters are updated synchronously. The specific implementation formula is as follows:
[0109]
[0110] Where, represents the global actor network percentage, c0 represents the global critic network percentage, and λ represents the learning rate of the global network update speed. This formula enables the global network to aggregate the exploration experience of all subnetworks, improving the generalization and stability of the strategy.
[0111] In A3C network training, this patent designs four fully connected layers for both the actor network and the critic network. The first two layers have 256 nodes each, and the last two layers have 128 nodes each. After training, the mean of the learned strategy output and the mean of the prior strategy can be combined to obtain the strategy mapping formula for the actor online network:
[0112] a t =μ(s t )
[0113] During actual flight tests, the drone's sensors collect flight status data in real time and transmit it to the onboard computer for calculation. The calculation results will be input into the trained Actor online network, thereby outputting flight commands in real time and realizing autonomous navigation.
[0114] When using traditional methods to decay the prior strategy, the prior strategy influence is likely to be too high in the later stages of training, thereby reducing the learning network's ability to optimize autonomously. Therefore, this patent designs a sinusoidal decay function to address the prior strategy decay problem in A3C. The formula is as follows:
[0115]
[0116] Where α0 represents the initial value of the prior strategy decay parameter, t is the current number of training times, and T is the total number of training times. This formula can make the prior strategy decay parameter increase rapidly according to the law of the increasing period of the sine function, thereby increasing the standard deviation of the prior strategy. In addition, in order to make the decay of the prior strategy more consistent with the state of the learning network, this patent designs a reward threshold. Before the reward value Z reaches a certain value, it is considered that the navigation success rate is low, and the traditional method is used to decay the prior strategy so that it can provide sufficient guidance to the model; after the reward value Z reaches a certain value, it is considered that the navigation success rate is high, and the sine decay function is used to decay the prior strategy so that the learning strategy takes the dominant position. The formula used is as follows:
[0117]
[0118] Among them C Z is a set reward threshold.
[0119] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0120] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0121] The present invention and its embodiments are described above. Such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by the above, and does not deviate from the purpose of the present invention, without inventive design, a structure and embodiment similar to the technical solution should fall within the scope of protection of the present invention.
Claims
1. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles, characterized by: The following steps are involved: S1. Establish an A3C network and set the state space and action space of the drone. The state space includes the drone's flight speed, distance to the target, angle with the target, distance to obstacles, and the drone's first-person perspective angle; the action space includes the drone's flight speed and deflection angle. S2. Design the learning strategy function and the prior strategy function. Both the learning strategy function and the prior strategy function are approximated by Gaussian distribution. S3. The learning strategy and the prior strategy are integrated to generate a behavioral strategy. The mean and standard deviation of the behavioral strategy are determined by the mean and standard deviation of the learning strategy and the prior strategy. S4. Design a sparse reward function to give rewards when the distance between the drone and the target is less than a preset threshold; S5. Construct a sinusoidal decay function and dynamically adjust the decay parameters of the prior strategy. The sinusoidal decay function is: Among them, α0 represents the initial value of the prior strategy attenuation parameter, t is the current training number, and T is the total training number. S6. Switch the attenuation mode according to the reward threshold. When the cumulative reward value is lower than the threshold, a fixed attenuation parameter is used. When the cumulative reward value reaches or exceeds the threshold, a sinusoidal attenuation function is used to adjust the attenuation parameter.
2. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The specific definitions of the state space and action space in step S1 are as follows: the state space takes the UAV flight speed, distance to the target, angle to the target, distance to the obstacle, and the UAV first-person viewing angle as state values, and the action space takes the UAV flight speed and deflection angle as action values.
3. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The Gaussian distribution of the learned policy function and the prior policy function in step S2 is expressed as: Among them, μ l (s) and σ l (s) is the mean and standard deviation of the learning strategy, μ p (s) and σ p (s) is the mean and standard deviation of the prior strategy.
4. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The mean and standard deviation of the behavior strategy in step S2 are calculated using the following formula:
5. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The sparse reward function in step S4 is defined as: Among them, (x t ,y t ) is the coordinate of the drone, (x d ,y d ) is the target coordinate, and C is the preset threshold.
6. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: In step S1, the A3C network includes a global actor-critic network and multiple sub-actor-critic networks. The sub-actor network and sub-critic network are used to generate learning strategies and evaluate state values, respectively.
7. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 6, characterized in that: The gradient update formula of the sub-Actor network is: Among them, R t (n) is the cumulative reward for n steps, E c (s t ) is the evaluation value of the Critic network.
8. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 6, characterized in that: The loss function of the sub-critic network is: J E =∑ t [R t (n)-E c (s t )] 2 。 9. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The adjustment of the prior strategy attenuation parameter in step S2 is achieved by the following formula: s p(t+1) (s)=σ p(t) (s)(1+αN) Where N is the number of interactions between the model and the environment, and α is the attenuation parameter.
10. A sinusoidal decay A3C algorithm for autonomous navigation training of unmanned aerial vehicles according to claim 1, characterized in that: The switching condition of the reward threshold in step S6 is: Among them, Z is the cumulative reward value, C Z is the preset reward threshold.