A method for controlling the posture balance of a UAV based on deep reinforcement learning

CN117170390BActive Publication Date: 2026-08-28HANGZHOU DIANZI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310273228.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2026-08-28
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

然而,由于各种外界条件(例如风力)对无人机机动性的影响难以实时估计,传统的控制方法无法实现对无人机的精准控制,因此无法满足如下对无人机控制的需求

Benefits of technology

[0039]采用上述技术方案,通过预训练与在线训练两个过程得到控制模型,能准确且高效地对无人机进行姿态控制,有效降低了无人机在自然环境下受强风或其他情况干扰时失控的风险,可以在复杂环境下对无人机姿态进行控制和矫正,提高无人机的安全性和实用性;另外,方法分为预训练与在线训练,还具备响应快、鲁棒性强,收敛速度快等特点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117170390B_ABST
    Figure CN117170390B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned plane posture balance control method based on deep reinforcement learning, comprising the following steps: S1, the m initial experience data of pre-acquisition is filled into experience pool;S2, the parameter of deep reinforcement learning network is initialized;S3, pre-train the initialized deep reinforcement learning network, obtain pre-training weight S4, online real-time training deep reinforcement learning network;S5, the control quantity of deep reinforcement learning network output controls unmanned plane.The method is used to control and correct the posture of unmanned plane in complex environment, improve the safety and practicality of unmanned plane.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control technology, specifically to a UAV attitude balance control method based on deep reinforcement learning. Background Technology

[0002] Unmanned aerial vehicles (UAVs) possess novel structural layouts and unique flight modes, utilizing aerodynamics to overcome their own weight. They can achieve hovering and fixed-path flight in the air and can be equipped with specialized instruments to perform specific functions. They have broad application prospects and research value. In harsh environments, safe and precise flight maneuverability is crucial for UAVs. However, because the impact of various external conditions (such as wind) on UAV maneuverability is difficult to estimate in real time, traditional control methods cannot achieve precise control of UAVs, thus failing to meet the following UAV control requirements. Summary of the Invention

[0003] The purpose of this invention is to address the shortcomings of existing technologies by proposing a deep reinforcement learning-based method for UAV attitude balance control, which can be used to control and correct the attitude of UAVs in complex environments, thereby improving the safety and practicality of UAVs.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] A method for attitude balance control of unmanned aerial vehicles based on deep reinforcement learning includes the following steps:

[0006] S1, The pre-collected data Initial experience data is collected and loaded into the experience pool;

[0007] S2. Initialize the parameters of the deep reinforcement learning network;

[0008] S3. The pre-trained, initialized deep reinforcement learning network obtains the pre-trained weights.

[0009] S4: Online real-time training of deep reinforcement learning networks;

[0010] S5 uses a deep reinforcement learning network to output control variables to control the drone.

[0011] Preferably, the initial empirical data is ,in, For the current drone attitude, To control the amount, For output value, This will be the next output state.

[0012] Preferably, the initialization parameters in step S2 include a weight vector. and target network weight vector .

[0013] Preferably, the pre-training method in step S3 is as follows:

[0014] S3-1, Set the delayed update base to... The number of pre-training iterations is ;

[0015]

[0016] in This represents the output function of a deep neural network;

[0017] S3-2, according to The size of the data in the experience pool is used to prioritize the data.

[0018] , which is the weighting coefficient of the empirical value;

[0019] S3-3, Sample n samples from the experience pool to obtain

[0020] S3-4, according to Calculate the loss and apply the gradient descent algorithm to... The weight vector is updated, where,

[0021]

[0022]

[0023] , For the control action corresponding to state s in the initial dataset,

[0024]

[0025] S3-5, When the number of iterations is the deferred update base When the value is an integer multiple, update Weight vector in ,make When the number of iterations reaches the set maximum number of iterations When the time comes, training is stopped and pre-training weights are obtained.

[0026] Preferably, the online training method for step S4 is as follows:

[0027] S4-1. Load pre-trained weights and set the desired output. Learning rate Delayed update base is Maximum number of iterations Initialize the online training experience pool and the iteration index. ;

[0028] S4-2, Setting the threshold Randomly generate a number ,if Then a random action is generated. Otherwise use Network generates control variables ,in ;

[0029] S4-3, Control quantity Inputs are fed into the control system, and the system outputs are collected. and the state vector at the next time step. ;

[0030] S4-4, Change the current state Current control quantity Current output The state of the next moment Form a set The data is stored in the priority experience replay pool. If the priority experience replay pool is full, the last data is discarded.

[0031] S4-5. Sample n samples from the experience pool to obtain And prioritize them;

[0032] S4-6, according to Calculate the loss and apply the gradient descent algorithm to... The weight vector is updated.

[0033]

[0034]

[0035] , For the control action corresponding to state s in the initial dataset,

[0036]

[0037] S4-7. Perform iterative training, where the number of iterations is equal to the delayed update base. When the value is an integer multiple, update Weight vector in ,make When the set maximum number of iterations is reached When the time is up, stop training and obtain the trained weight model.

[0038] This invention has the following characteristics and beneficial effects:

[0039] By adopting the above technical solution, a control model is obtained through two processes: pre-training and online training. This model can accurately and efficiently control the attitude of the UAV, effectively reducing the risk of the UAV losing control when it is disturbed by strong winds or other factors in natural environments. It can control and correct the attitude of the UAV in complex environments, improving the safety and practicality of the UAV. In addition, the method, which is divided into pre-training and online training, also has the characteristics of fast response, strong robustness, and fast convergence speed. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of the deep reinforcement learning structure in an embodiment of the present invention. Detailed Implementation

[0042] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0043] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0044] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0045] This invention provides a novel deep learning-based method for UAV attitude control, comprising the following steps:

[0046] Step (1) Deep reinforcement learning network pre-training, the specific implementation process is as follows:

[0047] Step 1: Transfer the pre-collected data Initial empirical data Loaded into the experience pool, including: Current state (drone attitude) To control the quantity (control the direction of motion) For output value, For the next output state;

[0048] Step 2: Randomly initialize the deep reinforcement learning network Weight vector in and target network weight vector Set the delayed update base to The number of pre-training iterations is ;

[0049]

[0050] in This represents the output function of a deep reinforcement learning network. Structure as Figure 1 As shown,

[0051] Step 3: According to The size of the data in the experience pool is used to prioritize the data.

[0052] , which is the weighting coefficient of the empirical value;

[0053] Step 4: Sample n samples from the experience pool to obtain

[0054] Step 5: According to Calculate the loss and apply the gradient descent algorithm to... The weight vector is updated. Where, ;

[0055]

[0056]

[0057] , For the control action corresponding to state s in the initial dataset,

[0058]

[0059] Step 6: When the number of iterations is equal to the delayed update cardinality When the value is an integer multiple, update Weight vector in ,make When the number of iterations reaches the set maximum number of iterations. When the time comes, training is stopped and pre-training weights are obtained.

[0060] After step (two) pre-training is completed, the following steps are performed: The specific process of online training is as follows:

[0061] Step 1: Load pre-trained weights and set the expected output. Learning rate Delayed update base is Maximum number of iterations Initialize the online training experience pool and the iteration index. ;

[0062] Step 2: Set the threshold Randomly generate a number. ,if Then a random action is generated. Otherwise use Network generates control variables ,in .

[0063] Step 3: Set the control quantity Inputs are fed into the control system, and the system outputs are collected. and the state vector at the next time step. .

[0064] Step 4: Set the current state Current control quantity Current output The state of the next moment Form a set The data is stored in the priority experience replay pool. If the priority experience replay pool is full, the last data is discarded.

[0065] Step 5: Sample n samples from the experience pool to obtain And prioritize them.

[0066] Step 6: According to Calculate the loss and apply the gradient descent algorithm to... The weight vector is updated.

[0067]

[0068]

[0069] , For the control action corresponding to state s in the initial dataset,

[0070]

[0071] Step 7: Perform iterative training, with the number of iterations equal to the delayed update base. When the value is an integer multiple, update Weight vector in ,make When the set maximum number of iterations is reached... When the time is up, training is stopped, and the trained weight model is obtained. This model can be directly deployed to the lower-level machine system for operation.

[0072] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments, including components, without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.

Claims

1. A method for attitude balance control of a UAV based on deep reinforcement learning, characterized in that, Includes the following steps: S1, The pre-collected data The initial experience data is loaded into the experience pool. ,in, For the current drone attitude, To control the amount, For output value, For the next output state; S2. Initialize the parameters of the deep reinforcement learning network; S3. Pre-trained and initialized deep reinforcement learning network: S3-1, Set the delayed update base to... The number of pre-training iterations is ; in This represents the output function of a deep reinforcement learning network; S3-2, according to The size of the data in the experience pool is used to prioritize the data. Weighting coefficients for empirical values; S3-3, Sample n samples from the experience pool to obtain S3-4. Calculate the loss function based on the sampled data; S3-5, When the number of iterations is the deferred update base When the value is an integer multiple, update Weight vector in ,make When the number of iterations reaches the set maximum number of iterations When the time comes, stop training and obtain the pre-trained weights; S4. Online real-time training of deep reinforcement learning networks: S4-1. Load pre-trained weights and set the desired output. Learning rate Delayed update base is Maximum number of iterations Initialize the online training experience pool and initialize the iteration index. ; S4-2, Setting the threshold Randomly generate a number ,if Then a random action is generated. Otherwise use Network generates control variables ,in ; S4-3, Control quantity Inputs are fed into the control system, and the system outputs are collected. and the state vector at the next time step. ; S4-4, Change the current state Current control quantity Current output The state in the next moment Form a set The data is stored in the priority experience replay pool. If the priority experience replay pool is full, the last data is discarded. S4-5. Sample n samples from the experience pool to obtain And prioritize them; S4-6. Calculate the loss function based on the sampled samples sorted by priority; S4-7. Perform iterative training, where the number of iterations is equal to the delayed update base. When the value is an integer multiple, update Weight vector in ,make When the set maximum number of iterations is reached When the time comes, stop training and obtain the trained weight model; S5 uses a deep reinforcement learning network to output control variables to control the drone.

2. The UAV attitude balance control method based on deep reinforcement learning according to claim 1, characterized in that, The initialization parameters in step S2 include randomly initializing the deep reinforcement learning network. Weight vector in and target network weight vector .

3. The UAV attitude balance control method based on deep reinforcement learning according to claim 2, characterized in that, In step S3-4, the loss function is calculated as follows: according to Calculate the loss and apply the gradient descent algorithm to... The weight vector is updated, where, , For the control action corresponding to state s in the initial dataset, 。 4. The UAV attitude balance control method based on deep reinforcement learning according to claim 2, characterized in that, The method for calculating the loss function in step S4-6 is as follows: according to Calculate the loss and apply the gradient descent algorithm to... Update the weight vector. , For the control action corresponding to state s in the initial dataset, 。

Citation Information

Patent Citations

  • Stable flight control method of multi-rotor unmanned aerial vehicle based on finite-time neurodynamics

    CN107368091A

  • Four-rotor unmanned aerial vehicle route following control method based on deep reinforcement learning

    CN110673620A