Wind disturbance compensation control method, device and system for quad-rotor unmanned aerial vehicle

By constructing a dynamic model and optimizing a deep network learning model, the motor thrust is dynamically adjusted to cope with wind disturbances, solving the problems of attitude instability and trajectory deviation of traditional quadrotor drones in complex environments, and improving the robustness and adaptability of the drone.

CN120669530APending Publication Date: 2025-09-19CHINA INST OF RADIO PROPAGATION
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510751621.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional quadrotor UAV control methods have difficulty dynamically adapting to model parameter perturbations caused by wind disturbances in complex environments, and lack the ability to actively perceive and model disturbance characteristics, resulting in attitude instability, trajectory deviation and control failure.

Method used

The flight trajectory under different wind conditions is tracked through a model predictive controller, a dynamic model is constructed, and an optimized deep network learning model is used to determine the attitude correction of the UAV. Combined with the basic thrust of the motor and the wind disturbance compensation force, the motor thrust is dynamically adjusted to cope with wind field changes.

Benefits of technology

It improves the robustness and adaptability of quadrotor drones in unknown or complex environments, solves the overshoot problem caused by high-frequency gusts, and achieves accurate compensation for external interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669530A_ABST
    Figure CN120669530A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle control, and discloses a wind disturbance compensation control method for a quad-rotor unmanned aerial vehicle, which comprises the following steps: tracking flight paths of the quad-rotor unmanned aerial vehicle under different wind conditions through a model prediction controller to obtain flight state data under different wind conditions; according to the flight state data and the kinetic equation under different wind conditions, a kinetic model of the four-rotor unmanned aerial vehicle is constructed, and a flight data set is obtained based on the kinetic model; determining the attitude correction of the unmanned aerial vehicle based on the optimized deep network learning model; according to the attitude correction amount of the unmanned aerial vehicle and the flight state data under different wind conditions, the motor basic thrust of the four-rotor unmanned aerial vehicle is obtained; and according to the basic thrust of the motor and the wind disturbance compensation force output by the optimized deep network learning model, obtaining the target thrust of the motor after wind disturbance compensation. The invention further discloses a wind disturbance compensation control device and system for the quad-rotor unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drone control technology, and for example, to a wind disturbance compensation control method, device, and system for a quadrotor drone. Background Art

[0002] Quadrotors, with their vertical takeoff and landing capabilities, maneuverability, and strong environmental adaptability, have been widely used in fields such as terrain exploration, disaster relief, and precision agriculture. However, in complex outdoor environments, time-varying wind disturbances can cause aircraft attitude instability, trajectory deviation, and even control failure. When flying close to the ground, turbulence caused by buildings or terrain exhibits strong nonlinearity, time-varying behavior, and spatial inhomogeneity. These issues pose significant challenges to traditional quadrotor control methods.

[0003] Classic control methods such as Proportional Derivative (PD) control rely on precise dynamic models, and their fixed gains make it difficult to dynamically adapt to model parameter perturbations caused by wind disturbances. Although Proportional Integral Derivative (PID) control can suppress steady-state errors through the integral term, it is prone to phase lag and overshoot when faced with high-frequency sudden wind disturbances. Although PID control is widely used in consumer-grade quadrotor products, actual parameter adjustment often relies on trial and error and human experience, and lacks a mechanism to systematically respond to external disturbances. Although the Linear Quadratic Regulator (LQR) can provide optimal control strategies under ideal conditions, its control gain matrix based on the linear system assumption is difficult to adapt to the strong nonlinear dynamic changes caused by wind disturbances, and lacks robustness guarantees for state estimation errors and model uncertainties. Its performance degrades significantly in unpredictable wind disturbance environments. Although Model Predictive Control (MPC) has feedforward compensation capabilities, the dynamic evolution of wind disturbances in the time domain that it predicts is difficult to accurately represent using finite-order models. Sudden changes in system state caused by wind disturbances can disrupt the consistency between the prediction model and the actual dynamics, leading to inaccurate rolling optimization. In short, traditional quadrotor UAV control methods treat wind disturbances as external interference and passively suppress them. They lack the ability to actively perceive and model disturbance characteristics, making it impossible to achieve feedforward-feedback coordinated compensation. Summary of the Invention

[0004] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0005] The embodiments of the present disclosure provide a wind disturbance compensation control method, device, and system for a quadrotor drone to improve the robustness and adaptability of the quadrotor drone when flying in unknown or complex environments.

[0006] In some embodiments, the method includes: tracking the flight trajectory of a quadrotor drone under different wind conditions through a model predictive controller to obtain flight status data under different wind conditions; constructing a dynamic model of the quadrotor drone based on the flight status data and dynamic equations under different wind conditions and obtaining a flight data set based on the dynamic model; determining the drone attitude correction amount based on an optimized deep network learning model; obtaining the basic motor thrust of the quadrotor drone based on the drone attitude correction amount and the flight status data under different wind conditions; obtaining the target motor thrust after compensating for the wind disturbance based on the motor basic thrust and the wind disturbance compensation force output by the optimized deep network learning model.

[0007] In some embodiments, a dynamic model of a quadrotor drone is constructed based on flight state data and dynamic equations under different wind conditions, and a flight data set is obtained based on the dynamic model, including: constructing a dynamic model equation for the quadrotor drone; obtaining different segments of polynomial trajectories based on the flight state data under different wind conditions; constructing an initial data set containing time series features based on the flight state data and timestamps corresponding to the different segments of the polynomial trajectories; wherein the flight state data includes the drone's position and attitude, velocity, angular velocity, acceleration, and angular acceleration, and the initial data set includes the total thrust of each motor and the three-axis torque; inputting the initial data set into the dynamic model equation to obtain the dynamic model of the quadrotor drone. wherein the dynamic model represents the mapping relationship between state quantities and wind disturbance compensation force under different wind conditions, and the state quantities include the drone's position and attitude, velocity, and angular velocity; integrating the dynamic models corresponding to the different segments of the polynomial trajectories to construct the flight data set.

[0008] In some embodiments, the drone attitude correction amount is determined based on an optimized deep network learning model, including: constructing a deep neural network learning model and iteratively optimizing the deep network learning model to obtain an optimized deep network learning model; obtaining a wind disturbance compensation force based on the optimized deep network learning model; determining the drone attitude correction amount based on the wind disturbance compensation force; wherein, the drone attitude correction amount is obtained after the wind disturbance compensation force is input into a compensator constructed based on the deep network learning model.

[0009] In some embodiments, the dynamic model is expressed as: y = f(Z, w) + ε, where y and f(Z, w) represent the wind disturbance compensation force and the non-interference wind disturbance compensation force, respectively. And Z∈R 2n , X represents the position and attitude of the drone, represents speed and angular velocity, and ε represents noise interference; construct a deep neural network learning model, including: obtaining samples under different wind disturbances based on flight status data under different wind conditions; solving the optimal representation function W corresponding to the minimum target error function best (Z k ) and latent variables Among them, the objective error function is: W(Z k ) represents the optimal representation function for characterizing the common characteristics under all wind conditions, c k represents the latent variable used to characterize the characteristics under different wind conditions, k represents the wind condition index, y k Represents the wind disturbance compensation force under the k-th wind condition; based on Configure f(Z,w) to build a neural network learning model; configure the adversarial meta-learning framework; where the adversarial meta-learning framework is expressed as V represents the discriminator used to predict the wind condition type of the input sample, and the discriminator output is the classification probability of each wind condition. L represents the classification loss function, a represents the parameter used to control the adversarial strength, and the loss function of the discriminator is the cross entropy loss.

[0010] In some embodiments, the deep network learning model is iteratively optimized to obtain an optimized deep network learning model, including: randomly extracting mutually exclusive batches of data sets from the flight data set as an adaptation set and a training set respectively; solving the optimal linear coefficient c corresponding to minimizing the square error of the adaptation set * (w); update φ(Z k ) to extract universal features that are independent of wind disturbance; and optimize the training set based on stochastic gradient descent.

[0011] In some embodiments, determining the attitude correction amount of the drone based on the wind disturbance compensation force includes: obtaining a compensation vector based on the vector sum of the acceleration vector corresponding to the wind disturbance compensator and the gravity acceleration vector; wherein the acceleration vector represents the ratio of the wind disturbance compensation force to the mass of the drone; constructing a compensation coordinate system based on the positive axis direction of the compensation vector; wherein the positive axis direction of the compensation vector represents the thrust direction of the motor, and the Z-axis direction of the compensation coordinate system is the positive axis direction of the compensation vector; obtaining a quaternion equation; wherein the quaternion equation is: θ represents the rotation coordinate of the fixed point in the world coordinate system; according to the compensation thrust vector g c With the quaternion equation, construct the quaternion vector q before rotation f and the rotated quaternion vector q0; where the compensation thrust vector g c represents the normalized compensation vector q f =[0,g c ], q0 = qq fq -1 ; According to the rotation matrix formula, obtain the rotation matrix R cw And determine the rotated quaternion vector as the attitude correction value of the drone; among them, the rotation matrix is ​​used to represent the attitude of the compensation coordinate system under the world coordinates, and the rotation matrix formula is: q0 = R cw q f ,

[0012] In some embodiments, the flight status data also includes the three-axis speed and relative position deviation. According to the UAV attitude correction amount and the flight status data under different wind conditions, the basic thrust of the motor of the quadrotor UAV is obtained, including: inputting the UAV attitude correction amount and the three-axis speed, angular velocity, and relative position deviation into the reinforcement learning model for reinforcement learning to obtain the basic thrust of the motors of the four motors; wherein, the reinforcement learning model is obtained by configuring the optimized deep network learning model in the following manner: based on the position error and attitude error, the basic thrust of the motors of the four motors, a reward function is constructed and the objective function is optimized; a dual-hidden layer Actor network is designed, and the ReLU activation function and the Gaussian strategy output layer are used to construct the strategy function π(a|s) and the Critic network is constructed in parallel; based on the off-policy learning method, the network parameters of the Actor network and the Critic network are optimized in parallel.

[0013] In some embodiments, the target thrust of the motor after compensating for wind disturbance is obtained based on the basic thrust of the motor and the wind disturbance compensation force output by the optimized deep network learning model, including: obtaining the component of the wind disturbance compensation force in the body coordinate system According to the basic thrust of each motor and The sum of the values ​​is used to determine the target thrust of each motor after compensating for wind disturbance.

[0014] In some embodiments, the device includes a processor and a memory storing program instructions, and the processor is configured to execute the aforementioned wind disturbance compensation control method for a quadrotor drone when running the program instructions.

[0015] In some embodiments, a quadrotor drone system includes: a quadrotor drone system body; and the aforementioned wind disturbance compensation control device for a quadrotor drone, installed on the quadrotor drone system body.

[0016] The wind disturbance compensation control method, device, and system for a quadrotor drone provided by the embodiments of the present disclosure can achieve the following technical effects:

[0017] Compared with traditional fixed-gain controllers such as PID and LQR, the embodiment of the present disclosure tracks the flight trajectory of a quadcopter drone under different wind conditions through a model predictive controller to obtain flight status data under different wind conditions, and then constructs a dynamic model based on the flight status data under different wind conditions and the dynamic equation, and obtains a flight data set based on the dynamic model. The embodiment of the present disclosure uses an optimized deep network learning model to determine the attitude correction of the drone, so that the quadcopter drone system can obtain the motor base thrust based on the attitude correction of the drone and the flight status data under different wind conditions, and combine the motor base thrust with the wind disturbance compensation force output by the optimized deep network learning model to obtain the motor target thrust after compensating for the wind disturbance. By dynamically adjusting the motor base thrust as a control compensation, it can cope with the ever-changing wind field environment, thereby improving the robustness and adaptability of the quadcopter drone in unknown or complex environments.

[0018] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,

[0020] Figure 1 is a schematic diagram of a wind disturbance compensation control method for a quadrotor drone provided by an embodiment of the present disclosure;

[0021] Figure 2 is a schematic diagram of another wind disturbance compensation control method for a quadrotor drone provided by an embodiment of the present disclosure;

[0022] Figure 3 is a schematic diagram of another wind disturbance compensation control method for a quadrotor drone provided by an embodiment of the present disclosure;

[0023] Figure 4 is a schematic diagram of another wind disturbance compensation control method for a quadrotor drone provided by an embodiment of the present disclosure;

[0024] Figure 5 is a schematic diagram of converting a world coordinate system into a compensation coordinate system provided by an embodiment of the present disclosure;

[0025] Figure 6 is a schematic diagram of a double-hidden-layer Actor network model provided by an embodiment of the present disclosure;

[0026] Figure 7 Schematic diagram of a wind disturbance compensation control device for a quadrotor drone provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] In the description and claims of the embodiments of the present disclosure, as well as in the accompanying drawings, the terms "first," "second," and the like are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe the embodiments of the present disclosure herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.

[0028] Unless otherwise stated, the term "plurality" means two or more.

[0029] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.

[0030] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0031] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0032] Combine Figure 1 As shown, the embodiment of the present disclosure provides a wind disturbance compensation control method for a quadrotor drone, comprising:

[0033] S01, the processor tracks the flight trajectory of the quadrotor drone under different wind conditions through a model prediction controller to obtain flight status data under different wind conditions.

[0034] S02, the processor constructs a dynamic model of the quadrotor drone according to the flight state data and dynamic equations under different wind conditions and obtains a flight data set based on the dynamic model.

[0035] S03, the processor determines the drone attitude correction amount based on the optimized deep network learning model.

[0036] S04, the processor obtains the basic thrust of the motor of the quadrotor drone based on the drone attitude correction value and the flight status data under different wind conditions.

[0037] S05: The processor obtains the target thrust of the motor after compensating for the wind disturbance based on the basic thrust of the motor and the wind disturbance compensation force output by the optimized deep network learning model.

[0038] The wind disturbance compensation control method for a quadrotor drone provided by the embodiment of the present disclosure, compared with traditional fixed-gain controllers such as PID and LQR, uses a model predictive controller to track the flight trajectory of the quadrotor drone under different wind conditions to obtain flight status data under different wind conditions, and then constructs a dynamic model based on the flight status data under different wind conditions and the dynamic equation, and obtains a flight data set based on the dynamic model. The embodiment of the present disclosure uses an optimized deep network learning model to determine the drone attitude correction value, so that the quadrotor drone system can obtain the motor base thrust based on the drone attitude correction value and the flight status data under different wind conditions, and combine the motor base thrust with the wind disturbance compensation force output by the optimized deep network learning model to obtain the motor target thrust after compensating for the wind disturbance. By dynamically adjusting the motor base thrust as control compensation, it can cope with the ever-changing wind field environment, thereby improving the robustness and adaptability of the quadrotor drone flying in unknown or complex environments.

[0039] In addition, the disclosed embodiment accurately estimates the impact of wind disturbance on the quadrotor drone through an optimized deep network learning model, calculates the drone attitude correction, and achieves accurate compensation for external interference, effectively solving the overshoot problem caused by high-frequency gusts that is difficult to solve with traditional control methods.

[0040] In the disclosed embodiment, the quadrotor drone repeatedly collected flight status data under different wind conditions to ensure data reliability and consistency of the results. Different wind conditions include different wind speeds. The wind speed range is greater than or equal to 0 m / s (meters per second) and less than or equal to 6.1 m / s.

[0041] In the specific experiment, a quadrotor drone used a model predictive controller to track random trajectories under different wind conditions. The quadrotor drone's single flight time was 3 minutes. The drone's position, velocity, attitude angle, motor speed, and corresponding PWM (Pulse Width Modulation) values ​​were collected during the flight. A set of k of these input-output pairs formed a dataset D, and different wind conditions were represented by w.

[0042] The random trajectory above represents a polynomial trajectory. Each polynomial trajectory consists of two waypoints: the starting position and the target position. The velocity, acceleration, and angular acceleration at the starting and ending waypoints are all constrained to 0. When the quadrotor reaches the ending waypoint of a trajectory, a new polynomial trajectory is automatically generated.

[0043] In the embodiment of the present disclosure, a seventh-order polynomial is used to describe a polynomial trajectory. The process of generating the polynomial trajectory is as follows:

[0044] Given the starting position p0 and the target position p f , the seventh-degree polynomial is expressed as:

[0045] p(t)=a0+a1t 1 +a2t 2 +a3t 3 +a4t 4 +a5t 5 +a6t 6 +a7t 7

[0046] Among them, a j Represents the polynomial coefficient, j = 0, 1, 2, ..., 7.

[0047] Solving the first-order derivative of the above septad polynomial can obtain the velocity v(t), solving the second-order derivative of the above septad polynomial can obtain the acceleration a(t), and solving the third-order derivative of the above septad polynomial can obtain the angular acceleration j(t):

[0048]

[0049]

[0050]

[0051] According to the constraints of the starting waypoint and the ending waypoint, we can get:

[0052]

[0053] Among them, p0 represents the starting position of each polynomial trajectory, t f Indicates the duration of the polynomial trajectory from the starting position to the target position.

[0054] In the disclosed embodiment, after obtaining different segments of polynomial trajectories, the relationship between the motor speed, torque, and motor thrust of the quadrotor drone is first defined:

[0055] T i =c1ω i 2

[0056]

[0057] Among them, T i represents the total thrust of the i-th motor, ω i represents the motor speed of the i-th motor, l represents the arm length of the quadrotor drone, μ represents the torque, c1 and c2 represent the generated force coefficient and z-axis moment coefficient respectively, and i = 1, 2, 3, 4.

[0058] Based on the flight status data under different wind conditions in step S01, an initial data set including time series features is obtained Among them, X represents the position and attitude of the drone, X represents the speed and angular velocity, Represents the acceleration and angular acceleration of the quadcopter, U represents the total thrust and three-axis torque of the four motors, and t represents the timestamp. Among them, U represents the following:

[0059] U=[T1,T2,T3,T4,μ x ,μ y ,μ z ]

[0060] In the above formula, μ x 、μ y 、μ z They represent the torque of the quadrotor drone along the x-axis, y-axis, and z-axis of the world coordinate system respectively.

[0061] Next, we introduce a dynamic model equation that describes the dynamic behavior of the aircraft:

[0062]

[0063] Where M(x) represents the symmetric positive definite inertia matrix, represents the Coriolis and centrifugal force matrix and is used to describe the inertial force due to motion, g represents the gravitational acceleration vector, U represents the control quantity defined above, represents the external wind disturbance compensation force and is characterized by an equation related to the dynamic behavior of the quadrotor UAV system, which includes external disturbances related to the wind condition w.

[0064] The initial dataset Substitute into the above dynamic model equation to solve the external wind disturbance compensation force Obtain the dynamic model of the quadrotor drone. The dynamic model is expressed as follows:

[0065] y=f(Z,w)+ε

[0066] Where y represents the wind disturbance compensation force.

[0067] Due to the existence of sensor noise and numerical errors, the external wind disturbance compensation force Also contains noise. Therefore, the interference-free wind disturbance compensation force is defined as f(Z,w), and Z∈R 2n , ε represents noise interference.

[0068] Finally, the dynamic equations obtained from each polynomial trajectory are integrated to construct the flight dataset D:

[0069]

[0070] Among them, Z k Represents the state quantity containing timestamp under the k-th wind condition w k Represents the kth wind condition, which is used to reflect the impact of different wind conditions on wind disturbance compensation force, and the output variable y k Represents the estimated value of the wind disturbance compensation force including noise.

[0071] Based on the above principles, step 02 in the above embodiment is described as follows:

[0072] Optionally, combined Figure 2 As shown, the processor constructs a dynamic model of the quadrotor drone based on the flight status data and dynamic equations under different wind conditions and obtains a flight data set based on the dynamic model, including:

[0073] S11, the processor constructs the dynamic model equations of the quadrotor drone.

[0074] In this step, the kinetic model equation is: Where M(x) represents the symmetric positive definite inertia matrix, represents the Coriolis and centrifugal force matrix and is used to describe the inertial force generated by motion, g represents the gravitational acceleration vector, U represents the control quantity defined above, represents the external wind disturbance compensation force, which includes the external disturbance related to the wind condition w.

[0075] S12: The processor obtains different segments of polynomial trajectories based on the flight status data under different wind conditions.

[0076] In this step, each polynomial trajectory can be represented by the aforementioned polynomial.

[0077] At S13, the processor constructs an initial dataset containing time series features based on the flight status data and timestamps corresponding to the different segments of the polynomial trajectory. The flight status data includes the drone's position and attitude, velocity, angular velocity, acceleration, and angular acceleration. The initial dataset includes the total thrust of each motor and the torque on all three axes.

[0078] At step S14, the processor inputs the initial data set into the dynamic model equation to obtain a dynamic model of the quadrotor drone. The dynamic model represents the mapping relationship between state variables and wind disturbance compensation forces under different wind conditions. The state variables include the drone's position, attitude, velocity, and angular velocity.

[0079] In this step, the dynamic model is expressed as: y = f(Z,w) + ε, where y and f(Z,w) represent the wind disturbance compensation force and the non-interference wind disturbance compensation force, respectively. And Z∈R 2n , X represents the position and attitude of the drone, represents speed and angular velocity, and ε represents noise interference.

[0080] In this step, the model input of the dynamic model is the state quantity (i.e., the UAV position and UAV attitude, speed and angular velocity) under different wind conditions with time series characteristics and wind disturbance conditions (i.e., actual wind disturbance conditions), and the model output of the dynamic model is the wind disturbance compensation force.

[0081] In S15 , the processor integrates the dynamic models corresponding to the different segments of the polynomial trajectory to construct a flight data set.

[0082] In this step, the flight data set is specifically represented as Among them, Z k Represents the state quantity under the kth wind condition w k represents the kth wind condition, which is used to reflect the impact of different wind conditions on the wind disturbance compensation force, y k Represents the estimated value of the wind disturbance compensation force including noise.

[0083] In this way, the embodiment of the present disclosure constructs a dynamic model of the quadrotor drone based on the collected flight status data and combined with the dynamic model equation of the quadrotor drone, and then integrates the dynamic models corresponding to different segments of polynomial trajectories to construct a flight data set, which can accurately represent the actual flight behavior of the quadrotor drone under different wind conditions.

[0084] In the disclosed embodiment, the optimal representation function W(Z) and a set of latent variables (c1...c k ), so that for any wind condition w, we can find the latent variable c(w) so that W(Z)c(w) can well approximate f(Z,w).

[0085] Here, W(Z) is a function that maps a 2n-dimensional state vector to an n×h-dimensional space, c k It is the linear coefficient of h dimension, which is used to adjust the features extracted by W(Z) and is independent of each other.

[0086] The definition of the objective error function S is:

[0087]

[0088] Among them, k represents the wind condition index, i represents the i-th sample under the k-th wind condition (the total number of samples under the k-th wind condition is N k ).y k represents the estimated value of wind disturbance compensation force including noise, Zk Represents the state quantity under different wind disturbance conditions

[0089] The optimal representation W(Z) is to solve the following optimization problem:

[0090]

[0091] Solve the minimum value of the prediction error function to find the optimal representation function W(Z k ) and the set of latent variables related to wind disturbance (c1...c k ). Here (c1...c k ) is a common characteristic under all wind conditions, and the linear coefficient c k Indicates characteristics under different wind conditions.

[0092] In the above optimization process, the inherent domain shift of the state space caused by the change of wind conditions has brought significant difficulties to the training of the deep network learning model. During the data collection phase, the quadcopter produced very different trajectories under different wind conditions, resulting in the state distribution Z corresponding to various wind conditions. k There are obvious differences. This shift in state distribution may cause the neural network W to tend to simply memorize the state distribution characteristics under various wind conditions, rather than truly learning the relationship mapping between wind conditions and dynamics. At this time, the dynamic changes {f(Z1,w1),f(Z2,w2),...,f(Z K ,w K )} may be mainly due to the difference in state distribution rather than the actual wind parameters {w1,w2,...,w K}, which will cause the deep network learning model to overfit and fail to learn the true wind condition-invariant representation W, thus affecting the generalization ability of the deep network learning model.

[0093] To solve this problem, the following adversarial meta-learning framework is proposed:

[0094]

[0095] Where V represents an additional deep neural network that acts as a discriminator to predict which of the K wind conditions the input sample comes from, L represents the classification loss function, a represents the parameter that controls the strength of the adversarial and a ≥ 0, k represents the wind condition index, and i is the i-th sample of the k-th wind condition. The goal of V is to predict index k directly from W(Z); the goal of W is to approximate the label At the same time, it makes the work of V more difficult. In other words, V represents a learning regularizer that is used to remove the environmental information contained in W. The output of V is a K-dimensional vector that represents the classification probability of K types of wind conditions.

[0096] This disclosed embodiment employs an adversarial meta-learning framework to train a deep neural network learning model, ensuring optimal performance under varying wind conditions. This approach transcends the limitations of traditional model predictive control within the prediction time domain and can better handle sudden system state changes caused by wind disturbances.

[0097] Next, define the above loss function. The loss function used by the discriminator V is the cross entropy loss, and the formula is as follows:

[0098]

[0099] Among them, if k=i, then α ki =1; otherwise α ki =0,α ki represents the indicator function, h(φ(Z k )) i represents the probability of the discriminator predicting the i-th wind condition;

[0100] After initialization, the inherent state variable Z changes with external wind disturbances, which brings challenges to deep learning. The deep neural network φ(Z) may memorize the distribution of Z under different state conditions, so that the differences in the dynamic change set are reflected in the distribution rather than directly in the wind condition set w.

[0101] The following explains the principle of iterative optimization of the constructed deep neural network learning model:

[0102] After building the deep neural network learning model, the model parameters are optimized through adaptation steps, training steps and regularization steps in sequence to achieve the purpose of rapid adaptation and precise control.

[0103] At the beginning of each iteration of the deep neural network learning model, a dataset D is randomly selected from the flight dataset D containing multiple different wind conditions. k , then from D k Randomly extract two mutually exclusive batches as adaptation set A and training set B. Where D={D1,D2,...D k}.

[0104] The adaptation step solves the least squares problem as a function of φ on the adaptation set B. The training step updates the representation φ learned on the training set B according to the best linear coefficients solved from the adaptation step. The regularization step updates the discriminator V on the training set as follows:

[0105] First, an adaptation step is performed to find an optimal set of linear coefficients by minimizing the squared error on the adaptation set A. The adaptation set A is used to solve the least squares problem to find an optimal set of linear coefficients c * (w), satisfies the following formula:

[0106]

[0107] Next, a training step is performed to improve the model's generalization capabilities, enabling the learned representation φ(Z) to extract universal features independent of wind disturbances. By updating φ(Z) using stochastic gradient descent (SGD) on the training set, leveraging the aforementioned adversarial meta-learning framework and incorporating a regularization term to counter the classification capabilities of the discriminator V, φ(Z) learns a more robust feature representation that is invariant across wind disturbances.

[0108] Then, a regularization step is performed to decide whether to update the discriminator V in each iteration with a certain probability β (0<β≤1). If it is decided to update, SGD is also used to optimize it on the training set B to improve its accuracy in classifying wind disturbance speed.

[0109] The loss function only focuses on the classification performance of the discriminator V:

[0110]

[0111] Among them, L represents the cross entropy loss, which is used to measure the degree of inconsistency between the wind condition label k predicted by the discriminator V and the actual wind condition label.

[0112] In this way, the performance of the deep neural network learning model in distinguishing different wind disturbance conditions can be enhanced, and φ(Z) can be promoted to learn more denoised and abstract features, effectively increasing the model's distinguishing ability.

[0113] Finally, the above three steps are continuously performed for iterative optimization to obtain an optimized deep network learning model and enhance the performance of the deep neural network learning model.

[0114] The following explains the principle of determining the drone's attitude correction based on the wind disturbance compensation force obtained using the optimized deep network learning model:

[0115] Design the compensator as follows:

[0116] After obtaining the optimized deep network learning model, the wind disturbance compensation force y can be estimated in real time. The estimated wind disturbance compensation force can be used to calculate the acceleration vector generated by the external wind disturbance:

[0117]

[0118] Among them, a est Represents the acceleration vector generated by the wind disturbance compensation force.

[0119] In order to perform disturbance compensation, the acceleration vector a generated by wind disturbance is combined est And the gravity acceleration vector g, define a new compensation vector to describe the acceleration formed by the combined action of wind disturbance compensation force and gravity:

[0120] a com =a est +g

[0121] a com It represents the acceleration caused by the combined effect of wind disturbance compensation force and gravity.

[0122] Next, define a compensation thrust vector. When the quadcopter's body coordinate system and the world coordinate system coincide, the external wind disturbance compensation force causes the quadcopter to tilt to a certain extent. This new compensation thrust vector is the direction of the drone's motor thrust in this state. The compensation thrust vector is expressed as follows:

[0123]

[0124] Among them, g c Represents the compensation thrust vector, whose positive axis is along the thrust direction of the quadrotor UAV's motor.

[0125] Thus, a compensation coordinate system can be obtained. The z-axis direction of the compensation coordinate system is g c direction.

[0126] In order to transform the posture of the quadcopter body in the world coordinate system into the compensation coordinate system, it is necessary to construct the rotation matrix R cw , the rotation matrix represents the posture of the compensation coordinate system in the world coordinate system.

[0127] In three-dimensional space, any rotation of coordinates about a fixed point is equivalent to a rotation of an angle θ around a fixed axis passing through the fixed point. c The rotation matrix R cw , using vector z i Representing the z-axis direction of the original world coordinate system, the quaternion equation is:

[0128]

[0129] Where x represents the unit vector of the rotation axis, defined as:

[0130] x=g c ×zi

[0131] Among them, z i Indicates the z-axis direction of the world coordinate system, and "×" represents the cross product, that is, x is perpendicular to g c The plane formed by zi.

[0132] Define the quaternion vector before rotation and the quaternion vector after rotation, and write the relationship between the two:

[0133] q f =[0,g c ]

[0134] q0=qq f q -1

[0135] Next, substitute q into the following formula:

[0136]

[0137] We can find:

[0138] q0=R cw q f

[0139] Thus, the relationship between the world coordinate system and the compensation coordinate system can be obtained, and the rotated quaternion vector q0 can be determined as the attitude correction value of the drone.

[0140] Based on the above principles, step 03 in the above embodiment is described as follows:

[0141] Optionally, combined Figure 3 As shown, the processor determines the drone's attitude correction based on an optimized deep network learning model, including:

[0142] S21, the processor builds a deep neural network learning model and iteratively optimizes the deep network learning model to obtain an optimized deep network learning model.

[0143] In step S23, the processor obtains wind disturbance compensation force based on the optimized deep network learning model.

[0144] At step S24, the processor determines a correction value for the drone's attitude based on the wind disturbance compensation force. The wind disturbance compensation force is input into a compensator built based on a deep network learning model to obtain the drone's attitude correction value.

[0145] Thus, the disclosed embodiment also features a feedforward-feedback synergy mechanism, which, combined with a deep network learning method, not only performs feedforward compensation based on real-time wind disturbance compensation force estimation, but also utilizes a deep network learning model to optimize and output motor thrust and torque control commands appropriate for the current flight state. By adopting a "perception-estimation-compensation" closed-loop architecture, the stability and control accuracy of the quadrotor drone system can be significantly enhanced.

[0146] Optionally, the processor builds a deep neural network learning model, including:

[0147] The processor obtains samples under different wind disturbances based on the flight status data under different wind conditions.

[0148] The processor solves the optimal representation function W that satisfies the minimum objective error function best (Z k ) and latent variables Among them, the objective error function is: W(Z k ) represents the optimal representation function for characterizing the common characteristics under all wind conditions, c k represents the latent variable used to characterize the characteristics under different wind conditions, k represents the wind condition index, y k represents the wind disturbance compensation force under the kth wind condition.

[0149] Processor based on Configure f(Z,w) to build a neural network learning model.

[0150] Processor configuration adversarial meta-learning framework. The adversarial meta-learning framework is represented as V represents the discriminator used to predict the wind condition type of the input sample, and the discriminator output is the classification probability of each wind condition, L represents the classification loss function, and a represents the parameter used to control the adversarial strength.

[0151] The loss function of the above discriminator is cross entropy loss, which is expressed as: Among them, if k=i, then α ki =1; otherwise α ki =0,α ki represents the indicator function, h(φ(Z k )) i represents the probability of the discriminator predicting the i-th wind condition.

[0152] Optionally, the processor iteratively optimizes the deep network learning model to obtain an optimized deep network learning model, including:

[0153] The processor randomly extracts mutually exclusive batches of data sets from the flight data set as the adaptation set A and the training set B respectively.

[0154] The processor solves the optimal linear coefficient c that satisfies the minimization of the square error of the adaptation set. * (w).

[0155] The processor updates φ(Z k ) to extract universal features that are not related to wind disturbance.

[0156] The processor optimizes the training set B based on stochastic gradient descent. In this step, when optimizing the training set B, the loss function only focuses on the classification performance of the discriminator, so the loss function uses the cross entropy loss, which is expressed as:

[0157] In this way, the performance of the deep neural network learning model in distinguishing different wind disturbance conditions can be enhanced, and φ(Z) can be promoted to learn more denoised and more abstract features, effectively increasing the model's distinguishing ability.

[0158] Optionally, combined Figure 4 and Figure 5 As shown, the processor determines the drone attitude correction value based on the wind disturbance compensation force, including:

[0159] S31: The processor obtains a compensation vector based on the vector sum of the acceleration vector corresponding to the wind disturbance compensator and the gravity acceleration vector, wherein the acceleration vector represents the ratio of the wind disturbance compensation force to the mass of the drone.

[0160] S32: The processor constructs a compensation coordinate system based on the positive axis direction of the compensation vector, wherein the positive axis direction of the compensation vector represents the thrust direction of the motor, and the Z axis direction of the compensation coordinate system is the positive axis direction of the compensation vector.

[0161] S33: The processor obtains a quaternion equation. The quaternion equation is: θ represents the rotation coordinate of the fixed point in the world coordinate system.

[0162] S34, the processor compensates the thrust vector g c With the quaternion equation, construct the quaternion vector q before rotation f And the rotated quaternion vector q0. Among them, the compensation thrust vector g c represents the normalized compensation vector q f =[0,g c ], q0 = qq f q -1 .

[0163] S35, the processor obtains the rotation matrix R according to the rotation matrix formula cw And determine the rotated quaternion vector as the drone attitude correction value.

[0164] Among them, the rotation matrix is ​​used to represent the posture of the compensation coordinate system under the world coordinate system. The rotation matrix formula is: q0 = R cw q f ,

[0165] In this way, the transformation of the quadcopter's body posture in the world coordinate system to the compensation coordinate system is achieved.

[0166] The following explains the principle of obtaining the basic thrust of the quadcopter's motor based on the drone's attitude correction and flight status data under different wind conditions:

[0167] The standard reinforcement learning framework consists of a learning agent interacting with an environment that follows a Markov decision process (MDP). The MDP consists of the tuple Definition, where S represents the state space and A represents the action space. represents the state transition probability of the environment, r represents the reward function, ρ0 represents the initial state distribution, and γ represents the discount factor.

[0168] The goal is to find a policy π(a∈A|S) that maximizes the cumulative discounted reward:

[0169]

[0170] In RL using deep neural networks, the agent follows the policy π(a|s;θ) = Pr(a|s;θ). The state space S is the wind-compensated attitude quaternion, three-axis velocity, angular velocity, and relative position deviation, and the action space A is the motor thrust values ​​of the four motors. The reward function is set as:

[0171] r=-(0.2e p +0.1e q +0.05μ i )

[0172] Among them, e p Represents the position error, e q represents the attitude error, μ i represents the motor thrust value of the i-th motor, i = 1, 2, 3, 4. It should be noted that the relative position deviation is usually represented by a vector.

[0173] This reward function takes various factors into consideration. Including an attitude error term in the reward function encourages the quadrotor to maintain or quickly return to the desired attitude. This improves the aircraft's stability and reduces unnecessary sway or tilt. The position error term helps the aircraft reach its desired location more accurately, enhancing its navigation capabilities. The motor thrust command is introduced as a penalty term to limit energy consumption.

[0174] Reinforcement learning models are trained in simulation using an off-police training method, which records the quadcopter's state, action, probability, and reward in each episode. Off-police training can leverage data generated by one policy (μ) to learn a different policy (π). During data collection, the quadcopter is randomly launched within a 2-meter cubic space. The training data includes the quadcopter's state, action, probability, and reward. Each episode consists of 200 steps within a two-second flight time.

[0175] Then, the following algorithm is used to update the strategy:

[0176]

[0177]

[0178] Where V is the state value, which can be used to evaluate the quality of the current strategy. H is the advantage function, which is defined as follows:

[0179]

[0180]

[0181] H π (s,a)=Q π (s,a)-V π (s)

[0182] Here, H π (s,a) represents the state V where a specific action a is selected π (s) and expected value Q π The advantage function can be used to determine whether the selected action is appropriate relative to the strategy π. Based on the above two equations, the stochastic gradient descent method is used to optimize the objective function, and we can get:

[0183]

[0184]

[0185] To approximate the function π(a|s) for proposing actions and V(s) for predicting state values, the two neural network equations used are as follows:

[0186]

[0187]

[0188] in, represents the state of the i-th hidden layer in the optimized deep network learning model with width j, represents the state of the i-th output layer in the optimized deep network learning model with width j, π(a|s) represents the policy function, V(s) represents the state, It represents the composition of functions, which is used to describe the hierarchical structure of the optimized deep network learning model, and is passed from the input to the left in sequence starting from the last side function. In π(a|s), the last side function is In V(s), the final side function is

[0189] After the above steps, a reinforcement learning model can be obtained. The input of the model is the attitude quaternion, three-axis velocity, angular velocity and relative position deviation. The output of the model is the basic motor thrust of the four motors.

[0190] Based on the above principles, step 04 in the above embodiment is described as follows:

[0191] Optionally, the flight status data also includes the three-axis speed and relative position deviation. The processor obtains the basic thrust of the quadcopter motor based on the drone attitude correction value and the flight status data under different wind conditions, including:

[0192] The processor inputs the drone's attitude correction value and the three-axis speed, angular velocity, and relative position deviation into the reinforcement learning model for reinforcement learning to obtain the basic motor thrust of the four motors.

[0193] The reinforcement learning model is obtained by configuring the optimized deep network learning model in the following way:

[0194] The processor constructs a reward function and optimizes the objective function based on the position error, attitude error, and the basic motor thrust of the four motors.

[0195] The processor designs a dual hidden layer Actor network, uses the ReLU activation function and the Gaussian strategy output layer to construct the policy function π(a|s), and constructs the Critic network in parallel. The Critic network uses a 128-node hidden layer structure. The designed dual hidden layer Actor network can be referenced. Figure 6 .

[0196] The processor iteratively optimizes the network parameters of the Actor network and the Critic network in parallel based on the off-policy learning method.

[0197] In this way, the embodiment of the present disclosure uses the drone attitude correction, three-axis speed, angular velocity, and relative position deviation obtained by wind disturbance compensation as state space inputs, and the basic motor thrust of the four motors as action space outputs, realizing the mapping from environmental perception to control execution. The embodiment of the present disclosure designs a dual-hidden layer Actor network, adopts the ReLU activation function and the Gaussian strategy output layer to approximate the strategy function π(a|s), and constructs a Critic network in parallel, adopts a 128-node hidden layer structure, accurately estimates the state value function, and provides a value benchmark for strategy optimization. The embodiment of the present disclosure is also based on the off-policy learning method, by parallel execution of flight status data collection and network parameter update process, iteratively optimizes the network parameters of the Actor network and the Critic network, and ensures that the learning process converges stably to the optimal control strategy. Based on this, the embodiment of the present disclosure adopts the state space-action space construction of the reinforcement learning model, the design of the reward function and optimization target, and the implementation of proximal strategy optimization training, etc., to achieve high system efficiency and good safety while ensuring the flight stability of the quadcopter drone.

[0198] Optionally, the processor constructs a reward function based on the position error and the attitude error and the basic thrust of the four motors, including: r = -(0.2e p +0.1e q +0.05μ i ). Among them, e p Represents the position error, e q represents the attitude error, μ i Represents the motor thrust value of the i-th motor, i=1,2,3,4.

[0199] The processor optimizes the objective function, including: the processor optimizes the objective function based on a stochastic gradient descent method.

[0200] The following explains the principle of obtaining the target thrust of the motor after compensating for wind disturbance based on the motor's basic thrust and the wind disturbance compensation force output by the optimized deep network learning model:

[0201] After steps S01 to S05, an optimized deep network learning model for estimating the external wind disturbance compensation force and a reinforcement learning model for realizing mapping from the state space to the action space can be obtained.

[0202] In a windy environment, the deep network learning model has estimated the wind disturbance compensation force y, and then used y to calculate the drone's attitude correction. This was then input into the reinforcement learning model to obtain the basic motor thrust value T for the four motors. i Next, calculate the motor target thrust using the following formula:

[0203]

[0204] Among them, T comi represents the target thrust of the i-th motor after compensating for wind disturbance, Represents the component of the wind disturbance compensation force in the body coordinate system.

[0205] Optionally, the processor obtains a target thrust of the motor after compensating for wind disturbance based on the basic thrust of the motor and the wind disturbance compensation force output by the optimized deep network learning model, including:

[0206] The processor obtains the component of the wind disturbance compensation force in the body coordinate system

[0207] The processor calculates the motor basic thrust and The sum of the values ​​is used to determine the target thrust of each motor after compensating for wind disturbance.

[0208] Combine Figure 7 As shown, an embodiment of the present disclosure provides a wind disturbance compensation control device 70 for a quadrotor drone, comprising a processor 700 and a memory 701. Optionally, the device 70 may further include a communication interface 702 and a bus 703. The processor 700, the communication interface 702, and the memory 701 may communicate with each other through the bus 703. The communication interface 702 may be used for information transmission. The processor 700 may call the logic instructions in the memory 701 to execute the wind disturbance compensation control method for a quadrotor drone of the above embodiment.

[0209] In addition, the logic instructions in the memory 701 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0210] Memory 701, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 700 executes the program instructions / modules stored in memory 701 to perform functional applications and data processing, thereby implementing the wind disturbance compensation control method for a quadrotor drone in the above-mentioned embodiments.

[0211] The memory 701 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 701 may include high-speed random access memory and non-volatile memory.

[0212] The disclosed embodiment provides a quadrotor drone system, comprising: a quadrotor drone system body, and the above-mentioned wind disturbance compensation control device 100 for the quadrotor drone. The wind disturbance compensation control device 100 for the quadrotor drone is installed on the quadrotor drone system body. The installation relationship described here is not limited to placement inside the quadrotor drone system body, but also includes installation connections with other components of the quadrotor drone system, including but not limited to physical connections, electrical connections or signal transmission connections, etc. Those skilled in the art will understand that the wind disturbance compensation control device 100 for the quadrotor drone can be adapted to a feasible quadrotor drone system body, thereby realizing other feasible embodiments.

[0213] The embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned wind disturbance compensation control method for a quadrotor drone. The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

Claims

1. A wind disturbance compensation control method for a quadrotor drone, characterized in that: include: The flight trajectory of the quadrotor drone under different wind conditions is tracked by the model predictive controller to obtain the flight status data under different wind conditions; Based on the flight status data and dynamic equations under different wind conditions, a dynamic model of the quadrotor UAV is constructed and a flight data set is obtained based on the dynamic model; Determine the drone attitude correction based on the optimized deep network learning model; The basic thrust of the quadrotor UAV's motor is obtained based on the UAV's attitude correction and flight status data under different wind conditions; According to the basic thrust of the motor and the wind disturbance compensation force output by the optimized deep network learning model, the target thrust of the motor after compensating for the wind disturbance is obtained.

2. The method according to claim 1, characterized in that Based on the flight status data and dynamic equations under different wind conditions, a dynamic model of the quadrotor UAV is constructed and a flight data set is obtained based on the dynamic model, including: Construct the dynamic model equations of the quadrotor drone; According to the flight status data under different wind conditions, different polynomial trajectories are obtained; Based on the flight status data and timestamps corresponding to different segments of the polynomial trajectory, an initial dataset containing time series features is constructed. The flight status data includes the drone's position and attitude, speed, angular velocity, acceleration, and angular acceleration. The initial dataset includes the total thrust of each motor and the three-axis torque. Input the initial data set into the dynamic model equation to obtain the dynamic model of the quadrotor drone. The dynamic model represents the mapping relationship between state variables and wind disturbance compensation force under different wind conditions. The state variables include the drone's position, attitude, velocity, and angular velocity. The dynamic models corresponding to different segments of polynomial trajectories are integrated to construct a flight dataset.

3. The method according to claim 2, characterized in that Based on the optimized deep network learning model, the drone attitude correction is determined, including: Construct a deep neural network learning model and iteratively optimize the deep network learning model to obtain an optimized deep network learning model; Obtain wind disturbance compensation force based on the optimized deep network learning model; The attitude correction value of the UAV is determined according to the wind disturbance compensation force; wherein, the attitude correction value of the UAV is obtained after the wind disturbance compensation force is input into the compensator constructed based on the deep network learning model.

4. The method according to claim 3, characterized in that The dynamic model is expressed as: y = f(Z,w) + ε, where y and f(Z,w) represent the wind disturbance compensation force and the non-interference wind disturbance compensation force respectively. And Z∈R 2n , X represents the position and attitude of the drone, represents speed and angular velocity, and ε represents noise interference; a deep neural network learning model is constructed, including: Based on the flight status data under different wind conditions, samples under different wind disturbances are obtained; Solve the optimal representation function W that satisfies the minimum objective error function best (Z k ) and latent variables Among them, the objective error function is: W(Z k ) represents the optimal representation function for characterizing the common characteristics under all wind conditions, c k represents the latent variable used to characterize the characteristics under different wind conditions, k represents the wind condition index, y k represents the wind disturbance compensation force under the k-th wind condition; based on Configure f(Z,w) to build a neural network learning model; Configure the adversarial meta-learning framework; where the adversarial meta-learning framework is expressed as V represents the discriminator used to predict the wind condition type of the input sample, and the discriminator output is the classification probability of each wind condition. L represents the classification loss function, a represents the parameter used to control the adversarial strength, and the loss function of the discriminator is the cross entropy loss.

5. The method according to claim 4, characterized in that Iteratively optimize the deep network learning model to obtain an optimized deep network learning model, including: Randomly extract mutually exclusive batches of data sets from the flight data set as the adaptation set and training set respectively; Solve the optimal linear coefficient c that satisfies the minimization of the square error of the adaptation set * (w); Using the adversarial meta-learning framework to update φ(Z k ) to extract universal features that are not related to wind disturbance; Optimize the training set based on stochastic gradient descent.

6. The method according to claim 3, characterized in that According to the wind disturbance compensation force, determine the attitude correction of the UAV, including: The compensation vector is obtained based on the vector sum of the acceleration vector corresponding to the wind disturbance compensator and the gravity acceleration vector. The acceleration vector represents the ratio of the wind disturbance compensation force to the mass of the UAV. According to the positive axis direction of the compensation vector, a compensation coordinate system is constructed; wherein the positive axis direction of the compensation vector represents the thrust direction of the motor, and the Z axis direction of the compensation coordinate system is the positive axis direction of the compensation vector; The quaternion equation is obtained; wherein the quaternion equation is: θ represents the rotation coordinate of the fixed point in the world coordinate system; According to the compensation thrust vector g c With the quaternion equation, construct the quaternion vector q before rotation f and the rotated quaternion vector q0; where the compensation thrust vector g c represents the normalized compensation vector q f =[0,g c ], q0 = qq f q -1 ; According to the rotation matrix formula, the rotation matrix R is obtained cw And determine the rotated quaternion vector as the attitude correction value of the drone; among them, the rotation matrix is ​​used to represent the attitude of the compensation coordinate system under the world coordinates, and the rotation matrix formula is: q0 = R cw q f , 7. The method according to claim 4, characterized in that The flight status data also includes the three-axis speed and relative position deviation. Based on the UAV attitude correction and the flight status data under different wind conditions, the basic thrust of the quadcopter motor is obtained, including: Input the drone's attitude correction value and the three-axis speed, angular velocity, and relative position deviation into the reinforcement learning model for reinforcement learning to obtain the basic motor thrust of the four motors; The reinforcement learning model is obtained by configuring the optimized deep network learning model in the following way: Based on the position error, attitude error, and basic thrust of the four motors, a reward function is constructed and the objective function is optimized; Design a dual-hidden layer Actor network, use the ReLU activation function and Gaussian strategy output layer to construct the strategy function π(a|s) and build the Critic network in parallel Based on the off-policy learning method, the network parameters of the Actor network and the Critic network are optimized in parallel and iteratively.

8. The method according to any one of claims 1 to 7, characterized in that Based on the motor's base thrust and the wind disturbance compensation force output by the optimized deep network learning model, the motor's target thrust after wind disturbance compensation is obtained, including: Obtain the component of the wind disturbance compensation force in the body coordinate system According to the basic thrust of each motor and The sum of the values ​​is used to determine the target thrust of each motor after compensating for wind disturbance.

9. A wind disturbance compensation control device for a quadrotor drone, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the wind disturbance compensation control method for a quadrotor drone according to any one of claims 1 to 8 when running the program instructions.

10. A quad-rotor drone system, characterized in that: include: Quadrotor UAV system body; The wind disturbance compensation control device for a quadrotor drone according to claim 9 is installed on the quadrotor drone system body.

Citation Information

Cited By

  • Unmanned aerial vehicle attitude stability control system in mountainous area high-altitude environment

    CN121028816A

  • Unmanned aerial vehicle residual force prediction network training and control method based on deep learning

    CN121302546A

  • Rotor unmanned aerial vehicle networking control method

    CN121418942A

  • A control method for networking of a rotor unmanned aerial vehicle

    CN121418942B