Multi-rotor unmanned aerial vehicle modeling method and system based on machine learning

Through adversarial imitation learning algorithms to learn the dynamic model of multi-rotor UAV from finite data, the problem that modeling assumptions in the existing technology is difficult to meet, and the modeling effect with higher accuracy and consistency is achieved, and the safety of flight control is enhanced.

CN120068642AActive Publication Date: 2025-05-30POLIXIR TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510204029.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing multi-rotor UAV modeling method is based on the Newton-Euler equation. Assuming that the rotor UAV is a uniformly symmetric rigid body, it is difficult to meet the actual conditions, resulting in a deviation in flight control results and poses safety hazards.

Method used

Using a machine learning-based method, through adversarial imitation learning algorithms, dynamic models are learned from finite rotor UAV flight data, and a high-fidelity digital model is directly constructed to avoid the dependence of constructing dynamic equations in advance.

Benefits of technology

It improves the accuracy and consistency of multi-rotor UAV modeling, reduces the requirements for data volume and data coverage, simplifies the construction process of dynamic models, and enhances the accuracy and safety of flight control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068642A_ABST
    Figure CN120068642A_ABST
Patent Text Reader

Abstract

The invention provides a multi-rotor unmanned aerial vehicle modeling method and system based on machine learning, and the method comprises the steps: S1, collecting flight data of a rotor unmanned aerial vehicle, completing the cleaning of early-stage data, and dividing a data set into a training set and a verification set; s2, dividing a state quantity and an action quantity from the data according to the flight data; s3, selecting a corresponding deep neural network model according to the dimensions of the state quantity and the action quantity, and setting input and output; s4, defining a training termination condition by using an adversarial imitation learning algorithm, and training the deep neural network model by using the training data set; and S5, evaluating the trained neural network model by using the verification set, and evaluating the generalization accuracy of the model. According to the method, the requirements on data volume and data coverage during data modeling of the rotor unmanned aerial vehicle can be reduced, and meanwhile, the process that personnel in the field in the professional direction need to analyze dynamic characteristics in advance and specifically construct a dynamic model and kinematics is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-rotor UAV modeling, and particularly to a multi-rotor UAV modeling method and system based on machine learning. Background Art

[0002] Due to the characteristics of simple structure, low price, greater flexibility, and the ability to take off and land vertically in a narrow space, multi-rotor UAVs have been widely used in recent years. Building a digital model of a rotor UAV also plays a key role in its development. For example, more accurate flight control can be achieved using more accurate modeling. General rotor UAV modeling methods are based on Newton-Euler equations to construct dynamic and kinematic models. During the modeling process, it is usually assumed that the rotor UAV is a uniformly symmetric rigid body, the geometric center coincides exactly with the center of gravity, the mass and moment of inertia are constant, and it is only affected by gravity and propeller thrust, etc. These assumptions are often difficult to meet under the limitations of actual conditions such as the materials of the aircraft, assembly accuracy, and dynamic stability (for example, low battery power will affect control). Therefore, during the actual flight of the rotor UAV, the flight control results may deviate, posing a safety hazard.

[0003] The system identification method can calibrate and correct the kinematic mechanism model constructed under ideal conditions by combining actual operation data. This method can improve the consistency between the corrected kinematic model and the true model of the rotor UAV. To complete the identification, this method requires prior construction of a basic model structure (i.e., a dynamic model with parameters), and also requires collection of various data in various scenarios, covering various flight postures and weather conditions. On the one hand, constructing a basic dynamic model still requires strong domain knowledge, and there may still be a large difference between the overly simplified dynamic model under ideal conditions and the actual model; on the other hand, it is not easy to collect comprehensive data in real scenarios.

[0004] In recent years, with the development of machine learning technology in the field of robotics, people in the field can use very limited data learning and deep neural networks to very efficiently imitate expert strategies for manipulators, robotic dogs, UAVs, etc. Therefore, due to the symmetry between control strategies and dynamic models, machine learning methods can also be used to directly automatically learn dynamic models from limited data without prior construction of a basic model structure. Summary of the Invention

[0005] To solve the above problems, the present invention discloses a multi-rotor UAV modeling method and system based on machine learning.

[0006] The specific solutions are as follows:

[0007] A multi-rotor UAV modeling method based on machine learning, comprising the following steps:

[0008] S1. Collect the flight data of the rotary-wing UAV, segment the trajectory fragments according to the maximum length or the task termination condition, complete the cleaning of the preliminary data, and divide the data set into a training set and a validation set;

[0009] S2. According to the flight data of the rotary-wing UAV and the basic domain knowledge, divide the state variables and action variables (usually the control stick amount, throttle amount, etc.) from the data;

[0010] S3. Select the corresponding deep neural network model according to the dimensions of the state variables and action variables of the rotary-wing UAV and set the input and output;

[0011] S4. Use the adversarial imitation learning algorithm, define the training termination condition, and use the training data set to train the deep neural network model defined in step S3. During the training process, use a partial subset of the training set to periodically evaluate the currently trained model, so as to realize the modeling of the rotary-wing UAV;

[0012] S5. Use the validation set to evaluate the trained neural network model and evaluate the generalization accuracy of the model.

[0013] Further, in step S1, after collecting a certain amount of flight data (for example, the flight data accumulated for 1 hour in windy and breezy environments), perform trajectory segmentation on the flight data. For example, divide it according to the maximum trajectory length not exceeding 10 minutes. There should be a certain amount of overlap between adjacent trajectories in the divided trajectories. At the same time, obtain the training set and the validation set by sampling according to a ratio of 7:3 or 8:2, etc.

[0014] Further, in step S2, the pose, speed, blade rotation speed, etc. are set as the state variable S, and the throttle amount, stick amounts in all directions, etc. are set as the action variable a.

[0015] Further, in step S3, select a multi-layer perceptron model and a residual module as the dynamic model of the rotary-wing UAV. Its input is the state S at the current moment and the executed action a, and the output is the state S' at the next moment. For example, set a 5-layer multi-layer perceptron, and set the number of units in the middle layer to [256, 256, 512, 512, 256], and add residual modules at every other layer.

[0016] Further, step S4 is specifically as follows:

[0017] S41. Establish a multi-layer perceptron neural network MLP as a discriminator to determine the credibility of a generated (S,a). The credibility value finally output by the discriminator is a real number between 0 and 1. The closer it is to 1, the more it resembles real data, and the closer it is to 0, the more it resembles generated data. Set the convergence condition (for example, the discriminator outputs close to 0.5 on both generated data and real data) and the maximum number of iterations (for example, 5000 times).

[0018] S42. Sample a batch of starting points of trajectories from the real data as initial points, and interact with the defined neural network dynamics model, that is, input the current (S,a) to obtain S’, where the action amount a is obtained from the historical data or an additional neural network is established to generate it, until a specific trajectory length is reached or the termination condition is met, generating a batch of trajectory sequences.

[0019] S43. Update the discriminator with (S,a) in the generated trajectory sequence and (S,a) in the real data trajectory sequence. Denote the real data set as D and the generated data set during training as D’. The update objective is as follows:

[0020]

[0021] where f is the discriminator, f(S,a), f((S,a) ′ ) represent the credibility values output by the discriminator on a single piece of real data and generated data respectively.

[0022] S44. Use the updated discriminator to score each pair of generated (S,a) as the single-step reward r, and update the neural network dynamics model using the reinforcement learning algorithm. The reinforcement learning algorithm can choose PPO, or SAC when in the continuous action space.

[0023] S45. Repeat steps S42 to S44 until the maximum number of iterations or the convergence condition is reached.

[0024] In the present invention, the corresponding aerodynamic model is learned from a limited amount of actual flight data of rotor UAVs, that is, the data modeling method in model construction. The previous data modeling methods mainly use the technical means of constructing a basic dynamic equation in advance, then leaving out some parameters and fitting these parameters using actual data. This largely depends on the correctness of the constructed equation and the rationality of the reserved parameters. Another pure data modeling method uses deep learning, that is, the supervised learning method in machine learning for modeling. To obtain better results, this method usually requires collecting a very large amount of data covering various states and actions, and has high requirements for the data collection process. Machine learning methods can avoid prior explicit modeling, especially when using the generative adversarial imitation method, a high-fidelity model can be restored from relatively limited data.

[0025] In step S5, the generalization accuracy of the model is evaluated, specifically including:

[0026] (1) The mean absolute error of multiple steps on all trajectories in the validation set;

[0027] (2) the absolute error and relative proportion of the last step on all trajectories in the validation set;

[0028] (3) Visual comparison of typical trajectories.

[0029] A multi-rotor UAV modeling system based on machine learning, comprising:

[0030] The data integration and cleaning unit is used to collect the flight data of the rotary-wing UAV, segment the trajectory segments according to the maximum length or mission termination conditions, complete the cleaning of the preliminary data, and divide the data set into a training set and a validation set;

[0031] The state-action definition unit is used to divide the state quantity and action quantity from the data according to the flight data of the rotorcraft UAV and basic domain knowledge;

[0032] A neural network model selection unit, used to select a corresponding deep neural network model according to the dimensions of the rotorcraft state quantity and the motion quantity and set the input and output;

[0033] The adversarial imitation learning unit uses the adversarial imitation learning algorithm to define the training termination conditions, uses the training data set to train the defined deep neural network model, and uses a subset of the training set to periodically evaluate the currently trained model during the training process, thereby achieving modeling of the rotorcraft drone;

[0034] The model evaluation and verification unit evaluates the trained neural network model by using the verification set and evaluates the generalization accuracy of the model.

[0035] The beneficial effects of the present invention are:

[0036] 1. The present invention proposes a universal rotorcraft UAV modeling method based on machine learning, which can improve the modeling accuracy of rotorcraft UAV while reducing the modeling difficulty, that is, reduce the requirements for data volume and data coverage when modeling rotorcraft UAV data. At the same time, it also simplifies the process of personnel in fields that require professional directions to analyze dynamic characteristics in advance and construct dynamic models and kinematics in a targeted manner.

[0037] 2. The present invention can directly use a deep neural network as the basic model structure without specifying a dynamic model in advance; it can also use equations in parts where aerodynamics are clear and a neural network model in unclear parts, and learn in combination with data and the learning method in this application. The model obtained by learning in this application is a digital model with high fidelity, which can be used to construct flight training simulations, train flight control strategies, or develop new types of rotor UAVs based on this model, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of the present invention.

[0039] Figure 2 It is a detailed process diagram of adversarial imitation learning in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be further clarified below in conjunction with the specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0041] The present invention proposes a multi-rotor UAV modeling method based on machine learning. Based on adversarial imitation learning, it learns a rotor UAV model from the flight data of the rotor UAV collected in advance. The flight data can be collected in some basic flight tasks, and usually only the basic data of several processes such as takeoff, landing, low, medium, and high-speed operations are included. As Figure 1 shown, the specific operation steps of the present invention are as follows:

[0042] 110. Obtain the flight data of the rotor UAV, segment the trajectory segments according to the maximum length or task termination conditions, complete the cleaning of the preliminary data, and divide the data set into a training set and a validation set;

[0043] 210. According to the flight data of the rotor UAV and basic domain knowledge, divide the state variables and action variables (usually control stick variables, throttle variables, etc.) from the data;

[0044] 220. Select a corresponding deep neural network model according to the dimensions of the state variables and action variables of the rotor UAV and set the input and output;

[0045] 230. Use the generative adversarial imitation learning algorithm to define the training termination conditions, and use the training data set to train the deep neural network model defined in step

[220] . During the training process, a partial subset of the training set can be used to periodically evaluate the currently trained model, so as to realize the modeling of the rotor UAV;

[0046] 310. Use the validation set to evaluate the trained neural network model and evaluate the generalization accuracy of the model.

[0047] Among them, after a certain amount of flight data is collected (for example, flight data accumulated for 1 hour in windy and breezy environments), the flight data is segmented by trajectory. For example, the segmentation is based on that the maximum trajectory length does not exceed 10 minutes, and there should be a certain amount of overlap between adjacent trajectories in the segmented trajectories. At the same time, the training set and the validation set are sampled according to a ratio such as 7:3 or 8:2, and the pose, speed, blade rotation speed, etc. are set as the state quantity S, and the throttle amount, stick amounts in each direction, etc. are set as the action quantity a. This process corresponds to steps

[110] and

[210] in the present invention. Then, a multi-layer perceptron model and a residual module are selected as the dynamics model of the rotor UAV. Its input is the state S at the current moment and the executed action a, and the output is the state S' at the next moment. For example, a 5-layer multi-layer perceptron is set, and the number of units in the middle layer is set to [256, 256, 512, 512, 256], and a residual module is added to every other layer. This process corresponds to step

[220] in the present invention. After the data set is ready and the neural network model is set up, the adversarial imitation learning training in step

[230] is started. The detailed process of adversarial imitation learning is as Figure 2 shown:

[0048]

A100

[0049]

A200

[0050]

A300

[0051]

[0052] where f is the discriminator, f(S, a), f((S, a) ′respectively represent the credibility scores output by the discriminator on a single real data and generated data.

[0053]

A400

[0054] Repeat

A200

A400

[0055] In step

[310] of the present invention, after the model training is completed, multiple evaluation metrics can be used to evaluate the model on the validation set:

[0056] (1) The average absolute error over multiple steps on all trajectories in the validation set;

[0057] (2) The absolute error and relative ratio of the last step on all trajectories in the validation set;

[0058] (3) Visual comparison of typical trajectories.

[0059] The optimal model after model evaluation can be used as the rotor UAV dynamics model constructed in the present invention.

[0060] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions formed by any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A multi-rotor drone modeling method based on machine learning, characterized in that: The following steps are involved: S1. Collect the flight data of the rotary-wing UAV, divide the trajectory segments according to the maximum length or mission termination conditions, complete the cleaning of the preliminary data, and divide the data set into a training set and a validation set; S2, according to the flight data of the rotor UAV, dividing the state quantity and the action quantity from the data; S3, selecting a corresponding deep neural network model according to the dimensions of the rotorcraft state quantity and motion quantity and setting the input and output; S4. Using the adversarial imitation learning algorithm, defining the training termination condition, using the training data set to train the deep neural network model defined in step S3, and using a subset of the training set to periodically evaluate the currently trained model during the training process, thereby achieving modeling of the rotorcraft drone; S5. Use the validation set to evaluate the trained neural network model and assess the generalization accuracy of the model.

2. The multi-rotor drone modeling method based on machine learning according to claim 1 is characterized in that: There are overlapping parts between adjacent trajectories in the trajectories divided in step S1; in step S2, the posture, speed, and blade speed are set as the state quantity S, and the throttle amount and the stick amount in each direction are set as the action amount a; in step 3, a multi-layer perceptron model and a residual module are selected as the rotor UAV dynamics model, whose input is the state quantity S at the current moment and the executed action quantity a, and the output is the state quantity S' at the next moment.

3. The multi-rotor drone modeling method based on machine learning according to claim 2 is characterized in that: The step S4 is specifically as follows: S41. Establish a multi-layer perceptron neural network MLP as a discriminator to determine the credibility of a certain generation (S, a). The credibility of the final output of the discriminator is a real number between 0 and 1. The closer it is to 1, the more it resembles real data, and the closer it is to 0, the more it resembles generated data. Set the convergence condition and the maximum number of cycles. S42, sampling a batch of trajectory starting points from real data as initial points, interacting with the defined neural network dynamics model, that is, inputting the current (S, a), obtaining S', where the action amount a is obtained from historical data or generated by an additional neural network, until a specific trajectory length is reached or the termination condition is met, generating a batch of trajectory sequences; S43. Update the discriminator with (S, a) in the generated trajectory sequence and (S, a) in the trajectory sequence of the real data. The real data set is recorded as D, and the generated data set in the training process is recorded as D'. The update target is as follows: Where f is the discriminator, f(S,a),f((S,a) ′ ) represent the credibility of the discriminator output on a single real data and generated data respectively; S44. Use the updated discriminator to score each generated pair of (S, a) as the single-step reward r, and use the reinforcement learning algorithm to update the neural network dynamics model. The reinforcement learning algorithm can be PPO, or SAC in the continuous action space; S45. Repeat steps S42 to S44 until the maximum number of cycles or the convergence condition is reached.

4. The multi-rotor drone modeling method based on machine learning according to claim 1 is characterized in that: In step S5, the generalization accuracy of the model is evaluated, specifically including: (1) The mean absolute error of multiple steps on all trajectories in the validation set; (2) the absolute error and relative proportion of the last step on all trajectories in the validation set; (3) Visual comparison of typical trajectories.

5. A multi-rotor drone modeling system based on machine learning, characterized in that: include: The data integration and cleaning unit is used to collect the flight data of the rotary-wing UAV, segment the trajectory segments according to the maximum length or mission termination conditions, complete the cleaning of the preliminary data, and divide the data set into a training set and a validation set; The state-action definition unit is used to divide the state quantity and action quantity from the data according to the flight data of the rotorcraft UAV and basic domain knowledge; A neural network model selection unit, used to select a corresponding deep neural network model according to the dimensions of the rotorcraft state quantity and the motion quantity and set the input and output; The adversarial imitation learning unit uses the adversarial imitation learning algorithm to define the training termination conditions, uses the training data set to train the defined deep neural network model, and uses a subset of the training set to periodically evaluate the currently trained model during the training process, thereby achieving modeling of the rotorcraft drone; The model evaluation and verification unit evaluates the trained neural network model by using the verification set and evaluates the generalization accuracy of the model.

Citation Information

Patent Citations

  • Unmanned aerial vehicle strong-robustness attitude control method based on deep reinforcement learning

    CN114237268A

  • Machine learning system with anti-fact intervention

    CN118871918A

  • Intelligently modifying digital calendars utilizing a graph neural network and reinforcement learning

    US20220343155A1