Unmanned aerial vehicle control method, system and equipment based on meta-learning and medium

Through the drone control method based on meta-learning, a task neural network is constructed and meta-parameter update process is performed. The multi-target control quantity fusion process is combined with feedback linearization control quantity and disturbance compensation quantity, which solves the problem of unmodeled dynamics in high-speed flight and realizes high-precision and robust drone control.

CN120233792APending Publication Date: 2025-07-01SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371553.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing UAV control methods are difficult to cope with unmodeled dynamics caused by sudden changes in air resistance during high-speed flight, resulting in a decrease in control accuracy and difficult to meet real-time requirements.

Method used

The drone control method based on meta-learning is adopted to achieve precise control of the drone by constructing a task neural network set, performing meta-parameter update processing, obtaining online data sets, performing model prediction contour control optimization processing, and combining feedback linearized control amounts and disturbance compensation amounts to achieve multi-target control amount fusion processing.

Benefits of technology

It improves the accuracy of drone control, can adapt to high-speed flight and robust control under uncertain disturbances, balances tracking accuracy and flight speed, and enhances the generalization ability of the untrained speed domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233792A_ABST
    Figure CN120233792A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle control method, system and device based on meta-learning, and a medium. The method comprises the following steps: constructing a task neural network set based on a preset speed domain; performing meta-parameter updating processing on the task neural network set to obtain global meta-parameters; acquiring an online data set, inputting the online data set into the unmanned aerial vehicle controller, and performing model prediction contour control optimization processing on the unmanned aerial vehicle controller based on the global meta parameters to obtain control output; performing multi-target control quantity fusion processing on the unmanned aerial vehicle controller based on the control output in combination with the feedback linearization control quantity and the disturbance compensation quantity to obtain an unmanned aerial vehicle control result; and performing control processing on the unmanned aerial vehicle based on the unmanned aerial vehicle control result, and updating the prediction time domain online to perform iterative optimization processing on the unmanned aerial vehicle control result until the unmanned aerial vehicle control task is finished. The embodiment of the invention can improve the control precision of the unmanned aerial vehicle, and can be widely applied to the technical field of unmanned aerial vehicle control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of unmanned aerial vehicle control, and in particular, to an unmanned aerial vehicle control method and system, an electronic device, and a storage medium based on meta-learning. Background Art

[0002] In related technologies, quadrotor unmanned aerial vehicle control methods, such as L1 adaptive control, nonlinear dynamic inversion, etc., need to rely on fixed dynamic models or data-driven models trained offline, and it is difficult to cope with the unmodeled dynamics caused by sudden changes in air resistance during high-speed flight, and it is difficult to meet the real-time requirements. In summary, the technical problems existing in related technologies need to be improved. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose an unmanned aerial vehicle control method, system, device, and medium based on meta-learning, which can improve the control accuracy of the unmanned aerial vehicle.

[0004] To achieve the above object, on the one hand, an embodiment of the present application proposes an unmanned aerial vehicle control method based on meta-learning, and the method includes:

[0005] Construct a task neural network ensemble based on a preset speed domain;

[0006] Perform meta-parameter update processing on the task neural network ensemble to obtain global meta-parameters;

[0007] Obtain an online data set and input it into the unmanned aerial vehicle controller, and perform model predictive profile control optimization processing on the unmanned aerial vehicle controller based on the global meta-parameters to obtain a control output;

[0008] Perform multi-objective control quantity fusion processing on the unmanned aerial vehicle controller based on the control output in combination with a feedback linearization control quantity and a disturbance compensation quantity to obtain an unmanned aerial vehicle control result;

[0009] Perform control processing on the unmanned aerial vehicle based on the unmanned aerial vehicle control result, and online update the prediction time domain to perform iterative optimization processing on the unmanned aerial vehicle control result until the unmanned aerial vehicle control task ends.

[0010] In some embodiments, the constructing a task neural network ensemble based on a preset speed domain includes the following steps:

[0011] Perform data acquisition processing on the unmanned aerial vehicle based on the preset speed domain to obtain a task data set, and the task data set includes state data and control data;

[0012] Perform data-driven modeling processing on the task data set to obtain the task neural network ensemble.

[0013] In some embodiments, the meta-parameter update process for the task neural network ensemble to obtain global meta-parameters includes the following steps:

[0014] Perform model parameter update processing on each neural network in the task neural network ensemble based on the stochastic gradient descent algorithm to obtain a local parameter set;

[0015] Perform error calculation processing on the task neural network ensemble based on the first loss function to obtain a prediction error set;

[0016] Update the meta-parameters based on the local parameter set and the prediction error set to obtain the global meta-parameters.

[0017] In some embodiments, the model predictive contour control optimization process for the UAV controller based on the global meta-parameters to obtain a control output includes the following steps:

[0018] Perform dynamic parameter adjustment processing on the UAV controller based on the global meta-parameters to obtain a learning rate, and the UAV controller adopts a model predictive contour control strategy;

[0019] Perform path tracking processing on the UAV based on the learning rate and the multi-objective cost function to obtain the control output.

[0020] In some embodiments, the dynamic parameter adjustment process for the UAV controller based on the global meta-parameters to obtain a learning rate includes the following steps:

[0021] Calculate the tracking error based on the online data set;

[0022] When the tracking error is greater than a preset threshold, perform function loss calculation processing on the UAV controller based on the global meta-parameters through a second loss function that fuses the tracking error and regularization constraints to obtain a control loss;

[0023] Dynamically adjust the learning rate of the UAV controller based on the control loss to obtain the learning rate.

[0024] In some embodiments, the path tracking process for the UAV based on the learning rate and the multi-objective cost function to obtain the control output includes the following steps:

[0025] Obtain the optimized contour error, path progress, and control input;

[0026] Perform adaptive adjustment processing on the contour weight based on the path progress to obtain a dynamic contour weight;

[0027] Construct a multi-objective cost function based on the optimized contour error, the path progress, the control input, and the dynamic contour weight;

[0028] Calculate the gradient of the multi-objective cost function based on the learning rate and update the parameters of the UAV controller to obtain an optimized controller;

[0029] Perform tracking control on the UAV based on the optimized controller and output the control output.

[0030] In some embodiments, the multi-objective control quantity fusion process of the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result includes the following steps:

[0031] Perform feedback linear control processing on the UAV to obtain the feedback linearization control quantity;

[0032] Calculate the control error of the UAV based on the control output to obtain the disturbance compensation quantity;

[0033] Perform weight assignment and fusion processing on the control output, the feedback linearization control quantity, and the disturbance compensation quantity to obtain the UAV control result.

[0034] To achieve the above object, another aspect of the embodiments of the present application proposes a UAV control system based on meta-learning, and the system includes:

[0035] The first module is used to construct a task neural network set based on a preset speed domain;

[0036] The second module is used to perform meta-parameter update processing on the task neural network set to obtain global meta-parameters;

[0037] The third module is used to obtain an online data set and input it into the UAV controller, and perform model predictive contour control optimization processing on the UAV controller based on the global meta-parameters to obtain a control output;

[0038] The fourth module is used to perform multi-objective control quantity fusion processing on the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result;

[0039] The fifth module is used to perform control processing on the UAV based on the UAV control result and iteratively optimize the UAV control result by online updating the prediction horizon until the UAV control task ends.

[0040] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned method is implemented.

[0041] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0042] The embodiments of the present application at least include the following beneficial effects: The present application provides a method, system, device and medium for controlling an unmanned aerial vehicle (UAV) based on meta-learning. This solution constructs a set of task neural networks based on a preset speed domain, and can strengthen the generalization ability across speed domains through data-driven modeling, providing a dynamic basis for model predictive contour control. In addition, this solution performs meta-parameter update processing on the set of task neural networks to obtain global meta-parameters; obtains an online data set and inputs it into the UAV controller, and optimizes the model predictive contour control of the UAV controller based on the global meta-parameters to obtain a control output, which can improve the generalization ability of the model for tasks with untrained speeds based on meta-learning and provide a fast adaptive dynamic model basis for online model predictive contour control. Moreover, this solution performs multi-objective control quantity fusion processing on the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result; controls the UAV based on the UAV control result, and iteratively optimizes the UAV control result by updating the prediction horizon online until the UAV control task ends. This solution can dynamically adjust the contour error weight of the model predictive contour control, thereby balancing the tracking accuracy and speed during high-speed flight, adapting to robust control under high-speed flight and uncertain disturbances, and improving the accuracy of UAV control. Description of the Drawings

[0043] Figure 1 is a flowchart of a method for controlling an unmanned aerial vehicle based on meta-learning provided by an embodiment of the present application;

[0044] Figure 2 is a model framework diagram of UAV control provided by an embodiment of the present application;

[0045] Figure 3 is a schematic structural diagram of a system for controlling an unmanned aerial vehicle based on meta-learning provided by an embodiment of the present application;

[0046] Figure 4 is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed Embodiments

[0047] In order to make the objectives, technical solutions and advantages of this application more clearly understood, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of systems and methods that are consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0048] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0049] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, at least one includes one, two or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0051] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0052] 1) Meta-learning refers to the algorithm's ability to summarize a strategy from past experience to help it quickly learn when facing new tasks. This is different from traditional machine learning methods, which usually rely on a large amount of data to train models, while meta-learning focuses on how to achieve efficient learning through a small amount of data. Meta-learning can be regarded as a process of "learning how to learn", that is, the model not only learns the laws of the task itself, but also learns how to use the knowledge of previous tasks to accelerate the learning process of the current task.

[0053] 2) Model Predictive Contour Control (MPCC) is an advanced control strategy that combines the advantages of Model Predictive Control (MPC) and contour control. MPC is an advanced control strategy that can solve the control problems of complex dynamic systems. By predicting the system behavior within a certain period of time in the future and optimizing based on this, the best control effect of the system can be achieved. Contour control focuses on the precise tracking of the path and the kinematic constraints of the robot.

[0054] In related technologies, there are control methods for quadrotor UAVs, such as L1 adaptive control, nonlinear dynamic inversion, etc. These methods rely on fixed dynamic models or data-driven models trained offline. However, in practical applications, it is found that these methods are difficult to cope with the unmodeled dynamics caused by sudden changes in air resistance during high-speed flight, resulting in difficulties in accurately controlling the UAV in unknown narrow spaces and complex environments, and affecting the control accuracy of the UAV.

[0055] Exemplarily, for example, when using traditional data-driven methods in UAV control, such as Gaussian process, residual learning, etc., the tracking error increases significantly in the untrained speed domain, and the computational complexity is high, making it difficult to meet the real-time requirements and resulting in insufficient model generalization ability. In related technologies, the controller based on feedback linearization is sensitive to instantaneous wind field disturbances and is prone to trajectory deviation, especially in the near-ground effect or inter-aircraft turbulence scenarios, there are potential safety hazards. In addition, although Model Predictive Control (MPC) can optimize trajectory tracking, the fixed weight parameters are difficult to balance the contour error and progress during high-speed flight, resulting in speed loss or tracking accuracy decline during turning.

[0056] In view of this, in the embodiments of the present application, a drone control method, system, device, and medium based on meta-learning are provided. This solution constructs a set of task neural networks based on a preset speed domain, and can strengthen the generalization ability across speed domains through data-driven modeling, providing a dynamic basis for model predictive contour control. In addition, this solution obtains global meta-parameters by performing meta-parameter update processing on the set of task neural networks; obtains an online data set and inputs it into the drone controller, and optimizes the model predictive contour control of the drone controller based on the global meta-parameters to obtain a control output, which can improve the generalization ability of the model for untrained speed tasks based on meta-learning and provide a fast-adaptive dynamic model basis for online model predictive contour control. Moreover, this solution performs multi-objective control quantity fusion processing on the drone controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the drone control result; controls the drone based on the drone control result, and iteratively optimizes the drone control result by updating the prediction horizon online until the drone control task ends. This solution can dynamically adjust the contour error weight of the model predictive contour control, thereby balancing the tracking accuracy and speed during high-speed flight, adapting to robust control under high-speed flight and uncertain disturbances, and improving the accuracy of drone control.

[0057] A drone control method based on meta-learning provided by the embodiments of the present application relates to the field of information technology. The drone control method based on meta-learning provided by the embodiments of the present application can be applied to a terminal, or to a server, or can be software running on a terminal or a server. The drone is controlled through the terminal, server, or software. The control method can be sent to the drone controller through the terminal, so that the drone controller controls the drone. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing a drone control method based on meta-learning, etc., but is not limited to the above forms.

[0058] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0059] Figure 1 is an alternative flowchart of a method for controlling an unmanned aerial vehicle based on meta-learning provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S101 to S105.

[0060] Step S101, construct a set of task neural networks based on a preset speed domain;

[0061] Step S102, perform meta-parameter update processing on the set of task neural networks to obtain global meta-parameters;

[0062] Step S103, obtain an online dataset and input it to the unmanned aerial vehicle controller, and perform model predictive contour control optimization processing on the unmanned aerial vehicle controller based on the global meta-parameters to obtain a control output;

[0063] Step S104, perform multi-objective control quantity fusion processing on the unmanned aerial vehicle controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain an unmanned aerial vehicle control result;

[0064] Step S105, perform control processing on the unmanned aerial vehicle based on the unmanned aerial vehicle control result, and online update the prediction horizon to perform iterative optimization processing on the unmanned aerial vehicle control result until the unmanned aerial vehicle control task ends.

[0065] Steps S101 to S105 shown in the embodiments of the present application construct a speed-dependent dynamic model through a meta-learning framework, model different speed domains as independent learning tasks, and thus obtain a set of task neural networks, that is, neural networks corresponding to different preset speed domains are modeled. Then, according to the set of task neural networks, meta-parameter update processing is performed, and the model parameters can be updated in real time by combining the online incremental learning strategy to achieve the generalization ability for un-trained speed domains. Among them, the meta-parameters belong to the meta-learning framework and are used to generate the prior of the dynamic model across speed domains. Its optimization goal is to improve the generalization ability of the model for un-trained speed tasks and provide a fast-adaptive dynamic model basis for online MPCC. In the embodiments of the present application, by obtaining the online data set and inputting it into the UAV controller, and optimizing the model predictive profile control of the UAV controller based on the global meta-parameters, continuous optimization of the model predictive control can be achieved through real-time data stream processing and dynamic parameter update, so as to obtain the control output output by the UAV controller. The embodiments of the present application integrate the model predictive profile control output, feedback linearization control and disturbance compensation, and can drive the UAV controller through multi-objective control fusion and attitude execution, improving the accuracy of UAV control. Finally, through closed-loop iteration and re-optimization, the prediction horizon is updated in real time and model fine-tuning is triggered until the task ends.

[0066] One of the above technical solutions has the following advantages or beneficial effects: In the embodiments of the present application, a speed-dependent dynamic model is constructed through a meta-learning framework, different speed domains are modeled as independent learning tasks, and the model parameters are updated in real time by combining the online incremental learning strategy to achieve the generalization ability for un-trained speed domains. In addition, in the embodiments of the present application, through high-level model predictive profile control (MPCC), the control input is optimized based on the meta-learning dynamic model, and by dynamically adjusting the profile error weight, the tracking accuracy and flight speed can be balanced. The low-level controller combines feedback linearization and radial basis neural networks, and can approximate the un-modeled disturbance online, improving the adaptive ability of the system to sudden airflow and turbulence. Experiments show that in the embodiments of the present application, under high speed (1-10 m / s) and dynamic wind field interference, the tracking error is significantly lower than that of traditional control strategies, which is suitable for the precise control of UAVs in unknown narrow spaces and complex environments.

[0067] In step S101 of some embodiments, constructing the set of task neural networks based on the preset speed domain includes the following steps:

[0068] Performing data acquisition processing on the UAV based on the preset speed domain to obtain a task data set, where the task data set includes state data and control data;

[0069] Performing data-driven modeling processing based on the task data set to obtain the set of task neural networks.

[0070] In the embodiments of the present application, data collection and task-based modeling are performed based on a preset multi-speed domain, and a task-specific neural network model can be constructed based on the preset speed domain. In a feasible embodiment, the unmanned aerial vehicle performs uniform flight at preset speed domains (1 m / s, 3 m / s, 5 m / s), and records 12-dimensional state data (three-dimensional position p x , p y , p z , attitude angle linear velocity v x , v y , v z , angular velocity p, q, r) and 4-dimensional motor thrust control inputs F1, F2, F3, F4, and noise interference is eliminated through data cleaning and normalization processing (normalized to the interval [-1, 1]). Subsequently, each speed domain is modeled as an independent task Construct a task-specific neural network f NN : The input layer receives 12-dimensional state and 4-dimensional control signals, is mapped through a hidden layer with 64-32-32 nodes (ReLU activation), and outputs a 6-dimensional state increment Δx'. In the embodiments of the present application, the loss function of each task-specific neural network incorporates the mean square error and the L2 regularization term, and the expression of the first loss function is as follows:

[0071]

[0072] In the formula, x is the current state, x' is the measured state at the next moment, u is the control output, and the regularization coefficient λ offline = 0.01, is the task data set, and θ i is the model parameter.

[0073] One of the above technical solutions has the following advantages or beneficial effects: The embodiments of the present application strengthen the generalization ability across speed domains through data-driven modeling, providing a dynamic basis for model predictive control.

[0074] In some embodiments, the step of performing meta-parameter update processing on the task neural network ensemble to obtain global meta-parameters includes the following steps:

[0075] Performing model parameter update processing on each neural network in the task neural network ensemble based on the stochastic gradient descent algorithm to obtain a local parameter set;

[0076] Performing error calculation processing on the task neural network ensemble based on the first loss function to obtain a prediction error set;

[0077] Performing update processing on the meta-parameters based on the local parameter set and the prediction error set to obtain the global meta-parameters.

[0078] In the embodiments of the present application, for each task neural network in the established task neural network ensemble, stochastic gradient descent (SGD) is used to update local parameters, and the expression of the update algorithm is as follows:

[0079]

[0080] In the formula, the learning rate α = 0.01, iterate 5 times (k = 0, 1..., 4), and use the task-specific dataset to optimize the parameter θ i , the parameters of each task neural network after optimization can be obtained, and a local parameter set can be obtained to enable the model to quickly fit the dynamic characteristics of a specific speed domain.

[0081] Among them, the first loss function is the loss function of the task neural network, which is used to calculate the average prediction error of each task on the independent validation set to obtain a prediction error set, so as to update the meta-parameters based on the local parameter set and the prediction error set to obtain global meta-parameters. The meta-parameter φ belongs to the meta-learning framework and is used to generate a prior of the dynamic model across speed domains. Its optimization goal is to improve the generalization ability of the model for untrained speed tasks and provide a fast-adaptive dynamic model basis for online MPCC. The formula for updating the global meta-parameters is as follows:

[0082]

[0083] In the formula, is the local parameter after the i-th task is iterated by the inner-layer stochastic gradient descent (SGD), the learning rate β = 0.001, is the independent validation set of each task. In the embodiments of the present application, to overcome the local optimum defect of traditional single-layer optimization, a second-order optimization technique is introduced in the outer-layer gradient calculation to explicitly model the dependence relationship between the inner-layer parameter θ i and the meta-parameter φ. The finally generated meta-model φ* has both task adaptability and generalization, providing core support for the online fast parameter tuning of subsequent model predictive contour control (MPCC).

[0084] One of the above technical solutions has the following advantages or beneficial effects: In the embodiments of the present application, the generalization ability of the model for untrained speed tasks is improved through meta-learning, and a fast-adaptive dynamic model basis is provided for online MPCC, improving the accuracy of UAV control.

[0085] In some embodiments, the model predictive contour control optimization process is performed on the UAV controller based on the global meta-parameters to obtain a control output, including the following steps:

[0086] Performing dynamic parameter adjustment processing on the UAV controller based on the global meta-parameters to obtain a learning rate, where the UAV controller adopts a model predictive contour control strategy;

[0087] Performing path tracking processing on the UAV based on the learning rate and the multi-objective cost function to obtain the control output.

[0088] In the embodiments of the present application, continuous optimization of model predictive control is achieved through real-time data stream processing and dynamic parameter update. Among them, the UAV control adopts a model predictive contour control strategy, and the embodiments of the present application optimize the model predictive contour control strategy through dynamic parameter update. The UAV controller can perform incremental learning based on the tracking error. Incremental learning is a machine learning method that allows the model to continuously learn from new data instead of retraining the entire model from scratch. Its goal is to solve the problem of catastrophic forgetting in model training, that is, when training on a new task, the performance on the old task will decline. The characteristic of incremental learning is that it can continuously learn new knowledge from new samples and retain the knowledge that has been learned. Moreover, the embodiments of the present application achieve robust tracking of complex paths through a dynamic weight adjustment strategy and multi-objective cost function design, and drive the underlying attitude controller by fusing MPCC output, feedback linearization control, and disturbance compensation.

[0089] One of the above technical solutions has the following advantages or beneficial effects: In the embodiments of the present application, the meta-parameters can quickly adapt to flight environment disturbances through dual-scale adaption, and through multi-objective control fusion and attitude execution, the system can approximately approach unmodeled disturbances online, improving the system's adaptive ability to sudden airflow and turbulence and enhancing the accuracy of UAV control.

[0090] In some embodiments, the performing dynamic parameter adjustment processing on the UAV controller based on the global meta-parameters to obtain a learning rate includes the following steps:

[0091] Calculating the tracking error according to the online data set;

[0092] When the tracking error is greater than a preset threshold, performing function loss calculation processing on the UAV controller based on the global meta-parameters through a second loss function that fuses the tracking error and regularization constraints to obtain a control loss;

[0093] Performing dynamic adjustment processing on the learning rate of the UAV controller based on the control loss to obtain the learning rate.

[0094] In the embodiments of the present application, continuous optimization of model predictive control is achieved through real-time data stream processing and dynamic parameter update. Real-time data stream processing collects sensor data (including position, attitude, and control input) at a period of 50 ms to construct an online data set When the tracking error ||e|| > 0.01m, the incremental learning mechanism is triggered. The incremental parameter update adopts stochastic gradient descent (SGD) with a dynamically adjusted learning rate. Its loss function combines the tracking error and regularization constraints. That is, the expression of the second loss function is as follows:

[0095]

[0096] In the formula, the regularization coefficient λ online = 0.001, and the learning rate η online is dynamically adjusted according to the error change: if the error decreases for 3 consecutive iterations, then η ← f ↑ η (f ↑ = 1.1); if the error increases, then η ← f ↓ η (f ↓ = 0.9), and the initial value of η init = 0.01. According to the calculated control loss, the learning rate can be dynamically adjusted to obtain the learning rate of the UAV controller. Among them, the learning rate is a key hyperparameter in machine learning and deep learning, which controls the step size of the model parameters moving in each update. An appropriate learning rate can accelerate the training speed and improve the model performance; an overly high learning rate may cause the model to diverge or oscillate, while an overly low learning rate may cause slow convergence or getting stuck in a local optimum.

[0097] One technical solution in the above technical solutions has the following advantages or beneficial effects: Through dual-scale adaption, data triggering, and learning rate adjustment in the embodiments of the present application, the balance between the model update speed and stability is achieved. While suppressing overfitting, the meta-parameters can quickly adapt to the flight environment disturbance, improving the accuracy of UAV control.

[0098] In some embodiments, the path tracking process of the UAV based on the learning rate and the multi-objective cost function to obtain the control output includes the following steps:

[0099] Obtain the optimized contour error, path progress, and control input;

[0100] Based on the path progress, perform an adaptive adjustment process on the contour weight to obtain a dynamic contour weight;

[0101] Based on the optimized contour error, the path progress, the control input, and the dynamic contour weight, construct a multi-objective cost function;

[0102] Based on the gradient of the multi-objective cost function calculated by the learning rate, update the parameters of the UAV controller to obtain an optimized controller;

[0103] Based on the optimized controller, the UAV is tracked and controlled to output the control output.

[0104] In the embodiment of the present application, the optimized contour error is calculated through the error between the model predictive contour control output and the actual contour output by the UAV controller, that is, the vertical deviation between the actual flight trajectory of the UAV and the desired path, which reflects the path tracking accuracy. The path progress is the completion percentage of the UAV along the predetermined path or the ratio of the remaining distance to time, and the control input is the drive signal of the UAV actuator (motor, servo). Based on the path progress, the contour weight can be adaptively adjusted to obtain the dynamic contour weight. The calculation formula of the dynamic contour weight is as follows:

[0105]

[0106] In the formula, q(κ(β k )) represents the dynamic contour weight, β k represents the path progress, q nom = 0.8 is the reference weight, P g,j is the path inflection point position, and the covariance matrix ∑ = diag(0.1, 0.1, 0.1) controls the adjustment intensity. In the turning area (curvature κ > 0.5), the weight is increased to 1.2 to enhance the anti-disturbance ability, and in the straight section, it is restored to q nom to reduce energy consumption. Based on the optimized contour error, path progress, control input, and dynamic contour weight, a multi-objective cost function can be constructed. The expression of the multi-objective cost function is as follows:

[0107]

[0108] In the formula, J MPCC represents the model predictive contour control output, the dynamic contour weight q(κ k ) is driven by the path curvature κ k , ρ = 0.2 is the progress weight, R is the control input penalty matrix, e c is the optimized contour error, β k is the path progress, and u k is the control input. Finally, based on the learning rate, the gradient of the multi-objective cost function is calculated to update the parameters of the UAV controller to obtain the optimized controller. Based on the optimized UAV controller, the UAV can be tracked and controlled, and the control output can be obtained through the output of the UAV controller.

[0109] One of the above technical solutions has the following advantages or beneficial effects: In the embodiment of the present application, the model predictive contour control (MPCC) is optimized by dynamically adjusting the contour error weight, which can balance the tracking accuracy and speed during high-speed flight, thereby improving the accuracy of UAV control.

[0110] In some embodiments, the multi-objective control quantity fusion processing is performed on the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result, including the following steps:

[0111] Performing feedback linear control processing on the UAV to obtain the feedback linearization control amount;

[0112] Performing control error calculation processing on the UAV based on the control output to obtain the disturbance compensation amount;

[0113] The control output, the feedback linearization control amount and the disturbance compensation amount are subjected to weight distribution fusion processing to obtain the UAV control result.

[0114] In the embodiment of the present application, the feedback linearization control in the drone control is a control method that converts the originally complex nonlinear system into a linear system through nonlinear state feedback and coordinate transformation, thereby simplifying the controller design and improving the control performance, and the feedback linearization control quantity can be output. Among them, the feedback linearization control law converts the nonlinear dynamic model into a double integrator form, and the conversion formula is as follows:

[0115]

[0116] Where x is the state vector of the UAV, v is the virtual control input generated by feedback linearization, is disturbance compensation, which is obtained through control error. Finally, the UAV control result is obtained by weight distribution fusion processing of control output, feedback linearization control quantity and disturbance compensation quantity. Among them, the weight distribution fusion formula is shown as follows:

[0117]

[0118] In the formula, u total Represents the final output result, u MPCC represents the model prediction contour control output, u FL represents the feedback linearization control quantity, This weight allocation strategy balances the long-term optimization capability of MPCC (60%), the fast response of feedback linearization (30%), and the disturbance compensation of RBFNN (10%).

[0119] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application combines feedback linearization and disturbance compensation through a controller, which can approximate unmodeled disturbances online, improve the system's adaptability to sudden airflow and turbulence, and is suitable for precise control of drones in unknown narrow spaces and complex environments.

[0120] Next, in combination with specific application examples, the solutions of the embodiments of the present application will be introduced and described in detail:

[0121] The embodiments of the present application can be applied to the application scenario of controlling an unmanned aerial vehicle (UAV). The control method of the embodiments of the present application is implemented through a controller, so as to perform tasks such as trajectory tracking on the UAV. Please refer to Figure 2 , in the embodiments of the present application, a task neural network is constructed through a preset speed domain (V = 1m / s, 3m / s, 5m / s) and inner-loop update is performed. The inner-loop update is performed on the parameters θ n of the task neural network n by means of the corresponding task data set D. And through acquisition, the meta-learning framework can be based on the stochastic gradient descent algorithm (SGD) and based on the real-time latest state parameters x new and control quantity parameters u new of the UAV to form a function sequence f(x new , u new ; φ) to perform outer-loop update on the meta-learning parameters, that is, to obtain the control error of the UAV to update the meta-parameters to obtain new meta-parameters φ new , and adaptively dynamically adjust the weight of the model predictive contour control (MPCC) through the control error, optimize the model predictive contour control through the meta-learning parameters φ, and combine the preset path P des to perform multi-objective control fusion and attitude execution, output the final control result u to the UAV, and the UAV will output the latest state parameters x' new to iteratively optimize the model parameters until the control task of the UAV ends. Specifically, in the embodiments of the present application, a speed-dependent dynamic model is constructed through the meta-learning framework, different speed domains are modeled as independent learning tasks, and the model parameters are updated in real time in combination with the online incremental learning strategy to achieve the generalization ability for un-trained speed domains. The high-level model predictive contour control (MPCC) optimizes the control input based on the meta-learning dynamic model, and balances the tracking accuracy and flight speed by dynamically adjusting the contour error weight. The low-level controller combines feedback linearization and a radial basis neural network to online approximate the unmodeled disturbance and improve the adaptive ability of the system to sudden airflow and turbulence.

[0122] Please refer to Figure 3 , the embodiments of the present application also provide a UAV control system based on meta-learning, which can implement the above-mentioned UAV control method based on meta-learning. The system includes:

[0123] A first module 301, configured to construct a task neural network set based on a preset speed domain;

[0124] A second module 302, configured to perform meta-parameter update processing on the task neural network set to obtain global meta-parameters;

[0125] The third module 303 is configured to obtain an online data set and input it into the UAV controller, and perform model predictive contour control optimization processing on the UAV controller based on the global meta-parameters to obtain a control output;

[0126] The fourth module 304 is configured to perform multi-objective control quantity fusion processing on the UAV controller based on the control output in combination with a feedback linearization control quantity and a disturbance compensation quantity to obtain a UAV control result;

[0127] The fifth module 305 is configured to perform control processing on the UAV based on the UAV control result, and online update the prediction horizon to perform iterative optimization processing on the UAV control result until the UAV control task ends.

[0128] It can be understood that the content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented in the present system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0129] The embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned UAV control method based on meta-learning. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0130] It can be understood that the content in the above method embodiments is applicable to the present device embodiment. The functions specifically implemented in the present device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0131] Please refer to Figure 4 , Figure 4 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0132] A processor 401, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0133] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute a method for controlling an unmanned aerial vehicle based on meta-learning according to an embodiment of the present application;

[0134] The input / output interface 403 is used to implement information input and output;

[0135] The communication interface 404 is used to implement communication and interaction between this device and other devices. Communication can be achieved through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.);

[0136] The bus 405 transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);

[0137] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 achieve communication connections with each other inside the device through the bus 405.

[0138] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for controlling an unmanned aerial vehicle based on meta-learning.

[0139] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0140] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0141] A drone control method, system, device and medium based on meta-learning provided by an embodiment of the present application. This solution constructs a set of task neural networks based on a preset speed domain, and can strengthen the generalization ability across speed domains through data-driven modeling, providing a dynamic basis for model predictive contour control. In addition, this solution obtains global meta-parameters by performing meta-parameter update processing on the set of task neural networks; obtains an online data set and inputs it into the drone controller, and optimizes the model predictive contour control of the drone controller based on the global meta-parameters to obtain a control output, which can improve the generalization ability of the model for untrained speed tasks based on meta-learning and provide a fast-adaptive dynamic model basis for online model predictive contour control. Moreover, this solution performs multi-objective control quantity fusion processing on the drone controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the drone control result; controls the drone based on the drone control result, and iteratively optimizes the drone control result by updating the prediction horizon online until the drone control task ends. This solution can dynamically adjust the contour error weight of the model predictive contour control, thereby balancing the tracking accuracy and speed during high-speed flight, adapting to robust control under high-speed flight and uncertain disturbances, and improving the accuracy of drone control.

[0142] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0143] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0144] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations.

[0146] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0147] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0148] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical or other forms.

[0149] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0152] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A UAV control method based on meta-learning, characterized in that: The method comprises the following steps: Constructing a set of task neural networks based on a preset speed domain; Performing meta-parameter updating processing on the task neural network set to obtain global meta-parameters; Obtaining an online data set as input to a UAV controller, and performing model prediction contour control optimization processing on the UAV controller based on the global meta-parameters to obtain a control output; Based on the control output combined with the feedback linearization control amount and the disturbance compensation amount, a multi-objective control amount fusion process is performed on the UAV controller to obtain a UAV control result; The drone is controlled based on the drone control result, and the drone control result is iteratively optimized in an online updated prediction time domain until the drone control task is completed.

2. The method according to claim 1, characterized in that The step of constructing a task neural network set based on a preset speed domain includes the following steps: Performing data collection and processing on the UAV based on the preset speed domain to obtain a task data set, wherein the task data set includes state data and control data; Data-driven modeling is performed based on the task data set to obtain the task neural network set.

3. The method according to claim 1, characterized in that The step of performing meta-parameter updating processing on the task neural network set to obtain global meta-parameters includes the following steps: Based on the stochastic gradient descent algorithm, the model parameters of each neural network in the task neural network set are updated to obtain a local parameter set; Performing error calculation processing on the task neural network set based on the first loss function to obtain a prediction error set; The meta-parameters are updated based on the local parameter set and the prediction error set to obtain the global meta-parameters.

4. The method according to claim 1, characterized in that: The step of performing model prediction profile control optimization processing on the UAV controller based on the global meta-parameters to obtain a control output comprises the following steps: Based on the global meta-parameters, the UAV controller is dynamically adjusted to obtain a learning rate, and the UAV controller adopts a model prediction profile control strategy; The UAV is subjected to path tracking processing based on the learning rate and the multi-objective cost function to obtain the control output.

5. The method according to claim 4, characterized in that The method of dynamically adjusting the parameters of the drone controller based on the global meta-parameters to obtain a learning rate comprises the following steps: Calculate the tracking error according to the online data set; When the tracking error is greater than a preset threshold, a second loss function that fuses the tracking error and the regularization constraint is used to perform function loss calculation processing on the UAV controller based on the global meta-parameters to obtain a control loss; The learning rate of the drone controller is dynamically adjusted based on the control loss to obtain the learning rate.

6. The method according to claim 4, characterized in that The performing path tracking processing on the UAV based on the learning rate and the multi-objective cost function to obtain the control output comprises the following steps: Get optimized contour error, path progress and control input; Adaptively adjusting the contour weight based on the path progress to obtain a dynamic contour weight; Constructing a multi-objective cost function based on the optimized contour error, the path progress, the control input and the dynamic contour weight; Calculating the gradient of the multi-objective cost function based on the learning rate to update the parameters of the drone controller to obtain an optimized controller; The UAV is tracked and controlled based on the optimization controller, and the control output is obtained as an output.

7. The method according to any one of claims 1 to 6, characterized in that: The method of performing multi-objective control quantity fusion processing on the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result includes the following steps: Performing feedback linear control processing on the UAV to obtain the feedback linearization control amount; Performing control error calculation processing on the UAV based on the control output to obtain the disturbance compensation amount; The control output, the feedback linearization control amount and the disturbance compensation amount are subjected to weight distribution fusion processing to obtain the UAV control result.

8. A UAV control system based on meta-learning, characterized in that: The system comprises: The first module is used to construct a task neural network set based on a preset speed domain; The second module is used to perform meta-parameter update processing on the task neural network set to obtain global meta-parameters; The third module is used to obtain an online data set and input it into a UAV controller, and perform model prediction contour control optimization processing on the UAV controller based on the global meta-parameters to obtain a control output; The fourth module is used to perform multi-objective control quantity fusion processing on the UAV controller based on the control output combined with the feedback linearization control quantity and the disturbance compensation quantity to obtain the UAV control result; The fifth module is used to control the drone based on the drone control result, and iteratively optimize the drone control result in an online updated prediction time domain until the drone control task is completed.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.