Intelligent scheduling method and system for drones based on knowledge-driven meta-learning device

By constructing a meta-task environment and introducing physical guidance items based on a knowledge-driven meta-learning device, the problems of data dependency and computing resource constraints in UAV intelligent scheduling are solved, and the rapid adaptation and efficient scheduling of UAV models in dynamic environments are achieved.

CN120031361BActive Publication Date: 2025-09-09SYST OVERALL RES INST INST OF SYST ENG ACAD OF MILITARY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510520931.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-09-09
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing meta-learning technology has problems in drone intelligent scheduling, such as strong data dependence, insufficient generalization ability and computing resource constraints, making it difficult to quickly adapt to new mission scenarios in dynamic and complex environments.

Method used

A knowledge-driven meta-learning device is used to optimize meta-parameters through a meta-training module, construct a meta-task environment covering multiple tasks, introduce physical guidance terms to optimize the loss function, and use unlabeled data for self-training, combining forward propagation and backpropagation to update the parameter set.

Benefits of technology

The drone intelligent scheduling model was able to quickly optimize its performance under data scarcity conditions, which improved the model's generalization and adaptability, shortened training time, and improved meta-learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031361B_ABST
    Figure CN120031361B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for intelligent scheduling of unmanned aerial vehicles (UAVs) based on a knowledge-driven meta-learning device. Specifically, the method comprises: deploying and initializing a knowledge-driven meta-learning device, including meta-training and target retraining modules; the meta-training module determines the meta-learning objectives of the UAV intelligent scheduling model and optimizes meta-parameters, performing weight estimation through linear regression and measuring the similarity between different feature representations through a similarity metric; constructing a meta-task environment covering multiple tasks to learn cross-task update rules, using feature domain basis matrices to form orthogonal multi-domain feature representations; introducing a physical guidance term into the meta-learning loss function to align the output with physical laws and optimize the meta-learning objective function; and the target retraining module performs self-training using unlabeled data, updating the parameter set to obtain the final model, and directing and optimizing the collaborative operation of multiple UAVs. This invention can rapidly optimize the performance of the UAV intelligent scheduling model even when data is scarce.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning and artificial intelligence technology, and in particular to a method and system for intelligent scheduling of unmanned aerial vehicles (UAVs) based on a knowledge-driven meta-learning device. Background Art

[0002] Drone scheduling often faces complex dynamic environments such as weather changes, sudden tasks, and multi-machine collaboration. Traditional methods rely on predefined rules or models trained on a single task, which makes it difficult to quickly adapt to new scenarios. Meta-learning can improve the generalization ability of the scheduling system through cross-task learning.

[0003] Meta-learning, a method for "learning how to learn," has demonstrated strong potential in scenarios such as few-shot learning and transfer learning. Its core concept is to learn model parameters or optimization rules across multiple tasks to improve adaptability to new tasks. In practical applications, meta-learning is often closely related to few-shot learning, particularly when data is scarce or new tasks are involved, enabling efficient learning with a small number of examples.

[0004] However, existing meta-learning technologies still have the following significant problems in applications related to drone intelligent scheduling: (1) Strong data dependence: Current meta-learning methods usually require a large amount of labeled data and diverse task samples to construct the meta-task environment. However, in many practical scenarios, data collection and labeling are expensive, especially in dynamic and complex environments, where obtaining sufficient training data is almost impossible. (2) Insufficient generalization ability: The generalization ability of most meta-learning methods depends on manually defined task distributions. When the model faces unseen tasks or drastically changing environments, its adaptability is significantly reduced. Especially in new task scenarios, due to data scarcity, the model performance often fails to meet expectations. (3) Computational resource constraints: Meta-learning frameworks usually require frequent gradient calculations and meta-parameter adjustments, resulting in high computational costs. In embedded devices or resource-constrained environments, existing methods have difficulty achieving real-time processing and rapid decision-making.

[0005] To address the above issues, researchers have attempted to improve the adaptability of meta-learning in drone intelligent scheduling through model lightweighting and few-sample optimization techniques. However, these methods often have the following defects in the actual application of drone intelligent scheduling: 1. Simplifying model functions or reducing the dimension of input signals will significantly limit the human-machine intelligent scheduling model's ability to handle complex tasks; 2. Failure to fully utilize prior knowledge in the scene means that the performance of the human-machine intelligent scheduling model in the target scene is still limited. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for intelligent scheduling of drones based on a knowledge-driven meta-learning device with short training time and high efficiency, so that the model can efficiently adapt to new mission scenarios, quickly optimize the performance of the target model through a small amount of training data, and achieve rapid convergence without the need for large-scale labeled data.

[0007] The technical solution to achieve the purpose of the present invention is: a method for intelligent scheduling of drones based on a knowledge-driven meta-learning device, comprising the following steps:

[0008] Step 1: Deploy and initialize the knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module;

[0009] Step 2: The meta-training module determines the meta-learning objectives of the UAV intelligent scheduling model and optimizes the meta-parameters, estimates the weights through linear regression, and measures the similarity between different feature representations through similarity metrics;

[0010] Step 3: Construct a meta-task environment covering multiple tasks, so that the UAV intelligent scheduling model can learn cross-task update rules and use the feature domain basis matrix to form an orthogonal multi-domain feature expression;

[0011] Step 4: Introduce a physical guidance term into the meta-learning loss function to align the output of the UAV intelligent scheduling model with physical laws and optimize the meta-learning objective function.

[0012] Step 5: The target retraining module uses unlabeled data to self-train the UAV intelligent scheduling model, combines forward propagation and backpropagation to update the parameter set, and obtains the final UAV intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple UAVs.

[0013] A UAV intelligent scheduling system based on a knowledge-driven meta-learning device is provided. The system is used to implement the UAV intelligent scheduling method based on a knowledge-driven meta-learning device. The system includes the first to fifth units, and the functions of each module are as follows:

[0014] The first unit is used to deploy and initialize a knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module;

[0015] In the second unit, the meta-learning objectives of the UAV intelligent scheduling model are determined through the meta-training module, and the meta-parameters are optimized. The weights are estimated through linear regression, and the similarity between different feature representations is measured through similarity metrics.

[0016] Unit 3: By building a meta-task environment covering multiple tasks, the UAV intelligent scheduling model learns cross-task update rules and uses the feature domain basis matrix to form an orthogonal multi-domain feature expression;

[0017] Unit 4: By introducing a physical guidance term into the meta-learning loss function, the output of the UAV intelligent scheduling model is aligned with physical laws and the meta-learning objective function is optimized.

[0018] In the fifth unit, the drone intelligent scheduling model is self-trained with unlabeled data through the target retraining module, and the parameter set is updated by combining forward propagation and backpropagation to obtain the final drone intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple drones.

[0019] Compared with the existing technology, the present invention has the following significant advantages: (1) By constructing a universal meta-parameter update rule, the intelligent scheduling of drones can efficiently adapt to new mission scenarios. Even when data is scarce, the performance of the target model can be quickly optimized with a small amount of training data, and the model has strong generalization ability and rapid adaptability; (2) In the target mission scenario, rapid convergence can be achieved without the need for large-scale labeled data, which shortens the training time and improves the efficiency of meta-learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of the intelligent scheduling method of drones based on the knowledge-driven meta-learning device of the present invention. DETAILED DESCRIPTION

[0021] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] The present invention provides a method for intelligently scheduling drones based on a knowledge-driven meta-learning device, comprising the following steps:

[0023] Step 1: Deploy and initialize the knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module;

[0024] Step 2: The meta-training module determines the meta-learning objectives of the UAV intelligent scheduling model and optimizes the meta-parameters, estimates the weights through linear regression, and measures the similarity between different feature representations through similarity metrics;

[0025] Step 3: Construct a meta-task environment covering multiple tasks, so that the UAV intelligent scheduling model can learn cross-task update rules and use the feature domain basis matrix to form an orthogonal multi-domain feature expression;

[0026] Step 4: Introduce a physical guidance term into the meta-learning loss function to align the output of the UAV intelligent scheduling model with physical laws and optimize the meta-learning objective function.

[0027] Step 5: The target retraining module uses unlabeled data to self-train the UAV intelligent scheduling model, combines forward propagation and backpropagation to update the parameter set, and obtains the final UAV intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple UAVs.

[0028] As a specific example, the deployment and initialization of the knowledge-driven meta-learning device described in step 1, which includes a meta-training module and a target retraining module, is as follows:

[0029] The meta-training module is used to define the meta-learning objectives, build the meta-learning environment, and generate general meta-parameter update rules through training. The meta-training module can process multimodal input data and optimize the meta-objective function in multi-task scenarios;

[0030] The target retraining module is used to adapt to the target task scenario based on the data samples and the knowledge obtained by the meta-training module. The target retraining module forms a target model by adjusting model parameters and rules.

[0031] Furthermore, the meta-learning device is deployed on a general-purpose computer device, a computing platform, an embedded device or an edge computing device. According to the device type, the deployment methods are divided into three types: stand-alone deployment, cluster deployment and edge inference deployment. The stand-alone deployment method is suitable for general-purpose computer devices and embedded devices, and can be directly run after the software environment is installed. The cluster deployment method is suitable for computing platforms or cloud environments, and a distributed framework is used to achieve multi-task parallel processing. The edge inference deployment method converts the optimized lightweight model into a format suitable for edge computing devices, and performs real-time inference through edge computing devices.

[0032] As a specific example, the meta-training module in step 2 determines the meta-learning objective of the UAV intelligent scheduling model and optimizes the meta-parameters, performs weight estimation through linear regression, and measures the similarity between different feature representations through similarity metrics, as follows:

[0033] Step 2.1: The meta-training module determines the meta-learning objective of the UAV intelligent scheduling model and defines the meta-learning objective as optimizing the meta-parameters , meta-learning goal The formula is:

[0034]

[0035] in, is the meta-parameter of the UAV intelligent scheduling model, Indicates the neural network The weights of the layers, represents the layer number of the neural network, and Indicates the total number of layers of the neural network; Indicates the task number, Indicates in The meta-learning objective function on each task is concretized as a loss function or performance indicator; It means finding the mathematical expectation of each task in the task distribution, that is, taking the average over multiple tasks;

[0036] Step 2.2: Gradient optimization of meta-parameters. In the inner loop, use Model parameter set for the current task Perform multiple iterations and updates. It is a meta-parameter learned across tasks, used to initialize and guide Rapid adaptation and optimization in specific tasks, the formula is:

[0037]

[0038] in, is the control parameter The learning rate for updating the step size; represents the iteration round, For the The meta-parameters after iterations (such as the weight vector of the neural network) represent the meta-parameter status of the current round; For the The meta-parameters after iterations represent the results of updating the current meta-parameters according to the gradient information;

[0039] is the meta-learning objective function, which is a loss function used to measure the performance of the current task. The input includes the model parameter set of the current task. And sample data for training tasks , which determines how the meta parameters are optimized. It is the sample data used for training in the current task or scenario. It may be a small amount of labeled data or unlabeled data. It is an important input for the model to learn specific task features.

[0040] Step 2.3: Perform weight estimation to capture the relationship between task features and introduce the weight estimation formula for linear regression:

[0041]

[0042] in, is the optimal weight vector; Representation Label and feature representation The degree of difference between is the regularization term; is the regularization coefficient, which is a hyperparameter used to control the regularization term The influence of ,prevents the model from overfitting on the training data, thereby improving the generalization ability; It is the parameter vector in linear regression, which is used to learn the mapping relationship between task features and output labels. In this scenario, it represents the weight coefficient of the relationship between the fitted feature representation and the label. Indicates the sample index, used to refer to the A sample, such as a sample in the meta-task environment, the index can be traversed within the training set or support set; For the The sample in the neural network Feature representation extracted by the layer;

[0043] Step 2.4: Introduce similarity metrics into the meta-learning objective function to measure the similarity between different feature representations. The formula is:

[0044]

[0045] in, is the optimal weight vector and feature representation The result of linear combination; The cosine distance function is used to measure the angle difference between two vectors, reflecting their directional similarity. The smaller the value, the more similar they are. In meta-learning, it is used to evaluate the directional consistency between the model prediction output and the true label. The test sample number is used to identify different test samples, usually taken from the query set outside the support set; For test samples The true label vector represents the first The expected output of the samples; For test samples In the neural network The feature representation extracted by the layer comes from the neural network The output of the layer is used to represent the high-dimensional abstract features of the sample;

[0046] As a specific example, step 3 constructs a meta-task environment covering multiple tasks, enabling the UAV intelligent scheduling model to learn cross-task update rules and use the feature domain basis matrix to form an orthogonal multi-domain feature expression, as follows:

[0047] Step 3.1: Diversify the task distribution and build a meta-task environment covering multiple tasks, so that the UAV intelligent scheduling model can learn the update rules across tasks and build a feature matrix using multi-dimensional perception data. , generate multiple subspaces by eigendecomposition:

[0048]

[0049] in, is the subspace of the feature matrix, representing the characteristics of the environment domain, network domain, and behavior domain respectively; Represents the feature matrix is a two-dimensional matrix in the real field with Line and Column, the dimension in the real number space is Matrix of Indicates the number of sampled features, matrix Each row in corresponds to a sample; Indicates the number of tasks or time steps; if applied to a multi-task scenario, then Indicates the number of different tasks involved; if it is a time series feature, it can also be interpreted as the number of time segments or the number of task evolution stages;

[0050] Step 3.2: Multi-domain feature representation: the basis matrices of each feature domain are all unit matrices, forming an orthogonal multi-domain feature expression, making the features between tasks discriminative and combinable.

[0051] As a specific example, in step 4, a physical guidance term is introduced into the meta-learning loss function to align the output of the drone intelligent scheduling model with physical laws and optimize the meta-learning objective function. The details are as follows:

[0052] Step 4.1: Meta-learning objective Expanding further, the expression is as follows:

[0053]

[0054] in, It's on a mission The loss function on Representation based on meta parameters The model is trained on data Training conducted on Indicates that after a gradient update, in the test data The losses, including is the learning rate used for gradient updates; Represents the distribution of meta-tasks Sampling task Finally, the expected value of the objective function is calculated to measure the average performance of the meta-learning model on all training tasks; Represents the balance coefficient of the inner and outer layer loss functions, which is used to adjust the relative weight of the loss in the training phase and the loss in the test phase in the optimization objective; Represents a pair parameter Calculate the gradient to reflect the directional update information of the current model on the training loss function; Indicates a task The training set loss function on ;

[0055] Step 4.2: Meta-learning loss function Introducing physical guidance items to align the output of the UAV intelligent scheduling model with physical laws;

[0056] Physics-guided loss function The form is:

[0057]

[0058] in, is the data-driven loss term used by the model during training; It is the physical guidance loss term introduced by the model to align the model output with the physical laws; It is the weight of data-driven loss, and its value is adjusted according to the needs of actual application; is the weight of each physical guidance loss item. Different priorities are given to different physical guidance items by adjusting the weights. Represents the index number of the physical guidance loss term, traversing all physical constraint types to be introduced; Represents the total number of physical guided loss terms, which is used to control the number or type of physical laws introduced by the model;

[0059] Step 4.3, loss matrix expansion:

[0060] and Contains multiple loss values ​​used to describe the error of the UAV intelligent scheduling model on different data samples or tasks. They are all in vector or matrix form. Therefore, the physical guidance loss function is expanded as follows:

[0061]

[0062] in, represents the first The loss value, represents the first The first loss loss value; , is the total number of loss values;

[0063] Step 4.4. Finally, the meta-learning objective is written as:

[0064]

[0065] in, Represents the meta-task set Sampling task After that, the task The overall loss function on is used to perform expectation operations to evaluate the average generalization performance of the model in the task distribution; Indicates that in the task Based on the training data and the current model meta-parameters Calculated data-driven loss terms, such as mean squared error, cross entropy, etc. Indicates that in the task The physical guidance loss term on is used to align the model output with specific physical laws (such as conservation laws, differential equations, motion models, etc.), thereby enhancing physical consistency and interpretability; It represents the balance coefficient of the loss function between the inner and outer layer tasks, and is used to control the proportion of the generalization effect of the model on the test set in the objective function before and after a gradient update.

[0066] As a specific example, the target retraining module in step 5 self-trains the drone intelligent scheduling model using unlabeled data, combines forward propagation and backpropagation to update the parameter set, and obtains the final drone intelligent scheduling model, as follows:

[0067] Step 5.1: Collect data from the target task scenario and perform forward propagation;

[0068] Step 5.2: Customize the error signal, normalize the error signal, and calculate the back propagation weight parameters;

[0069] Step 5.3: Calculate the forward weight update parameters and bias update parameters.

[0070] Step 5.4: Combine the forward weight update parameters, bias update parameters, and backpropagation weight update parameters of all layers into a parameter set for output.

[0071] As a specific example, step 5.1 collects data from the target task scene and performs forward propagation as follows:

[0072] Step 5.1.1: The model receives limited unlabeled data As input, data is collected from the target task scenario;

[0073] Step 5.1.2, Data Passed to each hidden layer of the model in turn, Input data in the layer By using the forward weight and bias Perform linear transformation to generate an inactive intermediate representation :

[0074]

[0075] Step 5.1.3, through batch normalization To normalize:

[0076]

[0077] in, Represents the normalized , represents the batch normalization function;

[0078] Step 5.1.4, use The activation function will Converted to the activated output :

[0079]

[0080] The activation function sets negative values ​​to zero and only keeps positive values, thus introducing nonlinearity.

[0081] As a specific example, the custom error signal described in step 5.2 is normalized and the back propagation weight parameters are calculated as follows:

[0082] Step 5.2.1, Custom error signal definition:

[0083] In each layer, Layer The neuron in The local error signal generated by the neuron position The calculation method is:

[0084]

[0085] in, represents the custom error signal for the top layer, yes The derivative of the activation function, where It is in The intermediate representations that are not activated in the layer; The derivative of shows how the error changes with the inactive output Change, if the activation function is ,but The derivative of is a step function; represents element-wise multiplication, The result is equivalent to adjusting the custom error signal to the error amplitude suitable for the current layer; and is the meta-parameter used to propagate the error signal; It comes from the upper layer, The error signal of the neurons in the layer is Sum and calculate the error contribution of all upper layer units to the current neuron, and then calculate the update of each layer weight through the convolutional neural network module;

[0086] Step 5.2.2, Error signal normalization:

[0087] Calculate the normalized error signal , ensuring that the amplitude of error signals at different levels is consistent during propagation, avoiding gradient explosion or vanishing problems. The calculation formula is:

[0088]

[0089] in, Indicates in Layer The neuron pairs The original error signal generated by samples; for Normalized error signal; It is the normalization dimension correction coefficient, which is used to adjust the contribution of error signals of different dimensions to the overall normalized amplitude to avoid imbalance of error signal amplitude in high-dimensional space; Indicates in Layer The neuron pairs The original error signal generated by the samples is understood as the error signal tensor A component corresponding to each sample;

[0090] Step 5.2.3, back propagation weight calculation:

[0091] The error signal is propagated back through the weights Passed to the upper layer neurons, the calculation formula is:

[0092]

[0093] in, Indicates the Layer The neuron in The total error signal received at the neuron position is The weighted summary of all relevant error information of the layer; Indicates the Layer The neuron in The local error signal generated by the position of each neuron is often used as the error transfer term in back propagation, indicating the sensitivity of the neuron to the loss function. Indicates the Layer from the The neuron transmits to the The forward connection weights between neurons.

[0094] As a specific example, the forward weight update parameters and bias update parameters are calculated as follows:

[0095] Step 5.3.1, forward weight update parameters:

[0096] Using Functions Calculate weight updates:

[0097]

[0098] Among them, the function Used in the back propagation process of the neural network to calculate the update amount of the weight; function Receiving input data , intermediate activation value , pre-activation value , model parameters and meta parameters In each layer, the function uses a setting function based on the local part of the neuron, combining the forward propagation signal and the error signal from the back propagation to determine how to calculate the update; Indicates the The weight update matrix of the layer neural network, that is, the gradient direction adjustment value calculated for the layer in this round of backpropagation, is used to update the weight of the layer;

[0099] Step 5.3.2, weight normalization:

[0100] Calculate the update , and normalize between layers to avoid gradient explosion or disappearance problems. The formula is:

[0101]

[0102] in, Indicates the Layer The normalized forward weight update corresponding to the neurons is used to replace the unnormalized As the final updated value; Indicates the Layer The unnormalized forward weight update corresponding to each neuron; Indicates the The number of neurons in the layer is the number of neurons in the current The input layer of the layer is used to calculate the total number of connection weights; Indicates the The number of neurons in the layer, that is, the output dimension and number of channels of the current layer; Indicates the Layer The neuron corresponding to Input channel and The updated value of the single connection weight between the output channels is composed of part of; Indicates the The index of the layer neuron, used to traverse all input channels; Indicates the The index of the layer neurons, used to traverse all output channels; The neurons connected Layer and All channels between layers;

[0103] Step 5.3.3, weight update formula:

[0104] To make the final weight update smooth, each weight update plane is linearly combined to generate the final update:

[0105]

[0106] in, is the learning rate for weight updates; Indicates the Layer The original forward weight matrix at the iteration, which represents the current weight state before the update; Indicates the Layer The updated forward weight matrix at the iteration is the result of fusing the old parameters and the update amount;

[0107] Step 5.3.4, bias update parameters:

[0108] Calculate the bias update to avoid all units entering the linear state and ensure that the network learns nonlinear relationships. The initial bias update is obtained by batch aggregation over the hidden dimension:

[0109]

[0110] in, For the The initial bias update value of the layer is obtained by aggregating the feature responses of the samples in the batch; Indicates the first samples; is the low-rank feature dimension index, the total number is ; is the weight vector in the low-rank readout module, where Indicates the The weights on the dimensions, For the Layer The sample in The hidden response on the channel; the whole formula represents the low-rank feature A weighted average is performed within the data to generate a stable bias adjustment base value;

[0111] represents the batch size, represents hidden dimensions; is the term related to activation and is multiplied by weight; Used to constrain bias updates to ensure that the bias does not cause the entire layer to enter an overly linear state:

[0112]

[0113] in, Indicates the The bias update amount of the layer after constraint processing, Indicates the initial bias update vector The bias update value corresponding to each neuron is used to calculate the average bias trend.

[0114] As a specific example, the forward weight update parameters, bias update parameters, and backpropagation weight update parameters of all layers described in step 5.4 are combined into a parameter set for output, as follows:

[0115] Will As a function set, it includes the first layer to the Forward weight update of the layer , bias update and backpropagation weight updates , expressed as:

[0116]

[0117] The present invention also provides a drone intelligent scheduling system based on a knowledge-driven meta-learning device, which is used to implement the drone intelligent scheduling method based on a knowledge-driven meta-learning device. The system includes first to fifth units, and the functions of each module are as follows:

[0118] The first unit is used to deploy and initialize a knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module;

[0119] In the second unit, the meta-learning objectives of the UAV intelligent scheduling model are determined through the meta-training module, and the meta-parameters are optimized. The weights are estimated through linear regression, and the similarity between different feature representations is measured through similarity metrics.

[0120] Unit 3: By building a meta-task environment covering multiple tasks, the UAV intelligent scheduling model learns cross-task update rules and uses the feature domain basis matrix to form an orthogonal multi-domain feature expression;

[0121] Unit 4: By introducing a physical guidance term into the meta-learning loss function, the output of the UAV intelligent scheduling model is aligned with physical laws and the meta-learning objective function is optimized.

[0122] In the fifth unit, the drone intelligent scheduling model is self-trained with unlabeled data through the target retraining module, and the parameter set is updated by combining forward propagation and backpropagation to obtain the final drone intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple drones.

[0123] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0124] Example

[0125] Combine Figure 1 This embodiment provides a method for intelligently scheduling drones based on a knowledge-driven meta-learning device, comprising the following steps:

[0126] Step 1: Deploy and initialize the knowledge-driven meta-learning device, including the meta-training module and the target retraining module, as follows:

[0127] The meta-training module is used to define the meta-learning objectives, build the meta-learning environment, and generate general meta-parameter update rules through training. This module can process multimodal input data and optimize the meta-objective function in multi-task scenarios;

[0128] The target retraining module is used to quickly adapt to the target task scenario based on limited data samples and the knowledge obtained by the meta-training module. This module can adjust model parameters and rules to form an efficient target model;

[0129] The meta-learning device is deployed on a general-purpose computer device, computing platform, embedded device or edge computing device. Depending on the device type, the deployment mode is divided into three types: stand-alone deployment, cluster deployment and edge inference deployment.

[0130] The stand-alone deployment method is applicable to general-purpose computers and embedded devices, and can be run directly after installing the necessary software environment;

[0131] The cluster deployment method is suitable for high-performance computing platforms or cloud environments, and uses distributed frameworks such as Horovod and DeepSpeed ​​to achieve multi-task parallel processing;

[0132] The edge inference deployment method converts the optimized lightweight model into a format suitable for edge devices, such as TensorRT and ONNX, and performs real-time inference through edge devices.

[0133] Step 2: The meta-training module determines the meta-learning target and performs pre-training as follows:

[0134] Step 2.1. Determine the meta-learning goal and define the meta-learning goal as optimizing meta-parameters , the target formula is:

[0135]

[0136] in, It is the meta-parameter of the UAV intelligent scheduling model, including the weight matrix or tensor of the multi-layer network; Indicates in The meta-learning objective function on each task can be concretized as a loss function or performance indicator;

[0137] Step 2.2: Gradient optimization of meta-parameters. In the inner loop, use Model parameter set for the current task Perform multiple iterative updates, the formula is:

[0138]

[0139] in, is the control parameter The learning rate for updating the step size;

[0140] Step 2.3: Weight estimation. To capture the relationship between task features, the weight estimation formula for linear regression is introduced:

[0141]

[0142] in, is the optimal weight vector; Representation Label and feature representation The degree of difference between is a regularization term used to prevent the model from overfitting;

[0143] Step 2.4: Similarity metric. Similarity metric is introduced into the meta-objective function to measure the similarity between different feature representations. The formula is:

[0144]

[0145] in, is the optimal weight obtained previously Task characteristics Linear combination results.

[0146] Step 3: Build a meta-task environment covering multiple tasks to ensure that the model can learn universal update rules across tasks and use the basis matrices of each feature domain to form an orthogonal multi-domain feature expression to ensure that the features between tasks are well distinguishable and combinable. The details are as follows:

[0147] Step 3.1: Diversify the task distribution and build a meta-task environment covering multiple tasks to ensure that the model can learn universal update rules across tasks and use multi-dimensional perception data to build a feature matrix , generate multiple subspaces by eigendecomposition:

[0148]

[0149] in, is the subspace of the feature matrix, representing the characteristics of the environment domain, network domain, and behavior domain respectively;

[0150] Step 3.2: Multi-domain feature representation. The basis matrices of each feature domain are all unit matrices, forming an orthogonal multi-domain feature expression to ensure that the features between tasks have good distinguishability and combination, such as the "magic cube" structure. Each slice or combination of the feature matrix corresponds to the feature representation at a different time or position.

[0151] Step 4: Introduce a physical guidance term into the meta-learning loss function to align the model output with physical laws and optimize the meta-learning objective function. The details are as follows:

[0152] Step 4.1: Further expand the meta-learning objective and express it as follows:

[0153]

[0154] in, It's on a mission The loss function on Representation based on meta parameters The model is trained on data The training conducted on Indicates that after a gradient update, in the test data The losses, including is the learning rate used for gradient updates;

[0155] Step 4.2: Add the meta-learning loss function Introducing physical guidance terms to align model outputs with physical laws. When constructing physical guidance terms based on known physical equations, nonlinear differential equations containing high-order differential terms are often involved. Introducing a physical guidance loss function avoids directly solving the numerical solution of the differential equation and instead extracts physical information by balancing the equation. By performing a weighted combination of multiple loss functions, a more diverse range of loss functions can be constructed, thereby improving the adaptability and physical consistency of the model.

[0156] The physical guidance loss function is in the form of:

[0157]

[0158] in, It is the data-driven loss term mainly used by the model during training. It can be a standard loss function based on the data label, such as cross entropy loss in classification problems or mean squared error in regression problems; It is a physics-guided loss term introduced by the model to align the model output with physical laws. In meta-learning or other scientific applications, these loss terms are constructed based on known physical equations, such as differential equations, to help the model optimize under the constraints of physical laws. is the weight of the data-driven loss. This value can be adjusted according to the needs of the actual application. For example, in scenarios where more attention is paid to the performance of the model on the data, this value can be increased. is the weight of each physics-guided loss term. By adjusting these weights, different priorities can be given to different physics-guided terms to ensure that the model complies with physical laws.

[0159] Step 4.3, loss matrix expansion:

[0160] and Contains multiple loss values ​​used to describe the model's error on different data samples or tasks, all in vector or matrix form, so the physical guidance loss function can be expanded as:

[0161]

[0162] in, represents the first The loss value, represents the first The first loss values;

[0163] Step 4.4. Finally, the meta-learning objective can be written as:

[0164]

[0165] Step 5: The target retraining module trains the model on a small amount of unlabeled data, combines forward propagation and backpropagation to update the parameter set, and uses the meta-parameters generated in the meta-training phase to quickly optimize the model to adapt it to the needs of the target scenario. The details are as follows:

[0166] Step 5.1: Collect data from the target task scenario and perform forward propagation, as follows:

[0167] Step 5.1.1: The model receives limited unlabeled data As input, data is collected from the target task scenario;

[0168] Step 5.1.2, Data Passed to each hidden layer of the model in turn , at each layer , enter data By using the forward weight and bias Perform linear transformation to generate an inactive intermediate representation :

[0169]

[0170] Step 5.1.3, through batch normalization To normalize:

[0171]

[0172] Step 5.1.4, use The activation function will Converted to the activated output :

[0173]

[0174] Negative values ​​are set to zero and only positive values ​​are retained, thus introducing nonlinearity. During the backward propagation of the inner loop, the model calculates the error signal of each layer. And adjust the forward weight , thereby gradually optimizing the feature representation. Based on a custom error definition at the top level, so it does not depend on data labeling.

[0175] Step 5.2: Customize the error signal, normalize the error signal, and calculate the back propagation weight parameters as follows:

[0176] Step 5.2.1, Custom error signal definition:

[0177] At each layer, the error signal is calculated as:

[0178]

[0179] in, represents the custom error signal for the top layer, yes The derivative of the activation function, where It is in The intermediate representations that are not activated in the layer; The derivative of shows how the error changes with the inactive output Change, if the activation function is ,but The derivative of is a step function; represents element-wise multiplication, The result is equivalent to adjusting the custom error signal to the error amplitude suitable for the current layer; and is the meta-parameter used to propagate the error signal; It comes from the upper layer, The error signal of the neurons in the layer is Sum and calculate the error contribution of all upper layer units to the current neuron, and then calculate the update of each layer weight through the convolutional neural network module;

[0180] Step 5.2.2, Error signal normalization:

[0181] Calculate the normalized error signal , the calculation formula is:

[0182]

[0183] Ensure that the amplitude of error signals at different levels is consistent during propagation to avoid gradient explosion or vanishing problems;

[0184] Step 5.2.3, back propagation weight calculation:

[0185] The error signal is propagated back through the weights Passed to the upper layer neurons, the calculation formula is:

[0186]

[0187] Step 5.3: Calculate the forward weight update parameters and bias update parameters as follows:

[0188] Step 5.3.1, forward weight update parameters:

[0189] Using Functions Calculate weight updates:

[0190]

[0191] function Receiving input data , intermediate activation value , pre-activation value , model parameters and meta parameters In each layer, the function uses a specific function based on the local part of the neuron, which may include low-rank readout, such as LowRR, which combines the forward propagation signal and the error signal from the backpropagation to determine how the update is calculated; low-rank readout projects high-dimensional features to low dimensions to simplify the calculation of weight updates.

[0192] Step 5.3.2, weight normalization:

[0193] Function computes the update , and normalize between layers to avoid gradient explosion or vanishing problems; the update formula for the forward weight is:

[0194]

[0195] Step 5.3.3, weight update formula:

[0196] To make the final weight update smooth, each weight update plane is linearly combined to generate the final update:

[0197]

[0198] in, is the learning rate for weight updates;

[0199] Step 5.3.4, bias update parameters:

[0200] Bias updates are calculated to prevent all units from entering the linear regime, thus ensuring that the network learns nonlinear relationships. The initial bias updates are obtained by batch aggregation over the hidden dimension:

[0201]

[0202] represents the batch size, represents hidden dimensions; is the term related to activation and is multiplied by weight; Used to constrain bias updates to ensure that the bias does not cause the entire layer to enter an overly linear state:

[0203]

[0204] Step 5.4: Combine the forward weight update parameters, bias update parameters, and backpropagation weight update parameters of all layers into a parameter set for output, as follows:

[0205] ComputeDeltaWeight is a function set, including the first layer to the Forward weight update of the layer , bias update and backpropagation weight updates , expressed as:

[0206]

[0207] By constructing universal meta-parameter update rules, this paper enables the intelligent drone scheduling model to efficiently adapt to new mission scenarios. Even in data-scarce environments, it can rapidly optimize the performance of the target model using a small amount of training data. In the target mission scenario, rapid convergence is achieved without the need for large amounts of labeled data, significantly reducing training time.

[0208] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for intelligent scheduling of drones based on a knowledge-driven meta-learning device, characterized in that: The following steps are involved: Step 1: Deploy and initialize the knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module; Step 2: The meta-training module determines the meta-learning objectives of the UAV intelligent scheduling model and optimizes the meta-parameters, estimates the weights through linear regression, and measures the similarity between different feature representations through similarity metrics; Step 3: Construct a meta-task environment covering multiple tasks, so that the UAV intelligent scheduling model can learn cross-task update rules and use the feature domain basis matrix to form an orthogonal multi-domain feature expression, as follows: Step 3.1: Diversify the task distribution and build a meta-task environment covering multiple tasks, so that the UAV intelligent scheduling model can learn the update rules across tasks and build a feature matrix using multi-dimensional perception data. , generate multiple subspaces by eigendecomposition: ; in, is the subspace of the feature matrix, representing the characteristics of the environment domain, network domain, and behavior domain respectively; Represents the feature matrix is a two-dimensional matrix in the real field with Line and List; Indicates the number of sampled features, matrix Each row in corresponds to a sample; Indicates the number of tasks or time steps; Step 3.2: Multi-domain feature representation: the basis matrices of each feature domain are all identity matrices, forming an orthogonal multi-domain feature representation, making the features between tasks discriminative and combinable; Step 4: Introduce a physical guidance term into the meta-learning loss function to align the output of the UAV intelligent scheduling model with physical laws and optimize the meta-learning objective function. Step 5: The target retraining module uses unlabeled data to self-train the UAV intelligent scheduling model, combines forward propagation and backpropagation to update the parameter set, and obtains the final UAV intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple UAVs.

2. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 1, characterized in that: Deploy and initialize the knowledge-driven meta-learning device described in step 1, which includes a meta-training module and a target retraining module, as follows: The meta-training module is used to define the meta-learning objectives, build the meta-learning environment, and generate general meta-parameter update rules through training. The meta-training module can process multimodal input data and optimize the meta-objective function in multi-task scenarios; The target retraining module is used to adapt to the target task scenario based on the data samples and the knowledge obtained by the meta-training module. The target retraining module forms a target model by adjusting model parameters and rules.

3. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 2, characterized in that: The meta-training module described in step 2 determines the meta-learning objectives of the UAV intelligent scheduling model and optimizes the meta-parameters. It estimates the weights through linear regression and measures the similarity between different feature representations through similarity metrics, as follows: Step 2.1: The meta-training module determines the meta-learning objective of the UAV intelligent scheduling model and defines the meta-learning objective as optimizing the meta-parameters , meta-learning goal The formula is: ; in, is the meta-parameter of the UAV intelligent scheduling model, Indicates the neural network The weights of the layers, represents the layer number of the neural network, and Indicates the total number of layers of the neural network; Indicates the task number, Indicates in The meta-learning objective function on each task is concretized as a loss function or performance indicator; It means finding the mathematical expectation of each task in the task distribution, that is, taking the average over multiple tasks; Step 2.2: Gradient optimization of meta-parameters, using Model parameter set for the current task Perform multiple iterative updates, the formula is: ; in, is the control parameter The learning rate for updating the step size; represents the iteration round, For the Meta parameters after iterations; For the Meta parameters after iterations; is the meta-learning objective function, which is a loss function used to measure the performance of the current task. The input includes the model parameter set of the current task. And sample data for training tasks ; Step 2.3: Perform weight estimation to capture the relationship between task features and introduce the weight estimation formula for linear regression: ; in, is the optimal weight vector; Representation Label and feature representation The degree of difference between is the regularization coefficient, is a regularization term used to prevent the model from overfitting; is the parameter vector in linear regression, which is used to learn the mapping relationship between task features and output labels; Indicates the sample index, used to refer to the samples; For the The sample in the neural network Feature representation extracted by the layer; Step 2.4: Introduce similarity metrics into the meta-learning objective function to measure the similarity between different feature representations. The formula is: ; in, is the optimal weight vector and feature representation The result of linear combination; is the cosine distance function, which is used to measure the angle difference between two vectors; The test sample number is used to identify different test samples; For test samples The true label vector of For test samples In the neural network The feature representation extracted by the layer.

4. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 3, characterized in that: In step 4, we introduce a physical guidance term into the meta-learning loss function, align the output of the UAV intelligent scheduling model with physical laws, and optimize the meta-learning objective function as follows: Step 4.1: Meta-learning objective Expanding further, the expression is as follows: ; in, It's on a mission The loss function on Representation based on meta parameters The model is trained on data Training conducted on Indicates that after a gradient update, in the test data The losses, including is the learning rate used for gradient updates; Represents the distribution of meta-tasks Sampling task After that, the expected value of the objective function is calculated; Represents the balance coefficient of the inner and outer layer loss functions, which is used to adjust the relative weight of the loss in the training phase and the loss in the test phase in the optimization objective; Represents a pair parameter Calculate the gradient to reflect the directional update information of the current model on the training loss function; Indicates a task The training set loss function on ; Step 4.2: Meta-learning loss function Introducing physical guidance items to align the output of the UAV intelligent scheduling model with physical laws; Physics-guided loss function The form is: ; in, is the data-driven loss term used by the model during training; It is the physical guidance loss term introduced by the model to align the model output with the physical laws; It is the weight of data-driven loss, and its value is adjusted according to the needs of actual application; is the weight of each physical guidance loss item. Different priorities are given to different physical guidance items by adjusting the weights. Represents the index number of the physical guidance loss term, traversing all physical constraint types to be introduced; Represents the total number of physical guided loss terms, which is used to control the number or type of physical laws introduced by the model; Step 4.3, loss matrix expansion: and Contains multiple loss values ​​used to describe the error of the UAV intelligent scheduling model on different data samples or tasks. They are all in vector or matrix form. Therefore, the physical guidance loss function is expanded as follows: ; in, represents the first The loss value, represents the first The first loss loss value; , is the total number of loss values; Step 4.

4. Finally, the meta-learning objective is written as: ; in, Represents the meta-task set Sampling task After that, the task Perform expectation operation on the overall loss function; Indicates that in the task Based on the training data and the current model meta-parameters The calculated data-driven loss term; Indicates that in the task The physical guided loss term on .

5. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 4, characterized in that: The target retraining module in step 5 self-trains the UAV intelligent scheduling model through unlabeled data, combines forward propagation and backpropagation to update the parameter set, and obtains the final UAV intelligent scheduling model, as follows: Step 5.1: Collect data from the target task scenario and perform forward propagation; Step 5.2: Customize the error signal, normalize the error signal, and calculate the back propagation weight parameters; Step 5.3: Calculate the forward weight update parameters and bias update parameters. Step 5.4: Combine the forward weight update parameters, bias update parameters, and backpropagation weight update parameters of all layers into a parameter set for output.

6. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 5, characterized in that: Step 5.1 collects data from the target task scene and performs forward propagation as follows: Step 5.1.1: The model receives limited unlabeled data As input, data is collected from the target task scenario; Step 5.1.2, Data Passed to each hidden layer of the model in turn, Input data in the layer By using the forward weight and bias Perform linear transformation to generate an inactive intermediate representation : ; Step 5.1.3, through batch normalization To normalize: ; in, Represents the normalized , represents the batch normalization function; Step 5.1.4, use The activation function will Converted to the activated output : ; The activation function sets negative values ​​to zero and only keeps positive values, thus introducing nonlinearity.

7. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 6, characterized in that: The custom error signal described in step 5.2 is normalized, and the back propagation weight parameters are calculated as follows: Step 5.2.1, Customize the error signal: In each layer, Layer The neuron in The local error signal generated by the neuron position The calculation method is: ; in, represents the custom error signal for the top layer, yes The derivative of the activation function, where It is in The intermediate representations that are not activated in the layer; The derivative of shows how the error changes with the inactive output Change, if the activation function is ,but The derivative of is a step function; represents element-wise multiplication, The result is equivalent to adjusting the custom error signal to the error amplitude suitable for the current layer; and is the meta-parameter used to propagate the error signal; It comes from the upper layer, The error signal of the neurons in the layer is Sum and calculate the error contribution of all upper layer units to the current neuron, and then calculate the update of each layer weight through the convolutional neural network module; Step 5.2.2, Error signal normalization: Calculate the normalized error signal , the calculation formula is: ; in, Indicates in Layer The neuron pairs The original error signal generated by samples; for Normalized error signal; is the normalized dimension correction coefficient; Indicates in Layer The neuron pairs The original error signal generated by samples; Step 5.2.3, back propagation weight calculation: The error signal is propagated back through the weights Passed to the upper layer neurons, the calculation formula is: ; in, Indicates the Layer The neuron in The total error signal received at each neuron location; Indicates the Layer The neuron in The local error signal generated by each neuron position; Indicates the Layer from the The neuron transmits to the The forward connection weights between neurons.

8. The method for intelligent scheduling of drones based on a knowledge-driven meta-learning device according to claim 7, characterized in that: The forward weight update parameters and bias update parameters are calculated as described in step 5.3 as follows: Step 5.3.1, forward weight update parameters: Using Functions Calculate weight updates: ; Among them, the function Used in the back propagation process of the neural network to calculate the update amount of the weight; function Receiving input data , intermediate activation value , pre-activation value , model parameters and meta parameters In each layer, the function uses a setting function based on the local part of the neuron, combining the forward propagation signal and the error signal from the back propagation to determine how to calculate the update; Indicates the The weight update matrix of the layer neural network; Step 5.3.2, weight normalization: Calculate the update , and normalize between layers, the formula is: ; in, Indicates the Layer The normalized forward weight update corresponding to the neurons is used to replace the unnormalized As the final updated value; Indicates the Layer The unnormalized forward weight update corresponding to each neuron; Indicates the The number of neurons in the layer is the current The input layer of the layer; Indicates the The number of neurons in the layer; Indicates the Layer The neuron corresponding to Input channel and The updated value of the single connection weight between the output channels is composed of part of; Indicates the The index of the layer neuron, used to traverse all input channels; Indicates the The index of the layer neurons, used to traverse all output channels; The neurons connected Layer and All channels between layers; Step 5.3.3, weight update formula: To make the final weight update smooth, each weight update plane is linearly combined to generate the final update: ; in, is the learning rate for weight updates; Indicates the Layer The original forward weight matrix at the iteration, which represents the current weight state before the update; Indicates the Layer The updated forward weight matrix at the iteration is the result of fusing the old parameters and the update amount; Step 5.3.4, bias update parameters: Compute the bias update. The initial bias update is obtained by batch aggregation over the hidden dimension: ; in, For the The initial bias update value of the layer is obtained by aggregating the feature responses of the samples in the batch; Indicates the first samples; is the low-rank feature dimension index, the total number is ; is the weight vector in the low-rank readout module, where Indicates the The weights on the dimensions, For the Layer The sample in Hidden responses on channels; represents the batch size, represents hidden dimensions; For constrained bias updates: ; in, Indicates the The bias update amount of the layer after constraint processing, Indicates the initial bias update vector The bias update value corresponding to each neuron is used to calculate the average bias trend.

9. A UAV intelligent scheduling system based on a knowledge-driven meta-learning device, characterized in that: The system is used to implement the drone intelligent scheduling method based on the knowledge-driven meta-learning device according to any one of claims 1 to 8. The system includes the first to fifth units, and the functions of each module are as follows: The first unit is used to deploy and initialize a knowledge-driven meta-learning device, which includes a meta-training module and a target retraining module; In the second unit, the meta-learning objectives of the UAV intelligent scheduling model are determined through the meta-training module, and the meta-parameters are optimized. The weights are estimated through linear regression, and the similarity between different feature representations is measured through similarity metrics. Unit 3: By building a meta-task environment covering multiple tasks, the UAV intelligent scheduling model learns cross-task update rules and uses the feature domain basis matrix to form an orthogonal multi-domain feature expression; Unit 4: By introducing a physical guidance term into the meta-learning loss function, the output of the UAV intelligent scheduling model is aligned with physical laws and the meta-learning objective function is optimized. In the fifth unit, the UAV intelligent scheduling model is self-trained with unlabeled data through the target retraining module, and the parameter set is updated by combining forward propagation and backpropagation to obtain the final UAV intelligent scheduling model, which is used to command and optimize the collaborative operation of multiple UAVs.

Citation Information

Patent Citations

  • Radar target constraint element learner intelligent identification method

    CN114839616A

  • Metalearning-based hybrid unmanned aerial vehicle multi-task demand power prediction method

    CN117933473A