Multi-UAV Vision Recognition Method, System, Device and Storage Medium
Through the multi-drone visual recognition method, the local data set is preprocessed and prompted by each drone in the drone cluster. Combined with the multi-grained prompt parameter aggregation of the server, the problem of low visual recognition accuracy of the drone is solved, and efficient model training and recognition effects are achieved.
Patent Information
- Application Number
- CN202510426948.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-07
AI Technical Summary
A single drone in a drone cluster is limited by local training data and computing resources, and cannot effectively train the recognition model, resulting in low visual recognition accuracy.
Through the multi-drone visual recognition method, each drone in the drone cluster is used to preprocess the local data set, the target data set is obtained, and the original recognition model is prompted based on the target data set, the candidate recognition model and local prompt parameters are obtained, and the candidate recognition model and local prompt parameters are sent to the server for multi-grained prompt parameters aggregation, and finally the target recognition model is obtained for visual recognition.
It effectively improves the visual recognition accuracy of drone, solves the problem of limited computing resources and training data of a single drone, and improves model training efficiency.
Smart Images

Figure CN119942386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a multi-unmanned aerial vehicle visual recognition method, system, device, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence and wireless communication technologies, unmanned aerial vehicles are increasingly widely used in fields such as emergency rescue and agricultural monitoring. These unmanned aerial vehicles need to perform real-time target recognition and classification during mission execution, generating a large amount of visual data and computational requirements. However, due to the limited computational resources of unmanned aerial vehicles, the complex and variable mission environment, and the high requirements for real-time performance, traditional centralized visual recognition solutions face severe challenges.
[0003] Although edge computing provides new ideas for real-time visual task processing of unmanned aerial vehicles, the following problems still exist in practical applications: The training data locally available on unmanned aerial vehicles is usually relatively limited and difficult to support end-to-end training of complex visual models; the computational resources of unmanned aerial vehicles are severely limited and difficult to support online training and inference of large-scale deep models. Therefore, in the current unmanned aerial vehicle cluster, a single unmanned aerial vehicle is restricted by local training data and computational resources and cannot effectively train an identification model, resulting in low visual recognition accuracy. Summary of the Invention
[0004] The main objective of the present invention is to provide a multi-unmanned aerial vehicle visual recognition method, system, device, and storage medium, aiming to solve the technical problem in the prior art that a single unmanned aerial vehicle in an unmanned aerial vehicle cluster is restricted by local training data and computational resources and cannot effectively train an identification model, resulting in low visual recognition accuracy.
[0005] To achieve the above objective, the present invention provides a multi-unmanned aerial vehicle visual recognition method. The multi-unmanned aerial vehicle visual recognition method is applied to each unmanned aerial vehicle in an unmanned aerial vehicle cluster. The unmanned aerial vehicle cluster includes multiple unmanned aerial vehicles, and each unmanned aerial vehicle is communicatively connected to a server. The multi-unmanned aerial vehicle visual recognition method includes:
[0006] Preprocess the local data set collected during the mission to obtain a target data set;
[0007] Perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model;
[0008] Send the local prompt parameters to the server and obtain multi-granularity prompt parameters feedback by the server based on the local prompt parameters. The multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each unmanned aerial vehicle;
[0009] Aggregate the multi-granularity prompt parameters to obtain fused prompt parameters;
[0010] Based on the fused prompt parameters, perform prompt training on the candidate recognition model to obtain a target recognition model, and input the target dataset into the target recognition model for visual recognition.
[0011] Optionally, the original recognition model is a multi-modal pre-trained model; the performing prompt training on the original recognition model based on the target dataset to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model includes:
[0012] Initialize the original recognition model to obtain initial prompt parameters;
[0013] Obtain a prompt template structure that matches the task type corresponding to the drone;
[0014] Based on the local dataset, the prompt template structure, and the initial prompt parameters, perform prompt training on the original recognition model to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model.
[0015] Optionally, the multi-granularity prompt parameters include fine-grained prompt parameters, medium-grained prompt parameters, and coarse-grained prompt parameters; the aggregating the multi-granularity prompt parameters to obtain fused prompt parameters includes:
[0016] Calculate the attention scores of the multi-granularity prompt parameters to obtain a fine-grained attention score, a medium-grained attention score, and a coarse-grained attention score:
[0017] ;
[0018] ;
[0019] ;
[0020] Wherein, represents the fine-grained prompt parameter, represents the medium-grained prompt parameter, represents the coarse-grained prompt parameter, 、 and respectively represent weight matrices, 、 and respectively represent bias vectors, is the attention vector, represents the fine-grained attention score, represents the medium-grained attention score, Represents the coarse-grained attention score;
[0021] Normalize the fine-grained attention score, the medium-grained attention score, and the coarse-grained attention score to obtain a normalized score:
[0022] ;
[0023] Wherein, and respectively represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-grained attention score, represents the normalized coarse-grained attention score;
[0024] Aggregate the fine-grained prompt parameter, the medium-grained prompt parameter, and the coarse-grained prompt parameter based on the normalized score to obtain a fused prompt parameter:
[0025] ;
[0026] Wherein, is the fused prompt parameter.
[0027] In addition, to achieve the above object, the present invention also proposes a multi-UAV vision recognition method. The multi-UAV vision recognition method is applied to a server, and the server is communicatively connected to each UAV in the UAV cluster. The UAV cluster includes multiple UAVs. The multi-UAV vision recognition method includes:
[0028] Obtain the local prompt parameters sent by each UAV. The local prompt parameters are obtained by the UAV through prompt training of the original recognition model based on the local dataset;
[0029] Construct an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters. The affinity matrix includes the target affinity values between different task types;
[0030] Construct a task relationship graph according to the affinity matrix;
[0031] Determine the multi-grained prompt parameters of each task type based on the task relationship graph, and send the multi-grained prompt parameters to each UAV, so that each UAV performs prompt training on the original recognition model based on the multi-grained prompt parameters.
[0032] Optionally, the constructing an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters includes:
[0033] Aggregate the local hint parameters of drones with the same task type based on the task types corresponding to each drone:
[0034] ;
[0035] Among them, represents the set of drones performing tasks, represents the data volume of the local data set of drone ; represents the local hint parameter of drone ; represents the aggregated local hint parameter;
[0036] Calculate the initial affinity values between different task types according to the aggregated local hint parameters:
[0037] ;
[0038] Among them, represents the initial affinity value, represents the aggregated local hint parameter of the th task type, represents the aggregated local hint parameter of the th task type, represents transpose;
[0039] Update the initial affinity value based on historical affinity information to obtain candidate affinity values:
[0040] ;
[0041] Among them, represents the candidate affinity value, represents historical affinity information, is the sliding average coefficient, used to adjust the retention degree of historical affinity information;
[0042] Perform temperature scaling on the candidate affinity value to obtain the target affinity value:
[0043] ;
[0044] Among them, represents the target affinity value, represents the temperature parameter, used to adjust the smoothness of the task similarity distribution;
[0045] Construct an affinity matrix based on the target affinity values between different task types.
[0046] Optionally, constructing a task relationship graph according to the affinity matrix includes:
[0047] Determining an affinity screening threshold based on the mean and standard deviation values of the affinity matrix:
[0048] ;
[0049] Wherein, is the affinity screening threshold, is the mean value, is the standard deviation value, is the threshold adjustment coefficient;
[0050] Screening the affinity matrix according to the affinity screening threshold to determine the target relationship values in the affinity matrix that are higher than the affinity screening threshold;
[0051] Constructing an edge set based on the target relationship values and constructing a node set according to the task types;
[0052] Constructing a task relationship graph based on the edge set and the node set.
[0053] Optionally, after determining the multi-granularity prompt parameters for each task type based on the task relationship graph and sending the multi-granularity prompt parameters to each drone, it further includes:
[0054] Returning to execute the step of obtaining the local prompt parameters sent by each drone, and determining the difference value between the local prompt parameters and the multi-granularity prompt parameters;
[0055] Judging whether the difference value meets the convergence condition, and the convergence condition includes that the difference value is less than a preset threshold;
[0056] If the convergence condition is met, stop executing the step of constructing the affinity matrix based on the task types corresponding to each drone and the local prompt parameters;
[0057] If the convergence condition is not met, continue to execute the step of constructing the affinity matrix based on the task types corresponding to each drone and the local prompt parameters.
[0058] In addition, to achieve the above object, the present invention further provides a multi-drone vision recognition system, and the multi-drone vision recognition system includes: a plurality of drones and a server, and each drone is communicatively connected to the server;
[0059] The drone is used for preprocessing the local data set collected during the task process to obtain a target data set;
[0060] The drone is also used to perform prompt training on the original recognition model based on the target dataset to obtain a candidate recognition model and local prompt parameters of the candidate recognition model;
[0061] The drone is also used to send the local prompt parameters to the server;
[0062] The server is used to obtain the local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on a local dataset;
[0063] The server is also used to construct an affinity matrix based on the task types corresponding to each drone and the local prompt parameters, where the affinity matrix includes target affinity values between different task types;
[0064] The server is also used to construct a task relationship graph according to the affinity matrix;
[0065] The server is also used to determine multi-granularity prompt parameters for each task type based on the task relationship graph and send the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters;
[0066] The drone is also used to obtain the multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each drone;
[0067] The drone is also used to aggregate the multi-granularity prompt parameters to obtain fused prompt parameters;
[0068] The drone is also used to perform prompt training on the candidate recognition model based on the fused prompt parameters to obtain a target recognition model, and input the target dataset into the target recognition model for visual recognition.
[0069] In addition, to achieve the above object, the present application also proposes a multi-drone visual recognition device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the multi-drone visual recognition method as described above.
[0070] In addition, to achieve the above object, the present application also proposes a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the multi-drone visual recognition method as described above are implemented.
[0071] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the multi-UAV vision recognition method described above.
[0072] In the present invention, the UAV preprocesses the local data set collected during the task process to obtain a target data set, and performs prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model, and sends the local prompt parameters to the server; the server obtains the local prompt parameters sent by each UAV. The local prompt parameters are obtained by the UAV performing prompt training on the original recognition model based on the local data set. An affinity matrix is constructed based on the task type and local prompt parameters corresponding to each UAV. The affinity matrix includes target affinity values between different task types. A task relationship graph is constructed based on the affinity matrix, and multi-granularity prompt parameters of each task type are determined based on the task relationship graph, and the multi-granularity prompt parameters are sent to each UAV so that each UAV performs prompt training on the original recognition model based on the multi-granularity prompt parameters; the UAV obtains the multi-granularity prompt parameters fed back by the server based on the local prompt parameters. The multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each UAV. The multi-granularity prompt parameters are aggregated to obtain fusion prompt parameters, and the candidate recognition model is trained by prompt based on the fusion prompt parameters to obtain a target recognition model, and the target data set is input into the target recognition model for vision recognition; since the present invention sends the multi-granularity prompt parameters obtained by processing the local prompt parameters sent by each UAV in the UAV cluster to each UAV for model training through the server, the problem of limited computing resources and training data of a single UAV is effectively solved, the model training efficiency of the UAV locally is effectively improved, and the UAV vision recognition accuracy is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0074] Figure 1 It is a schematic structural diagram of a multi-UAV vision recognition device in the hardware operating environment related to the solution of the embodiment of the present invention;
[0075] Figure 2 It is a schematic flowchart of the first embodiment of the multi-UAV vision recognition method of the present invention;
[0076] Figure 3Schematic flowchart of the second embodiment of the multi-UAV vision recognition method of the present invention;
[0077] Figure 4 Block diagram of the structure of the first embodiment of the multi-UAV vision recognition system of the present invention.
[0078] The realization, functional features and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0079] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0080] Referring to Figure 1 , Figure 1 Schematic diagram of the structure of the multi-UAV vision recognition device in the hardware operating environment involved in the embodiment solution of the present invention.
[0081] As Figure 1 shown, the multi-UAV vision recognition device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless-fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage system independent of the foregoing processor 1001.
[0082] Those skilled in the art can understand that Figure 1 the structure shown in
[0083] does not constitute a limitation on the multi-UAV vision recognition device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Figure 1 As
[0084] shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a multi-UAV vision recognition program. Figure 1In the multi-UAV vision recognition device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the multi-UAV vision recognition device of the present invention can be arranged in the multi-UAV vision recognition device, and the multi-UAV vision recognition device calls the multi-UAV vision recognition program stored in the memory 1005 through the processor 1001 and executes the multi-UAV vision recognition method provided by the embodiments of the present invention.
[0085] An embodiment of the present invention provides a multi-UAV vision recognition method, referring to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the multi-UAV vision recognition method of the present invention.
[0086] In this embodiment, the multi-UAV vision recognition method is applied to each UAV in a UAV cluster. The UAV cluster includes multiple UAVs, and each UAV is communicatively connected to a server. The multi-UAV vision recognition method includes the following steps:
[0087] Step S10: Preprocess the local data set collected during the task to obtain a target data set.
[0088] It should be noted that the execution subject of this embodiment can be a UAV device with functions of data processing, network communication, and program running, such as the central controller of a UAV. Hereinafter, the UAV is taken as an example to illustrate this embodiment and the following embodiments.
[0089] It should be noted that the local data set can be image data collected by the UAV during the task and the corresponding label information of the image data. The preprocessing can be to process the local data set by means of normalization, standardization, data cleaning, data augmentation, etc.
[0090] In some embodiments, in the system initialization stage, first define the set of UAVs in the system as , where N represents the total number of UAVs. Each UAV executes a specific vision recognition task , and the task set is denoted as , where M represents the total number of different task types. For the local data set of each UAV, it is expressed as:
[0091] ;
[0092] where represents the data volume, represents the image data, represents the corresponding label information.
[0093] In some embodiments, for each local UAV dataset , this embodiment designs a complete data preprocessing process. First, image normalization is performed. The image data is normalized through the formula where μ and σ represent the mean and standard deviation of the image respectively. Subsequently, to enhance the generalization ability of the model, data augmentation strategies are implemented on the normalized images, including operations such as random cropping and horizontal flipping, to generate an augmented dataset . Finally, the processed dataset is divided into mini-batches of size B to obtain the target dataset, thereby preparing for subsequent batch training.
[0094] Step S20: Perform prompt training on the original recognition model based on the target dataset to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model.
[0095] It should be noted that the original recognition model can be a pre-trained multi-modal large model deployed locally on the UAV. The original recognition model may include an image encoder and a text encoder. During the local training phase, the parameters of the pre-trained image encoder and text encoder are frozen.
[0096] In a specific implementation, the UAV can generate initial prompt parameters based on the target dataset and the task type executed by the UAV, and input the initial prompt parameters and the target dataset into the original recognition model to perform prompt parameter training on the original recognition model, obtain a candidate recognition model, and obtain the local prompt parameters of the candidate recognition model.
[0097] Step S30: Send the local prompt parameters to the server and obtain the multi-granularity prompt parameters feedback by the server based on the local prompt parameters.
[0098] It should be noted that the multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each UAV.
[0099] In some embodiments, the server obtains the local prompt parameters sent by each UAV. The local prompt parameters are obtained by the UAV performing prompt training on the original recognition model based on the local dataset. An affinity matrix is constructed based on the task type corresponding to each UAV and the local prompt parameters. The affinity matrix includes the target affinity values between different task types. A task relationship graph is constructed based on the affinity matrix. The multi-granularity prompt parameters of each task type are determined based on the task relationship graph, and the multi-granularity prompt parameters are sent to each UAV so that each UAV performs prompt training on the original recognition model based on the multi-granularity prompt parameters.
[0100] Furthermore, in order to improve the model performance and accurately obtain the local prompt parameters, in some embodiments, the original recognition model is a multi-modal pre-trained model, and the above step S30 may include:
[0101] Step S31: Initialize the original recognition model to obtain initial prompt parameters;
[0102] Step S32: Obtain a prompt template structure that matches the task type corresponding to the drone;
[0103] Step S33: Perform prompt training on the original recognition model based on the local dataset, the prompt template structure, and the initial prompt parameters to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model.
[0104] It should be noted that the original recognition model is a multi-modal pre-trained model. For example, the original recognition model can be a CLIP (Contrastive Language–Image Pretraining) model.
[0105] In a specific implementation, the drone can deploy a pre-trained CLIP model locally as a basic feature extractor. This model mainly includes two core components: an image encoder and a text encoder . Among them, the image encoder adopts the Vision Transformer architecture, receives an image input with a dimension of 3×H×W, and outputs a d-dimensional feature vector; the text encoder adopts the Transformer architecture, sets the maximum token length to L, and also outputs a d-dimensional feature vector. During the local training phase, the parameters of the pre-trained image encoder and the text encoder are frozen.
[0106] It should be noted that for each task a learnable prompt template structure is designed, and its form is:
[0107] ;
[0108] The prompt parameters among them are learnable continuous vectors and are initialized using a normal distribution . This template design allows the model to adapt to different visual recognition requirements by learning task-specific prompt parameters while maintaining the possibility of cross-task knowledge sharing. The dimension of each prompt parameter vector matches the text feature dimension of the CLIP model, ensuring the consistency of the feature space.
[0109] It is understandable that during the training of the prompt parameters, the local model receives two types of inputs: the image data collected by the drone and the corresponding prompt parameters . The image encoder and the text encoder extract features from the inputs respectively to obtain feature vectors and . The feature fusion module uses a multi-layer perceptron (MLP) for feature fusion: , where represents the feature concatenation operation. The classification prediction is implemented through a fully connected layer: , where W is the weight matrix and b is the bias term. The final class probability is calculated through the softmax function: .
[0110] The loss function design includes three components: the classification loss , the contrastive loss and the regularization loss . The expression of the total loss function is:
[0111] ;
[0112] Among them, the classification loss adopts the cross-entropy form:
[0113] ;
[0114] The contrastive loss is used to enhance the discriminability of the feature representation:
[0115] ;
[0116] The regularization loss controls the norm of the prompt parameters:
[0117] ;
[0118] Among them is a learnable parameter, represents the total number of classes. At the same time, we introduce the regularization loss to prevent overfitting. By adjusting the weight coefficients and , the contributions of the losses can be balanced.
[0119] Step S40: Aggregate the multi-granularity prompt parameters to obtain the fused prompt parameters.
[0120] It should be noted that the multi-granularity parameters are prompt parameters of multiple granularities, and the multi-granularity prompt parameters may include fine-grained prompt parameters, medium-grained prompt parameters, and coarse-grained prompt parameters.
[0121] In some embodiments, the server sends hint parameters of three granularities to the corresponding drones, and the adaptive fusion module in the drones uses an attention mechanism to fuse the hint parameters of different granularities to obtain fused hint parameters.
[0122] Further, in order to effectively fuse multi-granularity parameters, in some embodiments, the multi-granularity hint parameters include fine-granularity hint parameters, medium-granularity hint parameters, and coarse-granularity hint parameters. The above step S40 may include:
[0123] Step S41: Calculate the attention scores of the multi-granularity hint parameters to obtain a fine-granularity attention score, a medium-granularity attention score, and a coarse-granularity attention score.
[0124] It should be noted that the server sends hint parameters of three granularities to the corresponding drones, and the local adaptive fusion module of the drones uses an attention mechanism to fuse the hint information of different granularities. First, calculate the attention scores of the hint parameters of each granularity. The formula for calculating the attention score is as follows:
[0125] ;
[0126] ;
[0127] ;
[0128] Where represents the fine-granularity hint parameter, represents the medium-granularity hint parameter, represents the coarse-granularity hint parameter, 、 and respectively represent weight matrices, 、 and respectively represent bias vectors, is the attention vector, represents the fine-granularity attention score, represents the medium-granularity attention score, represents the coarse-granularity attention score.
[0129] Step S42: Normalize the fine-granularity attention score, the medium-granularity attention score, and the coarse-granularity attention score to obtain normalized scores.
[0130] It should be noted that the attention scores can be normalized by the function, referring to the following formula:
[0131] ;
[0132] Among them, and respectively represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-grained attention score, represents the normalized coarse-grained attention score;
[0133] Step S403: Aggregate the fine-grained prompt parameter, the medium-grained prompt parameter, and the coarse-grained prompt parameter based on the normalized scores to obtain a fused prompt parameter.
[0134] It should be noted that, based on the normalized scores, the prompt parameters of each granularity are aggregated, and is used as the input prompt parameter of the local text encoder The aggregation refers to the following formula:
[0135] ;
[0136] Among them, is the fused prompt parameter.
[0137] Step S50: Perform prompt training on the candidate recognition model based on the fused prompt parameter to obtain a target recognition model, and input the target dataset into the target recognition model for visual recognition.
[0138] It should be noted that the drone performs prompt training on the candidate recognition model based on the fused prompt parameter, updates the local prompt parameter based on the prompt training result, sends the updated local prompt parameter to the server, and the server determines whether to converge based on the received local prompt parameter (for example, determines whether the difference value between the local prompt parameters sent by the drone in two consecutive rounds is less than a preset threshold. If it is less than the preset threshold, it is determined to converge). If it is determined to converge, the processing of the local prompt parameter sent by the drone is stopped, and a model convergence message is fed back to the drone. After receiving the model convergence message, the drone performs prompt training on the candidate recognition model based on the updated local prompt parameter to obtain a target recognition model, and inputs the target dataset into the target recognition model for visual recognition.
[0139] In some embodiments, the drone swarm and the server implement the update of the local recognition model through federated learning, including the following steps:
[0140] Step 1: Parameter initialization;
[0141] In the initial stage of federated learning, the system initializes the prompt parameter , CLIP model parameters, and global prompt vectors. In the initial round, initialize the local prompt parameters of each client to the same value: .
[0142] Step 2: Model parameter broadcasting;
[0143] At the beginning of each iteration, the server broadcasts multi-granularity prompt parameters to all drone clients, including fine-grained prompt parameters , medium-grained prompt parameters and coarse-grained prompt parameters .
[0144] Step 3: Local training and parameter uploading;
[0145] The client receives the multi-granularity prompt parameters given by the server, and the local client freezes the parameters of the image encoder and the text encoder parameters, and updates other learnable parameters. Each drone client optimizes the loss function based on the local dataset , where:
[0146] ;
[0147] Subsequently, upload the updated prompt parameters and the current task of the client to the server.
[0148] Step 4: Server constructs the task affinity relationship graph;
[0149] 1. Calculate the task affinity matrix A, where each element is calculated by cosine similarity:
[0150] ;
[0151] 2. Update the affinity value using the moving average mechanism:
[0152] ;
[0153] 3. Construct a weighted undirected graph , and the edge weights are calculated by softmax after temperature scaling:
[0154] ;
[0155] 4. Apply the PageRank algorithm to optimize the edge weights:
[0156] ;
[0157] Step 5: Server calculates multi-granularity prompt parameters;
[0158] Based on the constructed relationship graph, for each task we define three neighbor sets with different hop counts , and , and calculate the multi-granularity prompt parameters for each task.
[0159] Repeat steps 3 - 5 until the model converges. The convergence condition is that the change in the global prompt parameters for two consecutive rounds is less than a preset threshold .
[0160] In this embodiment, by preprocessing the collected local dataset, a target dataset is obtained. Based on the target dataset, the original recognition model is trained with prompts to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model. The local prompt parameters are sent to the server, and the multi-granularity prompt parameters feedback by the server based on the local prompt parameters are obtained. The multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each drone in the drone cluster; the candidate recognition model is trained with prompts based on the fused prompt parameters obtained after aggregating the multi-granularity prompt parameters to obtain a target recognition model, and the target dataset is input into the target recognition model for visual recognition, effectively solving the problem of limited computing resources and training data of a single drone, and greatly improving the visual recognition accuracy of the drone.
[0161] Refer to Figure 3 , Figure 3 which is the schematic flowchart of the second embodiment of the multi-drone visual recognition method of the present invention.
[0162] In this embodiment, the multi-drone visual recognition method is applied to a server, and the server is communicatively connected to each drone in the drone cluster. The drone cluster includes multiple drones. The multi-drone visual recognition method includes:
[0163] Step S100: Obtain the local prompt parameters sent by each drone.
[0164] It should be understood that the execution subject of this embodiment can be a server with data processing, network communication, and program running functions, such as a computer, or a terminal electronic device capable of implementing the above functions. Hereinafter, the server is taken as an example to illustrate this embodiment and the following embodiments.
[0165] It should be noted that the local prompt parameters are obtained by the drone training the original recognition model with prompts based on the local dataset.
[0166] In some embodiments, the drone sends the locally trained local prompt parameters and the task type corresponding to the drone to the server, and the server aggregates the local prompt parameters of drones of the same type according to the task type.
[0167] Step S200: Construct an affinity matrix based on the task type corresponding to each drone and the local prompt parameters.
[0168] It should be noted that the affinity matrix includes the target affinity values between different task types.
[0169] In a specific implementation, the server can determine the affinity values between different task types by calculating the cosine similarity between different task types.
[0170] Furthermore, in order to accurately calculate the affinity values between different task types, in some embodiments, the above step S200 may include:
[0171] Step S2001: Aggregate the local prompt parameters of drones with the same task type based on the task type corresponding to each drone;
[0172] Step S2002: Calculate the initial affinity values between different task types according to the aggregated local prompt parameters;
[0173] Step S2003: Update the initial affinity values based on the historical affinity information to obtain candidate affinity values;
[0174] Step S2004: Perform temperature scaling on the candidate affinity values to obtain the target affinity values;
[0175] Step S2005: Construct an affinity matrix based on the target affinity values between different task types.
[0176] It should be noted that the server can classify the local prompt parameters of the drones based on the task type and aggregate the local prompt parameters of the same type:
[0177] ;
[0178] Among them, represents the set of drones performing tasks, represents the drone the data volume of the local dataset of, represents the drone the local prompt parameters of, represents the aggregated local prompt parameters;
[0179] It is understandable that the server calculates the affinity matrix between different types of tasks based on the aggregated local hint parameters. Each element in the matrix is obtained through cosine similarity calculation:
[0180] ;
[0181] Among them, represents the initial affinity value, represents the aggregated local hint parameter of the th task type, represents the aggregated local hint parameter of the th task type, represents transpose;
[0182] It should be noted that in order to improve the stability of affinity calculation, the server can adopt a moving average mechanism to update the affinity value:
[0183] ;
[0184] Among them, represents the candidate affinity value, represents the historical affinity information, is the moving average coefficient, used to adjust the retention degree of historical affinity information;
[0185] It should be noted that the temperature scaling of the candidate affinity value refers to the following formula:
[0186] ;
[0187] Among them, represents the target affinity value, represents the temperature parameter, used to adjust the smoothness of the task similarity distribution.
[0188] Step S300: Construct a task relationship graph according to the affinity matrix.
[0189] It should be noted that the task relationship graph can be a weighted undirected graph. The task relationship graph includes multiple nodes, each node represents a different task type, and each target affinity value in the affinity matrix serves as the edge connecting the nodes.
[0190] Furthermore, in order to accurately construct the task relationship graph, in some embodiments, the above step S300 may include:
[0191] Step S3001: Determine an affinity screening threshold based on the mean and standard deviation values of the affinity matrix;
[0192] Step S3002: Screen the affinity matrix according to the affinity screening threshold to determine the target relationship values in the affinity matrix that are higher than the affinity screening threshold;
[0193] Step S3003: Construct an edge set based on the target relationship values and construct a node set according to the task type;
[0194] Step S3004: Construct a task relationship graph based on the edge set and the node set.
[0195] It should be noted that in this embodiment, an adaptive threshold mechanism is designed to dynamically adjust the determination criterion of affinity. The affinity screening threshold is determined with reference to the following formula:
[0196] ;
[0197] where is the affinity screening threshold, is the mean value, is the standard deviation value, is the threshold adjustment coefficient, which is used to control the strictness of the affinity screening threshold.
[0198] It can be understood that the server can screen the affinity values based on the affinity screening threshold, only retain those higher than the threshold, and use the affinity values in the affinity matrix that are higher than the affinity screening threshold as the target relationship values, with reference to the following formula:
[0199] ;
[0200] It should be understood that based on the screened affinity matrix, the server constructs a weighted undirected graph G=(V, E), where the node set V corresponds to the task set; the edge set E only contains the connections with non-zero affinity values; the weights of the edges are normalized using softmax, that is, the affinity matrix is normalized:
[0201] ;
[0202] where N ( i ) represents the set of all nodes connected to node i, represents the normalized affinity matrix.
[0203] Step S400: Determine the multi-granularity hint parameters of each task type based on the task relationship graph, and send the multi-granularity hint parameters to each drone, so that each drone performs hint training on the original recognition model based on the multi-granularity hint parameters.
[0204] It should be noted that the server determines the node path lengths and node relationships between different task types based on the task relationship graph, constructs task neighbor sets for different task types based on the node path lengths and node relationships. The task neighbor sets include one-hop neighbor sets, two-hop neighbor sets, and three-hop neighbor sets, and performs hint aggregation based on the task neighbor sets to obtain multi-granularity hint parameters.
[0205] In a specific implementation, the server defines neighbor sets with three different hop counts for different task types - the one-hop neighbor set contains nodes directly connected to task i; the two-hop neighbor set is reachable through intermediate nodes; the three-hop neighbor set . Then two intermediate nodes need to be passed through. To smooth the influence of different hop counts, we introduce a distance attenuation factor , where d is the hop count, ∈(0,1) is the attenuation rate. The influence strength of node j on task i can be expressed as:
[0206] ;
[0207] where is the shortest path length between nodes, represents the weight value of the edge in the task relationship graph, that is, the normalized affinity matrix.
[0208] After obtaining task sets with different granularities, we design a hierarchical hint aggregation mechanism. For task , its fine-grained hint parameters are obtained by weighted averaging of one-hop neighbors:
[0209] ;
[0210] The medium-grained hint parameters consider tasks reachable in two hops:
[0211] ;
[0212] The coarse-grained hint parameters aggregate all task information within three hops:
[0213] ;
[0214] Furthermore, in order to dynamically update the hint parameters of the UAV and converge under reasonable conditions, in some embodiments, after the above step S400, the following is further included:
[0215] Step S500: Return to execute the step of obtaining the local hint parameters sent by each UAV, and determine the difference value between the local hint parameters and the multi-granularity hint parameters;
[0216] Step S600: Determine whether the difference value meets the convergence condition, where the convergence condition includes that the difference value is less than a preset threshold;
[0217] Step S700: If the convergence condition is met, stop executing the step of constructing the affinity matrix based on the task types corresponding to each drone and the local hint parameters;
[0218] Step S800: If the convergence condition is not met, continue to execute the step of constructing the affinity matrix based on the task types corresponding to each drone and the local hint parameters.
[0219] It should be noted that the drone performs hint training on the candidate recognition model based on the fusion hint parameters, updates the local hint parameters based on the hint training results, and sends the updated local hint parameters to the server. The server determines whether to converge based on the received local hint parameters (for example, determines whether the difference value between the local hint parameters sent by the drone in two consecutive rounds is less than a preset threshold. If it is less than the preset threshold, it is determined to converge). If it is determined to converge, stop processing the local hint parameters sent by the drone and feedback a model convergence message to the drone. After receiving the model convergence message, the drone performs hint training on the candidate recognition model based on the updated local hint parameters to obtain a target recognition model, and inputs the target data set into the target recognition model for visual recognition.
[0220] In this embodiment, by obtaining the local hint parameters sent by each drone, where the local hint parameters are obtained by the drone performing hint training on the original recognition model based on the local data set, constructing an affinity matrix based on the task types corresponding to each drone and the local hint parameters, where the affinity matrix includes the target affinity values between different task types, constructing a task relationship graph based on the affinity matrix, determining the multi-granularity hint parameters of each task type based on the task relationship graph, and sending the multi-granularity hint parameters to each drone, so that each drone performs hint training on the original recognition model based on the multi-granularity hint parameters, effectively avoiding the problem that the model parameters cannot be aggregated due to task differences between drones. By aggregating the hint parameters of each drone in the drone cluster, the adaptability of the recognition model in the drone to different tasks is effectively improved, and the recognition accuracy of the drone is effectively improved.
[0221] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, on which a multi-drone visual recognition program is stored. When the multi-drone visual recognition program is executed by a processor, the steps of the multi-drone visual recognition method described above are implemented.
[0222] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0223] The above computer-readable storage medium can be included in the multi-UAV vision recognition device; it can also exist independently without being assembled into the multi-UAV vision recognition device.
[0224] In addition, an embodiment of the present invention also provides a computer program product, including a multi-UAV vision recognition program. When the multi-UAV vision recognition program is executed by a processor, it implements the steps of the multi-UAV vision recognition method described above.
[0225] The specific implementation manner of the computer program product of the present invention is basically the same as that of the above embodiments of the multi-UAV vision recognition method, and will not be elaborated here.
[0226] Refer to Figure 4 , Figure 4 which is the structural block diagram of the first embodiment of the multi-UAV vision recognition system of the present invention.
[0227] As Figure 4 shown, the multi-UAV vision recognition system proposed by the embodiment of the present invention includes: a plurality of UAVs 10 and a server 20, and each UAV is communicatively connected to the server;
[0228] The UAV 10 is used to preprocess the local data set collected during the task to obtain a target data set;
[0229] The drone 10 is further configured to perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model;
[0230] The drone 10 is further configured to send the local prompt parameters to the server;
[0231] The server 20 is configured to obtain the local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on a local data set;
[0232] The server 20 is further configured to construct an affinity matrix based on the task type corresponding to each drone and the local prompt parameters, where the affinity matrix includes target affinity values between different task types;
[0233] The server 20 is further configured to construct a task relationship graph according to the affinity matrix;
[0234] The server 20 is further configured to determine multi-granularity prompt parameters for each task type based on the task relationship graph, and send the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters;
[0235] The drone 10 is further configured to obtain the multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are obtained by the server processing the local prompt parameters sent by each drone;
[0236] The drone 10 is further configured to aggregate the multi-granularity prompt parameters to obtain fused prompt parameters;
[0237] The drone 10 is further configured to perform prompt training on the candidate recognition model based on the fused prompt parameters to obtain a target recognition model, and input the target data set into the target recognition model for visual recognition.
[0238] In this embodiment, a drone preprocesses a local data set collected during a task to obtain a target data set, and performs prompt training on an original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model, and sends the local prompt parameters to a server; the server obtains the local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set, constructs an affinity matrix based on the task type and local prompt parameters corresponding to each drone, the affinity matrix includes target affinity values between different task types, constructs a task relationship graph according to the affinity matrix, determines multi-granularity prompt parameters of each task type based on the task relationship graph, and sends the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters; the drone obtains the multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are obtained by the server processing the local prompt parameters sent by each drone, aggregates the multi-granularity prompt parameters to obtain fusion prompt parameters, performs prompt training on the candidate recognition model based on the fusion prompt parameters to obtain a target recognition model, and inputs the target data set into the target recognition model for visual recognition; since in this embodiment, the multi-granularity prompt parameters obtained by the server processing the local prompt parameters sent by each drone in the drone cluster are sent to each drone for model training, the problem that the computing resources and training data of a single drone are limited is effectively solved, the model training efficiency of the drone locally is effectively improved, and the visual recognition accuracy of the drone is greatly improved.
[0239] The multi-drone visual recognition system provided in this application adopts the multi-drone visual recognition method in the above embodiment and can solve the technical problems of multi-drone visual recognition. Compared with the prior art, the beneficial effects of the multi-drone visual recognition system provided in this application are the same as those of the multi-drone visual recognition method provided in the above embodiment, and other technical features in the multi-drone visual recognition system are the same as those disclosed in the method of the above embodiment and will not be elaborated here.
[0240] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not limit this.
[0241] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.
[0242] In addition, for the technical details not described in detail in this embodiment, reference may be made to the multi-UAV vision recognition method provided in any embodiment of the present invention, which will not be elaborated here.
[0243] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or system comprising that element.
[0244] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.
[0245] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0246] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A multi-UAV vision recognition method, characterized in that, The multi-UAV vision recognition method is applied to each UAV in a UAV cluster. The UAV cluster includes multiple UAVs, and each UAV is communicatively connected to a server. The multi-UAV vision recognition method includes: Preprocess the local data set collected during the task to obtain a target data set; Perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; Send the local prompt parameters to the server, and obtain multi-granularity prompt parameters feedback by the server based on the local prompt parameters. The multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each UAV; Aggregate the multi-granularity prompt parameters to obtain fused prompt parameters; Perform prompt training on the candidate recognition model based on the fused prompt parameters to obtain a target recognition model, and input the target data set into the target recognition model for vision recognition; The original recognition model is a multi-modal pre-trained model. The performing prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model includes: Initialize the original recognition model to obtain initial prompt parameters; Obtain a prompt template structure matching the task type corresponding to the UAV; Perform prompt training on the original recognition model based on the local data set, the prompt template structure, and the initial prompt parameters to obtain a candidate recognition model and local prompt parameters of the candidate recognition model.
2. The multi-UAV vision recognition method according to claim 1, wherein The multi-granularity prompt parameters include fine-grained prompt parameters, medium-grained prompt parameters, and coarse-grained prompt parameters; The aggregating the multi-granularity prompt parameters to obtain fused prompt parameters includes: Calculate the attention scores of the multi-granularity prompt parameters to obtain fine-grained attention scores, medium-grained attention scores, and coarse-grained attention scores; ; ; ; Among them, represents the fine-grained hint parameter, represents the medium-grained hint parameter, represents the coarse-grained hint parameter, 、 and represent the weight matrices respectively, 、 and represent the bias vectors respectively, is the attention vector, represents the fine-grained attention score, represents the medium-grained attention score, represents the coarse-grained attention score; Perform normalization processing on the fine-grained attention scores, the medium-grained attention scores, and the coarse-grained attention scores to obtain normalized scores; ; Among them, and respectively represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-grained attention score, represents the normalized coarse-grained attention score; Aggregate the fine-grained prompt parameters, the medium-grained prompt parameters, and the coarse-grained prompt parameters based on the normalized scores to obtain fused prompt parameters; ; Among them, is the fusion prompt parameter.
3. A multi-UAV vision recognition method, characterized in that The multi-UAV vision recognition method is applied to a server. The server is communicatively connected to each UAV in a UAV cluster. The UAV cluster includes multiple UAVs. The multi-UAV vision recognition method includes: Obtain the local prompt parameters sent by each UAV. The local prompt parameters are obtained by the UAV performing prompt training on the original recognition model based on a local data set; Construct an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters. The affinity matrix includes target affinity values between different task types; Construct a task relationship graph according to the affinity matrix; Determine the multi-granularity prompt parameters of each task type based on the task relationship graph, and send the multi-granularity prompt parameters to each UAV so that each UAV performs prompt training on the original recognition model based on the multi-granularity prompt parameters; Constructing an affinity matrix based on the task type corresponding to each unmanned aerial vehicle (UAV) and the local prompt parameters includes: Aggregating the local prompt parameters of UAVs with the same task type based on the task type corresponding to each UAV: ; Among them, represents the set of drones performing tasks, represents the drone the data volume of the local dataset, represents the drone the local hint parameter, represents the aggregated local hint parameter; Calculating the initial affinity values between different task types according to the aggregated local prompt parameters: ; Among them, represents the initial affinity value, represents the aggregated local prompt parameter of the th task type, represents the aggregated local prompt parameter of the th task type, represents transpose; Updating the initial affinity values based on historical affinity information to obtain candidate affinity values: ; Among them, represents the candidate affinity value, represents the historical affinity information, is the sliding average coefficient, which is used to adjust the retention degree of the historical affinity information; Performing temperature scaling on the candidate affinity values to obtain target affinity values: ; Among them, represents the target affinity value, represents the temperature parameter, which is used to adjust the smoothness of the task similarity distribution; Constructing an affinity matrix based on the target affinity values between different task types.
4. The multi-UAV vision recognition method according to claim 3, wherein Constructing a task relationship graph according to the affinity matrix includes: Determining an affinity screening threshold based on the mean and standard deviation of the affinity matrix: ; Among them, is the affinity screening threshold, is the mean value, is the standard deviation value, is the threshold adjustment coefficient; Screening the affinity matrix according to the affinity screening threshold to determine the target relationship values in the affinity matrix that are higher than the affinity screening threshold; Constructing an edge set based on the target relationship values and constructing a node set according to the task type; Constructing a task relationship graph based on the edge set and the node set.
5. The multi-UAV vision recognition method according to any one of claims 3 or 4, characterized in that, After determining the multi-granularity prompt parameters of each task type based on the task relationship graph and sending the multi-granularity prompt parameters to each UAV, it further includes: Returning to execute the step of obtaining the local prompt parameters sent by each UAV, and determining the difference value between the local prompt parameters and the multi-granularity prompt parameters; Judging whether the difference value meets the convergence condition, where the convergence condition includes that the difference value is less than a preset threshold; If the convergence condition is met, stop executing the step of constructing an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters; If the convergence condition is not met, continue to execute the step of constructing an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters.
6. A multi-UAV vision recognition system, characterized in that, The multi-UAV vision recognition system includes: multiple UAVs and a server, and each UAV is communicatively connected to the server; The UAV is used for preprocessing the local data set collected during the task process to obtain a target data set; The UAV is further used for performing prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and the local prompt parameters of the candidate recognition model; The UAV is further used for sending the local prompt parameters to the server; The server is used for obtaining the local prompt parameters sent by each UAV, and the local prompt parameters are obtained by the UAV performing prompt training on the original recognition model based on the local data set; The server is further used for constructing an affinity matrix based on the task type corresponding to each UAV and the local prompt parameters, and the affinity matrix includes the target affinity values between different task types; The server is further used for constructing a task relationship graph according to the affinity matrix; The server is further used for determining the multi-granularity prompt parameters of each task type based on the task relationship graph and sending the multi-granularity prompt parameters to each UAV, so that each UAV performs prompt training on the original recognition model based on the multi-granularity prompt parameters; The drone is also used to obtain the multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are processed by the server based on the local prompt parameters sent by each drone; The drone is also used to aggregate the multi-granularity prompt parameters to obtain the fused prompt parameters; The drone is also used to perform prompt training on the candidate recognition model based on the fused prompt parameters to obtain the target recognition model, and input the target dataset into the target recognition model for visual recognition; The original recognition model is a multi-modal pre-trained model; the drone is also used to initialize the original recognition model to obtain the initial prompt parameters; obtain the prompt template structure matching the task type corresponding to the drone; perform prompt training on the original recognition model based on the local dataset, the prompt template structure, and the initial prompt parameters to obtain the candidate recognition model and the local prompt parameters of the candidate recognition model; The server is also used to aggregate the local prompt parameters of the drones with the same task type based on the task types corresponding to each drone: ; Among them, represents the set of drones performing tasks, represents the drone the data volume of the local dataset, represents the drone the local prompt parameter, represents the aggregated local prompt parameter; Calculate the initial affinity values between different task types according to the aggregated local prompt parameters: ; Among them, represents the initial affinity value, represents the aggregated local prompt parameter of the th task type, represents the aggregated local prompt parameter of the th task type, represents transpose; Update the initial affinity values based on the historical affinity information to obtain the candidate affinity values: ; Among them, represents the candidate affinity value, represents the historical affinity information, is the sliding average coefficient, which is used to adjust the retention degree of the historical affinity information; Perform temperature scaling on the candidate affinity values to obtain the target affinity values: ; Among them, represents the target affinity value, represents the temperature parameter, which is used to adjust the smoothness of the task similarity distribution; Construct an affinity matrix based on the target affinity values between different task types.
7. A multi-UAV vision recognition device, characterized in that, The multi-drone visual recognition device includes: a memory, a processor, and a multi-drone visual recognition program stored on the memory and executable on the processor, where the multi-drone visual recognition program is configured to implement the multi-drone visual recognition method according to any one of claims 1 to 2 or 3 to 5.
8. A computer-readable storage medium, characterized in that, A multi-drone visual recognition program is stored on the computer-readable storage medium, and when the multi-drone visual recognition program is executed by the processor, it implements the multi-drone visual recognition method according to any one of claims 1 to 2 or 3 to 5.
Citation Information
Patent Citations
Energy-saving federation method under fault diagnosis based on cross-granularity and prompt learning
CN118333137A
Task target allocation method for uncertain arrival time characteristics of auxiliary tasks in unmanned aerial vehicle group
CN118966646A