Multi-unmanned aerial vehicle visual identification method, system and device and storage medium
By using multi-drone visual recognition methods in the drone cluster, using servers to process and aggregate local prompt parameters of the drone to generate multi-grained prompt parameters, the problem of low recognition accuracy of a single drone in the drone cluster is solved, and efficient model training and visual recognition are achieved.
Patent Information
- Application Number
- CN202510426948.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
A single drone in a drone cluster is limited by local training data and computing resources, and cannot effectively train the recognition model, resulting in low visual recognition accuracy.
By introducing multi-drone visual recognition methods into the drone cluster, the server uses the local prompt parameters sent by each drone in the drone cluster, generate multi-grained prompt parameters, and send them back to the drone for model training and visual recognition.
It effectively solves the problem of limited computing resources and training data of a single drone, improves the local model training efficiency of the drone, and greatly improves the visual recognition accuracy of the drone.
Smart Images

Figure CN119942386A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle technology, and in particular to a multi-unmanned aerial vehicle visual recognition method, system, device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence and wireless communication technology, drones are increasingly used in emergency rescue, agricultural monitoring and other fields. These drones need to perform real-time target recognition and classification when performing tasks, generating a large amount of visual data and computing requirements. However, due to the limited computing resources of drones, complex and changeable mission environments, and high real-time requirements, traditional centralized visual recognition solutions face severe challenges.
[0003] Although edge computing provides new ideas for real-time visual task processing of drones, the following problems still exist in practical applications: the local training data of drones is usually limited, making it difficult to support end-to-end training of complex visual models; the computing resources of drones are severely limited, making it difficult to support online training and reasoning of large-scale deep models. Therefore, a single drone in the current drone cluster is limited by local training data and computing resources, and cannot effectively train the recognition model, resulting in low visual recognition accuracy. Summary of the invention
[0004] The main purpose of the present invention is to provide a multi-UAV visual recognition method, system, device and storage medium, aiming to solve the technical problem in the prior art that a single UAV in a UAV cluster is limited by local training data and computing resources, and cannot effectively train the recognition model, resulting in low visual recognition accuracy.
[0005] To achieve the above object, the present invention provides a multi-UAV visual recognition method, which is applied to each UAV in a UAV cluster, wherein the UAV cluster includes multiple UAVs, each UAV is communicatively connected to a server, and the multi-UAV visual recognition method includes: Preprocess the local data set collected during the task to obtain the target data set; Performing prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; Sending the local prompt parameter to the server, obtaining a multi-granularity prompt parameter fed back by the server based on the local prompt parameter, wherein the multi-granularity prompt parameter is obtained by the server based on the local prompt parameter sent by each drone; Aggregating the multi-granularity prompt parameters to obtain fused prompt parameters; The candidate recognition model is trained based on the fusion prompt parameters to obtain a target recognition model, and the target data set is input into the target recognition model for visual recognition.
[0006] Optionally, the original recognition model is a multimodal pre-trained model; the prompt training of the original recognition model based on the target data set to obtain a candidate recognition model and a local prompt parameter of the candidate recognition model includes: Initialize the original recognition model and obtain initial prompt parameters; Obtaining a prompt template structure that matches the mission type corresponding to the drone; Prompt training is performed on the original recognition model based on the local data set, the prompt template structure and the initial prompt parameters to obtain a candidate recognition model and local prompt parameters of the candidate recognition model.
[0007] Optionally, the multi-granularity prompt parameters include fine-granularity prompt parameters, medium-granularity prompt parameters, and coarse-granularity prompt parameters; and aggregating the multi-granularity prompt parameters to obtain fused prompt parameters includes: Calculate the attention scores of the multi-granularity prompt parameters to obtain fine-grained attention scores, medium-grained attention scores, and coarse-grained attention scores: ; ; ; in, Indicates fine-grained prompt parameters, Indicates medium-granularity prompt parameters, Represents the coarse-grained prompt parameter, 、 and They represent the weight matrices, 、 and They represent the bias vectors, is the attention vector, represents the fine-grained attention score, represents the medium-granularity attention score, represents the coarse-grained attention score; The fine-grained attention score, the medium-grained attention score, and the coarse-grained attention score are normalized to obtain a normalized score: ; in, and They represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-granularity attention score, represents the normalized coarse-grained attention score; Aggregating the fine-grained prompt parameter, the medium-grained prompt parameter, and the coarse-grained prompt parameter based on the normalized score to obtain a fused prompt parameter: ; in, Hint parameters for fusion.
[0008] In addition, to achieve the above-mentioned purpose, the present invention further proposes a multi-UAV visual recognition method, which is applied to a server, the server is communicatively connected with each UAV in a UAV cluster, the UAV cluster includes multiple UAVs, and the multi-UAV visual recognition method includes: Acquire local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on an original recognition model based on a local data set; Building an affinity matrix based on the task type corresponding to each drone and the local prompt parameter, the affinity matrix including target affinity values between different task types; constructing a task relationship diagram according to the affinity matrix; Based on the task relationship diagram, multi-granularity prompt parameters of each task type are determined, and the multi-granularity prompt parameters are sent to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters.
[0009] Optionally, constructing an affinity matrix based on the mission type corresponding to each drone and the local prompt parameter includes: Based on the mission type corresponding to each drone, the local prompt parameters of drones with the same mission type are aggregated: ; in, represents the set of drones that perform the mission, Indicates drone The amount of data in the local dataset, Indicates the local prompt parameters of the drone, Represents the aggregated local prompt parameters; Calculate the initial affinity values between different task types based on the aggregated local prompt parameters: ; in, represents the initial affinity value, Indicates The aggregated local prompt parameters for each task type, Indicates The aggregated local prompt parameters for each task type, stands for transpose; The initial affinity value is updated based on the historical affinity information to obtain a candidate affinity value: ; in, represents the candidate affinity value, Represents historical affinity information, is the sliding average coefficient, Used to adjust the degree of retention of historical affinity information; The candidate affinity values are temperature scaled to obtain the target affinity value: ; in, represents the target affinity value, represents the temperature parameter, Used to adjust the smoothness of task similarity distribution; An affinity matrix is constructed based on the target affinity values between different task types.
[0010] Optionally, constructing a task relationship diagram according to the affinity matrix includes: Affinity screening thresholds were determined based on the mean and standard deviation values of the affinity matrix: ; in, is the affinity screening threshold, is the mean, is the standard deviation, is the threshold adjustment coefficient; Screening the affinity matrix according to the affinity screening threshold, and determining target relationship values in the affinity matrix that are higher than the affinity screening threshold; Building an edge set based on the target relationship value, and building a node set according to the task type; A task relationship graph is constructed based on the edge set and the node set.
[0011] Optionally, after determining the multi-granularity prompt parameters of each task type based on the task relationship graph and sending the multi-granularity prompt parameters to each drone, the method further includes: Returning to the step of obtaining the local prompt parameter sent by each UAV, determining the difference between the local prompt parameter and the multi-granularity prompt parameter; Determine whether the difference value satisfies a convergence condition, wherein the convergence condition includes that the difference value is less than a preset threshold; If the convergence condition is met, the step of constructing the affinity matrix based on the task type corresponding to each drone and the local prompt parameter is stopped; If the convergence condition is not met, the step of constructing an affinity matrix based on the task type corresponding to each drone and the local prompt parameter is continued.
[0012] In addition, to achieve the above-mentioned purpose, the present invention also proposes a multi-UAV visual recognition system, the multi-UAV visual recognition system comprising: a plurality of UAVs and a server, each UAV being communicatively connected to the server; The drone is used to pre-process the local data set collected during the mission to obtain the target data set; The drone is further used to perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; The drone is further used to send the local prompt parameter to the server; The server is used to obtain local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set; The server is further used to construct an affinity matrix based on the task type corresponding to each drone and the local prompt parameter, wherein the affinity matrix includes target affinity values between different task types; The server is further used to construct a task relationship graph according to the affinity matrix; The server is further configured to determine multi-granularity prompt parameters of each task type based on the task relationship diagram, and send the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters; The drone is further used to obtain multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone; The drone is further used to aggregate the multi-granularity prompt parameters to obtain fused prompt parameters; The drone is also used to perform prompt training on the candidate recognition model based on the fusion prompt parameters to obtain a target recognition model, and input the target data set into the target recognition model for visual recognition.
[0013] In addition, to achieve the above-mentioned objectives, the present application also proposes a multi-UAV visual recognition device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multi-UAV visual recognition method as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-UAV visual recognition method described above are implemented.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-UAV visual recognition method described above.
[0016] The present invention uses a drone to pre-process a local data set collected during a task, obtains a target data set, performs prompt training on an original recognition model based on the target data set, obtains a candidate recognition model and local prompt parameters of the candidate recognition model, and sends the local prompt parameters to a server; obtains local prompt parameters sent by each drone through a server, the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set, constructs an affinity matrix based on the task type corresponding to each drone and the local prompt parameters, the affinity matrix includes target affinity values between different task types, constructs a task relationship diagram based on the affinity matrix, determines multi-granularity prompt parameters of each task type based on the task relationship diagram, and sends the multi-granularity prompt parameters to each drone, so that each drone can determine the target affinity value between different task types based on the multi-granularity prompt parameters. The original recognition model is prompted for training; the multi-granularity prompt parameters fed back by the server based on the local prompt parameters are obtained through the drone, the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone, the multi-granularity prompt parameters are aggregated to obtain fused prompt parameters, the candidate recognition model is prompted for training based on the fused prompt parameters, the target recognition model is obtained, and the target data set is input into the target recognition model for visual recognition; because the present invention processes the local prompt parameters sent by each drone in the drone cluster through the server to obtain the multi-granularity prompt parameters and sends them to each drone for model training, the problem of limited computing resources and training data of a single drone is effectively solved, the local model training efficiency of the drone is effectively improved, and the visual recognition accuracy of the drone is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 It is a schematic diagram of the structure of a multi-UAV visual recognition device in a hardware operating environment involved in an embodiment of the present invention; Figure 2 This is a schematic diagram of the process of the first embodiment of the multi-UAV visual recognition method of the present invention; Figure 3 This is a schematic diagram of the flow chart of the second embodiment of the multi-UAV visual recognition method of the present invention; Figure 4 This is a structural block diagram of the first embodiment of the multi-UAV visual recognition system of the present invention.
[0019] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0021] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a multi-UAV visual recognition device in the hardware operating environment involved in an embodiment of the present invention.
[0022] like Figure 1 As shown, the multi-UAV visual recognition device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wireless-Fidelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk storage. The memory 1005 may also be a storage system independent of the aforementioned processor 1001.
[0023] Those skilled in the art will understand that Figure 1 The structure shown in does not constitute a limitation on the multi-UAV visual recognition device, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0024] like Figure 1 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a multi-UAV visual recognition program.
[0025] exist Figure 1 In the multi-UAV visual recognition device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the multi-UAV visual recognition device of the present invention can be set in the multi-UAV visual recognition device, and the multi-UAV visual recognition device calls the multi-UAV visual recognition program stored in the memory 1005 through the processor 1001, and executes the multi-UAV visual recognition method provided by the embodiment of the present invention.
[0026] The embodiment of the present invention provides a multi-UAV visual recognition method, referring to Figure 2 , Figure 2 Schematic diagram of the process of the first embodiment of the multi-UAV visual recognition method of the present invention.
[0027] In this embodiment, the multi-drone visual recognition method is applied to each drone in a drone cluster, the drone cluster includes multiple drones, each drone is communicatively connected to a server, and the multi-drone visual recognition method includes the following steps: Step S10: pre-process the local data set collected during the task to obtain the target data set.
[0028] It should be noted that the execution subject of this embodiment may be a drone device with data processing, network communication and program running functions, such as a central controller of a drone, etc. The following takes a drone as an example to illustrate this embodiment and the following embodiments.
[0029] It should be noted that the local data set may be the image data collected by the drone during the mission and the label information corresponding to the image data. Preprocessing may be to process the local data set by normalizing, standardizing, clarifying the data, enhancing the data, etc.
[0030] In some embodiments, during the system initialization phase, the set of drones in the system is first defined as , where N is the total number of drones. Each drone Perform specific visual recognition tasks , the task set is recorded as , where M represents the total number of different task types. For each drone’s local dataset, it is expressed as: ; in Indicates the amount of data, represents image data, Indicates the corresponding label information.
[0031] In some embodiments, for each drone local dataset , this embodiment designs a complete set of data preprocessing procedures. First, image standardization is performed, and the formula The image data is normalized, where μ and σ represent the mean and standard deviation of the image, respectively. Subsequently, in order to enhance the generalization ability of the model, data augmentation strategies are implemented on the standardized images, including random cropping, horizontal flipping and other operations, to generate an enhanced dataset. Finally, the processed data set is divided into mini-batches of size B to obtain the target data set, thus preparing for subsequent batch training.
[0032] Step S20: Perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model.
[0033] It should be noted that the original recognition model can be a pre-trained multimodal large model deployed locally on the drone, and the original recognition model can include an image encoder and a text encoder. In the local training phase, the parameters of the pre-trained image encoder and text encoder are frozen.
[0034] In a specific implementation, the drone can generate initial prompt parameters based on the target data set and the type of mission performed by the drone, and input the initial prompt parameters and the target data set into the original recognition model to train the prompt parameters of the original recognition model, obtain a candidate recognition model, and obtain the local prompt parameters of the candidate recognition model.
[0035] Step S30: Send the local prompt parameter to the server, and obtain the multi-granularity prompt parameter fed back by the server based on the local prompt parameter.
[0036] It should be noted that the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone.
[0037] In some embodiments, the server obtains local prompt parameters sent by each drone, the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set, an affinity matrix is constructed based on the task type corresponding to each drone and the local prompt parameters, the affinity matrix includes target affinity values between different task types, a task relationship graph is constructed according to the affinity matrix, multi-granularity prompt parameters of each task type are determined based on the task relationship graph, and the multi-granularity prompt parameters are sent to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters.
[0038] Furthermore, in order to improve the model performance and accurately obtain the local prompt parameters, in some embodiments, the original recognition model is a multimodal pre-trained model, and the above step S30 may include: Step S31: Initialize the original recognition model and obtain initial prompt parameters; Step S32: obtaining a prompt template structure matching the mission type corresponding to the drone; Step S33: performing prompt training on the original recognition model based on the local data set, the prompt template structure and the initial prompt parameters to obtain a candidate recognition model and local prompt parameters of the candidate recognition model.
[0039] It should be noted that the original recognition model is a multimodal pre-training model, for example, the original recognition model may be a CLIP (Contrastive Language–Image Pretraining) model.
[0040] In the specific implementation, the drone can deploy the pre-trained CLIP model locally as the basic feature extractor. The model mainly consists of two core components: image encoder and text encoder The image encoder uses the Vision Transformer architecture, receives image input of dimension 3×H×W, and outputs a d-dimensional feature vector; the text encoder uses the Transformer architecture, sets the maximum token length to L, and also outputs a d-dimensional feature vector. In the local training phase, freeze the pre-trained image encoder and text encoder Parameters.
[0041] It should be noted that for each task A learnable prompt template structure is designed, which is of the form: ;
[0042] The prompt parameters is a learnable continuous vector, using a normal distribution Initialization. This template design allows the model to adapt to different visual recognition requirements by learning task-specific prompt parameters while maintaining the possibility of cross-task knowledge sharing. The dimension of each prompt parameter vector matches the text feature dimension of the CLIP model, ensuring the consistency of the feature space.
[0043] It is understandable that during the prompt parameter training process, the local model receives two types of input: image data collected by the drone and the corresponding prompt parameters . Image Encoder and text encoder Extract features from the inputs and get feature vectors and The feature fusion module uses a multi-layer perceptron (MLP) for feature fusion: ,in Represents the feature concatenation operation. Classification prediction is achieved through the fully connected layer: , where W is the weight matrix and b is the bias term. The final category probability is calculated by the softmax function: .
[0044] The loss function design consists of three components: classification loss , contrast loss and regularization loss The total loss function expression is: ; Among them, the classification loss is in the form of cross entropy: ; Contrastive loss is used to enhance the discriminability of feature representation: ; The norm of the regularized loss control hint parameter: ; in is a learnable parameter, represents the total number of categories. At the same time, we introduce the regularization loss To prevent overfitting. By adjusting the weight coefficient and , which can balance the contribution of losses.
[0045] Step S40: Aggregate the multi-granularity prompt parameters to obtain fused prompt parameters.
[0046] It should be noted that the multi-granularity parameters are prompt parameters of multiple granularities, and the multi-granularity prompt parameters may include fine-granularity prompt parameters, medium-granularity prompt parameters, and coarse-granularity prompt parameters.
[0047] In some embodiments, the server sends prompt parameters of three granularities to the corresponding drones, and the adaptive fusion module in the drones uses an attention mechanism to fuse prompt parameters of different granularities to obtain fused prompt parameters.
[0048] Furthermore, in order to effectively fuse the multi-granularity parameters, in some embodiments, the multi-granularity prompt parameters include fine-granularity prompt parameters, medium-granularity prompt parameters and coarse-granularity prompt parameters, and the above step S40 may include: Step S41: Calculate the attention scores of the multi-granularity prompt parameters to obtain a fine-grained attention score, a medium-grained attention score, and a coarse-grained attention score.
[0049] It should be noted that the server sends the prompt parameters of three granularities to the corresponding drones, and the local adaptive fusion module of the drone uses the attention mechanism to fuse the prompt information of different granularities. First, the attention score of each granularity prompt parameter is calculated. The attention score calculation formula is as follows: ; ; ; in, Indicates fine-grained prompt parameters, Indicates medium-granularity prompt parameters, Represents the coarse-grained prompt parameter, 、 and They represent the weight matrices, 、 and They represent the bias vectors, is the attention vector, represents the fine-grained attention score, represents the medium-granularity attention score, Represents the coarse-grained attention score.
[0050] Step S42: normalize the fine-grained attention score, the medium-grained attention score, and the coarse-grained attention score to obtain a normalized score.
[0051] It should be noted that the attention score can be obtained by Function normalization, refer to the following formula: ; in, and They represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-granularity attention score, represents the normalized coarse-grained attention score; Step S403: Aggregate the fine-grained prompt parameter, the medium-grained prompt parameter and the coarse-grained prompt parameter based on the normalized score to obtain a fused prompt parameter.
[0052] It should be noted that the prompt parameters of each granularity are aggregated based on the normalized score. As a local text encoder Input prompt parameters, aggregation refers to the following formula: ; in, Hint parameters for fusion.
[0053] Step S50: Perform prompt training on the candidate recognition model based on the fusion prompt parameters to obtain a target recognition model, and input the target data set into the target recognition model for visual recognition.
[0054] It should be noted that the UAV performs prompt training on the candidate recognition model based on the fused prompt parameters, updates the local prompt parameters based on the prompt training results, and sends the updated local prompt parameters to the server. The server determines whether convergence is achieved based on the received local prompt parameters (for example, determining whether the difference between the local prompt parameters sent by two consecutive rounds of the UAV is less than a preset threshold, and if so, convergence is determined). If convergence is determined, the local prompt parameters sent by the UAV are stopped from being processed, and a model convergence message is fed back to the UAV. After receiving the model convergence message, the UAV performs prompt training on the candidate recognition model based on the updated local prompt parameters to obtain a target recognition model, and inputs the target data set into the target recognition model for visual recognition.
[0055] In some embodiments, the drone cluster and the server update the local recognition model through federated learning, including the following steps: Step 1: Parameter initialization; In the initial stage of federated learning, the system initializes prompt parameters , CLIP model parameters, and the global hint vector. In the initial round, the local hint parameters of each client are initialized to the same value: .
[0056] Step 2: Model parameter broadcasting; At the beginning of each iteration, the server broadcasts multi-granularity prompt parameters to all drone clients, including fine-grained prompt parameters , Medium-grained prompt parameters and coarse-grained hint parameters .
[0057] Step 3: Local training and parameter upload; The client receives the multi-granularity hint parameters given by the server, and the local client freezes the image encoder and text encoder parameters, and updates other learnable parameters, each drone client based on the local data set Optimizing the loss function ,in: ;
[0058] Then upload the updated prompt parameters And the client's current task to the server.
[0059] Step 4: The server builds a task affinity relationship diagram; 1. Calculate the task affinity matrix A, where each element is calculated by cosine similarity: ; 2. Update affinity value using sliding average mechanism: ; 3. Construct a weighted undirected graph , the edge weights are calculated by temperature-scaled softmax: ; 4. Apply PageRank algorithm to optimize edge weights: ; Step 5: The server calculates multi-granularity prompt parameters; Based on the constructed relationship graph, we Three neighbor sets with different hop counts are defined , and , and calculate the multi-granularity cue parameters for each task.
[0060] Repeat steps 3-5 until the model converges. The convergence condition is that the change of the global prompt parameter in two consecutive rounds is less than the preset threshold. .
[0061] This embodiment obtains a target data set by preprocessing the collected local data set, performs prompt training on the original recognition model based on the target data set, obtains a candidate recognition model and local prompt parameters of the candidate recognition model, sends the local prompt parameters to the server, obtains multi-granularity prompt parameters based on the local prompt parameter feedback of the server, and the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone in the drone cluster; prompt training is performed on the candidate recognition model based on the fused prompt parameters obtained by aggregating the multi-granularity prompt parameters to obtain a target recognition model, and the target data set is input into the target recognition model for visual recognition, which effectively solves the problem of limited computing resources and training data of a single drone and greatly improves the visual recognition accuracy of the drone.
[0062] refer to Figure 3 , Figure 3 Schematic diagram of the flow chart of the second embodiment of the multi-UAV visual recognition method of the present invention.
[0063] In this embodiment, the multi-drone visual recognition method is applied to a server, the server is communicatively connected with each drone in a drone cluster, the drone cluster includes multiple drones, and the multi-drone visual recognition method includes: Step S100: Obtain local prompt parameters sent by each drone.
[0064] It should be understood that the execution subject of this embodiment can be a server with data processing, network communication and program running functions, such as a computer, or a terminal electronic device capable of implementing the above functions. The following takes the server as an example to illustrate this embodiment and the following embodiments.
[0065] It should be noted that the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set.
[0066] In some embodiments, the drone sends the locally trained local prompt parameters and the mission type corresponding to the drone to the server, and the server aggregates the local prompt parameters of drones of the same type according to the mission type.
[0067] Step S200: constructing an affinity matrix based on the mission type corresponding to each drone and the local prompt parameter.
[0068] It should be noted that the affinity matrix includes target affinity values between different task types.
[0069] In a specific implementation, the server may determine the affinity values between different task types by calculating the cosine similarities between different task types.
[0070] Furthermore, in order to accurately calculate the affinity values between different task types, in some embodiments, the above step S200 may include: Step S2001: Aggregating local prompt parameters of drones of the same mission type based on the mission type corresponding to each drone; Step S2002: Calculate the initial affinity values between different task types according to the aggregated local prompt parameters; Step S2003: updating the initial affinity value based on historical affinity information to obtain a candidate affinity value; Step S2004: performing temperature scaling on the candidate affinity value to obtain a target affinity value; Step S2005: constructing an affinity matrix based on the target affinity values between different task types.
[0071] It should be noted that the server can classify the local prompt parameters of the drone based on the mission type and aggregate the local prompt parameters of the same type: ; in, represents the set of drones that perform the mission, Indicates drone The amount of data in the local dataset, Indicates the local prompt parameters of the drone, Represents the aggregated local prompt parameters; It can be understood that the server calculates the affinity matrix between different types of tasks based on the aggregated local prompt parameters. Each element in the matrix is calculated by cosine similarity: ; in, represents the initial affinity value, Indicates The aggregated local prompt parameters for each task type, Indicates The aggregated local prompt parameters for each task type, stands for transpose; It should be noted that in order to improve the stability of affinity calculation, the server can use a sliding average mechanism to update the affinity value: ; in, represents the candidate affinity value, Represents historical affinity information, is the sliding average coefficient, Used to adjust the degree of retention of historical affinity information; It should be noted that the temperature scaling of the candidate affinity values is performed according to the following formula: ; in, represents the target affinity value, represents the temperature parameter, Used to adjust the smoothness of task similarity distribution.
[0072] Step S300: constructing a task relationship diagram according to the affinity matrix.
[0073] It should be noted that the task relationship graph may be a weighted undirected graph, which includes a plurality of nodes, each node representing a different task type, and each target affinity value in the affinity matrix serves as an edge connecting each node.
[0074] Furthermore, in order to accurately construct the task relationship diagram, in some embodiments, the above step S300 may include: Step S3001: determining an affinity screening threshold based on the mean and standard deviation of the affinity matrix; Step S3002: screening the affinity matrix according to the affinity screening threshold, and determining target relationship values in the affinity matrix that are higher than the affinity screening threshold; Step S3003: constructing an edge set based on the target relationship value, and constructing a node set according to the task type; Step S3004: construct a task relationship graph based on the edge set and the node set.
[0075] It should be noted that this embodiment designs an adaptive threshold mechanism to dynamically adjust the affinity judgment standard, and the affinity screening threshold is determined by referring to the following formula: ; in, is the affinity screening threshold, is the mean, is the standard deviation, is the threshold adjustment coefficient, Used to control the stringency of the affinity screening threshold.
[0076] It is understandable that the server may filter the affinity values based on the affinity screening threshold, retain only the affinity values above the threshold, and use the affinity values in the affinity matrix that are above the affinity screening threshold as the target relationship values, referring to the following formula: ;
[0077] It should be understood that the server constructs a weighted undirected graph G=(V, E) based on the filtered affinity matrix, where the node set V corresponds to the task set; the edge set E contains only connections with non-zero affinity values; the edge weights are normalized using softmax, that is, the affinity matrix is normalized: ; in, represents the set of all nodes connected to node i, represents the normalized affinity matrix.
[0078] Step S400: determining multi-granularity prompt parameters for each task type based on the task relationship diagram, and sending the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters.
[0079] It should be noted that the server determines the node path length and node relationship between different task types based on the task relationship graph, and constructs a task neighbor set for each different task type based on the node path length and node relationship. The task neighbor set includes a one-hop neighbor set, a two-hop neighbor set, and a three-hop neighbor set. Prompt aggregation is performed based on the task neighbor set to obtain multi-granularity prompt parameters.
[0080] In the specific implementation, the server defines three neighbor sets with different hop counts for different task types based on the task relationship graph: Contains nodes directly connected to task i; two-hop neighbor set Reachable through intermediate nodes; three-hop neighbor set . It needs to go through two intermediate nodes. In order to smooth the impact of different hop numbers, we introduce the distance attenuation factor , where d is the number of hops, ∈(0,1) is the decay rate. The influence of node j on task i can be expressed as: ; in is the shortest path length between nodes, Represents the weight value of the edge in the task relationship graph, that is, the normalized affinity matrix.
[0081] After obtaining a set of tasks of different granularities, we designed a hierarchical prompt aggregation mechanism. , whose fine-grained hint parameter is obtained by weighted averaging of one-hop neighbors: ; The medium-granularity hint parameter considers tasks that are reachable within two hops: ; The coarse-grained hint parameter aggregates all task information within three hops: ;
[0082] Furthermore, in order to dynamically update the prompt parameters of the drone and converge under reasonable conditions, in some embodiments, after the above step S400, the following is further included: Step S500: returning to the step of obtaining the local prompt parameter sent by each UAV, and determining the difference value between the local prompt parameter and the multi-granularity prompt parameter; Step S600: determining whether the difference value satisfies a convergence condition, wherein the convergence condition includes that the difference value is less than a preset threshold; Step S700: If the convergence condition is met, then the step of constructing the affinity matrix based on the task type corresponding to each drone and the local prompt parameter is stopped; Step S800: If the convergence condition is not met, continue to execute the step of constructing an affinity matrix based on the task type corresponding to each drone and the local prompt parameter.
[0083] It should be noted that the UAV performs prompt training on the candidate recognition model based on the fused prompt parameters, updates the local prompt parameters based on the prompt training results, and sends the updated local prompt parameters to the server. The server determines whether convergence is achieved based on the received local prompt parameters (for example, determining whether the difference between the local prompt parameters sent by two consecutive rounds of the UAV is less than a preset threshold, and if so, convergence is determined). If convergence is determined, the local prompt parameters sent by the UAV are stopped from being processed, and a model convergence message is fed back to the UAV. After receiving the model convergence message, the UAV performs prompt training on the candidate recognition model based on the updated local prompt parameters to obtain a target recognition model, and inputs the target data set into the target recognition model for visual recognition.
[0084] This embodiment obtains local prompt parameters sent by each drone, wherein the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set, constructs an affinity matrix based on the task type corresponding to each drone and the local prompt parameters, wherein the affinity matrix includes target affinity values between different task types, constructs a task relationship graph according to the affinity matrix, determines multi-granularity prompt parameters for each task type based on the task relationship graph, and sends the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters, thereby effectively avoiding the problem that model parameters cannot be aggregated due to task differences between drones, and by aggregating the prompt parameters of each drone in the drone cluster, the adaptability of the recognition model in the drone to different tasks is effectively improved, and the recognition accuracy of the drone is effectively improved.
[0085] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, on which a multi-drone visual recognition program is stored. When the multi-drone visual recognition program is executed by a processor, the steps of the multi-drone visual recognition method described above are implemented.
[0086] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0087] The above-mentioned computer-readable storage medium may be included in the multi-UAV visual recognition device; or it may exist independently without being assembled into the multi-UAV visual recognition device.
[0088] In addition, an embodiment of the present invention further proposes a computer program product, including a multi-drone visual recognition program, which, when executed by a processor, implements the steps of the multi-drone visual recognition method described above.
[0089] The specific implementation manner of the computer program product of the present invention is basically the same as the various embodiments of the above-mentioned multi-UAV visual recognition method, and will not be repeated here.
[0090] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the multi-UAV visual recognition system of the present invention.
[0091] like Figure 4 As shown, the multi-UAV visual recognition system proposed in the embodiment of the present invention includes: a plurality of UAVs 10 and a server 20, each UAV being communicatively connected with the server; The drone 10 is used to pre-process the local data set collected during the mission to obtain the target data set; The drone 10 is further used to perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; The drone 10 is further used to send the local prompt parameter to the server; The server 20 is used to obtain local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set; The server 20 is further used to construct an affinity matrix based on the task type corresponding to each drone and the local prompt parameter, wherein the affinity matrix includes target affinity values between different task types; The server 20 is further used to construct a task relationship graph according to the affinity matrix; The server 20 is further configured to determine multi-granularity prompt parameters of each task type based on the task relationship diagram, and send the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters; The drone 10 is further used to obtain the multi-granularity prompt parameter fed back by the server based on the local prompt parameter, where the multi-granularity prompt parameter is obtained by the server based on the local prompt parameter sent by each drone; The drone 10 is further used to aggregate the multi-granularity prompt parameters to obtain fused prompt parameters; The drone 10 is further used to perform prompt training on the candidate recognition model based on the fusion prompt parameter to obtain a target recognition model, and input the target data set into the target recognition model for visual recognition.
[0092] In this embodiment, a local data set collected during a task is preprocessed by a drone to obtain a target data set, prompt training is performed on an original recognition model based on the target data set, a candidate recognition model and local prompt parameters of the candidate recognition model are obtained, and the local prompt parameters are sent to a server; the local prompt parameters sent by each drone are obtained by the server through prompt training of the original recognition model by the drone based on the local data set, an affinity matrix is constructed based on the task type corresponding to each drone and the local prompt parameters, the affinity matrix includes target affinity values between different task types, a task relationship diagram is constructed according to the affinity matrix, multi-granularity prompt parameters of each task type are determined based on the task relationship diagram, and the multi-granularity prompt parameters are sent to each drone, so that each drone can determine the target affinity value based on the multi-granularity prompt parameters. The original recognition model is prompted for training; the multi-granularity prompt parameters fed back by the server based on the local prompt parameters are obtained through the drone, the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone, the multi-granularity prompt parameters are aggregated to obtain fused prompt parameters, the candidate recognition model is prompted for training based on the fused prompt parameters, the target recognition model is obtained, and the target data set is input into the target recognition model for visual recognition; because this embodiment uses the server to process the local prompt parameters sent by each drone in the drone cluster to obtain the multi-granularity prompt parameters and send them to each drone for model training, thereby effectively solving the problem of limited computing resources and training data of a single drone, effectively improving the local model training efficiency of the drone, and greatly improving the drone's visual recognition accuracy.
[0093] The multi-drone visual recognition system provided by the present application adopts the multi-drone visual recognition method in the above-mentioned embodiment, which can solve the technical problem of multi-drone visual recognition. Compared with the prior art, the beneficial effects of the multi-drone visual recognition system provided by the present application are the same as the beneficial effects of the multi-drone visual recognition method provided by the above-mentioned embodiment, and other technical features of the multi-drone visual recognition system are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0094] It should be understood that the above is only an example and does not constitute any limitation on the technical solution of the present invention. In specific applications, technicians in this field can make settings as needed, and the present invention does not limit this.
[0095] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of the present invention. In practical applications, technicians in this field can select part or all of them according to actual needs to achieve the purpose of the present embodiment, and no limitation is made here.
[0096] In addition, for technical details not described in detail in this embodiment, reference may be made to the multi-UAV visual recognition method provided in any embodiment of the present invention, and will not be repeated here.
[0097] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0098] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0099] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0100] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multi-UAV visual recognition method, characterized in that: The multi-UAV visual recognition method is applied to each UAV in a UAV cluster, wherein the UAV cluster includes multiple UAVs, each UAV is connected to a server for communication, and the multi-UAV visual recognition method includes: Preprocess the local data set collected during the task to obtain the target data set; Performing prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; Sending the local prompt parameter to the server, obtaining a multi-granularity prompt parameter fed back by the server based on the local prompt parameter, wherein the multi-granularity prompt parameter is obtained by the server based on the local prompt parameter sent by each drone; Aggregating the multi-granularity prompt parameters to obtain fused prompt parameters; The candidate recognition model is trained based on the fusion prompt parameters to obtain a target recognition model, and the target data set is input into the target recognition model for visual recognition.
2. The multi-UAV visual recognition method according to claim 1, characterized in that: The original recognition model is a multimodal pre-trained model; the prompt training of the original recognition model based on the target data set to obtain a candidate recognition model and a local prompt parameter of the candidate recognition model includes: Initialize the original recognition model and obtain initial prompt parameters; Obtaining a prompt template structure that matches the mission type corresponding to the drone; Prompt training is performed on the original recognition model based on the local data set, the prompt template structure and the initial prompt parameters to obtain a candidate recognition model and local prompt parameters of the candidate recognition model.
3. The multi-UAV visual recognition method according to claim 2, characterized in that: The multi-granularity prompt parameters include fine-granularity prompt parameters, medium-granularity prompt parameters and coarse-granularity prompt parameters; The aggregating the multi-granularity prompt parameters to obtain fused prompt parameters includes: Calculate the attention scores of the multi-granularity prompt parameters to obtain fine-grained attention scores, medium-grained attention scores, and coarse-grained attention scores: ; ; ; in, Indicates fine-grained prompt parameters, Indicates medium-granularity prompt parameters, Represents the coarse-grained prompt parameter, 、 and They represent the weight matrices, 、 and They represent the bias vectors, is the attention vector, represents the fine-grained attention score, represents the medium-granularity attention score, represents the coarse-grained attention score; The fine-grained attention score, the medium-grained attention score, and the coarse-grained attention score are normalized to obtain a normalized score: ; in, and They represent the normalized scores, represents the normalized fine-grained attention score, represents the normalized medium-granularity attention score, represents the normalized coarse-grained attention score; Aggregating the fine-grained prompt parameter, the medium-grained prompt parameter, and the coarse-grained prompt parameter based on the normalized score to obtain a fused prompt parameter: ; in, Hint parameters for fusion.
4. A multi-UAV visual recognition method, characterized in that: The multi-UAV visual recognition method is applied to a server, the server is in communication connection with each UAV in a UAV cluster, the UAV cluster includes a plurality of UAVs, and the multi-UAV visual recognition method includes: Acquire local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on an original recognition model based on a local data set; Building an affinity matrix based on the task type corresponding to each drone and the local prompt parameter, the affinity matrix including target affinity values between different task types; constructing a task relationship diagram according to the affinity matrix; Based on the task relationship diagram, multi-granularity prompt parameters of each task type are determined, and the multi-granularity prompt parameters are sent to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters.
5. The multi-UAV visual recognition method according to claim 4, characterized in that: The constructing of an affinity matrix based on the mission type corresponding to each drone and the local prompt parameter includes: Based on the mission type corresponding to each drone, the local prompt parameters of drones with the same mission type are aggregated: ; in, represents the set of drones that perform the mission, Indicates drone The amount of data in the local dataset, Indicates the local prompt parameters of the drone, Represents the aggregated local prompt parameters; Calculate the initial affinity values between different task types based on the aggregated local prompt parameters: ; in, represents the initial affinity value, Indicates The aggregated local prompt parameters for each task type, Indicates The aggregated local prompt parameters for each task type, stands for transpose; The initial affinity value is updated based on the historical affinity information to obtain a candidate affinity value: ; in, represents the candidate affinity value, Represents historical affinity information, is the sliding average coefficient, Used to adjust the degree of retention of historical affinity information; The candidate affinity values are temperature scaled to obtain the target affinity value: ; in, represents the target affinity value, represents the temperature parameter, Used to adjust the smoothness of task similarity distribution; An affinity matrix is constructed based on the target affinity values between different task types.
6. The multi-UAV visual recognition method according to claim 5, characterized in that: The step of constructing a task relationship diagram according to the affinity matrix includes: Affinity screening thresholds were determined based on the mean and standard deviation values of the affinity matrix: ; in, is the affinity screening threshold, is the mean, is the standard deviation, is the threshold adjustment coefficient; Screening the affinity matrix according to the affinity screening threshold, and determining target relationship values in the affinity matrix that are higher than the affinity screening threshold; Building an edge set based on the target relationship value, and building a node set according to the task type; A task relationship graph is constructed based on the edge set and the node set.
7. The multi-UAV visual recognition method according to any one of claims 4 to 6, characterized in that: After determining the multi-granularity prompt parameters of each task type based on the task relationship graph and sending the multi-granularity prompt parameters to each drone, the method further includes: Returning to the step of obtaining the local prompt parameter sent by each UAV, determining the difference between the local prompt parameter and the multi-granularity prompt parameter; Determine whether the difference value satisfies a convergence condition, wherein the convergence condition includes that the difference value is less than a preset threshold; If the convergence condition is met, the step of constructing the affinity matrix based on the task type corresponding to each drone and the local prompt parameter is stopped; If the convergence condition is not met, the step of constructing an affinity matrix based on the task type corresponding to each drone and the local prompt parameter is continued.
8. A multi-UAV visual recognition system, characterized in that: The multi-UAV visual recognition system comprises: a plurality of UAVs and a server, each UAV being communicatively connected to the server; The drone is used to pre-process the local data set collected during the mission to obtain the target data set; The drone is further used to perform prompt training on the original recognition model based on the target data set to obtain a candidate recognition model and local prompt parameters of the candidate recognition model; The drone is further used to send the local prompt parameter to the server; The server is used to obtain local prompt parameters sent by each drone, where the local prompt parameters are obtained by the drone performing prompt training on the original recognition model based on the local data set; The server is further used to construct an affinity matrix based on the task type corresponding to each drone and the local prompt parameter, wherein the affinity matrix includes target affinity values between different task types; The server is further used to construct a task relationship graph according to the affinity matrix; The server is further configured to determine multi-granularity prompt parameters of each task type based on the task relationship diagram, and send the multi-granularity prompt parameters to each drone, so that each drone performs prompt training on the original recognition model based on the multi-granularity prompt parameters; The drone is further used to obtain multi-granularity prompt parameters fed back by the server based on the local prompt parameters, where the multi-granularity prompt parameters are obtained by the server based on the local prompt parameters sent by each drone; The drone is further used to aggregate the multi-granularity prompt parameters to obtain fused prompt parameters; The drone is also used to perform prompt training on the candidate recognition model based on the fusion prompt parameters to obtain a target recognition model, and input the target data set into the target recognition model for visual recognition.
9. A multi-UAV visual recognition device, characterized in that: The multi-UAV visual recognition device includes: a memory, a processor, and a multi-UAV visual recognition program stored in the memory and executable on the processor, wherein the multi-UAV visual recognition program is configured to implement the multi-UAV visual recognition method as described in any one of claims 1 to 3 or 4 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multi-drone visual recognition program, which, when executed by a processor, implements the multi-drone visual recognition method according to any one of claims 1 to 3 or 4 to 7.
Citation Information
Patent Citations
Unmanned aerial vehicle cooperative positioning system and method based on graph neural network
CN112414401A
Federal learning-based speech recognition method and system, and computer equipment
CN116343760A
Human body analysis network training method and device based on unmanned aerial vehicle image
CN116778528A
Energy-saving federation method under fault diagnosis based on cross-granularity and prompt learning
CN118333137A
Multi-unmanned aerial vehicle multi-target tracking method and device, computer equipment and storage medium
CN118570494A
Cited By
Unmanned aerial vehicle cooperative intelligent visual detection method and system for mine key facilities
CN121599992A
Mine key facility intelligent visual detection method and system coordinated by unmanned aerial vehicle
CN121599992B