Unmanned aerial vehicle detection method based on visual language model

Through the drone detection method based on visual language model, text features in drone pictures are automatically extracted and marked, and text and images are jointly modeled using deep learning algorithms and cosine similarity, which solves the inefficiency problem caused by manual annotation and achieves efficient and accurate drone target detection.

CN119963947APending Publication Date: 2025-05-09HANGZHOU QITAI ELECTRONIC TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510034941.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, unlabeled drone pictures need to be manually marked, resulting in a large workload and affecting the drone detection efficiency.

Method used

The drone detection method based on visual language model is adopted to automatically extract text features in the sample, automatically label the sample, and use deep learning algorithms and cosine similarity to realize joint modeling of text and images to automatically identify and detect drone targets.

Benefits of technology

It improves the efficiency and detection efficiency of pre-data processing and detection of drone detection, ensures the efficiency and accuracy of drone target detection, and reduces the need for manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963947A_ABST
    Figure CN119963947A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle detection method based on a visual language model, and the method specifically comprises the steps: obtaining a plurality of text image pair samples, and constructing corresponding label vectors; respectively extracting initial text features and initial image features of the text image pair sample, and constructing an initial feature sample set; based on a deep learning algorithm, constructing a text description sub-model and an image understanding sub-model, and associating the text description sub-model and the image understanding sub-model through cosine similarity to obtain an unmanned aerial vehicle detection model; training an unmanned aerial vehicle detection model based on the initial feature sample set; and establishing unified text description, and outputting an unmanned aerial vehicle detection result of the to-be-detected image in combination with the unmanned aerial vehicle detection model. According to the method, combined modeling of the text and the image is carried out by combining computer vision and natural language processing, the association relationship between text description and image features is trained and learned, and efficient and accurate detection of the unmanned aerial vehicle is realized on the basis of expanding training sample diversity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drone detection, and in particular to a drone detection method based on a visual language model. Background Art

[0002] With the rapid development of drone technology, drones are increasingly used in military reconnaissance, civil monitoring, traffic intersection monitoring, large-scale assembly inspections, fault detection, and information collection in disaster-stricken areas with complex terrain. Currently, drone target detection mostly relies on image processing and pattern recognition technology, and performs computer vision tasks such as target classification and detection through deep learning algorithms such as neural networks to achieve drone detection.

[0003] The process of training deep learning algorithm models such as neural networks is a complex and delicate task, involving multiple links, including data preparation, model definition, training and optimization. Among them, data preparation is a critical and time-consuming link, in which a large amount of data needs to be classified, labeled or feature extracted. The data processing speed and data accuracy directly affect the subsequent drone detection efficiency and detection accuracy.

[0004] For drone detection, the currently publicly available annotated drone image datasets include M3D-Real, GLMD, DronevsBird, Det-Fly, Anti-UAV-RGB, etc. The total number of images is only about 50,000, which is far from the amount of image information on the Internet. If the above annotated drone image datasets are used to perform subsequent model training, the subsequent drone detection accuracy cannot be guaranteed. In order to improve the accuracy of drone detection, it is necessary to add unlabeled images to the dataset for model training to increase the diversity of the training dataset. However, before adding unlabeled images, the collected unlabeled images need to be manually marked before subsequent model training can be performed. The manual marking method is too labor-intensive and will affect the efficiency of drone detection. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art that when expanding the training data set by unannotated pictures and improving the detection accuracy of drone detection by using a deep learning algorithm, the annotation work of unannotated pictures needs to be manually processed, the workload is large, and the drone detection efficiency is low. A drone detection method based on a visual language model is provided. For the acquired model training samples, even if there are unannotated pictures, text features in the samples can be automatically extracted to achieve automatic labeling of the samples. At the same time, deep learning algorithms and cosine similarity are used to achieve joint modeling of text and images. The correlation between text descriptions and image features is learned through model training sample training, thereby achieving automatic recognition and detection of drone targets, and ensuring the efficiency and accuracy of drone target detection.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] The drone detection method based on visual language model includes:

[0008] Get several text-image pair samples and construct corresponding label vectors;

[0009] Extract initial text features and initial image features of text-image pair samples respectively, and construct an initial feature sample set;

[0010] Based on the deep learning algorithm, a text description sub-model and an image understanding sub-model are constructed, and the text description sub-model and the image understanding sub-model are associated through cosine similarity to obtain a drone detection model.

[0011] Train the drone detection model based on the initial feature sample set;

[0012] Establish a unified text description, combine it with the drone detection model, and output the drone detection results of the image to be detected.

[0013] For the training samples of the drone detection model, a large number of text-image pair samples are obtained to ensure the diversity of its training samples, thereby improving the detection accuracy of subsequent models. And for the obtained text-image pair samples, the text description content in the sample can be automatically understood without manual annotation, which effectively improves the efficiency of pre-data processing for drone detection, thereby improving the detection efficiency of subsequent drone detection. And the deep learning algorithm and cosine similarity are used to realize the joint modeling of text and image, combining the technical advantages of computer vision and natural language processing. For the drone detection model combined with visual language, the relationship between text description and image features is learned through model training sample training. Without limiting the annotation of training samples and expanding the diversity of training samples, automatic recognition and annotation can also be performed to improve the training efficiency of the drone detection model, thereby realizing the automatic recognition and detection of drone targets, and ensuring the efficiency and accuracy of drone target detection.

[0014] Furthermore, the training of the drone detection model based on the initial feature sample set includes:

[0015] The initial text features and the initial image features are processed by the text description sub-model and the image understanding sub-model respectively to obtain text features and image features of the same length;

[0016] Combine the text features and image features after feature processing in pairs, calculate the corresponding cosine similarities respectively, and construct a cosine similarity matrix;

[0017] The cross entropy loss is calculated based on the cosine similarity matrix and the label vector. The parameters of the drone detection model are trained and optimized with the goal of minimizing the cross entropy loss.

[0018] Furthermore, the calculation formula of the cosine similarity is:

[0019]

[0020] Among them: Similarity cosine is the cosine similarity, A i is the corresponding component of the i-th text image pair sample in the text feature, B i is the corresponding component of the i-th text-image pair sample in the image feature, and n is the number of text-image pair samples.

[0021] Furthermore, the feature processing of the initial text features and the initial image features by the text description sub-model and the image understanding sub-model respectively includes:

[0022] Perform linear mapping on the input initial text features or initial image features to obtain the corresponding output vector;

[0023] Apply activation function to the output vector and input it into the fully connected layer;

[0024] After applying the activation function to the input, the fully connected layer of the output vector is randomly deactivated to obtain the deactivated feature vector;

[0025] The output vector and the deactivated feature vector are combined for normalization, and the corresponding text features or image features are output.

[0026] The extracted features are subjected to linear mapping, activation function processing, fully connected layer calculation, random inactivation, and normalization operations to ensure that the extracted features have the same length and similar feature distribution, thereby improving the stability and accuracy of the subsequent drone detection model.

[0027] Furthermore, the expression of the normalization process is:

[0028]

[0029] Where: X is the normalized text feature or image feature, Y is the output vector obtained after linear mapping, DropOut is the feature vector after random inactivation, E() is the mean function, Var() is the variance function, and ∈ is a preset constant.

[0030] Furthermore, the cross entropy loss is calculated based on the cosine similarity matrix and the label vector, including:

[0031] Extract the diagonal elements of the cosine similarity matrix;

[0032] Calculate the cross entropy of the obtained diagonal elements and each label component in the label vector respectively;

[0033] The cross entropy calculation results are averaged to obtain the cross entropy loss between the cosine similarity matrix and the label vector.

[0034] Furthermore, the calculation formula of the cross entropy is:

[0035]

[0036] Where p is the cosine similarity matrix, q is the label vector, H(p,q) is the cross entropy between the cosine similarity matrix and the cross entropy, p i is the diagonal element of the cosine similarity matrix of the combination of text features and image features corresponding to the i-th text-image pair sample, q ij is the j-th label component of the label vector corresponding to the i-th text-image pair sample, and n is the number of text-image pair samples.

[0037] Furthermore, the establishment of a unified text description, combined with a drone detection model, and output of a drone detection result of the image to be detected includes:

[0038] Extract the initial text features corresponding to the two text descriptions in the unified text description and the initial image features of the image to be detected, and input them into the drone detection model;

[0039] Perform feature processing on the corresponding initial text features and initial image features respectively through the text description sub-model and the image understanding sub-model to obtain a first text feature, a second text feature and a first image feature;

[0040] respectively calculating the cosine similarity between the first text feature and the first image feature and the cosine similarity between the second text feature and the first image feature;

[0041] The text description corresponding to the text feature with a larger output cosine similarity is the drone detection result of the image to be detected.

[0042] By calculating the cosine similarity between the image to be detected and the known text description features, accurate detection and recognition of drone targets can be achieved. This not only improves the accuracy and efficiency of detection, but also has strong robustness and adaptability. Even related pictures in complex scenes can achieve stable detection of drone targets.

[0043] Furthermore, the extracting of initial text features and initial image features of the text-image pair samples respectively to construct an initial feature sample set includes:

[0044] Based on the text feature encoder, feature extraction is performed on the text content in the text image sample to obtain initial text features;

[0045] Based on the image feature encoder, feature extraction is performed on the image content in the text image pair sample to obtain initial image features;

[0046] An initial feature sample set is constructed based on initial text features and initial image features.

[0047] Furthermore, the step of obtaining a plurality of text-image pair samples and constructing corresponding label vectors includes:

[0048] Acquire, by crawling or generating, a number of text-image pair samples including text descriptions and image contents containing drones and text descriptions and image contents not containing drones;

[0049] A corresponding label vector is constructed for each text-image pair sample. The label vector includes a label component that contains a drone and a label component that does not contain a drone, and each label component is assigned a value.

[0050] The beneficial effects of the present invention are:

[0051] (1) For the training samples of the drone detection model, a large number of text-image pair samples are obtained to ensure the diversity of its training samples, thereby improving the detection accuracy of the subsequent model. Moreover, for the obtained text-image pair samples, the text description content in the sample can be automatically understood without manual annotation, which effectively improves the efficiency of the pre-data processing of drone detection, thereby improving the detection efficiency of subsequent drone detection. In addition, the deep learning algorithm and cosine similarity are used to realize the joint modeling of text and image, combining the technical advantages of computer vision and natural language processing. For the drone detection model combined with visual language, the relationship between text description and image features is learned through model training sample training. Without limiting the annotation of training samples and expanding the diversity of training samples, automatic recognition and annotation can also be performed to improve the training efficiency of the drone detection model, thereby realizing the automatic recognition and detection of drone targets, and ensuring the efficiency and accuracy of drone target detection.

[0052] (2) The extracted features are subjected to linear mapping, activation function processing, fully connected layer calculation, random inactivation, and normalization operations to ensure that the extracted features have the same length and similar feature distribution, thereby improving the stability and accuracy of the subsequent drone detection model.

[0053] (3) Accurate detection and recognition of drone targets are achieved by calculating the cosine similarity between the image to be detected and the known text description features. This not only improves the accuracy and efficiency of detection, but also has strong robustness and adaptability. Even for related images in complex scenes, stable detection of drone targets can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of a process of the present invention;

[0055] Figure 2 It is an inference logic block diagram of drone detection for an image to be detected according to an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0057] Example:

[0058] Drone detection methods based on visual language models, such as Figure 1 As shown, including:

[0059] Get several text-image pair samples and construct corresponding label vectors;

[0060] Extract initial text features and initial image features of text-image pair samples respectively, and construct an initial feature sample set;

[0061] Based on the deep learning algorithm, a text description sub-model and an image understanding sub-model are constructed, and the text description sub-model and the image understanding sub-model are associated through cosine similarity to obtain a drone detection model.

[0062] Train the drone detection model based on the initial feature sample set;

[0063] Establish a unified text description, combine it with the drone detection model, and output the drone detection results of the image to be detected.

[0064] Taking into account the small amount of information in the existing annotated drone image datasets that can be applied to drone detection, it is difficult to ensure the detection accuracy of the trained model by training the drone detection model through such drone image datasets. Therefore, when selecting training samples for the drone detection model, this embodiment is not limited to annotated drone images, but simply collects images with relevant text descriptions to obtain corresponding text-image pair samples.

[0065] And considering that the collected samples do not necessarily have annotated content, it is not certain whether there are drones in the image content. Therefore, the label vector is set with label components that include drones and label components that do not include drones for training and learning of the drone detection model.

[0066] Specifically, obtain several text-image pair samples and construct corresponding label vectors, including:

[0067] Acquire, by crawling or generating, a number of text-image pair samples including text descriptions and image contents containing drones and text descriptions and image contents not containing drones;

[0068] A corresponding label vector is constructed for each text-image pair sample. The label vector includes a label component that contains a drone and a label component that does not contain a drone, and each label component is assigned a value.

[0069] Among them, the number of text image pair samples obtained in this embodiment is at least 10 million.

[0070] If you want to obtain relevant data through crawlers, you can choose public websites that allow crawling, such as academic databases, public image libraries, etc., and extract relevant content through relevant library parsing methods to obtain the corresponding text description content and image data.

[0071] If relevant data is obtained through generation, different text descriptions can be generated by creating templates containing drones and excluding drones by replacing keywords. Natural language generation models such as GPT-3 can also be used to generate relevant text descriptions based on prompt words. For the corresponding image data, corresponding images containing drones and excluding drones can be generated through algorithms such as generative adversarial networks.

[0072] A rich training sample library is built by collecting a large number of text-image pair samples. The text descriptions in these sample libraries contain drone-related keywords and descriptions, while the images correspond to scenes with and without drones. The training samples are then used to train the deep learning model for drone detection. The model can learn the relationship between text descriptions and image features, thereby achieving automatic recognition and detection of drone targets.

[0073] For each acquired text-image pair sample, a corresponding label vector Label is constructed. The label vector Label includes a component Label containing the drone. i1 and the component Label that does not include the drone i2 , and assign values ​​to the two components respectively, where Label i1 =1, Label i2 =0.

[0074] In order to realize automatic annotation of the acquired text image pair samples, the text description content of the text image pair samples is further identified, and the corresponding initial text features are obtained as annotation content, which together with the initial image features obtained according to the image content are used as training samples for the drone detection model.

[0075] Specifically, the initial text features and initial image features of the text-image pair samples are extracted respectively, and an initial feature sample set is constructed, including:

[0076] Based on the text feature encoder, feature extraction is performed on the text content in the text image sample to obtain initial text features;

[0077] Based on the image feature encoder, feature extraction is performed on the image content in the text image pair sample to obtain initial image features;

[0078] An initial feature sample set is constructed based on initial text features and initial image features.

[0079] In this embodiment, the Transformer encoder is used to extract features of the text content in the text image sample. Specifically, a pre-trained Transformer encoder (such as BERT) can be used to segment and embed the text content, and then input it into the Transformer encoder. Multi-layer processing is performed through a self-attention mechanism and a feedforward neural network, and finally the corresponding embedding vector is extracted from the output of the Transformer as the initial text feature of the text content.

[0080] At the same time, the Vision Transformer encoder is used to extract features from the image content in the text image sample to obtain the initial image features, and then the initial feature sample set is constructed in combination with the obtained initial text features.

[0081] The drone detection model is then trained based on the initial feature sample set, including:

[0082] The initial text features and the initial image features are processed by the text description sub-model and the image understanding sub-model respectively to obtain text features and image features of the same length;

[0083] Combine the text features and image features after feature processing in pairs, calculate the corresponding cosine similarities respectively, and construct a cosine similarity matrix;

[0084] The cross entropy loss is calculated based on the cosine similarity matrix and the label vector. The parameters of the drone detection model are trained and optimized with the goal of minimizing the cross entropy loss.

[0085] Among them, the text description sub-model and the image understanding sub-model both rely on neural networks, which can perform feature processing on the features in the initial feature sample set so that the corresponding feature vectors have the same length.

[0086] Specifically, feature processing is performed on the initial text features and the initial image features respectively by using the text description sub-model and the image understanding sub-model, including the following steps:

[0087] Perform linear mapping on the input initial text features or initial image features to obtain the corresponding output vector.

[0088] Among them, the expression of linear mapping is:

[0089] T=x·W 1 +b 1 ;

[0090] Among them, T is the output vector obtained after linear mapping, x is the initial text feature or initial image feature of the input, and W 1 is the mapping parameter that needs to be trained, b 1 is one of the preset vector offsets, and Y is the output vector obtained after linear mapping.

[0091] After applying the activation function to the output vector Y, it is input into the fully connected layer.

[0092] The activation function applied to the output vector Y is defined as:

[0093]

[0094] The error function erf applied in the activation function is defined as:

[0095]

[0096] Among them, Z is the output vector after applying the activation function, a is the variable of the error function erf, and t is the integral variable.

[0097] The output vector Z after applying the activation function is input into the fully connected layer. The function of the fully connected layer is defined as:

[0098] FC=Z·W 2 +b 2 ;

[0099] Among them, FC is the output vector of the fully connected layer, W 2 is the network parameter that needs to be trained, b 2 is one of the pre-set vector offsets.

[0100] Then, the fully connected layer of the output vector after applying the activation function to the input is randomly deactivated, that is, the values ​​of a part of the output vector FC are randomly set to 0. In this embodiment, the proportion of the output vector FC whose values ​​are set to 0 is set to 0.2, and then the deactivated feature vector DropOut is obtained.

[0101] Finally, the output vector Y and the inactivated feature vector DropOut are combined for normalization, and the corresponding text features or image features are output.

[0102] The expression of the normalization process is:

[0103]

[0104] Wherein: X is the text feature or image feature after normalization, Y is the output vector obtained after linear mapping, DropOut is the feature vector after random dropout, E() is the mean function, Var() is the variance function, ∈ is a preset constant, which is a very small number used to avoid the variance result being 0. In this embodiment, ∈ is set to 0.0000001.

[0105] After feature processing is completed, text features and image features of equal length can be obtained. Combined with the correspondence between the text description and the image content in the corresponding text-image sample pairs, the text features and image features are combined in pairs to calculate the corresponding cosine similarity.

[0106] The calculation formula of the cosine similarity is:

[0107]

[0108] Among them: Similarity cosine is the cosine similarity, A i is the corresponding component of the i-th text image pair sample in the text feature, B i is the corresponding component of the i-th text-image pair sample in the image feature, and n is the number of text-image pair samples.

[0109] The corresponding cosine similarity calculation results are expressed in the form of a cosine similarity matrix, and then the diagonal of the cosine similarity matrix is ​​taken out and combined with the label vector to calculate the cross entropy loss.

[0110] Specifically, the cross entropy loss is calculated based on the cosine similarity matrix and the label vector, including:

[0111] Extract the diagonal elements of the cosine similarity matrix;

[0112] Calculate the cross entropy of the obtained diagonal elements and each label component in the label vector respectively;

[0113] The cross entropy calculation results are averaged to obtain the cross entropy loss between the cosine similarity matrix and the label vector.

[0114] After obtaining the diagonal elements of the cosine similarity matrix, the cross entropy with each label component is calculated on the label vector, that is, the diagonal elements in the cosine similarity matrix are respectively compared with the component Label containing the drone. i1 and the component Label that does not include the drone i2 Calculate the corresponding cross entropy.

[0115] The calculation formula of the cross entropy is:

[0116]

[0117] Where p is the cosine similarity matrix, q is the label vector, H(p,q) is the cross entropy between the cosine similarity matrix and the cross entropy, p i is the diagonal element of the cosine similarity matrix of the combination of text features and image features corresponding to the i-th text-image pair sample, q ij is the j-th label component of the label vector corresponding to the i-th text-image pair sample, and n is the number of text-image pair samples.

[0118] Through the above cross entropy calculation formula, we can get the cosine similarity matrix and the component Label containing the drone i1 The cross entropy H 1 , and the component Label that does not contain the drone i2 The cross entropy H 0 .

[0119] By averaging the two, we can get the final cross entropy loss H:

[0120]

[0121] Based on the above feature processing flow, combined with the given initial feature sample set, and with the goal of minimizing the cross entropy loss, the mapping parameters and network parameters that need to be trained in the text description sub-model and the image understanding sub-model are optimized to complete the model training of the final drone detection model.

[0122] Considering that the drone detection results only include two conclusions: the presence of a drone and the absence of a drone, a unified text description is first constructed before drone detection, including a text description for the label where a drone exists, which is described as "there is a drone in the picture", and a text description for the label where a drone does not exist, which is described as "there is no drone in the picture".

[0123] On this basis, combined with the drone detection model, the drone detection results of the image to be detected are output, which specifically includes the following steps:

[0124] Extracting initial text features corresponding to the two text descriptions in the unified text description through a text feature encoder;

[0125] For any image to be detected to see whether a drone exists, the initial image features of the image to be detected are extracted through an image feature encoder;

[0126] The two initial text features are input into the drone detection model together with the initial image features of the image to be detected;

[0127] The text description sub-model and the image understanding sub-model perform feature processing on the two initial text features and the initial image features of the image to be detected, respectively, to obtain the first text feature A T , Second text feature A F and the first image feature B D ;

[0128] Calculate the first text feature A respectively T and the first image feature B D The cosine similarity between and the second text feature A F and the first image feature B D The cosine similarity between ;

[0129] The text description corresponding to the text feature with a larger output cosine similarity is the drone detection result of the image to be detected, that is, the detection result of whether a drone exists in the image to be detected is obtained.

[0130] Specifically, the reasoning logic for drone detection of the image to be detected in this embodiment is as follows: Figure 2 shown.

[0131] The above-described embodiment is only a preferred solution of the present invention and does not limit the present invention in any form. There are other variations and modifications without exceeding the technical solution described in the claims.

Claims

1. A drone detection method based on a visual language model, characterized in that: include: Get several text-image pair samples and construct corresponding label vectors; Extract initial text features and initial image features of text-image pair samples respectively, and construct an initial feature sample set; Based on the deep learning algorithm, a text description sub-model and an image understanding sub-model are constructed, and the text description sub-model and the image understanding sub-model are associated through cosine similarity to obtain a drone detection model. Train the drone detection model based on the initial feature sample set; Establish a unified text description, combine it with the drone detection model, and output the drone detection results of the image to be detected.

2. The drone detection method based on visual language model according to claim 1 is characterized in that: The training of the drone detection model based on the initial feature sample set includes: The initial text features and the initial image features are processed by the text description sub-model and the image understanding sub-model respectively to obtain text features and image features of the same length; Combine the text features and image features after feature processing in pairs, calculate the corresponding cosine similarities respectively, and construct a cosine similarity matrix; The cross entropy loss is calculated based on the cosine similarity matrix and the label vector. The parameters of the drone detection model are trained and optimized with the goal of minimizing the cross entropy loss.

3. The method for detecting drones based on a visual language model according to claim 2, characterized in that: The calculation formula of the cosine similarity is: Among them: Similarity cosine is the cosine similarity, A i is the corresponding component of the i-th text image pair sample in the text feature, B i is the corresponding component of the i-th text-image pair sample in the image feature, and n is the number of text-image pair samples.

4. The method for detecting drones based on a visual language model according to claim 2, characterized in that: The feature processing of the initial text features and the initial image features by the text description sub-model and the image understanding sub-model respectively includes: Perform linear mapping on the input initial text features or initial image features to obtain the corresponding output vector; Apply activation function to the output vector and input it into the fully connected layer; After applying the activation function to the input, the fully connected layer of the output vector is randomly deactivated to obtain the deactivated feature vector; The output vector and the deactivated feature vector are combined for normalization, and the corresponding text features or image features are output.

5. The method for detecting drones based on a visual language model according to claim 4, characterized in that: The expression of the normalization process is: Where: X is the normalized text feature or image feature, Y is the output vector obtained after linear mapping, DropOut is the feature vector after random inactivation, E() is the mean function, Var() is the variance function, and ∈ is a preset constant.

6. The method for detecting drones based on a visual language model according to claim 2, characterized in that: The cross entropy loss is calculated based on the cosine similarity matrix and the label vector, including: Extract the diagonal elements of the cosine similarity matrix; Calculate the cross entropy of the obtained diagonal elements and each label component in the label vector respectively; The cross entropy calculation results are averaged to obtain the cross entropy loss between the cosine similarity matrix and the label vector.

7. The method for detecting drones based on a visual language model according to claim 6, characterized in that: The calculation formula of the cross entropy is: Where p is the cosine similarity matrix, q is the label vector, H(p,q) is the cross entropy between the cosine similarity matrix and the cross entropy, p i is the diagonal element of the cosine similarity matrix of the combination of text features and image features corresponding to the i-th text-image pair sample, q ij is the j-th label component of the label vector corresponding to the i-th text-image pair sample, and n is the number of text-image pair samples.

8. The method for detecting drones based on a visual language model according to claim 2, characterized in that: The unified text description is established, combined with the drone detection model, and the drone detection result of the image to be detected is output, including: Extract the initial text features corresponding to the two text descriptions in the unified text description and the initial image features of the image to be detected, and input them into the drone detection model; The corresponding initial text features and initial image features are processed by the text description sub-model and the image understanding sub-model respectively to obtain the first text feature, the second text feature and the first image feature; the cosine similarity between the first text feature and the first image feature and the cosine similarity between the second text feature and the first image feature are calculated respectively; The text description corresponding to the text feature with a larger output cosine similarity is the drone detection result of the image to be detected.

9. The method for detecting drones based on a visual language model according to claim 1, characterized in that: The step of respectively extracting initial text features and initial image features of the text-image pair samples and constructing an initial feature sample set comprises: Based on the text feature encoder, feature extraction is performed on the text content in the text image sample to obtain initial text features; Based on the image feature encoder, feature extraction is performed on the image content in the text image pair sample to obtain initial image features; An initial feature sample set is constructed based on initial text features and initial image features.

10. The method for detecting drones based on a visual language model according to claim 1, characterized in that: The method of obtaining a plurality of text-image pair samples and constructing corresponding label vectors includes: Acquire, by crawling or generating, a number of text-image pair samples including text descriptions and image contents containing drones and text descriptions and image contents not containing drones; A corresponding label vector is constructed for each text-image pair sample. The label vector includes a label component that contains a drone and a label component that does not contain a drone, and each label component is assigned a value.