Unmanned aerial vehicle target identification method based on multi-modal information fusion
Through the drone target recognition method of multimodal information fusion, radar, optical and radio data are used to train and infer, the accuracy and robustness of small drone detection and recognition are solved, and efficient target recognition and feature completion are achieved.
Patent Information
- Application Number
- CN202510653283.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-04
AI Technical Summary
The detection and identification of small and medium-sized drones in the prior art are difficult, especially due to their small size, slow speed and low flight altitude, which leads to the low recognition rate of the existing single-modal detection methods, large recognition delay, and insufficient fusion of multimodal information.
UAV target recognition method based on multimodal information fusion is adopted, and through the joint training, processing and inference of radar, optical and radio multimodal data, rich prior knowledge is used to perform small sample inference, achieving high-quality fusion and feature completion of multimodal data.
It improves the accuracy and robustness of drone target recognition, can effectively deal with data loss and noise, and enhances the adaptability and recognition efficiency of the model.
Smart Images

Figure CN120257058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology applications, and specifically relates to a method for identifying unmanned aerial vehicle (UAV) targets based on multimodal information fusion. Background Art
[0002] With the development of the economy, there are more and more small UAVs, which has been increasing the pressure on current air traffic management work year by year. However, small UAVs often have a small volume, slow speed, and low flight altitude, so it is difficult to detect and identify such aircraft. Improving the detection ability of small targets in the low-altitude airspace has become an urgent problem to be solved in the field of air traffic management in recent years. Currently, most of the methods for identifying UAV targets use single means such as radar, optics, and radio for detection, with a low recognition rate and large recognition delay; a small number use multiple means for collaborative detection, but most research and applications are limited to the simple integration of information and do not achieve essential information fusion. Summary of the Invention
[0003] The purpose of the present invention is to propose a method for identifying UAV targets based on multimodal information fusion for the application in the above background art. Based on a multimodal large model, multi-modal UAV detection data such as radar, optics, and radio are input simultaneously for joint training, processing, and reasoning, and rich prior knowledge is used to achieve few-shot reasoning for downstream tasks, and the identification results of small UAV targets are returned.
[0004] The technical solution adopted by the present invention is as follows: A method for identifying UAV targets based on multimodal information fusion, comprising the following steps: Step 1, multimodal pre-training: Collect multi-modal UAV detection data, including radar detection data, UAV optical images, and radio monitoring signals, construct a pre-training task, and pre-train a multimodal large model from multiple dimensions; Step 2, model adaptation and fine-tuning: On the basis of pre-training, adopt instruction fine-tuning technology, and the multimodal large model adjusts its behavior by understanding and executing explicit task instructions, so that the multimodal large model learns fine-grained features and representations for specific tasks; Step 3, input multimodal data: For a specific UAV target identification task, input radar, optical, and radio multi-modal UAV detection data, and use the respective independent encoders in the multimodal large model to encode them respectively to obtain corresponding feature vectors; Step 4, multimodal feature fusion: Map the feature vectors of multiple single modalities to a common multimodal semantic space through a linear mapping layer, and perform alignment, splicing, and fusion to obtain the features jointly represented by each input modality; Step 5, output the target recognition result: Through the feature enhancement network, extract and enhance the fused features to obtain a global representation of the UAV features, and perform unified inference and prediction, and finally obtain the UAV classification and target recognition result.
[0005] Further, Step 1 specifically includes the following steps: Step 1.1, collect multi-modal UAV detection data, including radar detection data, UAV optical images, and radio monitoring signals; among them, the radar detection data is from the echo signal of the radar detecting the UAV, the UAV optical image comes from the real-time video or image during the UAV flight, and the radio monitoring signal is obtained by using radio frequency scanning technology to perform real-time monitoring, analysis, and direction finding on the civilian UAV remote control and video transmission signal frequency bands; Step 1.2, use the redundancy discovery method to discover data repetitions at different granularities and eliminate them; and use the classifier-based method to judge the data quality, identify and filter out the data that does not meet the quality requirements; Step 1.3, perform decoding, scaling, cropping, and normalization processing on non-text data objects, and convert them into vector data through feature extraction; Step 1.4, input the processed radar detection data, UAV optical images, and radio monitoring signals into the corresponding single-modal encoders in the multi-modal large model respectively to obtain their respective feature vectors, and then perform interaction and fusion on the feature vectors of each modality based on the class Transformer to achieve semantic space alignment and information fusion in the multi-modal, so that the multi-modal large model can simultaneously learn the features and deep-level correlation relationships of different modalities.
[0006] The present invention has the following advantages compared with the prior art: 1. Compared with the single-modal algorithm, multi-modal data such as radar, optics, and radio are used simultaneously for joint processing, training, and inference. Different modality information complements and promotes each other, describing the UAV features comprehensively from different dimensions, and improving the accuracy and usability of algorithm inference in the field of UAV target recognition.
[0007] 2. When using the multi-modal model for inference, it can effectively solve the problem of multi-modal data loss. Even if a certain modality of data is missing, it can be compensated by other modality data. At the same time, it also makes the multi-modal data integrate with high quality, better cope with data noise, and better mine the features of the target, making the model more robust.
[0008] 3. The algorithm selected in the present invention can achieve the integration of multi-modal information, with relatively high efficiency, high accuracy, strong adaptability, and strong technical feasibility through optimization processing. Description of the Drawings
[0009] Figure 1 It is a schematic flow chart of the UAV target recognition method based on multi-modal information fusion of the present invention. Specific embodiments
[0010] Next, in combination with Figure 1 and specific embodiments, the present invention will be further described.
[0011] Based on a multi-modal large model, the present invention combines multi-modal UAV detection data such as radar, optics, and radio. After data collection, cleaning, and preprocessing, unsupervised pre-training is carried out, and adaptation fine-tuning is performed according to the requirements of specific application fields. Then, for a specific UAV recognition task, the corresponding multi-modal data is input into the multi-modal large model, and through multi-modal feature fusion, the UAV is classified and recognized, and the UAV target recognition result is output. As Figure 1 shown, it mainly includes the following steps: Step 1, multi-modal pre-training: Collect a large amount of multi-modal UAV detection data, including radar detection data, UAV optical images, and radio monitoring signals, construct a pre-training task, and pre-train the multi-modal large model from multiple dimensions to enable it to have rich prior knowledge and strong reasoning ability. After training, the multi-modal large model has excellent few-shot reasoning ability.
[0012] The specific steps are as follows: Step 1.1, data collection: First, a variety of data sets need to be collected. By training a large number of samples, the multi-modal large model can continuously learn to extract target features, locate and classify targets. When the data scale is large and the types are rich, the accuracy of the model in detecting and recognizing targets is higher, the robustness of the model obtained after training is better, and the model has a better recognition effect on targets in different scenarios. Therefore, establishing an effective data set to store target feature information is of great significance for improving the efficiency of target detection.
[0013] The data of the multi-modal large model comes from experimental data and dedicated data sets, and the forms include radar detection data, UAV optical images, radio monitoring signals, etc. By adding some known high-quality data sets, the diversity of the data set is improved.
[0014] The radar detection data mainly comes from the echo signals of the radar detecting UAVs. By detecting and analyzing it, UAV target features are extracted, such as parameters such as radar cross section (RCS) and its statistical parameters, flight trajectory features, and micro-motion features. Based on the above features, the UAVs are identified and classified.
[0015] The optical images of drones mainly come from the real-time videos or images captured during drone flights. Through image processing methods, a small drone image dataset is constructed. The small drone images should cover special situations such as different moving backgrounds, different scales and angles, random positions, and occlusions.
[0016] The radio monitoring signal acquisition method is as follows: Using radio frequency scanning technology, real-time monitoring, analysis, and direction finding are carried out for the civilian drone remote control and video transmission signal frequency bands to obtain the drone control signal waveform, and it is compared with the drone control waveforms in the "waveform database" to determine whether there is a drone and identify its type.
[0017] Step 1.2, Data cleaning: Due to factors such as the limited resolution of drone detection means, the long flight distance of drones, and the complex moving background, the accuracy, real-time performance, detection accuracy, etc. of the collected data are all greatly affected. The noise existing in the collection process reduces the data quality and increases the difficulty of analyzing and processing data. Therefore, in addition to collecting high-quality data, the collected data also needs to be cleaned to reduce the impact of other irrelevant factors on target recognition.
[0018] By using the redundant discovery method, data repetitions at different granularities are discovered and eliminated; by using the classifier-based method, the data quality is judged, and low-quality data is identified and filtered out. Through the above data cleaning method, high-quality data can be constructed, greatly reducing the data scale required for large model training.
[0019] Step 1.3, Data preprocessing: The training of multi-modal large models requires a large number of modality-aligned data, but there is often a modality imbalance problem. The solution is: for non-text data objects such as pictures, videos, and signals, first perform appropriate decoding, scaling, cropping, and normalization processing, and then through feature extraction, convert them into vector data to ensure that different modality data can be correctly aligned, increase data diversity, and improve data quality.
[0020] Step 1.4, Multi-modal pre-training: Based on a large amount of multi-modal aligned data, a multi-modal large model is trained. Multiple single-modal encoders are respectively constructed. The processed radar detection data, drone optical images, and radio monitoring signals are respectively input into the corresponding single-modal encoders to obtain their respective feature vectors. Then, based on a class Transformer, the features of each modality are interacted and fused to achieve semantic space alignment and information fusion in the multi-modal space, enabling the multi-modal large model to simultaneously learn the features of different modalities and learn this deep association relationship, better mining the features of the target, and describing the drone target in an all-round and multi-dimensional manner.
[0021] Through sufficient pre-training on data in the field of large-capacity drones, more parameters and more complex structures are used to accurately represent the data distribution and complex features in the drone field, enabling cross-modal understanding and generation of multi-modal large models, and improving the generalization ability of the model for drone target recognition.
[0022] Step 2, Model adaptation and fine-tuning: After pre-training, the multi-modal large model can obtain the general ability to handle various tasks, but may not have accumulated enough knowledge for specific tasks or fields. In this case, fine-tuning the model on a smaller, domain-specific dataset allows the large model to precisely master professional knowledge and enhance its performance in a specific field to meet the requirements of downstream tasks. Utilize its efficient fine-tuning and few-shot learning modes to quickly adapt to complex scenario requirements and provide a feasible solution for drone target recognition.
[0023] Based on pre-training, using instruction fine-tuning technology, the multi-modal large model adjusts its behavior by understanding and executing explicit task instructions, enabling the multi-modal large model to learn fine-grained features and representations for specific tasks, further enhancing its adaptability and flexibility to different tasks, and improving its performance on specific tasks.
[0024] The strategies of pre-training and fine-tuning reflect a complementarity: the former provides the model with broad language understanding ability through self-supervised learning, while the latter ensures that the model, based on the former, performs supervised fine-tuning on a small amount of labeled data to achieve optimization for specific tasks. This complementary strategy greatly improves the generalization ability of the model in various types of processing tasks.
[0025] Step 3, Input multi-modal data: Based on the rich prior knowledge and powerful capabilities of the multi-modal large model, downstream tasks can be applied in a few-shot manner. For a specific drone target recognition task, input multi-modal drone detection data from radar, optics, and radio, and use the respective independent encoders in the multi-modal large model to encode them to obtain corresponding feature vectors.
[0026] To make more full use of the knowledge learned by the multi-modal large model in the pre-training stage, keeping the model framework, input form used in the inference and prediction stage consistent with those used in the pre-training stage will greatly enhance the ability of the multi-modal large model to mine and obtain prior knowledge in a specific field and achieve good application effects in downstream tasks.
[0027] Step 4, Multi-modal feature fusion: Multiple single-modal feature vectors are mapped to a common multi-modal semantic space through a linear mapping layer, and then aligned, concatenated, and fused to make full use of the correlation relationships between multi-modal data, so that the information carried by different modal data is fully supplemented and verified, and the overall features of the joint representation of each input modality are obtained for the next step of prediction or reasoning.
[0028] During the feature fusion process, the prior knowledge of the multi-modal large model is fully utilized to effectively solve the problem of the differences in the features of each modality and improve the overall recognition accuracy of the model. The cross-modal attention mechanism is used to enhance the attention to certain important features to further improve the prediction performance of the model.
[0029] Step 5, output the target recognition result: Through the feature enhancement network, the fused features are extracted and enhanced to obtain a global representation of the UAV features, and unified inference and prediction are performed to finally obtain the UAV classification and target recognition results.
Claims
1. A method for identifying drone targets based on multi-modal information fusion, characterized in that, Including the following steps: Step 1, Multimodal pre-training: Collect multimodal UAV detection data, including radar detection data, UAV optical images, and radio monitoring signals, construct pre-training tasks, and pre-train a multimodal large model from multiple dimensions; Step 2, Model adaptation and fine-tuning: On the basis of pre-training, adopt instruction fine-tuning technology. The multimodal large model adjusts its behavior by understanding and executing explicit task instructions, enabling the multimodal large model to learn fine-grained features and representations for specific tasks; Step 3, Input multimodal data: For a specific UAV target recognition task, input radar, optical, and radio multimodal UAV detection data, and encode them respectively using the respective independent encoders in the multimodal large model to obtain corresponding feature vectors; Step 4, Multimodal feature fusion: Map multiple single-modal feature vectors to a common multimodal semantic space through a linear mapping layer, and perform alignment, splicing, and fusion to obtain the features of the joint representation of each input modality; Step 5, Output target recognition results: Through a feature enhancement network, extract and enhance the fused features to obtain a global representation of UAV features, and perform unified inference and prediction to finally obtain UAV classification and target recognition results.
2. The method for identifying an unmanned aerial vehicle target based on multi-modal information fusion according to claim 1, wherein Step 1 specifically includes the following steps: Step 1.1, Collect multimodal UAV detection data, including radar detection data, UAV optical images, and radio monitoring signals; among them, the radar detection data comes from the echo signal of the radar detecting the UAV, the UAV optical image comes from the actual shooting video or image during UAV flight, and the radio monitoring signal is obtained by using radio frequency scanning technology to perform real-time monitoring, analysis, and direction finding on the civilian UAV remote control and video transmission signal frequency bands; Step 1.2, Adopt a redundancy discovery method to discover data repetitions at different granularities and eliminate them; and adopt a classifier-based method to judge data quality, identify and filter out data that does not meet the quality requirements; Step 1.3, Perform decoding, scaling, cropping, and normalization processing on non-text data objects, and convert them into vector data through feature extraction; Step 1.4, Input the processed radar detection data, UAV optical images, and radio monitoring signals into the corresponding single-modal encoders in the multimodal large model respectively to obtain their respective feature vectors, and then perform interaction and fusion on the feature vectors of each modality based on a Transformer-like architecture to achieve alignment and information fusion in the multimodal semantic space, enabling the multimodal large model to simultaneously learn the features and deep-level correlation relationships of different modalities.
Citation Information
Cited By
Audio-visual fusion unmanned aerial vehicle detection and positioning method
CN121165207A
Unmanned aerial vehicle identification method and system based on linkage of radar and wireless signal
CN122310240A
Unmanned aerial vehicle identification method and system based on radar and wireless signal linkage
CN122310240B