Unmanned aerial vehicle detection method based on infrared-visible light multi-mode fusion

By adopting a phased and progressive fusion network architecture, the problem of modal inconsistency between infrared and visible light images is solved, enabling efficient and accurate drone detection. It is suitable for edge devices and has good technology transferability and robustness.

CN121811286APending Publication Date: 2026-04-07CHONGQING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing drone detection technologies, the inconsistency between infrared and visible light image modes affects the information integration effect, resulting in low fusion efficiency. Furthermore, existing methods have high computational complexity, making them difficult to deploy on resource-constrained edge devices.

Method used

A UAV detection method based on infrared-visible multimodal fusion is designed. It adopts a phased progressive fusion network architecture, including a dual-branch feature extractor, a cross-modal feature alignment module, a refined feature fusion module, and a multi-scale feature fusion module. Feature alignment and fusion are performed through channel attention mechanism and gating mechanism to reduce the number of parameters and achieve lightweight design.

Benefits of technology

It achieves efficient and accurate drone detection, significantly reduces the number of parameters, and its lightweight design allows the model to be deployed to edge devices, meeting real-time detection requirements and exhibiting good technology transferability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811286A_ABST
    Figure CN121811286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-source information fusion, in particular to an infrared-visible light multi-modal fusion-based unmanned aerial vehicle detection method, which comprises the following steps of: acquiring infrared image and visible light image data, and constructing an accurately aligned infrared-visible light unmanned aerial vehicle data set; constructing a model based on a staged progressive fusion network architecture, and training the model by adopting a data set; the staged progressive fusion network architecture comprises a double-branch feature extractor, a cross-modal feature alignment module, a refined feature fusion module, a multi-scale feature fusion module and a detection output module. And processing the data by adopting the model to obtain an unmanned aerial vehicle detection result. The unmanned aerial vehicle detection method is based on the design of a staged progressive fusion network architecture, the cross-modal feature fusion is divided into two stages of initial alignment and later refinement, and the unmanned aerial vehicle detection method has the characteristics of light weight, high reasoning speed and high discrimination accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-source information fusion, in particular to a UAV detection method based on infrared-visible light multi-modal fusion. BACKGROUND

[0002] The development of UAV technology has brought new security risks, especially when the UAV enters the no-fly zone, engages in illegal activities or threatens public safety, how to timely and accurately detect and respond to these unauthorized UAVs has become a security problem that needs to be solved. At present, anti-UAV technology mainly relies on various sensor technologies such as radar, acoustics, radio frequency (RF) and visual sensors, and the main problem of existing UAV detection technology is concentrated on modal alignment.

[0003] In the prior art, the infrared and visible light images of the multi-modal UAV detection dataset may have differences in shooting time, angle and sensor position, and such inconsistency will affect the effective fusion between modes, thereby reducing the effect of information integration. While the multi-modal fusion method either adopts pixel-level fusion, fuses the images first and then detects, which is easy to lose modal specific information, or adopts simple feature-level fusion, which fails to fully utilize the complementary characteristics between different modes, resulting in low fusion efficiency. In addition, many advanced fusion methods use complex attention mechanisms or stack multiple fusion modules, resulting in large parameter quantity and high computational complexity, which is difficult to deploy to resource-constrained edge devices.

[0004] Therefore, there is an urgent need for a UAV detection method with fast discrimination speed, high accuracy and easy implementation. SUMMARY

[0005] Therefore, the present application discloses a UAV detection method based on infrared-visible light multi-modal fusion to solve the problems of the prior art, comprising:

[0006] S1, acquiring infrared image and visible light image data, and constructing an accurately aligned infrared-visible light UAV dataset;

[0007] S2, constructing a model based on a phased progressive fusion network architecture, and training the model using the dataset; the phased progressive fusion network architecture comprises a dual-branch feature extractor, a cross-modal feature alignment module, a refined feature fusion module, a multi-scale feature fusion module and a detection output module.

[0008] The dual-branch feature extractor is used for feature extraction of the infrared image and the visible light image respectively.

[0009] The cross-modal feature alignment module is used for aligning the infrared image features and the visible light image features, and the data processing is divided into two stages of shallow layer and deep layer.

[0010] The shallow stage learns the importance of different feature channels respectively by adopting a channel attention mechanism, and weights the importance to other modalities to obtain modal complementary features;

[0011] The deep stage projects the modal complementary features to a hidden state space, utilizes a gating mechanism for feature transmission, and respectively outputs the aligned infrared image features and visible light image features;

[0012] The fine-grained feature fusion module performs fine-grained feature fusion on the aligned infrared image features and visible light image features based on a MemAttention mechanism;

[0013] The multi-scale feature fusion module is used for up-sampling and fusing multi-scale low-resolution features;

[0014] The detection output module is used for target classification and bounding box prediction to generate a detection result containing the category and position of the target;

[0015] S3, a model is used to process data to obtain a UAV detection result.

[0016] The beneficial effects of the present application include:

[0017] A phased progressive fusion network architecture is designed, based on the design concept of "efficient global modeling + accurate local optimization", the cross-modal feature fusion is divided into two stages of initial alignment and later fine-tuning, the cross-modal feature alignment module can achieve excellent performance without multiple stacking, the parameter amount is about 18.7M, which is significantly lower than the parameter amount of existing technical solutions such as CFT 44.76M and ICAFusion 20.15M; the lightweight design and high inference speed enable the model to be deployed to edge devices, meeting the real-time detection requirements;

[0018] The method designed in the present application performs excellently in false detection and missed detection control, has good technical transferability, and provides a design idea for those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 It is a schematic diagram of the phased progressive fusion network architecture in the embodiments of the present application processing data;

[0020] Figure 2 It is a schematic diagram of the model based on the phased progressive fusion network architecture in Embodiment 2 of the present application;

[0021] Figure 3 It is a schematic diagram of the cross-attention RGB-IR enhanced Mamba module in Embodiment 2 of the present application;

[0022] Figure 4A schematic diagram of a local two-dimensional selective scanning module in Embodiment 2 of the present application;

[0023] Figure 5 A schematic diagram of a cross-modal refinement fusion Transformer module in Embodiment 2 of the present application;

[0024] Figure 6 A schematic diagram of a YOLO V5 detection head in Embodiment 2 of the present application. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical scheme, characteristics and advantages of the present application more clear, in order to make the technical personnel in the art better understand the technical scheme of the present application, the present application will be further described in detail below in combination with the drawings and examples.

[0026] Embodiment 1:

[0027] The present embodiment includes a method for detecting unmanned aerial vehicles based on infrared-visible light multi-modal fusion, comprising:

[0028] S1, obtain infrared image and visible light image data, and construct an accurate alignment infrared-visible light unmanned aerial vehicle data set.

[0029] The infrared image and visible light image data, in the present embodiment, the target scene is data collected by an unmanned aerial vehicle equipped with infrared and visible light binocular cameras. The collection process covers a variety of environmental conditions, including different scenes such as urban buildings, suburban mountains, dense grasslands and open sky, as well as different times and weather conditions such as sunny, cloudy, foggy and night. The target unmanned aerial vehicle includes different sizes of models to ensure the diversity of the data set.

[0030] Before constructing the data set, the image data is processed by image registration: due to the differences in resolution and viewing angle between the binocular sensor, the infrared image resolution is 640x512, and the visible light image resolution is 3840x2160, accurate registration is needed. First, the infrared image and the visible light image are cropped to the same size; manually select at least four obvious and easily identifiable feature points (usually located at the edge, corner or area with significant brightness change) in each pair of images; calculate the projection transformation matrix from the infrared image to the visible light image based on these feature points; finally, perform morphing processing on the infrared image based on the projection transformation, to realize the spatial alignment of the two images.

[0031] In the process of constructing the data set, the data annotation uses a professional annotation tool LabelImg to perform detailed annotation on the UAV target in each pair of infrared and visible light images. The annotation process is performed by multiple experienced annotators and is reviewed and corrected multiple times to ensure that the annotation data in each pair of images is consistent and has high accuracy. Finally, the entire data is divided into a training set, a validation set and a test set according to a ratio of 7:2:1.

[0032] S2, constructing a model based on a staged progressive fusion network architecture, and training the model using the data set; the staged progressive fusion network architecture includes a dual-branch feature extractor, a cross-modal feature alignment module, a refined feature fusion module, a multi-scale feature fusion module, and a detection output module. A schematic diagram of the staged progressive fusion network architecture processing data is shown in Figure 1 .

[0033] The dual-branch feature extractor is configured to extract features from the infrared image and the visible light image, respectively. The input image is first preprocessed by standardization and normalization, and then processed by several convolutional layers. The resolution of the feature map gradually decreases, and the number of channels gradually increases. Different scale target features are extracted, and finally the infrared features and the visible light features .

[0034] The cross-modal feature alignment module is configured to align the infrared image features and the visible light image features. The processing of the data is divided into two stages: a shallow stage and a deep stage.

[0035] In the shallow stage, a channel attention mechanism is used to learn the importance of different feature channels and weight the importance to other modalities to obtain complementary features of the modalities.

[0036] Specifically, the channel attention mechanism is used to automatically learn the importance of each feature channel, thereby dynamically enhancing the features that are helpful for the task while suppressing unimportant features. The formula for calculating the channel attention of the two modalities is as follows:

[0037]

[0038]

[0039] wherein, represents the importance of the visible light channel, represents the importance of the infrared channel, represents a sigmoid activation function, represents a fully connected layer, represents a Relu activation function, represents average pooling, represents an infrared feature, represents a visible light feature.

[0040] Further, the visible light channel importance is applied to the infrared modality, and the infrared channel importance is applied to the visible light modality, to capture the complementary features between the modalities, and the formula is:

[0041]

[0042]

[0043] wherein, and respectively represent the visible light modality complementary feature and the infrared modality complementary feature.

[0044] In the deep stage, the modality complementary features are projected to the hidden state space, and the feature transmission is performed by using a gating mechanism, and the infrared image features and the visible light image features after alignment are respectively output. The feature processing in the deep stage includes:

[0045] Step 1, performing preliminary feature extraction on the input features; including: sequentially performing normalization, linear transformation, depth separable convolution, and local two-dimensional selective scanning on the output features; and the formula is:

[0046]

[0047]

[0048] wherein, and respectively represent the further extracted visible light features and infrared features, represents local two-dimensional selective scanning; the local two-dimensional selective scanning module is a lightweight module design implemented by the present application to simultaneously capture local details and global context information, including a first branch and a second branch, the first branch performs standard four-way scanning on the input feature map; the second branch divides the input feature map into blocks, and performs independent four-way scanning in each block; the first branch scanning result and the second branch scanning result are fused by a hyperparameter as the local two-dimensional selective scanning output; compared with the traditional SS2D method, the local two-dimensional selective scanning module is more suitable for multi-modal fusion tasks, and can capture local detail information while maintaining global perception ability.

[0049] Step 2, calculating the fusion weight based on the bidirectional gating fusion mechanism, and the calculation formula is:

[0050]

[0051]

[0052] wherein and correspond to the weight of the visible light feature and the infrared feature respectively.

[0053] Step 3, fusion according to the fusion weight, the formula is:

[0054]

[0055]

[0056]

[0057]

[0058] wherein, and respectively represent the aligned infrared image feature and the visible light image feature. Such design enables the cross-modal feature alignment module to capture both local details and global context information, ensuring a more comprehensive and accurate feature fusion process.

[0059] The fine feature fusion module, based on the MemAttention mechanism, performs fine feature fusion on the aligned infrared image feature and the visible light image feature. In this embodiment, the fine feature fusion module includes the following data processing:

[0060] Step 1, element-level fusion of features of different modalities to form a unified feature representation, the formula is:

[0061]

[0062]

[0063]

[0064] wherein, the unified feature representation is used for subsequent attention calculation.

[0065] Step 2, attention calculation on the infrared modal feature and the visible light modal feature respectively, the formula is:

[0066]

[0067] wherein, denotes a scaling factor to prevent gradient vanishing problem; for each attention head , the attention vector is as follows:

[0068]

[0069]

[0070]

[0071]

[0072] wherein, , , denotes the infrared modal feature, denotes the visible light modal feature, denote the projection matrix of query, key and value respectively.

[0073] Step 3, add the attention results of the infrared modal feature and the visible light modal feature, and splice in the head dimension to obtain a multi-head attention output feature , the formula is:

[0074]

[0075]

[0076] wherein, and denote the attention of the infrared modal and the visible light modal respectively.

[0077] Step 4, residual connection is performed between the multi-head attention output feature and the visible light feature to obtain the output of the fine feature fusion module.

[0078]

[0079] wherein, denotes the linear transformation weight matrix of the multi-head attention output, which is used to map the dimension of to the target dimension. The asymmetric residual connection strategy effectively preserves the detail expression ability of the target feature based on the characteristic that the visible light image is rich in detailed information, thereby helping to improve the performance of the model.

[0080] The multi-scale feature fusion module is used for up-sampling and fusing the multi-scale low-resolution features.

[0081] The detection output module is used for target classification and bounding box prediction to generate a detection result containing the category and position of the target; in the detection output module, the network performs target classification and bounding box prediction according to the fused features, and generates the final detection result through convolution and full connection operations.

[0082] Further, in order to avoid multiple bounding boxes detecting the same target, a non-maximum suppression (NMS) algorithm is adopted, by calculating the overlap degree of the detection boxes, the detection box with the highest confidence is retained, and the detection result with high overlap degree and low confidence is removed. In the final output detection result, each result contains the category, confidence score and bounding box position of the target.

[0083] S3, processing the data by using the model to obtain a UAV detection result.

[0084] Embodiment 2

[0085] The embodiment includes a UAV detection method based on infrared-visible light multi-modal fusion, which is different from embodiment 1. In the embodiment, the model based on the phased progressive fusion network architecture is as shown in Figure 2 The double-branch feature extractor is realized based on a YOLO V5 detection head; the cross-modal feature alignment module is realized based on three cross-attention RGB-IR enhanced Mamba modules. Compared with the traditional attention mechanism, Mamba has a significant parameter efficiency advantage in processing high-resolution multi-modal features, and can realize the same or even better global modeling performance with fewer parameters; the fine feature fusion module is realized based on three cross-modal fine fusion Transformer modules.

[0086] The YOLO V5 detection head is as shown in Figure 6 The cross-attention RGB-IR enhanced Mamba module is as shown in Figure 3 The local two-dimensional selective scanning module contained in the cross-attention RGB-IR enhanced Mamba module is as shown in Figure 4 The cross-modal fine fusion Transformer module is as shown in Figure 5

[0087] Further, the model based on the phased progressive fusion network architecture in the embodiment is tested. On the self-built data set, the mAP50 of the model based on the phased progressive fusion network architecture reaches 97.4%, the mAP75 reaches 69.2%, and the mAP reaches 59.3%, which is superior to the existing methods. On the LLVIP pedestrian detection data set, the mAP50 reaches 97.4%, and the mAP reaches 65.2%, which creates a new optimal record. On the M3FD multi-class detection data set, the mAP50 reaches 87.1%, and the performance level of the People, Bus and Truck categories reaches the most advanced level. On the Vedai remote sensing image data set, the mAP50 reaches 75.6%, which fully demonstrates the good cross-domain generalization ability of the method. These experimental results fully prove the detection accuracy advantage of the method in different application scenarios.

[0088] ​Compared with the prior art, the most core difference of the technical solution of the present application is to propose a staged progressive fusion architecture, based on the design concept of "efficient global modeling + accurate local optimization", the cross-modal feature fusion is divided into two stages of initial alignment and later refinement. In addition, the cross-modal feature alignment module designed in the present application can achieve excellent performance without multiple stacking, with a parameter amount of about 18.7M, which is significantly lower than the parameter amount of the prior art solutions such as CFT 44.76M and ICAFusion 20.15M. The lightweight design and high inference speed (about 117 frames per second) make the model deployable to edge devices, meeting the real-time detection requirements.

[0089] The self-built data set adopted in the present application covers multi-scene training data of different backgrounds such as city, suburb, grassland and sky, and performs well under different scenes, different weather conditions and different target sizes, proving that the method designed in the present application has excellent robustness.

[0090] In addition, the present application performs excellently in false detection and missed detection control. Through multi-modal information fusion and fine feature extraction, the reliability of the detection system can be significantly improved, especially in challenging scenes such as complex background, target occlusion and small target detection, the method in the present application can still maintain excellent detection performance, and can be extended to multiple fields such as pedestrian detection, vehicle detection and remote sensing target detection, and has good technical transferability. Lightweight design and high inference speed make the system have the condition of practical deployment, and can be widely applied to airport security monitoring, important facility protection, border patrol and other practical scenes, meeting the requirements of real-time detection and rapid response.

[0091] Finally, it should be noted that the above only describes part of the embodiments of the present application, and for those skilled in the art, various changes, modifications, replacements and deformations of the embodiments can be made without departing from the principles and spirits of the present application, the protection scope of the present application is defined by the appended claims and their equivalents, and the above behaviors should be covered within the protection scope of the present application.

Claims

1. A method for detecting unmanned aerial vehicles (UAVs) based on infrared-visible multimodal fusion, characterized in that, include: S1. Acquire infrared and visible light image data, and construct a precisely aligned infrared-visible light UAV dataset; S2. Construct a model based on a phased progressive fusion network architecture and train the model using a dataset; The phased progressive fusion network architecture includes: a dual-branch feature extractor, a cross-modal feature alignment module, a refined feature fusion module, a multi-scale feature fusion module, and a detection output module; The dual-branch feature extractor is used to extract features from infrared images and visible light images respectively; The cross-modal feature alignment module is used to align infrared image features and visible light image features. The data processing is divided into two stages: shallow and deep. In the shallow stage, a channel attention mechanism is used to learn the importance of different feature channels and weight the importance to other modalities to obtain modal complementary features. In the deep stage, the modal complementary features are projected into the latent state space, and a gating mechanism is used for feature transfer to output the aligned infrared image features and visible light image features respectively. The refined feature fusion module, based on the MemAttention mechanism, performs refined feature fusion on aligned infrared image features and visible light image features. The multi-scale feature fusion module is used to upsample and fuse low-resolution features at multiple scales. The detection output module is used to perform target classification and bounding box prediction, generating detection results that include the target's category and location. S3. The model is used to process the data to obtain the UAV detection results.

2. The UAV detection method based on infrared-visible multimodal fusion according to claim 1, characterized in that, The process of constructing a precisely aligned infrared-visible UAV dataset involves image registration processing of the image data before dataset construction. This includes: cropping the infrared and visible light images to the same size; manually selecting at least four obvious and easily identifiable feature points in each pair of images; calculating the projection transformation matrix from the infrared image to the visible light image based on the feature points; and performing deformation processing on the infrared image based on the projection transformation to achieve spatial alignment of the two images.

3. The UAV detection method based on infrared-visible multimodal fusion according to claim 1, characterized in that, The deep-level stage involves feature processing including: Step 1: Perform preliminary feature extraction on the input features; Step 2: Calculate the fusion weights based on the bidirectional gating fusion mechanism; Step 3: Perform fusion based on fusion weights.

4. The UAV detection method based on infrared-visible multimodal fusion according to claim 3, characterized in that, The preliminary feature extraction includes: performing normalization, linear transformation, depthwise separable convolution, and local two-dimensional selective scanning on the output features sequentially; the formula is: ; ; in, and These represent the visible light and infrared features, respectively, after further extraction. Indicates local two-dimensional selective scanning. and These represent the complementary features of the visible light modes and the complementary features of the infrared modes output in the shallow stage, respectively. The local two-dimensional selective scanning includes a first branch and a second branch. The first branch performs a standard four-way scan on the input feature map; the second branch divides the input feature map into blocks and performs an independent four-way scan within each block; the scanning results of the first branch and the second branch are fused together using hyperparameters to serve as the local two-dimensional selective scanning output.

5. The UAV detection method based on infrared-visible multimodal fusion according to claim 1, characterized in that, The refined feature fusion module processes the data as follows: it performs attention calculations on the infrared modal features and the visible light modal features respectively; it adds the attention results of the infrared modal features and the visible light modal features and concatenates them in the head dimension to obtain the multi-head attention output features; it performs residual connection between the multi-head attention output features and the visible light features to obtain the output of the refined feature fusion module.

6. The UAV detection method based on infrared-visible multimodal fusion according to claim 1, characterized in that, The detection output module uses a non-maximum suppression algorithm to calculate the overlap of detection boxes, retains the detection boxes with the highest confidence, and removes detection results with high overlap and low confidence.

7. The UAV detection method based on infrared-visible multimodal fusion according to claim 1, characterized in that, The dual-branch feature extractor uses a YOLO V5 detection head.