A target detection method fusing visible light image detail features and infrared image contour features

CN122821089APending Publication Date: 2026-09-25ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610958909.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]当前发展的目标检测技术大多基于可见光图像或者红外图像等单一模态,但是单模态图像具有特征信息的局限性,例如可见光图像在光线良好时具有较好的纹理和细节信息,但光线偏暗或有干扰时则信息损失严重;而红外图像由于其特殊的成像原理,始终能够保持良好的轮廓特征,但是细节信息往往不足

Benefits of technology

本发明将非对称互补特征融合模块添加到YOLOv12双分支特征提取骨干网络的中间位置,充分利用可见光图像与红外图像各自的特点和优势,可见光图像色彩、纹理和细节信息丰富,可以为红外图像提供更多的细节信息;红外图像轮廓清晰可以为可见光图像提供更多的轮廓信息,非对称互补特征融合模块可以有效融合两种模态的有益信息,从而有效提高后续对目标的精准检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821089A_ABST
    Figure CN122821089A_ABST
Patent Text Reader

Abstract

The application provides a target detection method fusing visible light image detail features and infrared image contour features, and belongs to the technical field of target detection. On the basis of YOLOv12, a double-flow parallel feature extraction network embedding an asymmetric complementary feature fusion module is used to extract features of different modes. While extracting features of visible light and infrared images respectively, information interaction between two different mode features is realized by using the proposed asymmetric complementary feature fusion, so that the respective characteristics and advantages of visible light images and infrared images are fully utilized, thereby effectively improving the accurate detection of subsequent targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target detection technology, and in particular relates to a target detection method that integrates visible light image detail features and infrared image contour features. Background Technology

[0002] Object detection technology is one of the core downstream tasks of computer vision. It is the foundation for tasks such as instance segmentation and object tracking, and is currently widely used in many fields such as autonomous driving and intelligent monitoring.

[0003] Current target detection technologies are mostly based on single modalities such as visible light or infrared images. However, single-modal images have limitations in feature information. For example, visible light images have good texture and detail information in good lighting conditions, but information loss is severe in low light or with interference. Infrared images, due to their special imaging principle, can always maintain good contour features, but detail information is often insufficient. Target detection based on the fusion of visible light and infrared images is one of the hot development directions in the field of target detection, and many research results have been achieved. However, most of these efforts focus on achieving complementary fusion of features from the two modalities, with less emphasis on complementary fusion addressing the advantages and disadvantages of different modalities.

[0004] Convolutional Neural Networks (CNNs) are the core of deep learning-based object detection and a primary tool for image feature extraction. They possess excellent feature extraction capabilities, but their small receptive field makes it difficult to capture local or even global feature relationships. Transformers enable information interaction between global features, allowing the network to focus more on "meaningful" features and achieve better interaction between features, resulting in a larger receptive field. However, the computational complexity increases dramatically because they require calculating the relationships between every feature point. In recent years, large-kernel CNNs have developed rapidly, achieving a better balance between feature extraction and a large receptive field. Currently, some researchers have attempted to apply large-kernel CNNs to multimodal fusion detection, but they suffer from insufficient consideration of complementary information. They only consider the contour information provided by infrared images while ignoring the detail and texture information provided by visible light images, and they also ignore the information interaction between channels, leading to limitations in the fusion effect. Summary of the Invention

[0005] To address the problems of existing technologies, this invention proposes a target detection method that integrates detailed features from visible light images and contour features from infrared images. Based on YOLOv12, this method employs a dual-stream parallel feature extraction network to extract features from different modalities. While extracting features from both visible light and infrared images, it utilizes the proposed asymmetric complementary feature fusion to achieve information exchange between the features of the two different modalities.

[0006] To achieve the above objectives, the present invention provides a target detection method that integrates visible light image detail features and infrared image contour features, comprising: Step 1: Collect Data Set Collect a dataset of paired visible light and infrared images, and divide the dataset into training, validation, and test sets according to a preset ratio; Step 2: Construct an object detection model Constructing a visible light NCF module: The visible light NCF module first uses a shared convolutional kernel of size 5 to perform preliminary extraction and alignment of visible light and infrared features. Then, infrared features are further extracted using a convolutional kernel of size 7 with an expansion rate of d to extract contour features from the infrared image. The visible light and infrared features are convolved and normalized separately before being concatenated along the channel dimension. The concatenated features undergo spatial and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension before concatenation along the channel dimension. The concatenated features are then convolved and compared with the features before convolution and normalization. The visible light and infrared features are subjected to residual processing, and the two residual-processed features are merged to obtain the spatial dimension fusion feature. The channel dimension information fusion first performs channel dimension global max pooling and channel dimension global average pooling respectively, and then concatenates along the channel dimension. The concatenated features are then convolved and residual-processed with the visible light and infrared features before convolution normalization. The two residual-processed features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and convolved. The convolved features are then residual-fused with the initial visible light features. Building an infrared NCF module: The infrared NCF module first uses a shared convolutional kernel of size 5 to perform preliminary extraction and alignment of visible light and infrared features. Then, the infrared features are further extracted using a convolutional kernel of size 7 with an expansion rate of d to extract the contour features of the infrared image. The visible light and infrared features are then processed by aggregation convolution and convolutional normalization, respectively, before being concatenated along the channel dimension. The concatenated features undergo spatial and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension, then concatenates along the channel dimension. The concatenated features are then processed by convolution and normalized. The visible light and infrared features before processing are subjected to residual processing, and the two residual-processed features are merged to obtain the spatial dimension fusion feature. The channel dimension information fusion is first performed by channel dimension global max pooling and channel dimension global average pooling respectively, and then concatenated along the channel dimension. The concatenated features are then processed by convolution and residually processed with the visible light and infrared features before convolution normalization. The two residual-processed features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and processed by convolution. The convolution-processed features are then residually fused with the initial visible light features. Construct an asymmetric complementary feature fusion module that includes visible light and infrared branches: The visible light branch, from input to output, includes a batch normalization layer, a visible light NCF block, another batch normalization layer, and a multilayer perceptron. The input of the first batch normalization layer and the output residual of the visible light NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer perceptron are also processed. The infrared branch, from input to output, includes a batch normalization layer, an infrared NCF block, another batch normalization layer, and a multilayer sensor. The input of the first batch normalization layer and the output residual of the infrared NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer sensor are processed. The visible light NCF block consists of a fully connected layer, a GELU activation function layer, a visible light NCF module, and a fully connected layer, from input to output. The infrared NCF block consists of a fully connected layer, a GELU activation function layer, an infrared NCF module, and a fully connected layer, from input to output. A multilayer perceptron consists of a fully connected layer, a deep convolutional layer, a GELU activation function layer, and another fully connected layer, from input to output. The output characteristics of the visible light NCF block serve as another input to the infrared NCF block, and vice versa. Add an asymmetric complementary feature fusion module to the middle position of the YOLOv12 dual-branch feature extraction backbone network: Each branch, from input to output, sequentially includes a first C3K2 feature extraction module composed of two CBSs, a second C3K2 feature extraction module composed of one CBS, a first A2C2f feature extraction module composed of one CBS, and a first A2C2f feature extraction module composed of one CBS. The first NFC unit, consisting of two stacked asymmetric complementary feature fusion modules with an expansion rate of 5, is embedded between the first C3K2 feature extraction module and the second C3K2 feature extraction module of the two branches. The output features of the first C3K2 feature extraction modules of the two branches are input into the first NFC unit. The output features of the first NFC unit are merged with the output features of the first C3K2 feature extraction modules of the two branches and then input into the second C3K2 feature extraction module of each branch. The output features of the second C3K2 feature extraction modules of the two branches are merged and then output. The second NFC unit, consisting of four stacked asymmetric complementary feature fusion modules with an expansion rate of 4, is embedded between the second C3K2 feature extraction module and the first A2C2f feature extraction module of the two branches. The output features of the second C3K2 feature extraction modules of the two branches are input into the second NFC unit. The output features of the second NFC unit are merged with the output features of the second C3K2 feature extraction modules of the two branches and then input into the first A2C2f feature extraction module of each branch. The output features of the first A2C2f feature extraction modules of the two branches are merged and then output. A third NFC unit, consisting of three stacked asymmetric complementary feature fusion modules with an expansion rate of 3, is embedded between the first A2C2f feature extraction modules and the second A2C2f feature extraction modules of the two branches. The output features of the first A2C2f feature extraction modules of the two branches are input into the third NFC unit. The output features of the third NFC unit are merged with the output features of the first A2C2f feature extraction modules of the two branches and then input into the second A2C2f feature extraction modules of their respective branches. The output features of the second A2C2f feature extraction modules of the two branches are merged and then output. Step 3: Train the target detection model The target detection model is trained by inputting the training set. Step 4: Test the target detection model The object detection model trained in step 3 is validated and tested using a validation set. The model parameters are then adjusted and retrained to determine the final object detection model. Step 5: Detect the target The final target detection model is used to detect targets in the visible light and infrared image pairs to be detected.

[0007] By adopting the above technical solution, the present invention has the following beneficial effects: This invention adds an asymmetric complementary feature fusion module to the middle position of the YOLOv12 dual-branch feature extraction backbone network, making full use of the characteristics and advantages of visible light images and infrared images. Visible light images are rich in color, texture and detail information, which can provide more detail information for infrared images; infrared images have clear contours, which can provide more contour information for visible light images. The asymmetric complementary feature fusion module can effectively fuse the beneficial information of the two modalities, thereby effectively improving the accuracy of subsequent target detection. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic diagram of the visible light NCF module. Figure 2 This is a schematic diagram of the infrared NCF module. Figure 3 This is a schematic diagram of the structure of the asymmetric complementary feature fusion module; Figure 4 A schematic diagram of the structure of the YOLOv12 backbone network with embedded asymmetric complementary feature fusion modules. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] This invention provides a target detection method that fuses visible light image detail features and infrared image contour features, comprising: Step 1: Collect Data Set Collect a dataset of paired visible light and infrared images, and divide the dataset into training, validation, and test sets according to a preset ratio.

[0012] Step 2: Construct an object detection model like Figure 1As shown, construct a visible light NCF module: The visible light NCF module first uses a shared convolutional kernel of size 5 to perform preliminary extraction and alignment of visible light and infrared features, aiming to extract common similar features. Then, infrared features are further extracted using a convolutional kernel of size 7 and dilation rate d to extract contour features from the infrared image, supplementing the visible light image with more prominent contour features from the infrared image. The visible light and infrared features are convolved and normalized separately before being concatenated along the channel dimension. The concatenated features undergo spatial and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension before concatenating along the channel dimension. The concatenated features are then convolved. After processing, residual processing is performed on the visible light features and infrared features before convolution and normalization. The two residual-processed features are then merged to obtain the spatial dimension fusion feature. For channel dimension information fusion, channel dimension global max pooling and channel dimension global average pooling are performed separately, and then concatenated along the channel dimension. The concatenated features are then convolved and residual-processed on the visible light features and infrared features before convolution and normalization. The two residual-processed features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and convolved. The convolved features are then residual-fused with the initial visible light features to supplement the contour information of the visible light image.

[0013] like Figure 2 As shown, construct the infrared NCF module: The infrared NCF module first uses a shared convolutional kernel with a kernel size of 5 to perform preliminary extraction and alignment of visible light and infrared features. Then, the infrared features are further extracted using a convolutional kernel with a kernel size of 7 and an expansion rate of d to extract contour features from the infrared image. The visible light and infrared features are then subjected to aggregation convolution (to extract visible light detail features at different scales, thus supplementing the infrared image) and convolutional normalization, and then concatenated along the channel dimension. The concatenated features undergo spatial dimension information fusion and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension, and then concatenates along the channel dimension. The concatenated features are then further processed... After convolution processing, residual processing is performed on the visible light features and infrared features before convolution normalization. The two residual features are then merged to obtain the spatial dimension fusion feature. For channel dimension information fusion, channel dimension global max pooling and channel dimension global average pooling are performed separately, and then concatenated along the channel dimension. The concatenated features are then convolved and residually processed on the visible light features and infrared features before convolution normalization. The two residual features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and convolved. The convolved features are then residually fused with the initial visible light features to supplement the detailed features of the infrared image.

[0014] like Figure 3 As shown, an asymmetric complementary feature fusion module including visible light and infrared branches is constructed: The visible light branch, from input to output, includes a batch normalization layer, a visible light NCF block (which, from input to output, includes a fully connected layer, a GELU activation function layer, a visible light NCF module, and a fully connected layer), a batch normalization layer, and a multilayer perceptron (which, from input to output, includes a fully connected layer, a depth convolutional layer, a GELU activation function layer, and a fully connected layer). The input of the first batch normalization layer and the output residual of the visible light NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer perceptron are also processed.

[0015] The infrared branch, from input to output, includes a batch normalization layer, an infrared NCF block (which, from input to output, includes a fully connected layer, a GELU activation function layer, an infrared NCF module, and a fully connected layer), a batch normalization layer, and a multilayer perceptron (which, from input to output, includes a fully connected layer, a deep convolutional layer, a GELU activation function layer, and a fully connected layer). The input of the first batch normalization layer and the output residual of the infrared NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer perceptron are also processed.

[0016] The output characteristics of the visible light NCF block serve as another input to the infrared NCF block, and vice versa.

[0017] The purpose of the asymmetric complementary feature fusion module is to fully utilize the feature advantages of different modalities, and to achieve complementary advantages between the two modalities by leveraging the prominent contour advantages of infrared images and the prominent detail and texture advantages of visible light images.

[0018] like Figure 4 As shown, the asymmetric complementary feature fusion module is added to the middle position of the YOLOv12 dual-branch feature extraction backbone network: Each branch, from input to output, sequentially includes a first C3K2 feature extraction module composed of two CBSs, a second C3K2 feature extraction module composed of one CBS, a first A2C2f feature extraction module composed of one CBS, and a first A2C2f feature extraction module composed of one CBS.

[0019] The first NFC unit, consisting of two stacked asymmetric complementary feature fusion modules with an expansion rate of 5, is embedded between the first C3K2 feature extraction module and the second C3K2 feature extraction module of the two branches. The output features of the first C3K2 feature extraction modules of the two branches are input into the first NFC unit. The output features of the first NFC unit are merged with the output features of the first C3K2 feature extraction modules of the two branches and then input into the second C3K2 feature extraction module of each branch. The output features of the second C3K2 feature extraction modules of the two branches are merged to output feature P2.

[0020] The second NFC unit, consisting of four stacked asymmetric complementary feature fusion modules with an expansion rate of 4, is embedded between the second C3K2 feature extraction module and the first A2C2f feature extraction module of the two branches. The output features of the second C3K2 feature extraction modules of the two branches are input into the second NFC unit. The output features of the second NFC unit are merged with the output features of the second C3K2 feature extraction modules of the two branches and then input into the first A2C2f feature extraction module of each branch. The output features of the first A2C2f feature extraction modules of the two branches are merged and output as P3.

[0021] A third NFC unit, consisting of three stacked asymmetric complementary feature fusion modules with an expansion rate of 3, is embedded between the first A2C2f feature extraction modules and the second A2C2f feature extraction modules of the two branches. The output features of the first A2C2f feature extraction modules of the two branches are input into the third NFC unit. The output features of the third NFC unit are merged with the output features of the first A2C2f feature extraction modules of the two branches and then input into the second A2C2f feature extraction modules of their respective branches. The output features of the second A2C2f feature extraction modules of the two branches are merged and output as P4.

[0022] Step 3: Train the target detection model The target detection model is trained by inputting the training set. Step 4: Test the target detection model The object detection model trained in step 3 is validated and tested using a validation set. The model parameters are then adjusted and retrained to determine the final object detection model. Step 5: Detect the target The final target detection model is used to detect targets in the visible light and infrared image pairs to be detected.

[0023] The following specific examples are provided in conjunction with the above embodiments. It should be understood that the following specific examples are only illustrative of the specific implementation of the above embodiments and are not intended to limit the technical solutions of the above embodiments.

[0024] 1. Experimental setup 1.1 Dataset To verify the effectiveness of the designed module, the public dataset FLIR-aligned was used to validate the object detection method. The dataset encompasses both daytime and nighttime scenes, containing 4113 image pairs for training and 1029 image pairs for testing and validation. 40% of the image pairs represent low-light nighttime conditions, and 60% represent well-lit daytime conditions. The dataset mainly covers scenes such as roads and streets, and includes multiple categories such as "people," "vehicles," and "animals," demonstrating good generalization ability.

[0025] 1.2 Experimental Conditions Our object detection model is implemented using the PyTorch deep learning framework, with the latest self-attention-based YOLOv12 version as the base model. Furthermore, to ensure objectivity, all experiments were conducted under unified hardware and software conditions. The operating system used in this experiment was Ubuntu 22.04, the processor was an Intel Xeon E5 2698 (12 cores), a single NVIDIA GeForce RTX 4060 graphics card, 24GB of RAM, and CUDA and CUDNN versions 12.1 and 8.9.2, respectively. The AdamW optimizer was used to adaptively adjust the initial learning rate and momentum parameters. The training batch size was set to 16, the number of training epochs to 150, and all other hyperparameters and loss functions remained consistent with YOLOv12.

[0026] We use common evaluation metrics, such as mAP50, mAP75, and mAP50-95, to evaluate the accuracy of the detection algorithm, the size of the algorithm by the memory size and number of parameters, and the real-time performance by the inference speed and FPS.

[0027] 2. Experimental Results and Analysis The experimental results of our designed object detection model on the FLIR dataset are shown in Table 1. As can be seen from Table 1, our designed object detection model achieved significant results on the FLIR dataset. The results show that mAP50 improved by 7.9% and 1.8% respectively compared to using only visible light and infrared images, while mAP75 also achieved significant improvements of 18.5% and 4.5%. Even mAP50-95, which has the most stringent accuracy requirements for object detection, achieved 50.4%, significantly better than the 39.5% for visible light images and 48.1% for infrared images. Meanwhile, we used the Add and Concat methods in the YOLOv12 backbone feature extraction network to directly fuse visible light and infrared images. The experimental results show that our innovative fusion method has a significant accuracy advantage, with mAP50-95 outperforming the YOLOv12+Add and YOLOv12+Concat fusion methods by 5.7% and 5.3% respectively.

[0028] Table 1

[0029] 3. Comparative Experiment To further demonstrate the effectiveness and superiority of this invention, we compared it with the latest research results in this field on the FLIR dataset, and the comparison results are shown in Table 2. Our method significantly outperforms other methods in detection metrics such as mAP75 and mAP50-95. While our mAP50 is only 0.3% lower than the GM-DETR algorithm, our method is 11.5% and 4.6% higher in mAP75 and mAP50-95, respectively. Therefore, in terms of detection precision and accuracy alone, our innovative algorithm surpasses most current dual-light fusion-based detection algorithms, achieving state-of-the-art (SOTA) performance in this field.

[0030] Table 2

[0031] Meanwhile, to fully demonstrate the superiority of our innovative algorithm, we conducted a thorough comparison with existing methods in terms of model size, number of parameters, and GFLOPs. The comparison results are shown in Table 3. From the perspective of algorithm application, the research value of a model lies in its ability to be linked to practical applications. Our innovative method significantly outperforms other methods in terms of parameter count, model size, and detection accuracy, demonstrating its superiority. Furthermore, our innovative algorithm achieves a processing speed of 82.2 frames per second, fully meeting the needs of practical applications.

[0032] Table 3

[0033] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.

Claims

1. A target detection method that integrates visible light image detail features and infrared image contour features, characterized in that, include: Step 1: Collect Data Set Collect a dataset of paired visible light and infrared images, and divide the dataset into training, validation, and test sets according to a preset ratio; Step 2: Construct an object detection model Constructing a visible light NCF module: The visible light NCF module first uses a shared convolutional kernel of size 5 to perform preliminary extraction and alignment of visible light and infrared features. Then, infrared features are further extracted using a convolutional kernel of size 7 with an expansion rate of d to extract contour features from the infrared image. The visible light and infrared features are convolved and normalized separately before being concatenated along the channel dimension. The concatenated features undergo spatial and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension before concatenation along the channel dimension. The concatenated features are then convolved and compared with the features before convolution and normalization. The visible light and infrared features are subjected to residual processing, and the two residual-processed features are merged to obtain the spatial dimension fusion feature. The channel dimension information fusion first performs channel dimension global max pooling and channel dimension global average pooling respectively, and then concatenates along the channel dimension. The concatenated features are then convolved and residual-processed with the visible light and infrared features before convolution normalization. The two residual-processed features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and convolved. The convolved features are then residual-fused with the initial visible light features. Building an infrared NCF module: The infrared NCF module first uses a shared convolutional kernel of size 5 to perform preliminary extraction and alignment of visible light and infrared features. Then, the infrared features are further extracted using a convolutional kernel of size 7 with an expansion rate of d to extract the contour features of the infrared image. The visible light and infrared features are then processed by aggregation convolution and convolutional normalization, respectively, before being concatenated along the channel dimension. The concatenated features undergo spatial and channel dimension information fusion. Spatial dimension information fusion first performs global max pooling and global average pooling in the spatial dimension, then concatenates along the channel dimension. The concatenated features are then processed by convolution and normalized. The visible light and infrared features before processing are subjected to residual processing, and the two residual-processed features are merged to obtain the spatial dimension fusion feature. The channel dimension information fusion is first performed by channel dimension global max pooling and channel dimension global average pooling respectively, and then concatenated along the channel dimension. The concatenated features are then processed by convolution and residually processed with the visible light and infrared features before convolution normalization. The two residual-processed features are then merged to obtain the channel dimension fusion feature. The spatial dimension fusion feature and the channel dimension fusion feature are then merged and processed by convolution. The convolution-processed features are then residually fused with the initial visible light features. Construct an asymmetric complementary feature fusion module that includes visible light and infrared branches: The visible light branch, from input to output, includes a batch normalization layer, a visible light NCF block, another batch normalization layer, and a multilayer perceptron. The input of the first batch normalization layer and the output residual of the visible light NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer perceptron are also processed. The infrared branch, from input to output, includes a batch normalization layer, an infrared NCF block, another batch normalization layer, and a multilayer sensor. The input of the first batch normalization layer and the output residual of the infrared NCF block are processed and then input to the second batch normalization layer. The input of the second batch normalization layer and the output residual of the multilayer sensor are processed. The visible light NCF block consists of a fully connected layer, a GELU activation function layer, a visible light NCF module, and a fully connected layer, from input to output. The infrared NCF block consists of a fully connected layer, a GELU activation function layer, an infrared NCF module, and a fully connected layer, from input to output. A multilayer perceptron consists of a fully connected layer, a deep convolutional layer, a GELU activation function layer, and another fully connected layer, from input to output. The output characteristics of the visible light NCF block serve as another input to the infrared NCF block, and vice versa. Add an asymmetric complementary feature fusion module to the middle position of the YOLOv12 dual-branch feature extraction backbone network: Each branch, from input to output, sequentially includes a first C3K2 feature extraction module composed of two CBSs, a second C3K2 feature extraction module composed of one CBS, a first A2C2f feature extraction module composed of one CBS, and a first A2C2f feature extraction module composed of one CBS. The first NFC unit, consisting of two stacked asymmetric complementary feature fusion modules with an expansion rate of 5, is embedded between the first C3K2 feature extraction module and the second C3K2 feature extraction module of the two branches. The output features of the first C3K2 feature extraction modules of the two branches are input into the first NFC unit. The output features of the first NFC unit are merged with the output features of the first C3K2 feature extraction modules of the two branches and then input into the second C3K2 feature extraction module of each branch. The output features of the second C3K2 feature extraction modules of the two branches are merged and then output. The second NFC unit, consisting of four stacked asymmetric complementary feature fusion modules with an expansion rate of 4, is embedded between the second C3K2 feature extraction module and the first A2C2f feature extraction module of the two branches. The output features of the second C3K2 feature extraction modules of the two branches are input into the second NFC unit. The output features of the second NFC unit are merged with the output features of the second C3K2 feature extraction modules of the two branches and then input into the first A2C2f feature extraction module of each branch. The output features of the first A2C2f feature extraction modules of the two branches are merged and then output. A third NFC unit, consisting of three stacked asymmetric complementary feature fusion modules with an expansion rate of 3, is embedded between the first A2C2f feature extraction modules and the second A2C2f feature extraction modules of the two branches. The output features of the first A2C2f feature extraction modules of the two branches are input into the third NFC unit. The output features of the third NFC unit are merged with the output features of the first A2C2f feature extraction modules of the two branches and then input into the second A2C2f feature extraction modules of their respective branches. The output features of the second A2C2f feature extraction modules of the two branches are merged and then output. Step 3: Train the target detection model The target detection model is trained by inputting the training set. Step 4: Test the target detection model The object detection model trained in step 3 is validated and tested using a validation set. The model parameters are then adjusted and retrained to determine the final object detection model. Step 5: Detect the target The final target detection model is used to detect targets in the visible light and infrared image pairs to be detected.