Lightweight bimodal fusion system for RGB-infrared target detection

By introducing a multimodal fusion system into the YOLOv8 network, combining RGB and infrared image features, using infrared images to compensate for the missing RGB image texture, and adopting lightweight self-calibration convolution and CA attention mechanisms, the problems of low accuracy and insufficient computing resources of drone target detection in low-light environments are solved, and efficient lightweight target detection is achieved.

CN120431497APending Publication Date: 2025-08-05NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510554693.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing drone target detection technology has low accuracy in low-light environments, and traditional RGB single-mode detection methods are difficult to effectively identify targets. The YOLO model has too large parameters and high calculation costs, which limits its application in resource-constrained environments.

Method used

Based on the YOLOv8 object detection network, a multimodal fusion strategy was introduced. Through a dual-modal fusion system, the radiation characteristics of the infrared image are combined, and the missing texture of the RGB image is compensated by using the radiation characteristics of the infrared image. The lightweight self-calibration convolution and CA attention mechanism are used to optimize the object detection model.

Benefits of technology

It improves the success rate of target detection of drones in low-light environments, reduces computing resource consumption, and achieves high-precision lightweight target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431497A_ABST
    Figure CN120431497A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight bimodal fusion system for RGB-infrared target detection. The lightweight bimodal fusion system comprises a backbone network, a neck module and a detection head network, the backbone network comprises a first feature extraction branch, a second feature extraction branch, a first multi-modal feature fusion layer, a second multi-modal feature fusion layer and a third multi-modal feature fusion layer; the neck module comprises a first up-sampling layer, a third feature fusion layer, a fifth feature decomposition layer, a second up-sampling layer, a fourth feature fusion layer, a sixth feature decomposition layer, a tenth convolution layer, a fifth feature fusion layer, a first self-calibration convolution layer, an eleventh convolution layer, a sixth feature fusion layer and a second self-calibration convolution layer which are connected in sequence; and the detection head network comprises three detection head modules and is used for outputting a target detection result. According to the method, the unmanned aerial vehicle can detect clear target contour information in a low-light environment, so that the success rate of target detection of the unmanned aerial vehicle is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) vision and target detection, and in particular to a lightweight dual-modal fusion system for RGB-infrared target detection. Background Art

[0002] In recent years, the development of drone vision and target detection technology has become one of the key breakthroughs in many fields, and the application of drone vision and target detection technology has become increasingly extensive and important.

[0003] Traditional RGB single-modality target detection methods for drones can provide rich color information, effectively utilizing the color difference between the target and the background, thereby more accurately identifying and distinguishing the target. RGB images can also capture the target's texture details, helping to improve target detection accuracy. However, in low-light environments, the brightness and contrast of RGB images are significantly reduced, and the distinction between the target and the background becomes poor, resulting in a significant decrease in target detection accuracy. At the same time, to obtain sufficient brightness, the camera needs to increase the exposure time, which causes image blur and further reduces detection performance. Some key visual features of the target object may disappear completely, making it difficult for target detection algorithms based on RGB images to effectively extract features. Therefore, in drone target detection tasks in low-light environments, the success rate of traditional RGB single-modality target detection methods is low, making it difficult to effectively meet the requirements of drone target detection tasks. To address this issue, some researchers have proposed combining infrared (IR) and RGB images. For example, patent application CN119722492A discloses a deep learning-based method and system for fusion of infrared and visible light image perception enhancement. This invention proposes fusing the features of these two images to achieve detail preservation, brightness uniformity, and color rendering in low-light conditions. This approach is particularly suitable for applications requiring high nighttime perception. However, this invention primarily focuses on image fusion and is not suitable for real-time target recognition.

[0004] In recent years, with the continuous advancement of deep learning theory and significant increases in computing power, the application of deep learning-based object detection technology in industrial scenarios has greatly expanded. These industrial scenarios include industrial digital design, smart warehousing, and autonomous driving, and object detection technology provides strong technical support for them. YOLO (You Only Look Once), a classic single-stage object detection algorithm, stands out among many other object detection algorithms due to its unique advantages. YOLO's advantages are mainly reflected in the following aspects: First, it has extremely high real-time performance, completing object detection tasks in a short time, which is crucial for industrial scenarios that require rapid response. Second, YOLO's simple and efficient design enables efficient operation even on resource-limited devices, reducing hardware requirements. Furthermore, YOLO has multi-scale detection capabilities, effectively identifying objects of varying sizes, which is extremely useful in complex industrial environments. It also fully utilizes global contextual information to further improve detection accuracy. Finally, YOLO supports multi-task learning, allowing a single model to complete multiple tasks, improving the model's versatility and flexibility. However, as the requirements for object detection accuracy in industrial applications continue to increase, the recognition accuracy of traditional YOLO-based model architectures is gradually failing to meet current needs. At the same time, the problem of excessive number of YOLO model parameters and high computational cost has long existed, which to some extent limits its application in resource-constrained environments.

[0005] Existing technologies also use the YOLOV8 dual-branch model to simultaneously extract and fuse RGB and infrared image features. However, due to limited drone computing resources, the fusion performed by existing dual-branch models includes a large number of redundant features, increasing computing resource consumption while also compromising accuracy. Therefore, finding a lightweight and improved YOLO model architecture while improving object detection performance has become a pressing technical challenge. Summary of the Invention

[0006] The purpose of this invention is to provide a lightweight dual-modal fusion system for RGB-infrared target detection. This system improves upon the YOLOv8 target detection network by introducing a multimodal fusion strategy. By combining the advantages of RGB and infrared images through dual-modal fusion, it enables drones to detect clear target outlines in low-light environments, significantly improving the success rate of drone target detection. Compared to traditional target detection methods, this system boasts higher detection accuracy and fewer parameters, effectively completing drone target detection tasks and meeting the requirements for drone target detection in low-light environments.

[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is: A lightweight dual-modal fusion system for RGB-infrared target detection, comprising a backbone network, a neck module, and a detection head network; The backbone network includes a first feature extraction branch, a second feature extraction branch, a first multimodal feature fusion layer, a second multimodal feature fusion layer, and a third multimodal feature fusion layer; the first feature extraction branch and the second feature extraction branch respectively use convolution layers to extract features from the input infrared image and RGB image, and the obtained original infrared person-vehicle feature map and RGB person-vehicle feature map are respectively subjected to four-layer convolution and feature decomposition networks, and finally enter the corresponding spatial pyramid pooling layer for feature extraction again, and output the final infrared person-vehicle feature map and the final RGB person-vehicle feature map; the first multimodal feature fusion layer, the second multimodal feature fusion layer, and the third multimodal feature fusion layer respectively fuse the infrared image features and RGB image features obtained by the first three layers of cascaded feature decomposition networks, and then restore them to two independent branches of RGB and IR through channel segmentation and enter the next feature decomposition layer; the infrared image features and RGB image features obtained by the last three layers of cascaded feature decomposition networks are respectively obtained by ADD operation to obtain the first person-vehicle feature map, the second person-vehicle feature map, and the third person-vehicle feature map; The third person-vehicle feature map output by the backbone network feature extraction enters the neck module and is divided into two paths. One path undergoes a double-cascade upsampling, feature fusion, and feature decomposition layer, and is fused with the first person-vehicle feature map and the second person-vehicle feature map respectively before undergoing feature decomposition, outputting the fourth person-vehicle feature map and the fifth person-vehicle feature map of different sizes respectively; the fifth person-vehicle feature map is convolved with the fourth person-vehicle feature map in the feature fusion layer and output to the first self-calibration convolution layer to obtain the sixth person-vehicle feature map; the sixth person-vehicle feature map is convolved with the second person-vehicle feature map in the feature fusion layer and output to the second self-calibration convolution layer to output the seventh person-vehicle feature map; The detection head network includes three detection head modules, which respectively perform CA attention mechanism processing on the fifth person-vehicle feature map, the sixth person-vehicle feature map, and the seventh person-vehicle feature map and then output the target detection results.

[0008] Furthermore, the first feature extraction branch and the second feature extraction branch have the same structure, both including a first convolutional layer, a second convolutional layer, a first feature decomposition layer, a third convolutional layer, a second feature decomposition layer, a fourth convolutional layer, a third feature decomposition layer, a fifth convolutional layer, a fourth feature decomposition layer and a spatial pyramid pooling layer; Among them, the first convolution layer, the second convolution layer, the first feature decomposition layer, and the third convolution layer are connected in sequence. The third convolution layer is connected to the second feature decomposition layer through the first multimodal feature fusion layer. The second feature decomposition layer is connected to the fourth convolution layer. The fourth convolution layer is connected to the third feature decomposition layer through the second multimodal feature fusion layer. The third feature decomposition layer is connected to the fifth convolution layer. The fifth convolution layer is connected to the fourth feature decomposition layer through the third multimodal feature fusion layer. The fourth feature decomposition layer is connected to the spatial pyramid pooling layer. The first multimodal feature fusion layer, the second multimodal feature fusion layer, and the third multimodal feature fusion layer simultaneously receive the infrared image features and RGB image features sent by the two feature extraction branches, perform feature fusion on the two, use the thermal radiation features of the infrared image to compensate for the missing texture of the RGB image, use the details of the RGB image to optimize the target contour of the infrared image, and then restore them to two independent branches of RGB image and infrared image through channel segmentation to enter the next feature decomposition layer.

[0009] Furthermore, the multimodal feature fusion layer includes an infrared radiation feature extraction module, an RGB contour feature extraction module, a bimodal linear fusion module, a noise channel filtering module and a channel segmentation module; The infrared radiation feature extraction module processes the input infrared image features and extracts the infrared image thermal radiation features contained therein; the RGB contour feature extraction module processes the input RGB image features and extracts the target contour features contained therein; the bimodal linear fusion module fuses the infrared image thermal radiation features and the target contour features, uses the infrared image thermal radiation features to compensate for the missing texture of the RGB image, and uses the details of the RGB image to optimize the infrared image target contour; the noise channel filtering module filters the noise channel to filter out redundant noise; the channel segmentation module restores the fused features into two independent branches, RGB and infrared, through channel segmentation, and sends them to the corresponding feature decomposition layer respectively.

[0010] Furthermore, the feature decomposition layer includes a sixth convolutional layer, a segmentation layer, n bottleneck structure layers, a first feature fusion layer and a seventh convolutional layer; In the feature decomposition layer, the human and vehicle feature maps are sequentially convolved and segmented and then input into the n-cascade bottleneck structure. The output of each bottleneck structure is combined with the human and vehicle feature maps after the segmentation operation. Figure 1 Then, it enters the first feature fusion layer, outputs the fused human-vehicle feature map and performs feature extraction again through convolution.

[0011] Furthermore, the first feature fusion layer fuses feature maps of different levels through concatenation and element-by-element addition to obtain human and vehicle feature maps with feature information at different levels.

[0012] Furthermore, the spatial pyramid pooling layer includes an eighth convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a second feature fusion layer and a ninth convolutional layer; The feature map after convolution processing in the eighth convolutional layer is divided into two paths, one of which enters the third maximum pooling layer. Figure 1 One path directly enters the feature fusion layer, and the other path passes through the cascaded second maximum pooling layer and the first pooling layer and then enters the second feature fusion layer. In the feature fusion layer, the feature map after convolution processing by the eighth convolution layer, the feature map after the first maximum pooling, the feature map after the second maximum pooling, and the human and vehicle feature map after the third maximum pooling are spliced to output a multi-scale feature map, and then the feature is extracted through the ninth convolution layer to output the human and vehicle feature map with reduced size.

[0013] Furthermore, the neck module includes a first upsampling layer, a third feature fusion layer, a fifth feature decomposition layer, a second upsampling layer, a fourth feature fusion layer, a sixth feature decomposition layer, a tenth convolutional layer, a fifth feature fusion layer, a first self-calibration convolutional layer, an eleventh convolutional layer, a sixth feature fusion layer, and a second self-calibration convolutional layer, which are connected in sequence; The third person-vehicle feature map output by the backbone network feature extraction enters the neck module and is divided into two paths. One path enters the sixth feature fusion layer, and the other path enters the first upsampling layer for sampling, and then enters the third feature fusion layer and is melted with the second person-vehicle feature map before entering the fifth feature decomposition layer to obtain the fourth person-vehicle feature map; the fourth person-vehicle feature map is also divided into two paths. One path enters the fifth feature fusion layer, and the other path enters the second upsampling layer for sampling, and then enters the fourth feature fusion layer and is fused with the first person-vehicle feature map before entering the sixth feature decomposition layer to obtain the fifth person-vehicle feature map; the fifth person-vehicle feature map is convolved by the tenth convolution layer and then enters the fifth feature fusion layer to be fused with the fourth person-vehicle feature map. The fusion result enters the first self-calibration convolution layer to output the sixth person-vehicle feature map; the sixth person-vehicle feature map is convolved by the eleventh convolution layer and then enters the sixth feature fusion layer to be fused with the third person-vehicle feature map. The fusion result enters the second self-calibration convolution layer to output the seventh person-vehicle feature map.

[0014] Furthermore, the detection head module includes two independent detection branches: a position detection branch and a category detection branch; the input human and vehicle feature maps enter the two detection branches respectively, and output target position information and target category information.

[0015] Compared with the prior art, the present invention has the following beneficial effects: First, the lightweight dual-modal fusion system for RGB-infrared target detection of the present invention proposes a new dual-modal feature fusion layer. Taking into account the characteristics of RGB and infrared images, the backbone network only extracts the radiation characteristics of the IR image, assists the external contour features of the RGB image, and filters out noise signals through the noise filtering module, thereby effectively combining the advantages of RGB and infrared images, improving recognition accuracy and reducing resource consumption.

[0016] Second, the lightweight dual-modal fusion system for RGB-infrared target detection of the present invention proposes a new neck module based on the characteristic properties extracted by the dual-modal feature fusion layer, and further performs multi-scale fusion processing on the fusion features extracted by the dual-modal feature fusion layer. The self-calibration convolution module adaptively constructs long-distance spatial domain and channel correlation through self-correction operations, reducing the spatial and channel redundancy widely present in standard convolution, enabling drones to detect clear target contour information in low-light environments, thereby significantly improving the success rate of drone target detection and better meeting the needs of practical applications.

[0017] Third, the lightweight dual-modal fusion system for RGB-infrared target detection of the present invention, by introducing the attention mechanism and the SCConv lightweight module, not only further improves the accuracy of target detection, but also solves the problem of excessive number of parameters in the traditional Yolo model, providing new ideas and methods for the further development of UAV target detection technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the structure of the lightweight dual-modal fusion system for RGB-infrared target detection of the present invention.

[0019] Figure 2 This is the feature-level fusion framework diagram of RGB images and infrared images proposed in this invention.

[0020] Figure 3 This is a block diagram of the spatial pyramid pooling layer in Yolo-NightV of the present invention.

[0021] Figure 4 This is a block diagram of the convolutional layer principle in Yolo-NightV of the present invention.

[0022] Figure 5 This is a functional block diagram of the feature decomposition layer in Yolo-NightV of the present invention.

[0023] Figure 6 This is a principle block diagram of the multimodal feature fusion layer in Yolo-NightV of the present invention.

[0024] Figure 7This is a block diagram of the bottleneck structure principle in Yolo-NightV of the present invention.

[0025] Figure 8 This is a block diagram of the detection head principle in Yolo-NightV of the present invention.

[0026] Figure 9 This is a block diagram of the CA attention mechanism used in Yolo-NightV of the present invention.

[0027] Figure 10 This is a diagram showing the results of the ablation experiment performed in Yolo-NightV of the present invention. DETAILED DESCRIPTION

[0028] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.

[0029] The present invention discloses a lightweight dual-modal fusion system for RGB-infrared target detection, which includes a backbone network, a neck module and a detection head network.

[0030] join Figure 1 The backbone network includes a first feature extraction branch, a second feature extraction branch, a first multimodal feature fusion layer, a second multimodal feature fusion layer, and a third multimodal feature fusion layer; the first feature extraction branch and the second feature extraction branch respectively use convolution layers to extract features from the input infrared image and RGB image, and the obtained original infrared person-vehicle feature map and RGB person-vehicle feature map are respectively subjected to four-layer convolution and feature decomposition networks, and finally enter the corresponding spatial pyramid pooling layer for feature extraction again, and output the final infrared person-vehicle feature map and the final RGB person-vehicle feature map; the first multimodal feature fusion layer, the second multimodal feature fusion layer, and the third multimodal feature fusion layer respectively fuse the infrared image features and RGB image features obtained by the first three layers of cascaded feature decomposition networks, and then restore them to two independent branches of RGB and IR through channel segmentation and enter the next feature decomposition layer; the infrared image features and RGB image features obtained by the last three layers of cascaded feature decomposition networks are respectively obtained by ADD operation to obtain the first person-vehicle feature map, the second person-vehicle feature map, and the third person-vehicle feature map.

[0031] The third person-vehicle feature map output by the backbone network feature extraction enters the neck module and is divided into two paths. One path undergoes a double-cascade upsampling, feature fusion and feature decomposition layer, and is fused with the first person-vehicle feature map and the second person-vehicle feature map respectively before feature decomposition, outputting the fourth person-vehicle feature map and the fifth person-vehicle feature map of different sizes respectively; the fifth person-vehicle feature map is convolved with the fourth person-vehicle feature map in the feature fusion layer and output to the first self-calibration convolution layer to obtain the sixth person-vehicle feature map; the sixth person-vehicle feature map is convolved with the second person-vehicle feature map in the feature fusion layer and output to the second self-calibration convolution layer to output the seventh person-vehicle feature map.

[0032] The detection head network consists of three detection head modules, which respectively perform CA attention mechanism processing on the fifth person-vehicle feature map, the sixth person-vehicle feature map, and the seventh person-vehicle feature map and then output the target detection results.

[0033] When performing target detection tasks, data set preparation and preprocessing are crucial steps. This paper uses the corresponding database to load target detection image data, converts its annotation files into the txt format required by the Yolo algorithm, and divides them into training and test sets in a 4:1 ratio. This step provides the necessary data support for subsequent target detection model training and evaluation.

[0034] Data-level fusion, feature-level fusion, and decision-level fusion are three common data fusion methods. The advantage of data-level fusion is that it maintains the integrity and authenticity of the original data, enabling the fused data to provide a more accurate and comprehensive representation or estimate of the observed target. The advantage of feature-level fusion is that it reduces the amount of raw data to be processed, improving system processing speed and real-time performance. Furthermore, by extracting representative features, it can reduce the impact of noise and redundant information on system processing. The advantage of decision-level fusion is that it allows for flexible selection of sensor results, improving the system's fault tolerance. Furthermore, by enhancing the ability to accommodate multi-source, heterogeneous sensors, more complex decision-making processes can be implemented. Furthermore, decision-level fusion can reduce data transmission and storage requirements.

[0035] In order to ensure the accuracy of target detection while simplifying the model architecture, the present invention uses the feature-level fusion method, as shown in the attached framework. Figure 2 The ablation experimental results of each fusion method are shown in the appendix. Figure 10The visible RGB image shown provides rich color and texture information, making it suitable for environments with good lighting conditions. The infrared image, on the other hand, provides temperature information, making it suitable for nighttime or inclement weather conditions. By fusing these two images and leveraging their respective strengths, the model can understand targets from multiple dimensions, thereby enhancing its adaptability to environmental changes and improving the accuracy and robustness of target detection. The two images are fed into a feature extractor to generate intermediate features. These features contain important information about the target in their respective spectra. These intermediate features are then fused to combine the target information from both spectra. The fused features are then fed into a multi-scale detector head, which can detect targets at different scales to accommodate targets of varying sizes. The head ultimately outputs the target detection results, including information such as the target's location and category.

[0036] Applying this idea to the YoloV8 algorithm model and improving the entire model architecture has formed our dual-modal fusion model Yolo-NightV based on RGB and infrared images, as shown in the attached figure. Figure 1 As shown in Figure 2, the RGB and IR images are split into two paths and fed simultaneously into the backbone, neck, and head of YOLO-NightV for feature extraction and object detection. Since RGB images have three channel dimensions, while IR images have only one, we perform channel upscaling on the IR images to align them with the three dimensions of the RGB image.

[0037] First, the structure and working principle of the functional module involved in the present invention are described below.

[0038] The spatial pyramid pooling layer is shown in the attached Figure 3 As shown in the figure, the spatial pyramid pooling layer includes the eighth convolutional layer, the first maximum pooling layer, the second maximum pooling layer, the third maximum pooling layer, the second feature fusion layer and the ninth convolutional layer; the feature map after convolution processing in the eighth convolutional layer is divided into two paths, one of which enters the third maximum pooling layer, and the pooled feature map is Figure 1 One path directly enters the feature fusion layer, and the other path passes through the cascaded second maximum pooling layer and the first pooling layer and then enters the second feature fusion layer. In the feature fusion layer, the feature map after convolution processing by the eighth convolution layer, the feature map after the first maximum pooling, the feature map after the second maximum pooling, and the human and vehicle feature map after the third maximum pooling are spliced to output a multi-scale feature map, and then the feature is extracted through the ninth convolution layer to output the human and vehicle feature map with reduced size.

[0039] In the convolution operation, the image to be detected is first convolved to complete the preliminary feature extraction, and batch normalization and activation function operations are required to output the human and vehicle feature map, as shown in the attached figure. Figure 4 shown.

[0040] In the feature decomposition operation, the human and vehicle feature maps are sequentially convolved and segmented and then input into the n-cascade bottleneck structure. The output of each bottleneck structure is combined with the human and vehicle feature maps after the segmentation operation. Figure 1 Then, the fused human-vehicle feature map is output and subjected to convolution for feature extraction, as shown in the attached figure. Figure 5 In the bottleneck structure, the human-vehicle feature map undergoes a double-cascade convolution operation for feature extraction and is fused with the original human-vehicle feature map, as shown in the attached figure. Figure 7 shown.

[0041] In the feature fusion layer, feature maps at different levels are fused through concatenation and element-by-element addition, so that the model can use feature information at different levels to improve the detection accuracy of the target.

[0042] The multimodal feature fusion layer includes an infrared radiation feature extraction module, an RGB contour feature extraction module, a bimodal linear fusion module, a noise channel filtering module, and a channel segmentation module; the infrared radiation feature extraction module processes the input infrared image features and extracts the infrared image thermal radiation features contained therein. The RGB contour feature extraction module processes the input RGB image features and extracts the target contour features contained therein. The bimodal linear fusion module fuses the infrared image thermal radiation features and the target contour features; since RGB images retain details but have low accuracy under weak light conditions, IR images can capture thermal radiation information but lack target contour information, so the IR image thermal radiation features are used in the multimodal feature fusion layer to compensate for the missing texture of the RGB image, and the details of the RGB image are used to optimize the IR target contour. The noise channel filtering module filters the noise channel to remove redundant noise; the channel segmentation module restores the fused features into two independent branches, RGB and infrared, through channel segmentation, and sends them to the corresponding feature decomposition layer respectively. At this point, the features of the two branches have been integrated into each other's contextual information, providing a better feature alignment basis for subsequent deep fusion, such as Figure 6 shown.

[0043] In the backbone network, the image to be detected undergoes convolution for feature extraction, outputting a human / vehicle feature map. This map then passes through a four-tiered convolution and feature decomposition network before entering a spatial pyramid pooling layer for further feature extraction, outputting a human / vehicle feature map. The spatial pyramid pooling layer reduces the size of the feature map while retaining important spatial information, thereby reducing computational complexity and improving the model's generalization capabilities.

[0044] Specifically, the first and second feature extraction branches have the same structure, consisting of a first convolutional layer, a second convolutional layer, a first feature decomposition layer, a third convolutional layer, a second feature decomposition layer, a fourth convolutional layer, a third feature decomposition layer, a fifth convolutional layer, a fourth feature decomposition layer, and a spatial pyramid pooling layer. To unify the channel dimensions of the input image, the input layers of the two feature extraction branches are slightly different. A dimensionality increase module is included in the input layer of the first feature extraction branch to increase the channel dimensions of the IR image to align with the three dimensions of the RGB image. The first convolutional layer, the second convolutional layer, the first feature decomposition layer, and the third convolutional layer are connected in sequence. The third convolutional layer is connected to the second feature decomposition layer through the first multimodal feature fusion layer. The second feature decomposition layer is connected to the fourth convolutional layer. The fourth convolutional layer is connected to the third feature decomposition layer through the second multimodal feature fusion layer. The third feature decomposition layer is connected to the fifth convolutional layer. The fifth convolutional layer is connected to the fourth feature decomposition layer through the third multimodal feature fusion layer. The fourth feature decomposition layer is connected to the spatial pyramid pooling layer. The first, second, and third multimodal feature fusion layers simultaneously receive the infrared image features and RGB image features sent by the two feature extraction branches, perform feature fusion on the two, use the infrared image thermal radiation features to compensate for the missing texture of the RGB image, and use the details of the RGB image to optimize the target contour of the infrared image. Then, through channel segmentation, the RGB image and infrared image are restored into two independent branches and enter the next feature decomposition layer.

[0045] The third person-vehicle feature map output by the backbone network feature extraction is divided into two paths. One path undergoes a double-cascade of upsampling, feature fusion, and feature decomposition layers. After the two feature decompositions, the output is a person-vehicle feature map of different sizes, denoted as the fourth person-vehicle feature map and the fifth person-vehicle feature map. The fifth person-vehicle feature map undergoes convolution, is fused with the fourth person-vehicle feature map in the feature fusion layer, and is output to the first self-calibration convolution layer. The resulting image is denoted as the sixth person-vehicle feature map. The sixth person-vehicle feature map undergoes convolution, is fused with the third person-vehicle feature map in the feature fusion layer, and is output to the second self-calibration convolution layer. The output is denoted as the seventh person-vehicle feature map.

[0046] In the detection head network part, multi-scale human and vehicle feature maps are sequentially input into the CA attention mechanism and the detection head to detect the target. Specifically, the fifth human and vehicle feature map output by the dual cascade, the sixth human and vehicle feature map output by the first self-calibration convolution, and the seventh human and vehicle feature map output by the second self-calibration convolution are sequentially input into the CA attention mechanism and the detection head for target detection. In the detection head, the input human and vehicle feature maps are divided into two paths, wherein one path of human and vehicle feature maps passes through the convolution layer, the two-dimensional convolution layer and the detection frame loss in sequence to output the target position information, and the other path of human and vehicle feature maps passes through the convolution layer, the two-dimensional convolution layer and the classification loss in sequence to output the target category information, as shown in the attached figure. Figure 7 shown.

[0047] The principle block diagram of CA attention mechanism is as follows Figure 9 As shown in Figure 2, the CA attention mechanism is an advanced feature enhancement technology. In the Yolo-NightV model of the present invention, the CA attention mechanism is added before the detection head to enhance the quality of the feature map input to the detection head.

[0048] The quality of feature maps is improved through a series of operations. We first perform a one-dimensional global pooling operation on the input feature map, which helps capture global information in the image. Next, through a series of convolutions and activation functions, two direction-aware feature maps are generated. These two feature maps aggregate features along different directions, thereby capturing long-range dependencies while retaining precise location information. In the coordinate attention generation stage, the two direction-aware feature maps are merged using the concat operation to obtain the final attention map. This operation effectively integrates feature information from different directions, generating a comprehensive attention map that contains rich spatial and channel information. Finally, we apply the attention map to the input feature map through element-wise multiplication to achieve feature enhancement. This process enables the model to focus more on important feature regions while suppressing unimportant or noisy features, thereby improving model performance and accuracy. In this way, the added CA attention mechanism not only improves the quality of feature maps but also enhances the model's ability to recognize objects, especially in dense prediction tasks such as object detection and semantic segmentation.

[0049] The present invention also improves some convolution modules in the original YOLOv8 network into self-calibration convolution modules. Through self-correction operations, it adaptively constructs long-range spatial domain and channel correlations, reduces spatial and channel redundancy that is widely present in standard convolution, reduces computational costs and model storage, and improves the performance of the CNN model, achieving a lightweight model. Preferably, the present invention adds SCConv lightweight convolution modules at two self-calibration convolutions, which can further reduce the number of computational parameters while maintaining high accuracy in target detection.

[0050] The training set is used to train the bimodal target detection model to obtain the optimal bimodal target detection model. The training results are shown in the attached figure. Figure 10 The main data we observe include: precision, recall, mean average precision (mAP), and parameters.

[0051] Figure 10The top two rows show the results of training the images to be detected in the dataset using the original Yolov8 model architecture. It can be seen that under low-light conditions, the accuracy of single RGB image detection is difficult to reach a high standard. Although the detection accuracy of the single infrared modality has reached a high level, its operation parameters are extremely large, which greatly reduces the target detection rate. At the same time, when the algorithm is deployed on a drone, the computing power of the drone will also face huge challenges. Therefore, we tried three common modality fusion methods, as shown in the attached figure. Figure 10 As shown in rows 3-5, the training results of data-level fusion, feature-level fusion, and decision-level fusion are respectively displayed. Due to the need to balance detection accuracy and model simplicity, we finally chose feature-level fusion as the fusion method of our model architecture. Compared with the common RGB single-modal target detection algorithm, feature-level fusion increased the detection accuracy mAP value from 0.864 to 0.92, while the number of parameters was also greatly reduced. After determining the modality fusion method, we tried four different attention mechanisms, and the results obtained from the training are shown in the attached figure. Figure 10 As shown in rows 6-9 of the data, it is not difficult to see that after adding the CA attention mechanism, the mAP value of the model detection accuracy has achieved the greatest improvement, increasing to 0.942. At the same time, the number of operating parameters is also in a better range compared to the other three attention mechanisms. It can be seen that the CA attention mechanism helps the model better capture the correlation and dependency between channels, thereby improving the model's ability to understand the input data. In addition, the CA attention mechanism suppresses unimportant channels, reduces redundant information in the input data, and increases the model's attention to key features, which helps to reduce the model's computational complexity and improve the model's generalization ability. On the basis of determining the use of the CA attention mechanism, we compared four different loss functions. The results are shown in the attached figure. Figure 10 As shown in rows 10-13 of the , GIoU is selected due to its higher mAP value. Finally, we use the self-calibrated lightweight convolution module to lightweight the model. We add lightweight convolution modules to three places in the model, namely Figure 1 15, 18, and 21, and observe the results after training, as shown in the attached Figure 10 SCConv15_18_21 means that the lightweight convolution module is added in three places, and SCConv18_21 means that the lightweight convolution module is added in the attached Figure 1 The two self-calibration convolution modules are added, while SCConv21 indicates that only Figure 1 The self-calibration convolution module is added in the lower middle part. By comparing the three sets of data, it can be found that the number of model operation parameters is significantly reduced after adding the SCConv lightweight convolution module. When only adding the self-calibration convolution at two locations, the target detection accuracy is the highest. Therefore, in the final model Yolo-NightV, only the SCConv lightweight convolution module is added to the self-calibration convolution module.

[0052] Finally, the dual-modal fusion model Yolo-NightV based on RGB and infrared images of the present invention is obtained and compared with the unimproved Yolov8n algorithm, as shown in the attached figure. Figure 10 As shown in the two bottom rows, Yolo-NightV has a significant advantage in detection accuracy, computing parameters, and computing speed, proving the effectiveness and robustness of the model of the present invention and providing a new idea and paradigm for target detection under low-light conditions.

[0053] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0054] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A lightweight dual-modal fusion system for RGB-infrared target detection, characterized in that: The system includes a backbone network, a neck module and a detection head network; The backbone network includes a first feature extraction branch, a second feature extraction branch, a first multimodal feature fusion layer, a second multimodal feature fusion layer, and a third multimodal feature fusion layer; the first feature extraction branch and the second feature extraction branch respectively use convolution layers to extract features from the input infrared image and RGB image, and the obtained original infrared person-vehicle feature map and RGB person-vehicle feature map are respectively subjected to four-layer convolution and feature decomposition networks, and finally enter the corresponding spatial pyramid pooling layer for feature extraction again, and output the final infrared person-vehicle feature map and the final RGB person-vehicle feature map; the first multimodal feature fusion layer, the second multimodal feature fusion layer, and the third multimodal feature fusion layer respectively fuse the infrared image features and RGB image features obtained by the first three layers of cascaded feature decomposition networks, and then restore them to two independent branches of RGB and IR through channel segmentation and enter the next feature decomposition layer; the infrared image features and RGB image features obtained by the last three layers of cascaded feature decomposition networks are respectively obtained by ADD operation to obtain the first person-vehicle feature map, the second person-vehicle feature map, and the third person-vehicle feature map; The third person-vehicle feature map output by the backbone network feature extraction enters the neck module and is divided into two paths. One path undergoes a double-cascade upsampling, feature fusion, and feature decomposition layer, and is fused with the first person-vehicle feature map and the second person-vehicle feature map before undergoing feature decomposition. The fourth person-vehicle feature map and the fifth person-vehicle feature map of different sizes are output respectively. The fifth person-vehicle feature map is convolved and fused with the fourth person-vehicle feature map in the feature fusion layer and output to the first self-calibration convolution layer to obtain the sixth person-vehicle feature map; The sixth person-vehicle feature map is convolved and fused with the second person-vehicle feature map in the feature fusion layer and output to the second self-calibration convolution layer to output the seventh person-vehicle feature map; The detection head network includes three detection head modules, which respectively perform CA attention mechanism processing on the fifth person-vehicle feature map, the sixth person-vehicle feature map, and the seventh person-vehicle feature map and then output the target detection results.

2. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 1, characterized in that: The first feature extraction branch has the same structure as the second feature extraction branch, and both include a first convolutional layer, a second convolutional layer, a first feature decomposition layer, a third convolutional layer, a second feature decomposition layer, a fourth convolutional layer, a third feature decomposition layer, a fifth convolutional layer, a fourth feature decomposition layer, and a spatial pyramid pooling layer; Among them, the first convolution layer, the second convolution layer, the first feature decomposition layer, and the third convolution layer are connected in sequence. The third convolution layer is connected to the second feature decomposition layer through the first multimodal feature fusion layer. The second feature decomposition layer is connected to the fourth convolution layer. The fourth convolution layer is connected to the third feature decomposition layer through the second multimodal feature fusion layer. The third feature decomposition layer is connected to the fifth convolution layer. The fifth convolution layer is connected to the fourth feature decomposition layer through the third multimodal feature fusion layer. The fourth feature decomposition layer is connected to the spatial pyramid pooling layer. The first multimodal feature fusion layer, the second multimodal feature fusion layer, and the third multimodal feature fusion layer simultaneously receive the infrared image features and RGB image features sent by the two feature extraction branches, perform feature fusion on the two, use the thermal radiation features of the infrared image to compensate for the missing texture of the RGB image, use the details of the RGB image to optimize the target contour of the infrared image, and then restore them to two independent branches of RGB image and infrared image through channel segmentation to enter the next feature decomposition layer.

3. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 2, characterized in that: The multimodal feature fusion layer includes an infrared radiation feature extraction module, an RGB contour feature extraction module, a bimodal linear fusion module, a noise channel filtering module and a channel segmentation module; The infrared radiation feature extraction module processes the input infrared image features to extract the infrared image thermal radiation features contained therein; the RGB contour feature extraction module processes the input RGB image features to extract the target contour features contained therein; The dual-modal linear fusion module fuses the thermal radiation features of the infrared image and the target contour features, uses the thermal radiation features of the infrared image to compensate for the missing texture of the RGB image, and uses the details of the RGB image to optimize the target contour of the infrared image; the noise channel filtering module filters the noise channel to remove redundant noise; the channel segmentation module restores the fused features into two independent branches of RGB and infrared through channel segmentation, and sends them to the corresponding feature decomposition layer respectively.

4. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 1, characterized in that: The feature decomposition layer includes a sixth convolutional layer, a segmentation layer, n bottleneck structure layers, a first feature fusion layer and a seventh convolutional layer; In the feature decomposition layer, the human and vehicle feature maps are sequentially convolved and segmented before being input into an n-cascaded bottleneck structure. The output of each bottleneck structure, together with another human and vehicle feature map after the segmentation operation, enters the first feature fusion layer. The fused human and vehicle feature map is output and again convolved for feature extraction.

5. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 4, characterized in that: The first feature fusion layer fuses feature maps of different levels through concatenation and element-by-element addition to obtain human and vehicle feature maps with feature information at different levels.

6. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 1, characterized in that: The spatial pyramid pooling layer includes an eighth convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a second feature fusion layer and a ninth convolutional layer; The feature map after convolution processing by the eighth convolutional layer is divided into two paths, one of which enters the third maximum pooling layer. The feature map after pooling directly enters the feature fusion layer, and the other enters the second feature fusion layer after the cascaded second maximum pooling layer and the first pooling layer. In the feature fusion layer, the feature map after convolution processing by the eighth convolutional layer, the feature map after the first maximum pooling, the feature map after the second maximum pooling, and the human and vehicle feature map after the third maximum pooling are spliced to output a multi-scale feature map, and then feature extraction is performed through the ninth convolutional layer to output the human and vehicle feature map with reduced size.

7. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 1, characterized in that: The neck module includes a first upsampling layer, a third feature fusion layer, a fifth feature decomposition layer, a second upsampling layer, a fourth feature fusion layer, a sixth feature decomposition layer, a tenth convolutional layer, a fifth feature fusion layer, a first self-calibration convolutional layer, an eleventh convolutional layer, a sixth feature fusion layer and a second self-calibration convolutional layer connected in sequence; The third person-vehicle feature map output by the backbone network feature extraction enters the neck module and is divided into two paths. One path enters the sixth feature fusion layer, and the other path enters the first upsampling layer for sampling. Then, it is fused with the second person-vehicle feature map in the third feature fusion layer and enters the fifth feature decomposition layer to obtain the fourth person-vehicle feature map. The fourth person-vehicle feature map is also divided into two paths. One path enters the fifth feature fusion layer, and the other path enters the second upsampling layer for sampling. Then, it is fused with the first person-vehicle feature map in the fourth feature fusion layer and enters the sixth feature decomposition layer to obtain the fifth person-vehicle feature map. The fifth person-vehicle feature map is convolved by the tenth convolutional layer and then enters the fifth feature fusion layer to be fused with the fourth person-vehicle feature map. The fusion result enters the first self-calibration convolutional layer to output the sixth person-vehicle feature map. The sixth person-vehicle feature map is convolved by the eleventh convolutional layer and then enters the sixth feature fusion layer to be fused with the third person-vehicle feature map. The fusion result enters the second self-calibration convolutional layer to output the seventh person-vehicle feature map.

8. The lightweight dual-modal fusion system for RGB-infrared target detection according to claim 1, characterized in that: The detection head module includes two independent detection branches: a position detection branch and a category detection branch; the input human and vehicle feature maps enter the two detection branches respectively, and output target position information and target category information.

Citation Information

Patent Citations

  • Infrared and visible light image perception enhancement fusion method and system based on deep learning

    CN119722492A

Cited By

  • Infrared image target detection network

    CN121725342A