Infrared and visible light image fusion detection method and system

By constructing a fusion detection model and a two-stage training strategy, combining environmental perception and dynamic feature fusion, the problem of low performance of all-weather and all-scene object detection in traffic monitoring scenarios is solved, and real-time and efficient infrared and visible image fusion detection is achieved.

CN120298833APending Publication Date: 2025-07-11BEIJING SINOITS TECH

Patent Information

Application Number
CN202510283632.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing technology cannot achieve all-weather and full-scene target detection in traffic monitoring scenarios, especially in special scenarios such as light-free night, rainy days, fog, and snowy days, and the existing data set is not wide and unbalanced enough. The computing requirements of the fusion detection algorithm are high and cannot meet the real-time requirements.

Method used

A fusion detection model is built, a two-stage training strategy and a dual-stream fusion detection model based on YOLOv9 are adopted, and a dual-stream fusion detection model is based on YOLOv9. Combining the environment perception module and dynamic feature fusion module, image registration is carried out through CycleGAN and SIFT algorithms, fusion strategy is dynamically adjusted to solve the modal loss problem, and a dual cross attention Transformer interaction strategy is introduced to optimize feature fusion.

Benefits of technology

It improves the detection performance of the model in actual engineering applications, meets the needs of real-time and accuracy, solves the problems of modal loss and calculation complexity, and realizes object detection for all-weather and all-scene scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298833A_ABST
    Figure CN120298833A_ABST
Patent Text Reader

Abstract

The invention provides an infrared and visible light image fusion detection method and system, and the method comprises the following steps: constructing a fusion detection model based on a fusion data set; and performing image fusion detection by using the fusion detection model to obtain an image detection result. According to the technical scheme, the fusion strategy can be dynamically adjusted according to the actual scene, the problem of mode loss is solved, and the detection performance is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] As a core technology in the field of computer vision, object detection based on visible light images has reached a relatively high level. However, in practical applications, the environment is often open and dynamic. Relying solely on visible light images for object detection has obvious limitations. Especially in outdoor special scenarios such as insufficient light, bad weather (such as rain, snow, fog), and occlusion, the accuracy of traffic scene object detection based on visible light images drops significantly. Therefore, it is necessary to use models and algorithms to address these challenges. By integrating the characteristics of different sensors, the image information collected by multiple sensors is complementarily fused to achieve all-weather high-precision object detection.

[0003] As Figure 1 shown, due to its sensitive capture ability of thermal radiation, infrared images exhibit excellent object detection and recognition performance at night or under bad weather conditions, being unaffected by lighting conditions and weather changes. However, the resolution and detail expressiveness of infrared images are relatively weak, and the contrast is also low. In contrast, visible light images, with their rich color and texture information, as well as contrast and details that are more in line with human visual habits, can provide clear and vivid scene reproduction, but their imaging performance is limited under conditions such as low light, bad weather, and occlusion. Therefore, by leveraging the complementary characteristics between infrared and visible light images, the features of infrared and visible light images are fused to achieve all-weather and full-scene object detection.

[0004] 1. Object Detection Technology for Traffic Monitoring Scenarios Against the background of the rapid development of current intelligent transportation and autonomous driving technologies, traffic and road detection technologies have received extensive attention due to their important role in enhancing traffic safety, efficiency, and intelligence. Smart transportation and the Internet of Vehicles have become important areas for promoting the construction of new infrastructure. Because achieving accurate recognition of traffic road targets is crucial for the future development of smart transportation.

[0005] However, most current traffic road object detections are based on visible light and have poor performance in low light, at night, and in weather such as rain, snow, and fog, and cannot achieve accurate recognition in all-weather and full-scene scenarios.

[0006] 2. Background Introduction of Public Datasets in this Field As shown in Table 1, there are currently 13 public datasets related to the field of infrared and visible light fusion. From 2014 to 2022, researchers have proposed a total of 13 public datasets for the fusion of infrared images and visible light images. In terms of acquisition methods, these datasets mainly use the method of placing the equipment on the top of the vehicle for head-on acquisition. Few datasets use the vertical acquisition method of drones, and only the LLVIP dataset uses an oblique upward shooting method. In terms of acquisition scope, there is no dataset that fully covers different light intensities such as daytime, nighttime, and dark light, and different climates such as rainy days, snowy days, and foggy days. In addition, the existing acquisition fields are small, which is not conducive to the research of infrared and visible light image fusion detection. In terms of acquisition scenarios, there are no datasets for traffic road monitoring scenarios in existing research. In addition, with the rapid development of computer vision and hardware equipment in recent years, the original datasets are also in urgent need of updating to better meet current engineering applications.

[0007] Table 1 Public datasets of infrared and visible light images In summary, we found that there are three problems with the current public datasets that need to be solved: Lack of traffic monitoring scene datasets with a broad field of view; Some public data sets were collected a long time ago and are no longer compatible with current software and hardware development; Currently, there is no public dataset that fully covers “day-night-rain-snow-fog” and has balanced samples, which is very unfavorable for the research of infrared and visible light image fusion detection.

[0008] 3. Background of infrared and visible light image fusion detection technology 3.1 Fusion Method In order to make full use of the complementary features between visible light and infrared images, domestic and foreign scholars have conducted research on multimodal fusion detection technology, which can be mainly divided into three categories: pixel-level fusion detection, feature-level fusion detection, and decision-level fusion detection. Figure 2 shown.

[0009] (1) The multimodal detection algorithm based on pixel-level fusion aims to obtain a fusion image with prominent targets and rich textures from multiple registered images according to a certain fusion strategy, and build a detection network to determine whether the detection target exists in the fusion image. For example, Chen et al. designed a total variation fusion model to mine the complementary features between cross-modal images, and then input the fused image into the YOLOv3 network to perform the target detection task. Xie et al. constructed an encoder-decoder structure to autonomously learn heterogeneous deep features, and verified through experiments that multimodal fusion detection has stronger anti-interference and fault tolerance than single-modal detection algorithms. However, such methods must additionally design a fusion module, require multimodal images to be highly aligned, and generally have high network complexity.

[0010] (2) The multi-modal detection algorithm based on feature-level fusion uses a convolutional neural network to extract multi-source complementary features, designs a cross-modal feature fusion module to achieve modal interaction, and finally performs classification and regression tasks on the fused features. For example, Li et al. utilize the excellent global modeling ability of Graph Learning to learn the feature representation of multi-modal data, which can significantly improve the detection accuracy of the network in cases of overlap and occlusion. An Haonan et al. introduce a residual network to capture the complementary features between multi-sensor data to avoid the problem of poor detection performance of the algorithm due to insufficient feature information. Fang et al. propose a cross-modal Transformer module to synchronously perform intra-modal and inter-modal feature fusion to explore the potential interactions between heterogeneous data, which can significantly improve the detection performance of the algorithm.

[0011] (3) The multi-modal detection algorithm based on decision-level fusion independently completes the detection tasks for each modal data, and then aggregates the recognition results of multiple sensors to make a globally optimal decision. For example, Li et al. use a light perception module to assign weights to the detection results of different sensors, and thus output the final decision result using the master-slave detector structure. Geng et al. propose a decision-level fusion detection algorithm based on rule mining to solve the problem that the network cannot detect the target due to the failure of a single sensor in an uncertain information source scenario. However, the decision-level fusion method severs the strong correlation between multi-modal images, making it difficult to fully utilize the complementary features of multi-source information.

[0012] In 2019, Li further explored the mid-term fusion (feature fusion) strategy on Faster R-CNN, and proposed that feature fusion can significantly improve the detection accuracy compared with pixel-level fusion and decision-level fusion. In 2021, the multi-channel feature fusion module (MCFF) proposed by Cao was embedded into the YOLOv4 framework to comprehensively utilize visible light and infrared features under different lighting conditions, verifying the advantages of the feature fusion scheme. Therefore, feature-level fusion has become the default strategy for current multi-modal fusion.

[0013] 3.2 Background of Fusion Detection Technology In 2015, S. Hwang built a data acquisition system for a color camera and a thermal imager, captured a large number of aligned visible light and infrared image pairs, and developed a bimodal object detector based on these data, which combined ACF+T+HOG features and an Adaboost classifier, effectively improving the object detection accuracy. The experimental results show that the fusion model can achieve higher detection performance than the model trained only based on visible light, regardless of day or night.

[0014] With the widespread application of convolutional neural networks (CNNs) in computer vision, researchers have begun to explore infrared and visible light fusion detection methods based on deep learning. In terms of fusion strategies, in 2016, Wagner was the first to introduce the idea of multimodal feature fusion into a CNN object detection model, designed a two-branch network structure, and studied the effects of two strategies, early fusion (pixel-level fusion) and late fusion (decision-level fusion), on detection accuracy. Experiments showed that the decision-level fusion architecture performed better. Subsequently, in 2019, Li further explored the mid-term fusion (feature fusion) strategy on Faster R-CNN, proposing that feature fusion could significantly improve detection accuracy compared to pixel-level fusion and decision-level fusion. In 2021, the multi-channel feature fusion module (MCFF) proposed by Cao was embedded into the YOLOv4 framework to comprehensively utilize visible light and infrared features under different lighting conditions, verifying the advantages of the mid-term fusion scheme. Therefore, feature-level fusion has become the default strategy for current multimodal fusion Through the joint efforts of previous researchers, the impact of different fusion positions on the performance of fusion detection algorithms has also been clarified. Then, how to more effectively fuse the information of infrared and visible light modalities has become a new research direction. In 2023, Yan designed the Complementary Cross-modal Information Fusion Network (CCIFNet), proposed a cross-modal fusion mechanism, which can effectively fuse the feature information of different modalities and maintain the spatial relationship between modalities during the feature extraction stage. Zhang improved the utilization rate of multi-spectral features and enhanced the accuracy of object detection through the Feature Recursive Fusion Refinement Module. Xing proposed the multi-spectral detection model MS-DETR for the problems of modality misalignment and modality imbalance commonly existing in multi-spectral detection. This model consists of two modality-specific backbones, a Transformer encoder, and a modality fusion Transformer decoder. To effectively address the misalignment challenge between modality images, a loose-coupling fusion strategy was designed, which sparsely samples key points from multi-modal features independently and performs fusion with the help of adaptively learned attention weights. In addition, a novel instance-aware modality balance optimization strategy was introduced to accurately measure and adjust the contribution degree of each modality. Zhang deeply studied the potential impact of noise on detection performance when fusing different modality feature maps and found that enhancing feature contrast is the key to alleviating the misdetection phenomenon, and proposed a multi-spectral detection algorithm based on the object-aware fusion strategy. This algorithm can adaptively highlight the features closely related to the target, while effectively suppressing irrelevant background and noise features, thus generating more discriminative fusion features. This model achieves state-of-the-art performance while having comparable efficiency to similar algorithms. In 2024, Fu et al. proposed a fast single-stage detector YOLO-Adaptor to solve the problem of non-aligned visible and infrared object detection with complex deviations of translation, scaling, and rotation. It introduced a lightweight multi-modal adaptor to predict alignment parameters and confidence weights simultaneously. It showed good performance in weakly aligned infrared and visible light datasets and contributed to solving the problem that the research in this field overly relies on aligned image pairs. In the same year, Xiao et al. proposed the Generalized Multi-Spectral Detection Transformer (GM-DETR) model to deeply explore the diverse potential of data in this field and enhance the adaptability of the model in actual application scenarios. This model was specifically designed with a Modality-Specific Feature Interaction (MSFI) module to extract deep-level information from RGB and IR images. In addition, they also proposed a Cross-Modal Scale Feature Fusion (CMSF) module to integrate data of RGB and IR modalities. The CMSF module can perform multi-scale cross-modal information fusion, thus optimizing the feature extraction process and improving the performance and accuracy of the model in complex environments.

[0015] In summary, although domestic and foreign researchers have done a lot of research on infrared image and visible light image fusion detection, it still cannot meet the target detection tasks of all-weather and all-scenario traffic monitoring scenarios, especially in special scenes such as night without light, rainy days, foggy days, snowy days, etc. The detection performance is very low. Its main shortcomings include: a. Dataset There are currently three problems with the dataset in this field, as shown in Table 1 above: (1) There is a lack of traffic monitoring scene datasets with a broad field of view; (2) Some public data sets were collected a long time ago and are no longer compatible with current software and hardware development; (3) Currently, there is no public dataset that fully covers “day-night-rain-snow-fog” and has balanced samples, which is very unfavorable for the research on infrared and visible light image fusion detection.

[0016] b. Traffic target detection Currently, most traffic road target detection is based on visible light, and its performance is poor in scenes such as dim light, night, rain, snow, and fog.

[0017] c. Infrared and visible light image fusion detection technology Although many projects currently use infrared and visible light image fusion detection technology, the current technology has the following problems: (1) Some projects use decision-level fusion, which has high computing requirements and cannot meet the real-time requirements of engineering applications. It also breaks the strong correlation between multimodal images. (2) In the research of feature fusion, the variability of actual traffic scenes is not considered. Therefore, the present invention realizes real-time dynamic fusion according to the actual scene, and adopts different fusion strategies for different scenes and weather conditions.

[0018] (3) Currently, most of the research in this field focuses on the study of fusion strategies, ignoring the fact that there are few data that meet the training conditions in this field and the problem of mode loss in actual engineering. Therefore, the present invention proposes a two-stage training strategy.

[0019] (4) Currently, most technologies in this field consider the indicators after fusion. The network structure is complex and the computing requirements are high, while the real-time requirements in engineering are ignored. Summary of the invention

[0020] The present application provides an infrared and visible light image fusion detection method and system to improve the performance of image detection.

[0021] In a first aspect, a method for infrared and visible light image fusion detection is provided, comprising the following steps: Build a fusion detection model based on the fusion dataset; Use the fusion detection model to perform image fusion detection to obtain an image detection result.

[0022] In the above technical solution, by building a fusion detection model based on the fusion dataset; using the fusion detection model to perform image fusion detection to obtain an image detection result; it is possible to dynamically adjust the fusion strategy according to the actual scenario, solve the problem of modality loss, and better meet the actual engineering application in terms of model size, speed, accuracy, and reliability, and the detection performance has been significantly improved.

[0023] In a specific feasible implementation, it further includes: Build an infrared and visible light image fusion detection dataset to obtain the fusion dataset.

[0024] In a specific feasible implementation, it further includes: Design a training strategy for image fusion detection based on the fusion dataset to obtain a two-stage training strategy.

[0025] In a specific feasible implementation, use a registration algorithm based on infrared and visible light images to construct the fusion dataset.

[0026] In a specific feasible implementation, the registration algorithm based on infrared and visible light images specifically includes: Input the visible light image and its grayscale image into the CycleGAN network to generate a pseudo-infrared image; Use the SIFT algorithm to extract and match feature points of the pseudo-infrared image and the real infrared image; Derive a spatial transformation model between the visible light image and the infrared image according to the feature point matching relationship between the pseudo-infrared image and the real infrared image.

[0027] In a specific feasible implementation, the two-stage training strategy includes: First, in the first stage, use independent infrared and visible light data, mix the two modalities together without considering whether the paired infrared and visible light images are aligned, and then copy the same infrared and visible light data and train them simultaneously to obtain the weight result obtained from the first-stage training; In the second stage, formally enter the training of aligned / weakly aligned infrared and visible light image fusion detection, and use the weight result obtained from the first-stage training as the pre-training weight for the second stage.

[0028] In a specific feasible implementation, the fusion detection model includes a two-stream fusion detection model based on YOLOv9.

[0029] In a specific feasible implementation, the dual-stream fusion detection model includes: An environmental perception module for real-time judgment of the current lighting probability and calculation of the infrared and visible light modality weights; A dynamic feature fusion module for introducing a dual-cross attention Transformer interaction strategy and only using the auxiliary modality to solve for the correlation.

[0030] In a specific feasible implementation, the weight calculation formula of the environmental perception module is: W d = P d / (P d +P n ); W n =P n / (P d +P n ); Wherein, Wd represents the probability of the first modality; Wn represents the probability of the second modality; Pd represents the original predicted value or score of the first modality calculated through the environmental perception network; Pn represents the original predicted value or score of the second modality calculated through the environmental perception network.

[0031] Second, an infrared and visible light image fusion detection system is provided, including: A dataset module for constructing an infrared and visible light image fusion detection dataset to obtain a fusion dataset; A training strategy module for designing a training strategy for image fusion detection based on the fusion dataset to obtain a two-stage training strategy; A fusion detection model module for constructing a fusion detection model based on the fusion dataset; A fusion detection module for performing image fusion detection using the fusion detection model to obtain an image detection result.

[0032] In the above technical solution, by constructing a fusion detection model based on the fusion dataset and performing image fusion detection using the fusion detection model to obtain an image detection result, the fusion strategy can be dynamically adjusted according to the actual scenario, the problem of modality loss can be solved, and it better meets the actual engineering applications in terms of model size, speed, accuracy, and reliability, and the detection performance has been significantly improved. Brief Description of the Drawings

[0033] Figure 1 It is a visualization schematic diagram of paired infrared images and visible light images provided by an embodiment of the present application; wherein, the first column is the visible light image, and the second column is the infrared image; Figure 2Schematic diagram of the multi-modal fusion detection type provided by the embodiment of the present application; Figure 3 Specific flowchart of the infrared and visible light image fusion detection method provided by the embodiment of the present application; Figure 4 Schematic diagram of the structure of the data set acquisition device provided by the embodiment of the present application; Figure 5 Schematic diagram of partial visualization of the infrared and visible light image fusion detection data set constructed by the embodiment of the present application; Figure 6 Schematic flowchart of the infrared and visible light image registration algorithm provided by the embodiment of the present application; Figure 7 Schematic diagram of the structure of the infrared and visible light image dynamic fusion detection network provided by the embodiment of the present application; Figure 8 Schematic diagram of the structure of the lighting perception module provided by the embodiment of the present application; Figure 9 Schematic diagram of the infrared and visible light feature fusion structure provided by the embodiment of the present application; Figure 10 Flowchart of the infrared and visible light image fusion detection method provided by the embodiment of the present application; Figure 11 Structural block diagram of another infrared and visible light image fusion detection method provided by the embodiment of the present application. Detailed implementation mode

[0034] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present application will become more clearly defined.

[0035] The special term "exemplary" here means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0036] In addition, the technical features involved in different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0037] To facilitate understanding of the infrared and visible light image fusion detection method and system provided in the embodiment of the present application, its application scenario is first explained. The infrared and visible light image fusion detection method and system provided in the embodiment of the present application are used to improve the performance of image detection. Although domestic and foreign researchers have done a lot of research on infrared image and visible light image fusion detection, it is still unable to meet the target detection task of all-weather and all-scenario traffic monitoring scenes, especially in special scenes such as night without light, rainy days, foggy days, snowy days, etc. The detection performance is very low. Its shortcomings mainly include: a. In terms of data sets, there are currently the following three problems in the data sets in this field: (1) There is a lack of traffic monitoring scene data sets with a wide field of view; (2) Some public data sets were collected a long time ago and are no longer compatible with the current software and hardware development; (3) The current public data sets do not fully cover the "day-night-rain-snow-fog" and sample-balanced data sets, which is very unfavorable for the research on infrared and visible light image fusion detection. b. In terms of traffic target detection, most of the current traffic road target detection is based on visible light, and the performance is poor in scenes such as dark light, night, rain, snow and fog. c. In terms of infrared and visible light image fusion detection technology, although many projects currently use infrared and visible light image fusion detection technology, the current technology has the following problems: (1) Some projects use decision-level fusion, which has large computing requirements and cannot meet the real-time requirements of engineering applications, and also breaks the strong correlation between multi-modal images. (2) In the research on feature fusion, the variability of actual traffic scenes is not considered. Therefore, the present invention realizes real-time dynamic fusion according to the actual scene, and adopts different fusion strategies for different scenes and weather. (3) At present, most of the research in this field focuses on the research of fusion strategies, ignoring the fact that there are few data that meet the training conditions in this field and the problem of modality loss in actual engineering. Therefore, the present invention proposes a two-stage training strategy. (4) At present, most of the technologies in this field consider the indicators after fusion, the network structure is complex, the computing requirements are high, and the real-time requirements in engineering are ignored. For this reason, the embodiment of the present application provides an infrared and visible light image fusion detection method and system to improve the performance of image detection. The following is a detailed description of the embodiment in conjunction with specific drawings.

[0038] refer to Figures 1 to 11 , Figure 1 A visualization diagram of a pair of infrared images and visible light images provided in an embodiment of the present application; wherein the first column is a visible light image and the second column is an infrared image; Figure 2 A schematic diagram of a multimodal fusion detection type provided in an embodiment of the present application; Figure 3 A specific flow chart of the infrared and visible light image fusion detection method provided in the embodiment of the present application; Figure 4 A schematic diagram of the structure of a data set acquisition device provided in an embodiment of the present application; Figure 5Schematic diagram of partial visualization of the infrared and visible light image fusion detection dataset constructed for the embodiments of this application; Figure 6 Schematic flowchart of the infrared and visible light image registration algorithm provided for the embodiments of this application; Figure 7 Schematic diagram of the infrared and visible light image dynamic fusion detection network structure provided for the embodiments of this application; Figure 8 Schematic diagram of the structure of the lighting perception module provided for the embodiments of this application; Figure 9 Schematic diagram of the infrared and visible light feature fusion structure provided for the embodiments of this application; Figure 10 Flowchart of the infrared and visible light image fusion detection method provided for the embodiments of this application; Figure 11 Structural block diagram of another infrared and visible light image fusion detection method provided for the embodiments of this application.

[0039] In Figures 3 to 10 the embodiments of this application provide an infrared and visible light image fusion detection method, including the following steps: Based on the fusion dataset, construct a fusion detection model; Use the fusion detection model to perform image fusion detection to obtain an image detection result.

[0040] In the above technical solution, by constructing a fusion detection model based on the fusion dataset; using the fusion detection model to perform image fusion detection to obtain an image detection result; it is possible to dynamically adjust the fusion strategy according to the actual scenario, solve the problem of modality loss, and better meet the actual engineering application in terms of model size, speed, accuracy, and reliability, and the detection performance has been significantly improved.

[0041] In a specific feasible implementation, it further includes: Construct an infrared and visible light image fusion detection dataset to obtain the fusion dataset.

[0042] In a specific feasible implementation, it further includes: Based on the fusion dataset, design a training strategy for image fusion detection to obtain a two-stage training strategy.

[0043] In a specific feasible implementation, use a registration algorithm based on infrared and visible light images to construct the fusion dataset.

[0044] In a specific feasible implementation, the registration algorithm based on infrared and visible light images specifically includes: Input the visible light image and its grayscale image into the CycleGAN network to generate a pseudo-infrared image; Use the SIFT algorithm to extract and match feature points between the pseudo-infrared image and the real infrared image; Based on the feature point matching relationship between the pseudo-infrared image and the real infrared image, a spatial transformation model between the visible light image and the infrared image is derived.

[0045] In a specific implementable embodiment, the two-stage training strategy includes: First, in the first stage, independent infrared and visible light data are used. The two modalities are mixed together without considering whether the paired infrared and visible light images are aligned. Then, a copy of the same infrared and visible light data is made, and training is carried out simultaneously to obtain the weight results obtained from the first-stage training. In the second stage, formal training for the alignment / weak alignment of infrared and visible light image fusion detection is carried out, and the weight results obtained from the first-stage training are used as the pre-training weights for the second stage.

[0046] In a specific implementable embodiment, the fusion detection model includes a dual-stream fusion detection model based on YOLOv9.

[0047] In a specific implementable embodiment, the dual-stream fusion detection model includes: An environment perception module for judging the current lighting probability in real time and calculating the infrared and visible light modality weights. A dynamic feature fusion module for introducing a dual cross-attention Transformer interaction strategy and only using the auxiliary modality to solve the correlation.

[0048] In a specific implementable embodiment, the weight calculation formula of the environment perception module is: W d = P d / (P d +P n ); W n =P n / (P d +P n ); Where, Wd represents the probability of the first modality; Wn represents the probability of the second modality; Pd represents the original predicted value or score of the first modality calculated by the environment perception network; Pn represents the original predicted value or score of the second modality calculated by the environment perception network.

[0049] Specifically, referring to Figures 3 to 9 , to solve the problem that when only visible light is used for detection in the current traffic monitoring scenario, in the face of changing environments, the efficiency is poor and the working requirements of all-weather and full-scene cannot be achieved. The embodiment of the present application proposes a method for real-time dynamic fusion detection based on infrared and visible light images, as shown in Figure 3 , which includes: dataset construction, training strategy, and fusion detection network.

[0050] 1. Dataset Acquisition and Registration To address the problems existing in the publicly available datasets in the field of infrared and visible image fusion detection, the device shown in Figure 4 was used to construct an infrared and visible image fusion detection dataset ZS - BUPT with "richer scenes - wider fields of view - balanced samples" for traffic monitoring scenarios, as shown in Figure 5 .

[0051] Meanwhile, during the construction of the dataset, a registration algorithm based on infrared and visible images was proposed. As shown in Figure 6 , by constructing a Cycle - Consistent Generative Adversarial Network (CycleGAN), style transfer from visible images to the infrared image domain was achieved, thereby weakening the feature differences between multi - source image data and improving the accuracy of registration. The whole method is divided into the following three main stages: First, the visible image and its grayscale image are input into the CycleGAN network to generate a pseudo - infrared image. CycleGAN consists of two generators and two discriminators. Among them, the generator is responsible for migrating the visible image to the infrared domain and ensuring that the pseudo - infrared image maintains the original structural features; the discriminator improves the authenticity of the generated image through adversarial training. To ensure the stability of the style transfer process, the network introduces a cycle - consistency loss to constrain the pseudo - infrared image to be able to restore the original image after being transformed back to the visible domain.

[0052] Second, the Scale - Invariant Feature Transform (SIFT) algorithm is used to extract and match feature points from the pseudo - infrared image and the real infrared image. Specifically, the key points and descriptors of the pseudo - infrared image and the infrared image are extracted by SIFT, and the initially matched feature point pairs are screened using the nearest - neighbor ratio test. To further eliminate false matches, the homography matrix between the feature point pairs is calculated, thereby improving the robustness and accuracy of the matching.

[0053] Finally, according to the feature point matching relationship between the pseudo - infrared image and the real infrared image, the spatial transformation model between the visible image and the infrared image is derived. The feature points in the pseudo - infrared image are mapped back to the visible image through the homography matrix to achieve the unification of feature points in the visible and infrared domains, and the homography matrix is used to complete the registration of the infrared image to the visible image.

[0054] 2. Two - Stage Training Strategy Currently, in the research field of dual - spectrum image fusion detection, there are two problems with the training methods: (1) Limiting the infrared and visible fusion detection model to be trained only on Visible and Infrared images that are temporally and spatially aligned / weakly aligned greatly restricts the available training data for the model.

[0055] (2) Existing methods directly fuse infrared and visible light data during the training process. This training strategy enables the model to only process the ability of infrared and visible light fusion data, hindering the comprehensive understanding of each modal feature separately. Considering the actual application scenarios, when the multi-spectral detection model encounters modal loss in the input data, the model performance will significantly decline.

[0056] To solve the above two problems, a two-stage training strategy is designed by referring to other research fields. First, in the first stage, independent infrared and visible light data are used. The two modalities are mixed together without considering whether the paired infrared and visible light images are aligned. Then, a copy of the same data is made and sent into the network simultaneously to improve the network's learning of the understanding of single modalities. In the second stage, the training of the aligned / weakly aligned infrared and visible light image fusion detection is formally carried out, and the weight results obtained from the first stage training are used as the pre-training weights for the second stage.

[0057] 3. Infrared and Visible Light Image Fusion Detection Algorithm Based on YOLOv9, a two-stream fusion detection model is proposed, which inputs paired infrared and visible light image data simultaneously, and an environmental perception module and an Infrared and Rgb Feature Fusion (IRFF) module are proposed and embedded in two feature extraction branches.

[0058] As Figure 7 shown in the network structure, the infrared and visible light images are input simultaneously. The feature maps at three scales of C3, C4, and C5 are respectively fused through the IRFF module. During the fusion process, the environmental perception module calculates the illumination probability of the current image, and the weights of the two modalities are calculated through formulas (1) and (2), and the fusion weights are adjusted in real time to make the model pay more attention to clear modal features. Finally, three fusion feature maps at different scales are output and enter the Neck and Detection Head parts.

[0059] Environmental Perception Module: In the designed network, the fusion detection performance highly depends on the accuracy of environmental perception. As Figure 8 shown, the input of the environmental perception module is the visible light image, and the output is the probabilities of the two modalities, and W d and W n are calculated through the following formula. The environmental perception network consists of four convolutional layers, a global average pooling layer responsible for integrating illumination information, and two fully connected layers for calculating illumination probabilities.

[0060] W d = P d / (P d + P n ) (1) Wn =P n / (P d +P n ) (2) Infrared and Visible Light Feature Fusion (IRFF) Module: As Figure 9 shown, the infrared and visible light feature fusion module is mainly responsible for aggregating the feature information of the infrared and visible light modalities from both local and global perspectives. First, to meet the requirements of engineering applications for computational costs, the RIFF module first compresses the features of the two modalities using convolutional and pooling operations. Then, drawing on the idea of Vision Transformer, a Feature Interaction (Infrared and Rgb Feature Interaction - IRFI) module is designed, enabling a single modality to learn more complementary information from the auxiliary modality from a global perspective and overcoming the deficiencies in modeling the long-range dependencies of cross-modal features. During multiple iterations, enhanced infrared and visible light image feature maps are obtained. Finally, the enhanced infrared features and visible light features are fused by addition to obtain the fused feature map Ffused that aggregates the high-quality features of the two modalities, which is fed into the Neck and Head.

[0061] The real-time dynamic fusion algorithm for infrared and visible light images proposed in the present invention is verified on the ZS-BUPT dataset, and its performance is greatly improved compared with single-modal detection, demonstrating the value of the present invention in practical applications. Table 2 Experimental Verification Data Table of This Application In the above technical solution, the beneficial effects of the infrared and visible light image fusion detection method include: (1) For the traffic monitoring scenario, an infrared and visible light image fusion detection dataset of "richer scenes - wider field of view - balanced samples" is constructed, solving problems existing in all currently public datasets in this field, such as "lack of wide-field datasets for traffic monitoring scenarios", "public datasets do not fully cover white / night / rain / snow / fog scenarios", and "mismatch between shooting devices and the development of software and hardware in recent years".

[0062] (2) For the field of multi-source image fusion, a two-stage training strategy is proposed, solving the problems of less registration data for infrared and visible light image fusion detection training and the loss of a single modality in actual engineering applications.

[0063] (3) The real-time dynamic fusion detection algorithm for infrared and visible light images in all-weather and all-scene proposed by the present invention can dynamically adjust the fusion strategy according to the actual scene and solve problems such as modality loss. Compared with the prior art, it better meets the actual engineering applications in terms of model size, speed, accuracy, and reliability, including two innovation points: "environmental perception" and "dynamic feature fusion". Through experimental verification, the detection performance has been significantly improved compared with the current mainstream methods.

[0064] Environmental perception: Real-time judgment of the current lighting probability, calculation of the weights of the two modalities, and guidance of the attention degree of the fusion network to different modality features.

[0065] Dynamic feature fusion: Introduction of a double-cross attention Transformer interaction strategy, and only using the Q of the auxiliary modality to solve the correlation. On the premise of aggregating the features of infrared and visible light images from both local and global perspectives, it overcomes the deficiency in modeling the long-range dependence of cross-modal features and also ensures the requirements of load and memory.

[0066] In Figure 11 , the embodiments of the present application provide an infrared and visible light image fusion detection system, including: A dataset module for constructing an infrared and visible light image fusion detection dataset to obtain a fusion dataset; A training strategy module for designing a training strategy for image fusion detection based on the fusion dataset to obtain a two-stage training strategy; A fusion detection model module for constructing a fusion detection model based on the fusion dataset; A fusion detection module for using the fusion detection model to perform image fusion detection to obtain an image detection result.

[0067] In the above technical solution, by constructing a fusion detection model based on the fusion dataset; using the fusion detection model to perform image fusion detection to obtain an image detection result; it can dynamically adjust the fusion strategy according to the actual scene, solve the problem of modality loss, better meet the actual engineering applications in terms of model size, speed, accuracy, and reliability, and the detection performance has been significantly improved.

[0068] Those skilled in the art of the technical field know that the present application can be implemented as a system, a method, or a computer program product.

[0069] Accordingly, the present disclosure may be embodied in the following forms, namely: it may be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as "circuitry", "module" or "system". Additionally, in some embodiments, the present application may also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.

[0070] Any combination of one or more computer-readable media may be used. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0071] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations of the present application. Those of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application. On this basis, various substitutions and improvements can be made to the present application, all of which fall within the protection scope of the present application.

Claims

1. An infrared and visible light image fusion detection method, characterized in that, It includes the following steps: Based on the fused dataset, construct a fused detection model; Use the fused detection model to perform image fusion detection to obtain an image detection result.

2. The infrared and visible light image fusion detection method according to claim 1, wherein It also includes: Construct an infrared and visible light image fusion detection dataset to obtain the fused dataset.

3. The infrared and visible light image fusion detection method according to claim 2, wherein It also includes: Based on the fused dataset, design a training strategy for image fusion detection to obtain a two-stage training strategy.

4. The infrared and visible light image fusion detection method according to claim 3, wherein, Use a registration algorithm based on infrared and visible light images to construct the fused dataset.

5. The infrared and visible light image fusion detection method according to claim 4, characterized in that The registration algorithm based on infrared and visible light images specifically includes: Input the visible light image and its grayscale image into the CycleGAN network to generate a pseudo-infrared image; Use the SIFT algorithm to extract and match feature points of the pseudo-infrared image and the real infrared image; Based on the feature point matching relationship between the pseudo-infrared image and the real infrared image, deduce the spatial transformation model between the visible light image and the infrared image.

6. The infrared and visible light image fusion detection method according to claim 5, wherein The two-stage training strategy includes: First, in the first stage, use independent infrared and visible light data, mix the two modalities together without considering whether the paired infrared and visible light images are aligned, and then copy the same infrared and visible light data and perform training simultaneously to obtain the weight result obtained from the first-stage training; In the second stage, officially enter the training of aligned / weakly aligned infrared and visible light image fusion detection, and use the weight result obtained from the first-stage training as the pre-training weight for the second stage.

7. The infrared and visible light image fusion detection method according to claim 6, characterized in that The fused detection model includes a two-stream fused detection model based on YOLOv9.

8. The infrared and visible light image fusion detection method according to claim 7, wherein The two-stream fused detection model includes: An environment perception module for real-time judging the current illumination probability and calculating the infrared and visible light modality weights; A dynamic feature fusion module for introducing a double cross-attention Transformer interaction strategy and only using the auxiliary modality to solve the correlation.

9. The infrared and visible light image fusion detection method according to claim 8, wherein The weight calculation formula of the environment perception module is: W d = P d / (P d +P n ); W n =P n / (P d +P n ); Where, Wd represents the probability of the first modality; Wn represents the probability of the second modality; Pd represents the original prediction value or score of the first modality calculated by the environment perception network; Pn represents the original prediction value or score of the second modality calculated by the environment perception network.

10. An infrared and visible light image fusion detection system, characterized in that, It includes: A dataset module for constructing an infrared and visible light image fusion detection dataset to obtain a fused dataset; A training strategy module for designing a training strategy for image fusion detection based on the fused dataset to obtain a two-stage training strategy; A fused detection model module for constructing a fused detection model based on the fused dataset; A fused detection module for using the fused detection model to perform image fusion detection to obtain an image detection result.

Citation Information

Patent Citations

  • Infrared light and visible light image fusion method combining target detection

    CN116188342A

  • Image-based big data analysis method

    CN117333409A

  • Multispectral target detection method and device based on feature fusion and electronic equipment

    CN119445306A

Cited By

  • Abnormality detection method and device for fusion of visible light image and infrared image

    CN121280860A