Infrared and visible light image fusion network based on illumination intensity perception

Through the infrared and visible image fusion network based on light intensity perception, the problem that the prior art is difficult to retain high-frequency and global features under different lighting conditions is solved, and image quality improvement and semantic information are achieved.

CN119942296AInactive Publication Date: 2025-05-06HOHAI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411902698.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing infrared and visible image fusion technology processes images under different lighting conditions and complex environments, it is difficult to effectively retain high-frequency and global features, and it is difficult to provide detailed semantic information for downstream visual tasks.

Method used

An infrared and visible image fusion network based on light intensity perception is adopted, including an image registration module, a feature fusion module, a light perception module and a semantic extraction and injection module. Through the coordinated work of these modules, efficient fusion of images and semantic information extraction are achieved.

Benefits of technology

Under different lighting conditions, the high-frequency and global features in the fusion result can be better preserved, image quality can be improved, and semantic information extraction capabilities of downstream visual tasks can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942296A_ABST
    Figure CN119942296A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared and visible light image fusion network based on illumination intensity perception, belongs to an artificial intelligence technology, and aims to solve the problems that in an image fusion task, registration is often needed as a preprocessing link before fusion, and a registration effect also has direct influence on a final result of image fusion. Therefore, an automatic image registration algorithm based on common interested sub-regions is used as a preprocessing step before fusion. In addition, in consideration of most image fusion methods, fusion is directly carried out in a spatial domain, and an existing fusion method based on deep learning often carries out feature extraction and fusion on a single scale, so that rich multi-scale feature information in a source image is ignored and adaptability is lacked. Therefore, in the aspect of a fusion network, the invention provides a space-frequency domain feature fusion method for fusing Fourier transform guided fusion branches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the infrared and visible light image fusion technology, which aims to improve the quality, details and information content of the image, and is widely used in many fields, especially in low light, complex environment and night monitoring scenes. Specifically, it is an infrared and visible light image fusion network based on light intensity perception. Background Art

[0002] Infrared and visible light image fusion technology is an important pre-processing technology in many visual tasks. Due to the hardware limitations of imaging devices, sensors of a single type or a single setting are usually unable to fully characterize the imaging scene. For example, visible light images usually contain rich texture detail information, but are easily affected by extreme environments and occlusions and lose the target in the scene. In contrast, infrared sensors can effectively highlight pedestrians, vehicles and other prominent targets by capturing the thermal radiation information emitted by objects, but lack detailed descriptions of the scene. It is worth noting that sensors of different types or different optical settings usually contain a lot of complementary information, which also inspires people to integrate this complementary information into a single image. Therefore, image fusion technology came into being.

[0003] Infrared and visible light image fusion technology can overcome the limitations of a single sensor, improve the accuracy of target detection and recognition, improve image quality, etc., and has broad application prospects. However, due to the different external and internal parameters of the camera, it is impossible to directly capture aligned images for fusion. In addition, cross-channel binocular photos may be affected by jitter, delay, and special noise. In particular, infrared cameras may be severely disturbed by internal temperature and external thermal airflow. Therefore, in actual fusion, it is urgent to consider the misalignment problem between the input infrared and visible light images. Moreover, most image fusion methods are fusions performed directly in the spatial domain: such as intensity hue saturation (IHS) and principal component analysis (PCA) of pixel domain images. Although such fusion methods are simple and fast, their fusion effects are difficult to meet performance requirements and it is difficult to provide detailed semantic information for downstream advanced visual tasks.

[0004] In order to solve the problems existing in the above registration algorithm, it is considered to use an infrared and visible light image fusion network based on light intensity perception to better preserve the high-frequency and global features in the fusion results under different lighting conditions, while also preparing for the semantic information extraction of downstream visual tasks. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to provide an infrared and visible light image fusion network based on light intensity perception.

[0006] Technical solution: The infrared and visible light image fusion network based on light intensity perception described in the present invention includes an image registration module, a feature fusion module, a light perception module, and a semantic extraction and injection module. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is an overall schematic diagram of the network structure of the present invention;

[0008] Figure 2 This is a schematic diagram of the registration module of the present invention:

[0009] Figure 3 Schematic diagram of the cross-attention fusion module of the present invention.

Claims

1. An infrared and visible light image fusion network based on light intensity perception, characterized in that: The image fusion network includes an image registration module, an image fusion module, a target detection module, and a light intensity perception module; The specific steps of the fusion method are as follows: Step (1): automatically detect and register the common sub-regions of interest in the two images (target image and floating image), then obtain the transformation parameters from the sub-region registration, and finally apply the transformation parameters to the overall registration of the global image to achieve the final registration of the source images; Step (2): In the fusion stage, an adaptive fusion algorithm is used to extract spatial and frequency domain features from the two outputs of step (1), and the features are fused to obtain adaptive space-frequency fusion features. Step (3): Inject the semantic information extracted from the network learning of high-level visual tasks (such as object detection) into the decoding network through the semantic injection module; Step (4): Considering that light imbalance may affect information distribution, the light intensity perception network will adaptively fuse common and complementary information according to the light conditions, that is, determine the weight of the loss function based on light intensity; Step (5): Finally, back propagation is performed based on the loss function output of step (4) to guide the fusion network and the target detection network to extract features that are more consistent with the current environment and high-level task semantic information.

2. According to the image fusion network based on light intensity perception described in claim 1, it is characterized in that: In step (1), the input floating image is an infrared or visible light image that has been rotated and transformed to different degrees. It mainly includes three steps: first, in the preprocessing stage, a smoothing filter (median filter) is applied to both the source image and the target image to remove noise and other artifacts. If the input image is not grayscale, the original RGB value is replaced with a new grayscale value to convert it into a grayscale image, and then the grayscale image is converted into a binary image through a threshold. Then there is sub-region detection and segmentation. After binary conversion, two functions are applied to the floating image and the target image respectively, one of which divides them into sub-regions, and the other function is used to detect and segment the largest common sub-region in the two images. The mathematical expression is as follows: X(i)=F(I Ms ) i=1upto N Y(i)=F(I FT )i=1upto M C1=Max(X(i))for all i=1upto N C2=Max(Y(i))for all i=1upto M Where MS and FT represent the binary floating image and target image, respectively, and N and M are the sum of the number of all sub-regions of MS and FT. F is a function that uses thresholding and normalized cross correlation (NCC) to obtain all sub-regions, and the Max function is responsible for selecting the largest sub-region with a standard threshold in the floating image and the target image. Finally, a rigid transformation involving rotation and translation is applied to both regions. After estimating the parameters obtained from the registration of the common sub-regions, the parameters are applied to the entire floating image to make the image similar to the original target image.

3. According to the infrared and visible light image fusion network based on light intensity perception described in claim 1, it is characterized in that: The adaptive fusion algorithm in step (2) is a spatial-frequency domain feature fusion method with Fourier transform-guided fusion branches. This method uses a cross-attention fusion module CAF (as shown in Figure 3). This module uses a multi-head self-attention mechanism to adaptively fuse features from different modalities of the source image. The similarity between the two modalities is calculated by exchanging key vectors between the two modalities, and then the similarity score is used for adaptive fusion. This enhances the high-frequency features of different modalities and effectively alleviates the problem of information loss in the fused image. The specific steps are as follows: First, the multi-scale shallow features ψj of the input source image are extracted through the encoding network, and then the fusion features in the spatial domain are obtained from different scales through CAF. The expression is as follows: Considering that the frequency domain can provide complementary information for spatial domain fusion, the image features from different sources are transformed by Fourier transform to obtain the corresponding frequency domain features. The expression is as follows: The frequency domain features of adaptive fusion are obtained through the CAF block. The expression is as follows: After inverse Fourier transform, we get the frequency information. The expression is as follows: In view of the fact that frequency fusion features and spatial fusion features can complement and guide each other, the cross-attention of frequency domain fusion features and spatial domain fusion features are calculated respectively, and adaptive fusion is performed to obtain adaptive fusion features. The expression is as follows: Finally, the decoding network is used to reconstruct the multi-scale fusion features and obtain the fused image.

4. According to the infrared and visible light image fusion network based on light intensity perception as described in claim 1, it is characterized in that: In step (4), considering that illumination imbalance may affect information distribution, a method for determining the weight of the loss function based on perceived illumination intensity is proposed. The goal of the light intensity perception network is to estimate the illumination distribution of the scene. Its input is a visible light image and its output is an illumination probability. The network consists of four convolutional layers, a global average pool, and two fully connected layers. The 4×4 convolutional layer with a step size of 2 compresses and extracts illumination information. After this network, two probabilities (p1, p2<1) can be output as weights of the loss function. The greater the probability of daytime, the greater the loss of structure and content of the visible light image in the fused image, that is, the more sensitive it is to the preservation of the intensity information of the visible light image, and vice versa.

5. According to the infrared and visible light image fusion network based on light intensity perception described in claim 1, it is characterized in that ,In step (5), back propagation is performed based on the loss function output of step (4) to guide the behavior of steps (3) and (2). A new image fusion method is proposed, which combines image registration, environmental perception, image fusion and semantic requirements for high-level visual tasks into one framework.

Citation Information

Patent Citations

  • Unlocking method based on face image processing and intelligent lock

    CN110751757A

  • Image fusion method and device and storage medium

    CN114372948A

  • High-pressure sterilization cabinet temperature monitoring device based on infrared thermal imaging

    CN115468658A

  • High-order interaction synergistic visible light and infrared image fusion method

    CN118429199A

  • Infrared and visible light image lightweight fusion method based on spatial domain and frequency domain information

    CN118864266A