Infrared and visible image fusion method and device based on frequency domain decoupling

CN122243768BActive Publication Date: 2026-08-18CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610702829.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-18
Estimated Expiration
2046-05-21

AI Technical Summary

Technical Problem

然而,在实际部署过程中,融合系统常面临复杂多变的光照环境挑战,例如夜间或雨雾等恶劣天气条件下,可见光图像信号强度显著降低且噪声干扰严重;而在强日光照射、逆光场景或局部明暗对比剧烈的环境中,可见光图像又易出现过曝区域或细节信息缺失现象

Benefits of technology

利用可见光图像增强与红外图像互补的双路径架构,能够提高鲁棒性,提高融合图像的融合精度;而且,在可见光编码器分支引入光照增强与空间频域解耦机制:首先利用频域解耦机制将图像解耦为幅度谱与相位谱,在保持相位结构不变的前提下实现幅度分量的动态校正,从而在不破坏纹理细节的情况下校正光照,这样有效提升复杂光照场景下的融合图像质量,在保留红外热辐射目标的同时,能够根据环境光照恢复可见光纹理细节,然后通过从第一图像中提取全局光照因子和局部亮度特征,并生成光照权重图,使得融合过程能够动态感知并响应环境光照的变化,通过分别提取第一图像的多尺度空间域特征图和多尺度频域特征图,并将其与光照权重图进行加权融合,实现了对可见光模态的差异化处理,这种处理方式使得可见光图像的频域信息能够根据其在不同光照条件下的可靠性进行调整,从而在融合过程中更好地利用可见光图像的有效信息,有效解决了现有技术中融合系统在动态光照环境下模态失衡、缺乏环境光感知能力以及未能对可见光模态特有退化进行差异化处理的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122243768B_ABST
    Figure CN122243768B_ABST
Patent Text Reader

Abstract

The application relates to an infrared and visible light image fusion method and device based on frequency domain decoupling. The method extracts a multi-scale space domain feature map and a multi-scale frequency domain feature map of a first image, extracts a global light factor and a local brightness feature from the first image, generates a light weight map, and performs weighted fusion on the multi-scale space domain feature map and the multi-scale frequency domain feature map, so that the change of the ambient light can be dynamically perceived and responded, the visible light mode is differentially processed, the frequency domain information of the visible light image can be adjusted according to the reliability under different light conditions, the effective information of the visible light image can be better utilized in the fusion process, and the technical problems that the modal imbalance in the existing stage under the dynamic light environment, the lack of ambient light perception ability and the failure to differentially process the visible light mode specific degradation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for fusing infrared and visible light images based on frequency domain decoupling. Background Technology

[0002] Infrared and visible light image fusion technology aims to effectively integrate the thermal radiation information acquired by infrared sensors with the texture detail information captured by visible light sensors to generate fused images with richer information and clearer semantic understanding. This technology has significant value in key application scenarios such as autonomous driving, security monitoring, and military reconnaissance. However, in actual deployment, fusion systems often face challenges from complex and changing lighting environments. For example, at night or in adverse weather conditions such as rain and fog, the signal strength of visible light images is significantly reduced and noise interference is severe. In environments with strong sunlight, backlighting, or dramatic local contrast, visible light images are prone to overexposure or loss of detail. This information imbalance between infrared and visible light modes caused by dynamic lighting changes leads to insufficient stability of the fusion results during lighting transitions, severely limiting the practicality and robustness of the algorithm.

[0003] In existing technical solutions, spatial domain-based fusion methods generally rely on pre-set fixed feature extraction mechanisms and lack the ability to dynamically perceive changes in ambient light. This means that they can only maintain basic performance under specific lighting conditions, and the fusion quality deteriorates rapidly once the lighting conditions fluctuate. Although some studies have attempted to introduce frequency domain analysis to enhance feature representation, they mostly adopt a structurally symmetrical two-stream network framework, which fails to differentiate and optimize the brightness degradation and structural distortion problems unique to visible light modes under lighting fluctuations. Therefore, they cannot fundamentally alleviate the negative impact of modal imbalance. Summary of the Invention

[0004] To achieve the above objectives, a first aspect of the present invention provides a method for infrared and visible light image fusion based on frequency domain decoupling, the method comprising: Acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; The first image and the second image are fused based on a fusion model to obtain a fused image. The process of obtaining the fused image by the fusion model includes: Extract the global illumination factor and local brightness features of the first image from the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image.

[0005] The infrared and visible light image fusion method based on frequency domain decoupling provided in the first aspect of this application has at least the following beneficial effects: By utilizing a dual-path architecture that combines visible light image enhancement with infrared image complementarity, robustness and fusion accuracy of the fused images can be improved. Furthermore, illumination enhancement and spatial... Frequency domain decoupling mechanism: First, the image is decoupled into amplitude spectrum and phase spectrum using a frequency domain decoupling mechanism. Dynamic correction of amplitude component is achieved while maintaining the phase structure, thereby correcting illumination without destroying texture details. This effectively improves the quality of fused images in complex lighting scenarios. While preserving infrared thermal radiation targets, visible light texture details can be recovered according to ambient light. Then, global illumination factors and local brightness features are extracted from the first image, and an illumination weight map is generated, enabling the fusion process to dynamically sense and respond to changes in ambient light. By extracting multi-scale spatial domain feature maps and multi-scale frequency domain feature maps from the first image respectively, and then weighting and fusing them with the illumination weight map, differentiated processing of visible light modes is achieved. This processing method allows the frequency domain information of the visible light image to be adjusted according to its reliability under different lighting conditions, thereby better utilizing the effective information of the visible light image during the fusion process. This effectively solves the technical problems of modal imbalance, lack of ambient light perception capability, and failure to differentiate the degradation specific to visible light modes in existing fusion systems under dynamic lighting environments.

[0006] The second aspect of this application provides an infrared and visible light image fusion device based on frequency domain decoupling, the device comprising: An image acquisition module is used to acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; An image fusion module is used to fuse the first image and the second image based on a fusion model to obtain a fused image, wherein the process of obtaining the fused image by the fusion model includes: Extract the global illumination factor and local brightness features of the first image from the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image.

[0007] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the above-described infrared and visible light image fusion method based on frequency domain decoupling.

[0008] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described infrared and visible light image fusion method based on frequency domain decoupling.

[0009] It is understood that the beneficial effects of the second to fourth aspects compared with the related technologies are the same as the beneficial effects of the first aspect compared with the related technologies. Please refer to the relevant description in the first aspect above, which will not be repeated here. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic flowchart of an infrared and visible light image fusion method based on frequency domain decoupling provided in an embodiment of this application; Figure 2 This is a schematic diagram of the fusion model provided in this application for fusing visible light images and infrared images; Figure 3 This is a schematic diagram of the process for extracting the illumination weight map provided in an embodiment of this application; Figure 4 This is a schematic diagram of the process for extracting frequency domain features provided in an embodiment of this application; Figure 5 This is a structural diagram of an infrared and visible light image fusion device based on frequency domain decoupling provided in an embodiment of this application; Figure 6 This is a structural diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0013] Traditional infrared and visible light image fusion techniques employ fixed feature extraction and mapping strategies, lacking the ability to dynamically perceive ambient lighting conditions. This leads to information imbalance between modalities in extremely variable lighting scenarios. Specifically, visible light images exhibit weak signals and severe noise at night or in adverse weather conditions, while in strong light, backlight, or uneven brightness distribution scenarios, local overexposure or loss of detail occurs. This modal imbalance prevents fusion algorithms from effectively combining the thermal radiation information from infrared sensors with the texture detail information from visible light sensors, thereby reducing the information completeness and semantic understanding capabilities of the fused image.

[0014] To address the aforementioned technical problems, the following embodiments are provided: like Figure 1 This paper presents a method for fusing infrared and visible light images based on frequency domain decoupling, the method comprising: Step S110: Acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; Step S120: Based on the fusion model, the first image and the second image are fused to obtain a fused image. The process of obtaining the fused image using the fusion model includes: Extract the global illumination factor and local brightness features of the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image.

[0015] In step S110, image acquisition can be achieved in various ways. For example, it can be achieved by two independent image sensors capturing images synchronously at the same location and time, one sensor for visible light imaging and the other for infrared imaging. Alternatively, it can be obtained by reading registered visible light and infrared image pairs from a pre-stored image database.

[0016] In step S120, the fusion model can adopt a deep learning-based neural network model, which is designed to receive input images from two modalities and generate a fusion output through multi-layer processing.

[0017] Specifically, the process of obtaining the fused image by the fusion model includes extracting the global illumination factor and local brightness features of the first image, and generating an illumination weight map of the first image based on the global illumination factor and local brightness features.

[0018] Global illumination factors can be extracted by calculating the overall average pixel intensity of the first image, or by downsampling the image and calculating its average value. Local brightness features can be extracted by applying a series of predefined convolution kernels (such as Sobel and Laplacian operators) to capture the edge and texture information of the image, or by calculating the variance of local regions of the image to reflect brightness changes.

[0019] When generating the illumination weight map, linear interpolation or nonlinear mapping functions can be used to combine the extracted global illumination factors and local brightness features to generate a weight map with the same size as the image. The weight values ​​reflect the illumination reliability of the corresponding region.

[0020] For the extraction of spatial domain feature maps, multi-layer convolutional neural networks can be used. By stacking convolutional kernels and pooling layers (downsampling) of different sizes, features at different levels of abstraction can be extracted from the image step by step. For example, shallow networks can extract low-level features such as edges and textures, while deep networks can extract high-level features such as semantic information.

[0021] For extracting frequency domain feature maps, the first image can be converted into a frequency domain representation through Fourier transform, and then the frequency domain representation can be decomposed into multiple scales, for example, by using a frequency domain filter bank to obtain features in different frequency ranges.

[0022] Weighted fusion can be achieved by element-wise multiplying the illumination weight map with the spatial and frequency domain feature maps, or by dynamically adjusting the channel or spatial weights of the feature maps by using the weight map as input to an attention mechanism. For example, in regions with poor illumination conditions, the weight values ​​of the weight map may be lower, thereby reducing the contribution of the visible light features of the region, and vice versa.

[0023] The fusion process can employ various strategies, such as implementing a connection structure similar to U-net, enabling the model to adaptively select and integrate the most informative features from different modalities, thereby generating high-quality fused images.

[0024] The following example will provide a more detailed explanation of the above technical solution: Suppose an intelligent monitoring system is deployed at location A. The system needs to continuously and clearly capture images of the monitored area under various lighting conditions, including daytime, nighttime, and foggy or hazy conditions.

[0025] First, through its integrated visible light sensor and infrared sensor, it simultaneously acquires visible light images (first image) and infrared images (second image) of the monitored area. For example, at night, the visible light image may be very dark and noisy, while the infrared image clearly shows the thermal radiation outline of people or vehicles.

[0026] The fusion model extracts global illumination factors from visible light images, such as calculating the average brightness value of the entire image, to determine whether the current environment is under low light, normal light, or strong light conditions. At the same time, it also extracts local brightness features, such as identifying whether there are local overexposed or shadowed areas in the image, and the brightness changes in these areas. Based on this global and local illumination information, the system generates an illumination weight map.

[0027] Then, visible light and infrared images are processed in parallel. For visible light images, multi-scale spatial domain feature maps and multi-scale frequency domain feature maps are extracted. Spatial domain feature maps capture detailed information such as edges and textures of visible light images, while frequency domain feature maps capture the overall structure and frequency distribution of the image. These feature maps are extracted at different scales to ensure that various information from fine textures to macroscopic contours can be captured. At the same time, for infrared images, multi-scale infrared feature maps are also extracted. These feature maps mainly reflect the thermal radiation distribution and target contours of infrared images.

[0028] In the feature fusion stage, the method performs weighted fusion of the multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image with the previously generated illumination weight map to obtain a multi-scale fused feature map of the first image. For example, when a certain region of the visible light image is overexposed due to strong light, the weight value of the illumination weight map in that region will be lower, thereby reducing the contribution of visible light features from the overexposed region during fusion and avoiding the introduction of distorted visible light information into subsequent fusion. Conversely, in well-lit regions, the weight value is higher, and the contribution of visible light features is preserved. This adaptive weighting mechanism enables visible light features to be intelligently adjusted according to their own quality and reliability.

[0029] The multi-scale fusion feature map of the first image after illumination adjustment is fused with the multi-scale infrared feature map of the second image to obtain the final fused image. The final fused image contains the robust target information of the infrared image under poor lighting conditions, preserves the rich texture details of the visible light image as much as possible, and effectively suppresses the degradation of the visible light image caused by uneven illumination.

[0030] (1) This embodiment utilizes a dual-path architecture that combines visible light image enhancement with infrared image complementarity, which can improve robustness and the fusion accuracy of the fused image; (2) In this embodiment, illumination enhancement and spatial design are introduced into the visible light encoder branch. Frequency domain decoupling mechanism: First, the image is decoupled into amplitude spectrum and phase spectrum using a frequency domain decoupling mechanism. Dynamic correction of amplitude component is achieved while maintaining the phase structure, thereby correcting illumination without destroying texture details. This effectively improves the quality of fused images in complex lighting scenarios. While preserving infrared thermal radiation targets, visible light texture details can be recovered according to ambient light. Then, global illumination factors and local brightness features are extracted from the first image, and an illumination weight map is generated, enabling the fusion process to dynamically sense and respond to changes in ambient light. By extracting multi-scale spatial domain feature maps and multi-scale frequency domain feature maps from the first image respectively, and then weighting and fusing them with the illumination weight map, differentiated processing of visible light modes is achieved. This processing method allows the frequency domain information of the visible light image to be adjusted according to its reliability under different lighting conditions, thereby better utilizing the effective information of the visible light image during the fusion process. This effectively solves the technical problems of modal imbalance, lack of ambient light perception capability, and failure to differentiate the degradation specific to visible light modes in existing fusion systems under dynamic lighting environments. In some embodiments of this application, extracting the global illumination factor and local brightness features of the first image from the first image includes: Extract the luminance component from the first image.

[0031] The global illumination factor is extracted from the luminance component using a global average pooling layer and a multilayer perceptron.

[0032] Local brightness features are extracted from the brightness component using convolutional layers.

[0033] An illumination weight map of the first image is generated based on the global illumination factor and local brightness features of the first image, including: The global illumination factor and local brightness features of the first image are weighted, fused, and activated to obtain the illumination weight map of the first image.

[0034] Extracting the luminance component from the first image refers to obtaining information reflecting light intensity or brightness from the image. This step aims to decouple the image's luminance information from its color information, making the analysis of lighting characteristics more focused and accurate.

[0035] Global illumination factors are extracted from the luminance component using a global average pooling layer and a multilayer perceptron. The global illumination factor is a single or small number of values ​​describing the average luminance level or illumination intensity of the entire image or most of its area. The global average pooling layer is a neural network layer that averages the spatial dimension of each channel of the feature map into a single value. The multilayer perceptron is a feedforward neural network composed of multiple fully connected layers capable of learning complex nonlinear mappings. This combination aims to capture overall illumination information from the luminance component, compress spatial information through the global average pooling layer, and then learn the complex relationship between this information and the global illumination factor through the multilayer perceptron, thus obtaining a factor that represents the overall illumination level of the image.

[0036] Convolutional layers extract local brightness features from the brightness component. These local brightness features refer to the brightness variation patterns and detailed information in different regions of an image. Convolutional layers are a core component for feature extraction in deep learning. By sliding convolution kernels across the input data and performing convolution operations, they can capture local patterns and spatial hierarchical features. Convolutional layers can effectively capture detailed information such as local texture, edges, and brightness gradients in an image. This information is crucial for understanding brightness variations in different regions of an image and forms the basis for generating refined illumination weight maps.

[0037] Activation functions are non-linear functions used to introduce non-linearity, enabling neural networks to learn more complex patterns. The illumination weight map is a two-dimensional matrix of the same size as the image, where each pixel value represents the weight assigned to a location during the fusion process. These weights are dynamically adjusted based on the local and global illumination conditions of the image. Through weighted fusion, the overall information of global illumination and the detailed information of local brightness distribution can be organically combined, ensuring that the weight map reflects both the overall brightness trend and captures local detail changes. The activation function introduces non-linearity, allowing the generated weight map to adapt more flexibly and precisely to complex illumination conditions.

[0038] This embodiment first extracts the luminance component from the first image, decoupling the image's luminance and color information to lay the foundation for subsequent illumination analysis. Using a global average pooling layer and a multilayer perceptron, global illumination factors are efficiently extracted from the luminance component, thereby capturing the overall luminance level of the image. Simultaneously, the luminance component is processed through a convolutional layer to finely extract local luminance features, thereby capturing detailed information such as image texture, edges, and local luminance gradients. Finally, these global illumination factors and local luminance features are weighted, fused, and activated to generate an illumination weight map of the first image.

[0039] This embodiment employs a layered and refined illumination feature extraction and fusion mechanism, ensuring that the generated illumination weight map accurately reflects the overall illumination trend and local brightness details of the first image. The weight map is then used to weight and fuse the multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image. This allows the fusion process to intelligently adjust the contribution of visible light features according to the actual illumination conditions of the image, effectively suppressing information loss in overly bright or dark areas, enhancing the contrast and detail of the image, and significantly improving the visual quality and information integrity of the fused image.

[0040] In some embodiments of this application, the process of extracting the frequency domain feature map of the first image at the target scale includes: Obtain the first input; wherein, when the target scale is the largest scale among multiple scales, the first input is the visible light feature map extracted from the first image; when the target scale is not the largest scale, the first input is the fused feature map of the corresponding upper scale of the target scale.

[0041] The first input is converted into the frequency domain to obtain a preliminary frequency domain feature map.

[0042] The initial frequency domain feature map is decoupled into amplitude component feature map and phase component feature map.

[0043] The frequency channel descriptor is obtained by compressing the spatial dimension of the amplitude component feature map through a global average pooling layer.

[0044] The corrected amplitude component feature map is extracted from the frequency channel descriptor by a cascaded first 1x1 convolutional layer, an LReLU activation function, and a second 1x1 convolutional layer.

[0045] The amplitude component feature map and phase component feature map are fused and converted into the spatial domain to obtain the frequency domain feature map of the first image at the target scale.

[0046] The first input refers to the image feature representation used for frequency domain analysis, which can be a transformation of the original image or the output of an intermediate layer of a deep learning network.

[0047] When the target scale is the largest scale (highest resolution) among multiple scales, the first input is the visible light feature map extracted from the first image. This means that when processing the highest resolution of the image, features are directly extracted from the original visible light image as the starting point for frequency domain analysis. For example, features can be extracted directly from the visible light image using an initial convolutional layer or a series of convolutional layers to obtain its spatial feature representation at the largest scale; or, preprocessing steps, such as Gaussian filtering or mean filtering, can be used to obtain a smooth feature map of the visible light image as input.

[0048] When the target scale is not the maximum scale, the first input is the fused feature map of the corresponding upper scale of the target scale. This means that when processing lower resolution or finer-grained features of an image, the features that have been fused at the previous scale are used as the input of the current scale to achieve the gradual fusion and transfer of multi-scale information. For example, in a multi-scale feature pyramid structure, the first input of the current scale can be obtained by downsampling operations (such as max pooling or stride convolution) of the fused feature map of the previous scale (i.e., larger scale or higher resolution); or, feature resampling techniques can be used to adjust the fused feature map of the upper scale to the current target scale as input.

[0049] The first input is converted into the frequency domain to obtain a preliminary frequency domain feature map. The aim is to convert the spatial domain representation of the image into the frequency domain representation to reveal the frequency components of the image, such as texture and edge information. For example, this can be achieved through other frequency domain transformation methods such as Discrete Cosine Transform (DCT) or Wavelet Transform.

[0050] Decoupling the initial frequency domain feature map into amplitude component feature maps and phase component feature maps aims to decompose the complex representation of the frequency domain feature map into two independent components: amplitude and phase. The amplitude component typically represents the energy distribution and contrast information of the image, while the phase component contains image structure and positional information. For example, for a complex form of frequency domain feature map, the amplitude component feature map can be obtained by calculating the magnitude of each complex number, and the phase component feature map can be obtained by calculating the argument of each complex number; alternatively, neural network modules (such as two independent convolutional layers) can be used to learn and extract amplitude-related and phase-related information from the frequency domain feature map, respectively.

[0051] The amplitude component feature map is spatially compressed using a global average pooling layer to obtain a frequency channel descriptor. The aim is to perform global average pooling on the amplitude component feature map, compressing the spatial information of each channel into a single numerical value, thus obtaining a descriptor representing the overall frequency characteristics of the channel. For example, the average of all pixel values ​​in each channel of the amplitude component feature map can be calculated to obtain the average intensity value of the channel as the frequency channel descriptor; alternatively, a global max pooling layer or a global norm pooling layer can be used to compress the spatial dimension of the amplitude component feature map.

[0052] This study extracts corrected amplitude component feature maps from frequency channel descriptors using cascaded 1x1 convolutional layers and the LReLU activation function. The aim is to utilize a neural network structure to process the frequency channel descriptors, learning and generating a corrected amplitude component feature map. The 1x1 convolutional layers are used for feature interactions and dimensionality adjustment between channels, while the LReLU activation function introduces non-linearity to enhance the model's expressive power. For example, a 1x1 convolutional layer can first compress the number of channels in the frequency channel descriptor, the LReLU activation function introduces non-linearity, and then another 1x1 convolutional layer restores or adjusts the number of channels to the target dimension, thus obtaining the corrected amplitude component feature map. Alternatively, other activation functions such as ReLU, Sigmoid, or Tanh can be used.

[0053] The modified amplitude component feature map and phase component feature map are fused and converted into the spatial domain to obtain the frequency domain feature map of the first image at the target scale. The aim is to recombine the modified amplitude component feature map with the original phase component feature map and transform it back into the spatial domain through inverse frequency domain transformation, thereby obtaining the final spatial domain feature map with fine frequency information. For example, the modified amplitude component feature map and phase component feature map can be recombine into a complex form of frequency domain feature map, and then transformed back into the spatial domain through two-dimensional inverse discrete Fourier transform (2D-IDFT); or, a neural network module (such as a convolutional layer) can be used to directly learn how to fuse the modified amplitude component feature map and phase component feature map to generate a spatial domain feature map without explicit inverse frequency domain transformation.

[0054] This embodiment first flexibly selects the source of the first input based on the current processing scale, ensuring that initial features are directly obtained from the original visible light image at the maximum scale, while utilizing the fused features from the previous scale at non-maximum scales, thus achieving effective transmission and utilization of multi-scale information. Subsequently, the acquired first input is converted to the frequency domain to capture the image's frequency information. Crucially, the frequency domain feature map is decoupled into amplitude component feature maps and phase component feature maps. The amplitude component carries the image's energy and contrast information, while the phase component contains the image's structure and position information. To refine the amplitude information, a global average pooling layer is used to refine the amplitude component... Spatial dimension compression is performed on the amplitude feature map to obtain frequency channel descriptors, which summarize the overall characteristics of each frequency channel. Next, cascaded 1x1 convolutional layers and the LReLU activation function are used to process the frequency channel descriptors, thereby extracting the corrected amplitude component feature map. This correction mechanism effectively filters out redundant information, enhances the expression of key frequency components, and introduces nonlinear transformation to improve the discriminative power of the features. Finally, the corrected amplitude component feature map is re-fused with the original phase component feature map, and then transformed back to the spatial domain through inverse frequency domain transformation to obtain the final, refined frequency domain feature map of the first image at the target scale. Through this frequency domain decoupling and amplitude correction strategy, this embodiment can more accurately capture the frequency details of the image, avoiding information redundancy and noise interference that may result from direct frequency domain transformation. This provides high-quality frequency domain information for subsequent weighted fusion with spatial domain features, significantly improving the performance of the fusion model and the visual effect of the final fused image.

[0055] In some embodiments of this application, the process of extracting the infrared feature map of the second image at the target scale includes: Obtain the second input; wherein, if the target scale is the largest scale among multiple scales, the second input is the second image; if the target scale is not the largest scale, the second input is the infrared feature map of the corresponding upper scale of the target scale.

[0056] Local feature maps and global feature maps are extracted from the second input.

[0057] The infrared feature map of the second image at the target scale is obtained by stitching and fusing the local feature map and the global feature map.

[0058] In this embodiment, a module that runs in parallel with Transformer and Convolutional Neural Network (CNN) can be used. The CNN branch extracts local feature maps using stacked residual convolutional blocks, while the Transformer branch uses a multi-head self-attention mechanism to calculate global correlations in a specific channel dimension to obtain a global feature map. The features extracted by both are then spliced ​​and fused to obtain infrared features with local precision and global semantic consistency, which is beneficial for subsequent image fusion.

[0059] In some embodiments of this application, the training process of the fusion model includes: Obtain a first sample and a second sample. The first sample is a visible light image sample, and the second sample is an infrared image sample corresponding to the visible light image sample.

[0060] Extract the global illumination factor and local brightness features of the first sample, and generate the illumination weight map of the first sample based on the global illumination factor and local brightness features of the first sample.

[0061] The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are extracted respectively, and the multi-scale infrared feature map of the second sample is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are then weighted and fused with the illumination weight map of the first sample to obtain the multi-scale fused feature map of the first sample.

[0062] The multi-scale fusion feature map of the first sample is fused with the multi-scale infrared feature map of the second sample to obtain the output image of the fusion model.

[0063] The restored image is obtained by reconstructing the fused feature map of the first sample.

[0064] Based on the first sample and the restored image, construct the restoration loss.

[0065] The backpropagation of the fusion model is guided by the recovery loss to obtain the trained fusion model.

[0066] The training process of a fusion model refers to adjusting the internal parameters of the fusion model using specific algorithms and datasets, enabling it to learn the mapping relationship from the input image to the desired output image. This process optimizes the model's performance, allowing it to generate high-quality fused images in practical applications. The training process typically includes data preparation, forward propagation, loss calculation, backpropagation, and parameter updates. In unsupervised learning, the model learns by discovering the inherent structure of the data.

[0067] Unlike the inference phase of the aforementioned model, this embodiment reconstructs the restored image based on the fusion feature map (the smallest-scale fusion feature map) of the first sample. This step aims to inversely generate a restored image from the visible light feature representation within the fusion model. The fusion feature map is the representation of the visible light image after feature extraction and illumination weighting. The reconstruction process is typically achieved through a series of deconvolutional layers (or transposed convolutional layers) and upsampling operations, gradually restoring the abstract feature representation to the image pixel space. For example, a decoder network can be designed with the fusion feature map as input and the restored image of the same size as the first sample as its output.

[0068] Based on the first sample and the reconstructed image, a restoration loss is constructed. The restoration loss measures the difference between the reconstructed image and the original first sample, guiding the model to learn how to better preserve the details and structure of the visible light image. Loss functions typically include brightness loss, structure loss, and color loss. For example, brightness loss can use L1 or L2 norms to measure pixel intensity differences; structure loss can use the Structural Similarity Index (SSIM) or gradient loss to measure image structural similarity; and color loss measures the difference in color distribution between the reconstructed image and the original image. By minimizing the restoration loss, the model can learn to better preserve the inherent information of the visible light image during the fusion process.

[0069] The backpropagation of the fusion model is guided by the recovery loss to obtain the trained fusion model. Backpropagation is the core mechanism of neural network training. It calculates the error based on the loss function, calculates the gradient of the model parameters layer by layer through the chain rule, and updates these parameters using optimizers (such as Adam, SGD, etc.), thereby gradually improving the model performance. Here, the recovery loss serves as one of the bases for backpropagation, ensuring that the fusion model can effectively extract and retain key information from the visible light image while generating the fused image, avoiding the loss of important visible light details during the fusion process.

[0070] To ensure that the fusion model can fully preserve the details and structural information of the visible light image during the fusion process, this embodiment introduces a restoration mechanism. Specifically, based on the fusion feature map of the first sample, a reconstruction network is used to inversely generate a restored image. This restored image aims to restore the visual content of the first sample as much as possible. Subsequently, by comparing the first sample and the restored image, a restoration loss is constructed. The restoration loss quantifies the degree to which the model retains visible light image information during feature extraction and fusion. The restoration loss is used to guide the backpropagation process of the fusion model, and the model's internal parameters are adjusted through optimization algorithms such as gradient descent. In this way, the fusion model not only learns how to effectively combine infrared and visible light information, but more importantly, it is forced to learn how to extract and retain key features from the visible light image. Therefore, when generating the fused image, it can better maintain the texture, color, and structural details of the visible light image, avoid information loss, and thus improve the overall visual quality and information richness of the fused image.

[0071] In some embodiments of this application, before guiding the backpropagation of the fusion model based on the recovery loss to obtain the trained fusion model, the method further includes: Construct the fusion loss based on the second sample, the restored image, and the output image of the fusion model; The above-mentioned backpropagation of the fusion model based on recovery loss includes: Construct the total loss based on the recovery loss and fusion loss; The backpropagation of the fusion model is guided by the total loss, where the fusion loss includes: (1); (2); (3); in, For the loss of fusion, For strength loss, For gradient loss, The output image of the fusion model, To perform the maximum value operation, As the second sample, To restore the image, To of Norm value, To of Norm value, For the Sobel gradient operator, They are respectively for and gradient plot and Absolute value operation, These are the weights for the gradient loss.

[0072] The fusion loss is designed to evaluate the quality of the output image of the fusion model, especially its performance in combining infrared and visible light image information. By introducing the fusion loss, the model can be encouraged to focus not only on the reconstruction of the visible light image during training, but also to optimize the visual effect and information richness of the fused image.

[0073] This embodiment overcomes the limitations of a single restoration loss in optimizing the overall performance of fused images by introducing a fusion loss during the training process described above. Specifically, before the model guides backpropagation based on the restoration loss, a fusion loss is first constructed based on the second sample, the restored image, and the output image of the fusion model. The fusion loss, through a combination of intensity loss and gradient loss, quantifies the quality of the fused image from two dimensions: pixel intensity and image structure. The intensity loss ensures that the fused image can effectively inherit the significant brightness information of the two source images by comparing the pixel-wise maximum intensity value of the fused image with that of the infrared image and the restored image. The gradient loss, by comparing the gradient of the fused image with that of the infrared image and the restored image, enables the fused image to better preserve the edge and texture details in the source images.

[0074] The constructed fusion loss is combined with the recovery loss to form a total loss. This total loss serves as a more comprehensive optimization objective, guiding the fusion model through backpropagation. In this way, during training, the model not only learns how to accurately reconstruct the features of visible light images, but also learns how to effectively fuse the complementary information of infrared and visible light images to generate visually more natural and information-rich fused images. This dual loss mechanism enables the model to more comprehensively capture and integrate the advantages of different modalities of images, thereby significantly improving the overall quality and practicality of the fused images.

[0075] Most importantly, replacing the original first sample with a reconstructed image from visible light in the fusion loss significantly improves the fusion quality and model robustness under complex lighting conditions. The first sample suffers from abnormal brightness, missing texture, and noise interference in low-light, strong-light, and backlight scenarios. Directly participating in fusion supervision can lead to the model learning incorrect texture and intensity distributions, resulting in defects such as blurred details, structural distortion, or unclear targets in the fusion result. By incorporating the illumination-corrected reconstructed image into the intensity loss and gradient loss calculations, a stable and reliable visible light benchmark can be provided for the fusion process. This allows the fused image to inherit the clear texture, uniform brightness, and complete structural information of the reconstructed image, while preserving the saliency of the thermal radiation target in the infrared image. The design makes the fusion constraints more closely match the needs of real-world scenarios, avoiding the impact of lighting fluctuations on the fusion effect. This enables the model to output fusion results with rich details, clear targets, and complete scene semantics in all weather conditions and variable lighting environments, effectively improving the algorithm's practicality and generalization ability.

[0076] In some embodiments of this application, loss recovery includes: (4); (5); (6); (7); in, To recover the losses, For brightness recovery loss, To preserve structural losses, To maintain the weight of the loss in the structure, For color balance loss, The weighting of color balance loss, The number of non-overlapping local region blocks to divide the restored image into. For the first The average intensity value of each local region block, This is the preset brightness reference value; To restore the total number of pixels in the image, For the first The four-neighbor area of ​​a pixel To recover the first in the image 1 pixel To recover the first in the image 1 pixel For the first sample 1 pixel For the first sample 1 pixel; To restore the image in channels Global average intensity on To restore the image in channels Global average intensity on This is a combination of all channel pairs.

[0077] Restoration loss is a key metric used to measure the difference between the restored image and the original visible light image. Its design directly affects the accuracy of the fusion model in learning how to reconstruct the visible light image from the fusion features. It mainly includes: brightness restoration loss, structure preservation loss, and color equalization loss.

[0078] 1) Brightness restoration loss aims to ensure that the overall brightness distribution of the restored image is consistent with the preset brightness reference value E, and avoid the restored image being too bright or too dark. Its implementation usually involves dividing the image into multiple local region blocks and calculating the difference between the average intensity of each region block and the reference value.

[0079] 2) Structure preservation loss is used to ensure that the structural information of the restored image, such as edges and textures, is highly consistent with the original visible light image, thereby preventing blurring or structural distortion in the restored image. This can be achieved by comparing the local gradient differences between the restored image and the original image to quantify the degree of preservation of structural information.

[0080] 3) Color equalization loss aims to ensure that the color distribution of the restored image is natural and balanced, avoiding color cast or color imbalance. Its implementation usually involves calculating the difference between the global average intensity of different color channels in the restored image to promote the balance of the intensity of each channel.

[0081] This method refines the restoration loss into a weighted combination of brightness restoration loss, structure preservation loss, and color balance loss. This allows the fusion model to comprehensively evaluate the quality of the restored image from multiple dimensions during training. The brightness restoration loss guides the model to adjust the overall brightness of the restored image to meet preset visual standards, avoiding overexposure or underexposure. The structure preservation loss ensures that the model retains a high degree of detail in the edges, textures, and other information of the original visible light image during restoration, effectively preventing image blurring and structural distortion. The color balance loss ensures a natural and harmonious color distribution in the restored image, avoiding color cast. This multi-dimensional loss design, together with the fusion loss, constitutes the total loss and collaboratively guides the fusion model in backpropagation. This enables the model not only to generate high-quality fused images but also to internally reconstruct a restored image that is highly consistent with the original visible light image in terms of brightness, structure, and color. This comprehensive loss mechanism significantly improves the training effect of the fusion model, enabling it to better understand and process the complex information of visible light images, thereby outputting fused images with better visual effects and more complete information preservation.

[0082] like Figures 2 to 4One embodiment of this application provides a method for fusing infrared and visible light images based on frequency domain decoupling. This method includes: Step S210 (Training Phase): Constructing a system containing infrared images With visible light images The paired dataset was used for image registration, cropping, and normalization preprocessing. Step S220: Train the fusion model based on the dataset; The fusion model includes: an illumination assessment module, an encoder (including an infrared encoder branch and a visible light encoder branch), a visible light recovery module, and a decoder.

[0083] First, the visible light encoder branch extracts the input visible light image. The luminance component is then used to capture visible light images via a global average pooling layer and a multilayer perceptron (MLP). The system uses global illumination priors to generate global illumination factors; simultaneously, it extracts local brightness features from the image using convolutional layers, weights the global illumination factors and local brightness features, and generates an illumination weight map using a sigmoid activation function. .

[0084] Then, the infrared image With visible light images Features were extracted using the infrared encoder branch and the visible light encoder branch, respectively. The infrared encoder branch employs a parallel Transformer and CNN branch. The CNN branch extracts local detail features using stacked residual convolutional blocks, while the Transformer branch uses a multi-head self-attention mechanism to calculate global correlations across specific channel dimensions. The features extracted by both branches are then concatenated and fused to obtain an infrared feature map with both local finesse and global semantic consistency. .

[0085] Then, the visible light encoder branches (spatial domain branch and frequency domain branch) extract features from the visible light image. The structure of the spatial domain branch is the same as that of the infrared encoder branch, thus obtaining spatial domain features. The frequency domain branch first uses Fast Fourier Transform (FFT) to transform the visible light features to the frequency domain, decoupling them into amplitude components representing global illumination and phase components representing texture structure. The decoupled amplitude components are first input into a global average pooling layer to compress the spatial dimension, obtaining frequency channel descriptors. Then, cascaded 1×1 convolutions and LReLU activation functions are used to capture the nonlinear correlations between frequency channels and learn the brightness mapping relationship. At the same time, to ensure that structural distortion is not introduced while restoring brightness and to keep the phase unchanged, the corrected amplitude spectrum is combined with an Inverse Fourier Transform (IFFT) to transform it back to the spatial domain. Then, the residual connection operation is used for fusion to finally obtain the frequency domain features. .

[0086] Subsequently, spatial domain features were extracted in parallel. Frequency domain characteristics The two are weighted and fused using the illumination weight map to obtain the fused feature map. : (8); Then, the visible light restoration module will use the fused feature map output by the encoder. The restored image is generated using the decoder. .

[0087] Finally, the decoder employs an infrared saliency-guided mutual attention fusion strategy to integrate the extracted infrared features. With visible light characteristics The fusion process is performed in the intermediate fusion layer to obtain the fused features. That is, firstly... Spatial attention is extracted, a spatial attention map is generated, and the spatial attention map is used for... Perform weighting, and compare the weighted features with the original features. The images are stitched together and then fused using a 1×1 convolution to generate a fused image from the fused features.

[0088] This method utilizes recovery loss Monitor and recover losses Loss of brightness recovery Structural retention loss and color balance loss Weighted composition.

[0089] The fusion phase is monitored using fusion loss, which includes intensity loss. and gradient loss ; Finally, a composite total loss function containing fusion loss and recovery loss is constructed, and all parameters of the forward fusion network are updated through backpropagation until the model converges.

[0090] Loss function: (9); These are the weight parameters. As shown in formulas (1) to (3) above, they will not be elaborated here. As shown in formulas (4) to (7) above, they will not be elaborated here.

[0091] Step S230 (inference): Based on the fusion model, the first image and the second image are fused to obtain a fused image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image.

[0092] This embodiment has at least the following beneficial effects: (1) This embodiment utilizes a dual-path architecture that combines visible light image enhancement with infrared image complementarity, which can improve robustness and the fusion accuracy of the fused image; (2) In this embodiment, illumination enhancement and spatial design are introduced into the visible light encoder branch. Frequency domain decoupling mechanism: First, the global and local illumination priors are extracted through the illumination evaluation module to generate a weight map; then, the frequency domain decoupling mechanism is used to decouple the image into amplitude spectrum and phase spectrum. While keeping the phase structure unchanged, the amplitude component is dynamically corrected, thereby correcting the illumination without destroying the texture details. This effectively improves the quality of the fused image in complex illumination scenes. While preserving the infrared thermal radiation target, it can recover the visible light texture details according to the ambient illumination. (3) When constructing the loss function of the training model, the visible light restoration module is used to reconstruct the fused feature map to obtain the restored image. Then, the image is used for supervision as an intermediate supervision output to supervise the learning process of the visible light enhancement branch. Through brightness, structure and color loss constraints, the network is prompted to generate complete restored visible light features with lighting and texture, providing high-quality priors for subsequent feature fusion and ultimately improving the effect of the fused image. Moreover, in the fusion loss, the restored image is used to replace the original first sample, which can significantly improve the fusion quality and model robustness under complex lighting.

[0093] like Figure 5 As shown, in some embodiments of this application, an infrared and visible light image fusion device based on frequency domain decoupling is also provided, the device comprising: Image acquisition module 1001 is used to acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; Image fusion module 1002 is used to fuse a first image and a second image based on a fusion model to obtain a fused image. The process of obtaining the fused image by the fusion model includes: Extract the global illumination factor and local brightness features of the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image.

[0094] It should be noted that the infrared and visible light image fusion device based on frequency domain decoupling provided in this embodiment and the infrared and visible light image fusion method based on frequency domain decoupling described above are based on the same inventive concept. Therefore, the content of the infrared and visible light image fusion method based on frequency domain decoupling in the above embodiment is also applicable to the content of the infrared and visible light image fusion device based on frequency domain decoupling in this embodiment, and will not be repeated here.

[0095] like Figure 6 This application also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the infrared and visible light image fusion method based on frequency domain decoupling described above in this disclosure.

[0096] Electronic devices can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0097] The electronic devices according to embodiments of this application will now be described in detail.

[0098] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to perform the infrared and visible light image fusion method based on frequency domain decoupling according to the embodiments of the present invention.

[0099] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.

[0100] This invention also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described infrared and visible light image fusion method based on frequency domain decoupling.

[0101] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0102] The embodiments described in this invention are intended to more clearly illustrate the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.

[0103] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0106] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0107] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0108] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0112] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.

Claims

1. A method for fusing infrared and visible light images based on frequency domain decoupling, characterized in that, The method includes: Acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; The first image and the second image are fused based on a fusion model to obtain a fused image. The process of obtaining the fused image by the fusion model includes: Extract the global illumination factor and local brightness features of the first image from the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image; The process of extracting the frequency domain feature map of the first image at the target scale includes: Obtain a first input; wherein, when the target scale is the largest scale among multiple scales, the first input is a visible light feature map extracted from the first image; when the target scale is not the largest scale, the first input is a fused feature map of the corresponding upper-level scale of the target scale; The first input is converted into the frequency domain to obtain a preliminary frequency domain feature map; The preliminary frequency domain feature map is decoupled into an amplitude component feature map and a phase component feature map; The amplitude component feature map is spatially compressed using a global average pooling layer to obtain a frequency channel descriptor; The corrected amplitude component feature map is extracted from the frequency channel descriptor by cascading 1x1 convolutions and LReLU activation functions; By fusing the corrected amplitude component feature map and the phase component feature map and converting them into the spatial domain, the frequency domain feature map of the first image at the target scale is obtained. The training process of the fusion model includes: Obtain a first sample and a second sample; wherein the first sample is a visible light image sample, and the second sample is an infrared image sample corresponding to the visible light image sample; Extract the global illumination factor and local brightness features of the first sample from the first sample, and generate the illumination weight map of the first sample based on the global illumination factor and local brightness features of the first sample. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are extracted respectively, and the multi-scale infrared feature map of the second sample is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are then weighted and fused with the illumination weight map of the first sample to obtain the multi-scale fused feature map of the first sample. The multi-scale fusion feature map of the first sample is fused with the multi-scale infrared feature map of the second sample to obtain the output image of the fusion model; The restored image is reconstructed based on the fusion feature map of the first sample; Based on the first sample and the restored image, a restoration loss is constructed; The fusion model is backpropagated based on the recovery loss to obtain the trained fusion model.

2. The infrared and visible light image fusion method based on frequency domain decoupling according to claim 1, characterized in that, The step of extracting the global illumination factor and local brightness features of the first image from the first image includes: Extract the luminance component from the first image; The global illumination factor is extracted from the brightness component using a global average pooling layer and a multilayer perceptron. Local brightness features are extracted from the brightness components using convolutional layers; The step of generating an illumination weight map of the first image based on the global illumination factor and local brightness features of the first image includes: The global illumination factor and local brightness features of the first image are weighted, fused, and activated to obtain the illumination weight map of the first image.

3. The infrared and visible light image fusion method based on frequency domain decoupling according to claim 1, characterized in that, The process of extracting the infrared feature map of the second image at the target scale includes: Obtain the second input; wherein, when the target scale is the largest scale among multiple scales, the second input is the second image; when the target scale is not the largest scale, the second input is the infrared feature map of the corresponding upper scale of the target scale; Local feature maps and global feature maps are extracted from the second input; The infrared feature map of the second image at the target scale is obtained by stitching and fusing the local feature map and the global feature map.

4. The infrared and visible light image fusion method based on frequency domain decoupling according to claim 1, characterized in that, Before guiding the backpropagation of the fusion model based on the recovery loss to obtain the trained fusion model, the method further includes: Based on the second sample, the restored image, and the output image of the fusion model, a fusion loss is constructed; The step of guiding the backpropagation of the fusion model based on the recovery loss includes: Construct the total loss based on the recovery loss and the fusion loss; The fusion model is backpropagated based on the total loss; wherein the fusion loss includes: ; ; ; in, For the loss of fusion, For strength loss, The weights are the gradient loss values. For gradient loss, The output image of the fusion model, To perform the maximum value operation, As the second sample, To restore the image, for of Norm value, for of Norm value, For the Sobel gradient operator, They are respectively for and gradient plot and Absolute value operation.

5. The infrared and visible light image fusion method based on frequency domain decoupling according to claim 4, characterized in that, The recovery loss includes: ; ; ; ; in, To recover the losses, For brightness recovery loss, To preserve structural losses, To maintain the weight of the loss in the structure, For color balance loss, The weighting of color balance loss, The number of non-overlapping local region blocks to divide the restored image into. For the first The average intensity value of each local region block, This is the preset brightness reference value; To restore the total number of pixels in the image, For the first The four-neighbor area of ​​a pixel To recover the first in the image 1 pixel To recover the first in the image 1 pixel For the first sample 1 pixel For the first sample 1 pixel; To restore the image in channels Global average intensity on To restore the image in channels Global average intensity on This is a combination of all channel pairs.

6. An infrared and visible light image fusion device based on frequency domain decoupling, characterized in that, The device includes: An image acquisition module is used to acquire a first image and a second image; wherein the first image is a visible light image and the second image is an infrared image corresponding to the first image; An image fusion module is used to fuse the first image and the second image based on a fusion model to obtain a fused image, wherein the process of obtaining the fused image by the fusion model includes: Extract the global illumination factor and local brightness features of the first image from the first image, and generate the illumination weight map of the first image based on the global illumination factor and local brightness features of the first image. Multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are extracted respectively, and multi-scale infrared feature map of the second image is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first image are then weighted and fused with the illumination weight map of the first image to obtain the multi-scale fused feature map of the first image. The multi-scale fusion feature map of the first image is fused with the multi-scale infrared feature map of the second image to obtain a fused image; The process of extracting the frequency domain feature map of the first image at the target scale includes: Obtain a first input; wherein, when the target scale is the largest scale among multiple scales, the first input is a visible light feature map extracted from the first image; when the target scale is not the largest scale, the first input is a fused feature map of the corresponding upper-level scale of the target scale; The first input is converted into the frequency domain to obtain a preliminary frequency domain feature map; The preliminary frequency domain feature map is decoupled into an amplitude component feature map and a phase component feature map; The amplitude component feature map is spatially compressed using a global average pooling layer to obtain a frequency channel descriptor; The corrected amplitude component feature map is extracted from the frequency channel descriptor by cascading 1x1 convolutions and LReLU activation functions; By fusing the corrected amplitude component feature map and the phase component feature map and converting them into the spatial domain, the frequency domain feature map of the first image at the target scale is obtained. The training process of the fusion model includes: Obtain a first sample and a second sample; wherein the first sample is a visible light image sample, and the second sample is an infrared image sample corresponding to the visible light image sample; Extract the global illumination factor and local brightness features of the first sample from the first sample, and generate the illumination weight map of the first sample based on the global illumination factor and local brightness features of the first sample. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are extracted respectively, and the multi-scale infrared feature map of the second sample is extracted. The multi-scale spatial domain feature map and multi-scale frequency domain feature map of the first sample are then weighted and fused with the illumination weight map of the first sample to obtain the multi-scale fused feature map of the first sample. The multi-scale fusion feature map of the first sample is fused with the multi-scale infrared feature map of the second sample to obtain the output image of the fusion model; The restored image is reconstructed based on the fusion feature map of the first sample; Based on the first sample and the restored image, a restoration loss is constructed; The fusion model is backpropagated based on the recovery loss to obtain the trained fusion model.

7. An electronic device, characterized in that, include: At least one control processor and a memory for communicatively connecting to the at least one control processor; The memory stores instructions that can be executed by the at least one control processor, which, when executed, enables the at least one control processor to perform the infrared and visible light image fusion method based on frequency domain decoupling as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the infrared and visible light image fusion method based on frequency domain decoupling as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Denoising diffusion model driven texture enhanced infrared and visible light image fusion method and system

    CN119540071A

  • Frequency domain decoupling flare removal method

    CN120765521A