Image fusion method based on RFN-Nest network and illumination perception

By introducing the light-aware subnet and CBAM attention mechanism, the fusion layer training of the RFN-Nest network is optimized, and the information loss and texture details of image fusion under extreme lighting conditions are solved, and adaptive high-quality image fusion is achieved.

CN120495827APending Publication Date: 2025-08-15JILIN UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510983092.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing RFN-Nest network is prone to lose texture details and information when image fusion is fused under extreme lighting conditions, the fusion layer information is insufficient, and it fails to adaptively integrate meaningful information according to lighting conditions.

Method used

The light-aware subnet is introduced to estimate the light distribution, and the fusion layer network is trained through the light probability guidance, combining the CBAM attention mechanism and illuminance, auxiliary intensity, and texture loss function to optimize the information retention and texture details of the fusion layer.

Benefits of technology

Adaptive processing under different lighting conditions is realized, reducing information loss, retaining rich texture details and optimal intensity distribution, and improving image fusion quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495827A_ABST
    Figure CN120495827A_ABST
Patent Text Reader

Abstract

The invention discloses an image fusion method based on an RFN-Nest network and illumination perception, and the method comprises the steps: training an auto-encoder network and an illumination perception sub-network based on infrared and visible light images, and guiding the training of a fusion layer network through the illumination perception sub-network; inputting the infrared and visible light images to an encoder network, and extracting depth features of the infrared and visible light images on different scales; inputting depth features of different scales into each trained fusion layer network, performing multi-scale depth feature capture and feature splicing, and introducing a CBAM attention mechanism to perform double attention optimization to obtain fused depth features; inputting the multi-scale features into a decoder network, and reconstructing the fused multi-scale features; according to the invention, the illumination perception sub-network is introduced to realize adaptive processing of different illumination images; the fusion layer network is improved to reserve more information, information loss is reduced, and the performance of the network is improved by introducing the CBAM attention module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and more particularly to an image fusion method based on RFN-Nest network and illumination perception. Background Art

[0002] At present, multi-source information fusion technology overcomes the limitations of single sensors by integrating multimodal data and has become a key technology in computer vision. The fusion of infrared and visible light images is the most representative. Visible light provides texture details but is greatly affected by the environment, while infrared imaging is stable but has low resolution. The complementary advantages of the two can generate high-quality images with comprehensive information. This technology is widely used in remote sensing, medical treatment, intelligent navigation and other fields.

[0003] Image fusion algorithms based on deep learning can be divided into three categories: algorithms based on convolutional neural networks, based on autoencoders, and based on generative adversarial networks; the retention of gradient information in image fusion algorithms based on convolutional neural networks depends on manual adjustment, which may lead to the loss of texture structure; image fusion algorithms based on autoencoders: DenseFuse first applied autoencoders to the fusion of infrared and visible light, and NestFuse introduced multi-scale fusion and attention mechanisms on this basis to alleviate the problem of detail loss. RFN-Nest abandoned the manually designed fusion rules of DenseFuse and NestFuse, and innovatively adopted a separately trained fusion module to learn multi-level features. Feature fusion; Image fusion algorithm based on generative adversarial network. Since the balance between the generator and the discriminator is difficult to accurately control, the FusionGAN network performs poorly in retaining infrared target features. Some models proposed based on this include: The DDcGAN model introduces a visible light gradient discriminator and a low-resolution infrared discriminator, and adopts a dual discriminator structure to improve the multimodal feature learning ability; AttentionFGAN introduces a multi-scale attention mechanism to balance target and background features, but faces the challenge of high training complexity of the dual discriminator; The UIFGAN model extends GAN to multi-image fusion scenarios for the first time, breaking through the single-task limitation and promoting the evolution of technical boundaries in this field.

[0004] For the RFN-Nest method, although a learnable fusion strategy is designed without manual intervention, its modeling process still assumes that texture exists only in visible light images and does not consider lighting factors. This may lead to the loss of texture details when fusing images of low-light scenes. Moreover, due to the structural characteristics of its fusion layer, it does not adequately consider the problem of fusion layer information extraction, and there is still room for further improvement. Finally, in the training of the fusion layer loss function, this method only considers the background detail preservation loss and the target feature enhancement loss, without texture detail constraints. The fusion layer cannot adaptively integrate meaningful information according to lighting conditions.

[0005] Therefore, how to overcome the problem of detail loss under extreme lighting conditions, the problem of fusion layer information loss, and the problem of texture detail loss in image fusion is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0006] In view of this, the present invention provides an image fusion method based on RFN-Nest network and illumination perception to solve some of the technical problems mentioned in the background technology.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions: An image fusion method based on RFN-Nest network and illumination perception includes the following steps: S1. Based on infrared and visible light images, the codec is trained as an autoencoder network. The illumination perception subnetwork is trained based on visible light images. The illumination probability output by the illumination perception subnetwork guides the training of multiple fusion layer networks. S2. Input infrared and visible light images into the encoder network and extract deep features of infrared and visible light images at different scales. S3. Input the deep features of infrared and visible light images of different scales into the trained fusion layer network respectively, perform multi-scale deep feature capture and feature splicing, and introduce the CBAM attention mechanism for dual attention optimization to obtain the fused deep features; S4. Input the fused deep features into the decoder network and reconstruct the fused multi-scale features.

[0008] Preferably, in step S1, the training content of the autoencoder network is: The encoder and decoder are directly connected. The encoder network downsamples the input image to extract multi-scale depth features. The decoder reconstructs the input image with multi-scale depth features. The encoder-decoder network is trained without the fusion layer. The total loss function of the autoencoder network is: ; ; ; in, and represents the pixel loss and structural similarity loss between the input image I and the reconstructed image O, β is the trade-off parameter between pixel loss and structural similarity loss, is the F norm, It is a structural similarity measure used to quantify the structural similarity between two images.

[0009] Preferably, in step S1, the training process of the illumination perception sub-network is constrained by using cross entropy loss, specifically: ; Where m is the lighting label of the input image, is the illumination probability output by the illumination perception subnetwork, and σ is the softmax function, which normalizes the illumination probability to [0,1].

[0010] Preferably, in step S1, the loss function for training multiple fusion layer networks guided by the illumination probability output by the illumination perception sub-network is: L fusion =α1L ill +α2L aux +α3L tex; Among them, α1, α2, and α3 are trade-off parameters, and L ill is the illumination loss, L aux is the auxiliary strength loss, L tex For texture loss.

[0011] Preferably, the loss function of the fusion layer network includes illumination loss, auxiliary intensity loss, and texture loss; the specific contents are: The specific illumination loss is: ; ; ; in, and are the intensity losses of infrared and visible light images, respectively, K ir and K vi The illumination perception weights contributed by infrared images and visible light images, P x The output of the illumination perception sub-network is the probability that the illumination probability belongs to the night scene, P y is the probability of belonging to the daytime scene; The intensity loss of infrared and visible light images is specifically: ; ; Where H is the height of the input image, W is the width of the input image, is the L1 norm, I ir and I vi are the input infrared image and visible light image, I f is the output fused image; The auxiliary strength loss is specifically: ; in, represents an element-wise maximum selection; The texture loss is specifically: ; in, To measure the gradient operator of image texture information, the Sobel operator is used to calculate the gradient. It is an absolute value operation.

[0012] Preferably, in step S1, the trained illumination perception subnetwork input is a visible light image, and the output is illumination probability; the specific architecture includes: Four 4×4 convolutional layers with a stride of 2, each followed by an LReLU activation function without changing the size of the feature map, are used to compress spatial information and extract illumination information; A global average pooling layer to integrate lighting information; Two fully connected layers calculate the illumination probability distribution based on the integrated illumination information, including the probabilities of daytime scenes and nighttime scenes; The illumination distribution mechanism calculates the illumination perception weights representing the contributions of infrared images and visible light images based on the illumination probability.

[0013] Preferably, the specific content of step S3 is: S31. Input the infrared and visible light image depth features extracted by the encoder, perform a 1×1 convolution operation, and then pass it through four convolutional layers with different kernel sizes to capture multi-scale depth features. There is no pooling layer for each convolution operation. The kernel sizes of the convolutional layers are 1×1, 3×3, 5×5, and 7×7, respectively. S32. Concatenate the obtained features according to the channel size and pass the concatenated features through a 3×3 convolutional layer. S33. Introducing the CBAM attention mechanism module, which consists of a cascaded channel attention module and a spatial attention module. The feature map output by the convolutional layer is first weighted in the channel dimension, then enhanced in the spatial dimension, and finally outputs the feature representation after dual attention optimization. S34. The optimized features are passed through two 3×3 convolutional layers to obtain fused deep features and input into the decoder network.

[0014] Preferably, the specific content of the channel dimension weight adjustment is: The input feature map generates two spatially compressed vectors through parallel global maximum pooling and global average pooling, which capture the salient area and global statistical information respectively. The two vectors are then processed by a multi-layer perceptron (MLP) with shared parameters and superimposed element by element. They are normalized into a channel attention weight matrix using the Sigmoid function. Finally, the original feature map and the weight matrix are element-wise multiplied to obtain the input features of the spatial attention module.

[0015] Preferably, the specific content of spatial dimension feature enhancement is: The output features of the channel attention are subjected to maximum pooling and average pooling in the channel dimension to generate two single-channel spatial feature maps. The two channel features are then concatenated along the channel axis, the channel dimension is compressed to 1 through a 7×7 convolutional layer, and a spatial weight matrix is generated through Sigmoid. Finally, the original feature map and the spatial weight matrix are weighted and calculated to output the final optimized features.

[0016] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses an image fusion method based on the RFN-Nest network and illumination perception. On the basis of the RFN-Nest network structure, an illumination perception subnetwork is introduced to estimate illumination distribution and calculate illumination probability to guide fusion training and realize adaptive processing of images with different illuminations. The fusion layer network is improved so that the fusion layer can retain more information and reduce information loss. The CBAM attention module is introduced to improve the performance of the network. When training the fusion layer network, loss functions related to illumination loss, auxiliary intensity loss and texture loss are adopted, so that the fusion layer network can adaptively integrate meaningful information according to illumination conditions, and the fused image can retain rich texture details while maintaining the optimal intensity distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0018] Figure 1 A schematic diagram of an image fusion method based on RFN-Nest network and illumination perception provided by the present invention; Figure 2 Schematic diagram of the autoencoder network training framework provided by the present invention; Figure 3 A schematic diagram of the light perception sub-network architecture provided by the present invention; Figure 4 Schematic diagram of the fusion layer network training framework provided by the present invention; Figure 5 Schematic diagram of the CBAM attention module provided by the present invention; Figure 6 Schematic diagram of the channel attention module provided by the present invention; Figure 7 Schematic diagram of the spatial attention module provided by the present invention; Figure 8 Schematic diagram of the experimental results comparing the present invention with the RFN-Nest method. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] The embodiment of the present invention discloses an image fusion method based on RFN-Nest network and illumination perception, such as Figure 1 , including the following steps: S1. Based on infrared and visible light images, the codec is trained as an autoencoder network. The illumination perception subnetwork is trained based on visible light images. The illumination probability output by the illumination perception subnetwork guides the training of multiple fusion layer networks. S2. Input infrared and visible light images into the encoder network and extract deep features of infrared and visible light images at different scales. S3. Input the deep features of infrared and visible light images of different scales into the trained fusion layer network respectively, perform multi-scale deep feature capture and feature splicing, and introduce the CBAM attention mechanism for dual attention optimization to obtain the fused deep features; S4. Input the fused deep features into the decoder network and reconstruct the fused multi-scale features.

[0021] In order to further implement the above technical solutions, Figure 2 , step S1, the training content of the autoencoder network is: The encoder and decoder are directly connected. The encoder network downsamples the input image to extract multi-scale depth features. The decoder reconstructs the input image with multi-scale depth features. The encoder-decoder network is trained without the fusion layer. The total loss function of the autoencoder network is: ; ; ; in, and represents the pixel loss and structural similarity loss between the input image I and the reconstructed image O, β is the trade-off parameter between pixel loss and structural similarity loss, is the F norm, It is a structural similarity measure used to quantify the structural similarity between two images.

[0022] In order to further implement the above technical solution, step S1 adopts cross entropy loss The training process of the constrained light perception subnetwork is as follows: ; Where m is the lighting label of the input image, is the illumination probability output by the illumination perception subnetwork, and σ is the softmax function, which normalizes the illumination probability to [0,1].

[0023] In this embodiment, the illumination perception subnetwork is trained to generate the illumination probability corresponding to the input image. The cross entropy loss is used to constrain the training process of the illumination perception subnetwork, and the illumination probability can be used to obtain the illumination perception weight K contributed by the source image. ir and K vi .

[0024] In order to further implement the above technical solution, in step S1, the loss function of training multiple fusion layer networks guided by the illumination probability output by the illumination perception sub-network is: L fusion =α1L ill +α2L aux +α3L tex; Among them, α1, α2, and α3 are trade-off parameters, and L ill is the illumination loss, L aux is the auxiliary strength loss, L tex is texture loss; The fusion layer network is in L ill and L aux Under the guidance of L tex Under the guidance of , ideal texture details can be obtained.

[0025] In order to further implement the above technical solution, the loss function of the fusion layer network includes illumination loss, auxiliary intensity loss, and texture loss; the specific contents are: Illumination loss Illumination loss enables the fusion layer network to adaptively integrate meaningful information according to lighting conditions, specifically: ; ; ; in, and are the intensity losses of infrared and visible light images, respectively, K ir and Kvi The illumination perception weights contributed by infrared images and visible light images, P x The output of the illumination perception sub-network is the probability that the illumination probability belongs to the night scene, P y is the probability of belonging to the daytime scene; The intensity loss function quantifies the intensity consistency deviation between the fusion result and the source image in the pixel domain through a pixel-by-pixel comparison mechanism. The intensity loss of infrared and visible light images is specifically: ; ; Where H is the height of the input image, W is the width of the input image, is the L1 norm, I ir and I vi are the input infrared image and visible light image, I f is the output fused image; In this embodiment, the illumination perception weight K is obtained by the illumination perception sub-network. ir and K vi Adjust the intensity constraint of the fused image, that is, when the probability of night scene is high, K ir Greater than K vi , at this time, the intensity loss of the infrared image dominates, and vice versa (when the probability of judging that the input image is a daytime lighting scene is high, the fused image relatively retains more information in the visible light image; and when the probability of judging that the input image is a nighttime, i.e., low-light lighting scene is high, the fused image relatively retains more information in the infrared image); The illumination loss drives the fusion layer network to dynamically preserve the brightness information of the source image according to the lighting conditions. However, the fused image usually cannot maintain an optimal intensity distribution. Therefore, an auxiliary intensity loss is further introduced. The auxiliary strength loss is specifically: ; in, represents an element-wise maximum selection; In order to ensure that the fused image maintains the optimal intensity distribution while retaining rich texture details, texture loss is introduced; The texture loss is specifically: ; in, To measure the gradient operator of image texture information, the Sobel operator is used to calculate the gradient. It is an absolute value operation.

[0026] In order to further implement the above technical solutions, Figure 3In step S1, the trained illumination perception subnetwork inputs visible light images and outputs illumination probabilities; the specific architecture includes: Four 4×4 convolutional layers with a stride of 2, each followed by an LReLU activation function without changing the size of the feature map, are used to compress spatial information and extract illumination information; A global average pooling layer to integrate lighting information; Two fully connected layers calculate the illumination probability distribution based on the integrated illumination information, including the probabilities of daytime scenes and nighttime scenes; The illumination allocation mechanism calculates the illumination perception weights representing the contributions of infrared images and visible light images based on the illumination probability. Specifically, after obtaining the illumination probability, in order to use the illumination probability to construct the illumination loss to guide the training of the fusion network, a simple normalization function is used. This function can calculate the illumination perception weights representing the contributions of the source images based on the illumination probability.

[0027] In this embodiment, if Figure 4 When training the learnable fusion layer network, the weights of the codec and the illumination perception sub-network are fixed and remain unchanged; in this stage, the training first uses a fixed encoder network to extract multi-scale deep features from the source image; secondly, for each scale, a fusion layer is used to fuse these deep features; finally, the fused multi-scale features are input into the fixed decoder network to obtain the fused image; the training of the learnable fusion layer network is guided by illumination loss, auxiliary intensity loss, and texture detail loss function, where the illumination loss is calculated by the illumination perception weight.

[0028] In step S3, the infrared and visible light images are input into the trained encoder network, passing through a 1×1 convolution layer and four convolution blocks. Each convolution block contains two 3×3 convolution layers and a maximum pooling operator, which enables the encoder network to extract deep features at different scales.

[0029] In order to further implement the above technical solution, the specific content of step S3 is: S31. Input encoder to extract the depth features of infrared and visible light images at different scales and Input 4 fusion layer networks with the same structure respectively ( and represents the deep features of the nth scale extracted by the encoder network, n∈{1,2,3,4}), first, and Perform a 1×1 convolution operation, and then capture multi-scale deep features through four convolutional layers with different kernel sizes. There is no pooling layer for different convolution operations. The kernel sizes of the convolutional layers are 1×1, 3×3, 5×5, and 7×7 respectively. S32. Concatenate the obtained features according to the channel size and pass the concatenated features through a 3×3 convolutional layer. S33. Introducing the CBAM attention mechanism module, which consists of a cascaded channel attention module and a spatial attention module. The feature map output by the convolutional layer is first weighted in the channel dimension, then enhanced in the spatial dimension, and finally outputs the feature representation after dual attention optimization. S34. The optimized features are passed through two 3×3 convolutional layers to obtain the fused deep features And input into the decoder network.

[0030] In step S4, the decoder convolutional block structure of the decoder network is the same, consisting of two 3×3 convolutional layers, and the convolutional blocks are nested and connected, and then a 1×1 convolutional layer is used to obtain the final output fusion image. .

[0031] In this embodiment, if Figure 5 As a lightweight attention module, the CBAM attention mechanism module integrates both channel and spatial attention mechanisms, effectively strengthening the feature representation of key channels and regions in the feature map; in the infrared and visible light image fusion task, CBAM can help the model more accurately identify and extract important features in the source image. By enhancing the representation of these important features, CBAM can improve the quality and accuracy of the fused image; the CBAM module consists of cascaded channel attention and spatial attention modules. The feature map output by the convolutional layer is first adjusted by the channel dimension weight, and then the spatial dimension features are enhanced, and finally the feature representation after dual attention optimization is output.

[0032] In order to further implement the above technical solutions, Figure 6 , the specific content of channel dimension weight adjustment is: The input feature map is processed through parallel global maximum pooling and global average pooling to generate two spatially compressed vectors, capturing the salient region and global statistical information respectively. The two vectors are then processed element-by-element by a multi-layer perceptron (MLP) with shared parameters and normalized into a channel attention weight matrix using a sigmoid function. Finally, the original feature map is element-wise multiplied with the weight matrix to obtain the input features of the spatial attention module. Channel attention evaluates the importance of feature channels by compressing spatial dimensions (H×W→1×1). Maximum pooling strengthens local salient features, while average pooling preserves global distribution information. The two complement each other to improve feature expression capabilities. The channel attention mechanism can be expressed as: ; In order to further implement the above technical solutions, Figure 7,The specific content of spatial dimension feature enhancement is: The output features of the channel attention are subjected to channel-wise maximum pooling and average pooling to generate two single-channel spatial feature maps. The two channel features are then concatenated along the channel axis, compressed to 1 by a 7×7 convolutional layer, and then a spatial weight matrix is generated by Sigmoid. Finally, the original feature map and the spatial weight matrix are weighted to output the final optimized features. Spatial attention constructs a spatial position importance map by compressing the channel dimension (C→1). The convolutional layer can capture local spatial correlation across channels and enhance the model's ability to identify key areas. The spatial attention mechanism can be expressed as: ; Among them, σ is the sigmoid operation, f 7×7 For the convolution layer, a 7×7 convolution kernel works better than a 3×3 one.

[0033] Example 2

[0034] A light perception adaptive fusion system based on the RFN-Nest network, based on a light perception adaptive fusion method based on the RFN-Nest network, includes a model training module, a trained autoencoder network, a fusion layer network, and a light perception sub-network; the trained autoencoder network includes an encoder network and a decoder network; The model training module is used to train the codec as an autoencoder network based on infrared and visible light images, train the illumination perception subnetwork based on visible light images, and use the illumination probability output by the illumination perception subnetwork to guide the training of multiple fusion layer networks; Encoder network for extracting deep features of infrared and visible light images at different scales; The fusion layer network is used to capture and stitch the depth features of infrared and visible light images of different scales, and introduces the CBAM attention mechanism for dual attention optimization to obtain the fused depth features. The decoder network is used to input the fused deep features into the decoder network to fused multi-scale features.

[0035] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a light perception adaptive fusion method based on an RFN-Nest network.

[0036] A processing terminal includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the computer program, an illumination perception adaptive fusion method based on the RFN-Nest network is implemented.

[0037] Example 3

[0038] In this example, the visible light images and infrared images in the TNO dataset, MSRS dataset, and LLVIP dataset were used as input data, and the average gradient, standard deviation, spatial frequency, and information entropy were used as evaluation indicators to compare the RFN-Nest method with the present invention. The results are shown in the figure. Figure 8 ; The RFN-Nest method does not consider the influence of illumination factors during the modeling process. It assumes that texture only exists in visible light images. However, it ignores the fact that infrared images in low-light scenes contain richer textures than visible light images, resulting in the loss of some texture detail information and target information when fusing images under some extreme lighting conditions. In contrast, the present invention introduces an illumination perception subnetwork to guide the training of the fusion layer, so that the network can adaptively process input images with different illumination differently, that is, when the probability of judging that the input image is a daytime lighting scene is high, the fused image relatively retains more information in the visible light image. information; and when the probability of judging that the input image is a nighttime (i.e., low-light) illumination scene is high, the fused image relatively retains more information in the infrared image; moreover, the present invention improves the fusion layer network so that the fusion layer can retain more information, reduces information loss, and introduces the CBAM attention module to improve the performance of the network; in addition, when training the fusion layer network, the present invention adopts loss functions related to illumination loss, auxiliary intensity loss, and texture loss, so that the fusion layer network can adaptively integrate meaningful information according to the illumination conditions, and the fused image can retain rich texture details while maintaining the optimal intensity distribution.

[0039] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0040] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image fusion method based on RFN-Nest network and illumination perception, characterized in that: The following steps are involved: S1. Based on infrared and visible light images, the codec is trained as an autoencoder network. The illumination perception subnetwork is trained based on visible light images. The illumination probability output by the illumination perception subnetwork guides the training of multiple fusion layer networks. S2. Input infrared and visible light images into the encoder network and extract deep features of infrared and visible light images at different scales. S3. Input the deep features of infrared and visible light images of different scales into the trained fusion layer network respectively, perform multi-scale deep feature capture and feature splicing, and introduce the CBAM attention mechanism for dual attention optimization to obtain the fused deep features; S4. Input the fused deep features into the decoder network and reconstruct the fused multi-scale features.

2. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: Step S1, the training content of the autoencoder network is: The encoder and decoder are directly connected. The encoder network downsamples the input image to extract multi-scale depth features. The decoder reconstructs the input image with multi-scale depth features. The encoder-decoder network is trained without the fusion layer. The total loss function of the autoencoder network is: ; ; ; in, and represents the pixel loss and structural similarity loss between the input image I and the reconstructed image O, β is the trade-off parameter between pixel loss and structural similarity loss, is the F norm, It is a structural similarity measure used to quantify the structural similarity between two images.

3. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: Step S1: Use cross entropy loss to constrain the training process of the illumination perception sub-network, specifically: ; Where m is the lighting label of the input image, is the illumination probability output by the illumination perception subnetwork, and σ is the softmax function, which normalizes the illumination probability to [0,1].

4. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: In step S1, the illumination probability output by the illumination perception sub-network guides the training of multiple fusion layer networks with a loss function of: L fusion =α1L ill +α2L aux +α3L tex Among them, α1, α2, and α3 are trade-off parameters, and L ill is the illumination loss, L aux is the auxiliary strength loss, L tex For texture loss.

5. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: The loss function of the fusion layer network includes illumination loss, auxiliary intensity loss, and texture loss; the specific contents are: The specific illumination loss is: ; ; ; in, and are the intensity losses of infrared and visible light images, respectively, K ir and K vi The illumination perception weights contributed by infrared images and visible light images, P x The output of the illumination perception sub-network is the probability that the illumination probability belongs to the night scene, P y is the probability of belonging to the daytime scene; The intensity loss of infrared and visible light images is specifically: ; ; Where H is the height of the input image, W is the width of the input image, is the L1 norm, I ir and I vi are the input infrared image and visible light image, I f is the output fused image; The auxiliary strength loss is specifically: ; in, represents an element-wise maximum selection; The texture loss is specifically: ; in, To measure the gradient operator of image texture information, the Sobel operator is used to calculate the gradient. It is an absolute value operation.

6. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: In step S1, the trained illumination perception subnetwork takes visible light images as input and outputs illumination probabilities. The specific architecture includes: Four 4×4 convolutional layers with a stride of 2, each followed by an LReLU activation function without changing the size of the feature map, are used to compress spatial information and extract illumination information; A global average pooling layer to integrate lighting information; Two fully connected layers calculate the illumination probability distribution based on the integrated illumination information, including the probabilities of daytime scenes and nighttime scenes; The illumination distribution mechanism calculates the illumination perception weights representing the contributions of infrared images and visible light images based on the illumination probability.

7. The image fusion method based on RFN-Nest network and illumination perception according to claim 1, characterized in that: The specific content of step S3 is: S31. Input the infrared and visible light image depth features extracted by the encoder, perform a 1×1 convolution operation, and then pass it through four convolutional layers with different kernel sizes to capture multi-scale depth features. There is no pooling layer for each convolution operation. The kernel sizes of the convolutional layers are 1×1, 3×3, 5×5, and 7×7, respectively. S32. Concatenate the obtained features according to the channel size and pass the concatenated features through a 3×3 convolutional layer. S33. Introducing the CBAM attention mechanism module, which consists of a cascaded channel attention module and a spatial attention module. The feature map output by the convolutional layer is first weighted in the channel dimension, then enhanced in the spatial dimension, and finally outputs the feature representation after dual attention optimization. S34. The optimized features are passed through two 3×3 convolutional layers to obtain fused deep features and input into the decoder network.

8. The image fusion method based on RFN-Nest network and illumination perception according to claim 7, characterized in that: The specific content of channel dimension weight adjustment is: The input feature map generates two spatially compressed vectors through parallel global maximum pooling and global average pooling, which capture the salient area and global statistical information respectively. The two vectors are then processed by a multi-layer perceptron (MLP) with shared parameters and superimposed element by element. They are normalized into a channel attention weight matrix using the Sigmoid function. Finally, the original feature map and the weight matrix are element-wise multiplied to obtain the input features of the spatial attention module.

9. The image fusion method based on RFN-Nest network and illumination perception according to claim 8, characterized in that: The specific contents of spatial dimension feature enhancement are as follows: The output features of the channel attention are subjected to maximum pooling and average pooling in the channel dimension to generate two single-channel spatial feature maps. The two channel features are then concatenated along the channel axis, the channel dimension is compressed to 1 through a 7×7 convolutional layer, and a spatial weight matrix is generated through Sigmoid. Finally, the original feature map and the spatial weight matrix are weighted and calculated to output the final optimized features.

Citation Information

Patent Citations

  • Image tampering detection method based on adaptive spatial feature fusion

    CN115311204A

  • Multi-exposure image fusion method based on multi-scale auto-encoder

    CN115689962A

  • Self-attention progressive network and method for multi-source data fusion

    CN118279708A

  • Image color imitation method based on color migration model

    CN118967536A

  • Intrauterine device ultrasound image segmentation method and model construction method thereof

    CN119741316A