A multi-scene dual-band image fusion method, device and storage medium

By designing a multi-scene dual-band image fusion method, using a scene recognition network and an image fusion network, combining a multi-scene feature extraction module and an attention feature extraction module, the problems of single application scenarios and low fusion efficiency in the existing technology are solved, and high-quality image fusion for different scenarios are achieved.

CN118037561BActive Publication Date: 2025-05-09CHONGQING RES INST OF CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410077825.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-05-09
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

The existing infrared image and visible light image fusion method has a single application scenario, low fusion efficiency, and insufficient extraction of features in different scenes.

Method used

A multi-scene dual-band image fusion method is designed. By constructing an end-to-end network model including scene recognition network and image fusion network, the multi-scene feature extraction module and attention feature extraction module are used to adjust the loss function according to different scenarios to improve the efficiency of image fusion and the adequacy of feature extraction.

Benefits of technology

High-quality image fusion in different scenarios is realized, the efficiency of image fusion is improved, the ability to extract features of different scenarios is enhanced, and the impact on the fusion effect is reduced due to different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118037561B_ABST
    Figure CN118037561B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image fusion technology, and in particular, is an image fusion method, device, and storable medium. The present invention includes: preparing a data set, acquiring an infrared image and a visible light image; denoising, enhancing, and operating the acquired infrared image and visible light image; sending the pre-processed visible light image and infrared image to a scene recognition network to obtain the scene of the image, and then sending the infrared image and the visible light image to an image fusion network, performing different feature extraction, feature fusion, and feature reconstruction according to different scenes to obtain a fused image. The present invention can select a fusion strategy by judging the image scene, so that the fused image contains more information and is more in line with the visual effect of the human eye.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a multi-scene dual-band image fusion method, device and storable medium. Background Art

[0002] Image fusion is a technology that combines image information from different sensors or different modalities to obtain more comprehensive image information. Image fusion technology has become crucial in various fields, such as medical imaging, security monitoring and autonomous driving. Infrared images mainly capture the thermal radiation of target objects, so they can be used to detect temperature differences of objects and are not affected by light. Images can still be captured under low light or completely dark conditions. In addition, infrared radiation can penetrate certain substances, such as smoke, haze and some materials, so it can also provide useful information in bad weather. The color and brightness of visible light images are consistent with human eye perception and have wide applications in medical diagnosis and security monitoring, but the quality of their images is affected by lighting conditions. Fewer details may be captured under low light conditions. Therefore, we propose a multi-scene dual-band image fusion method to solve the above problems.

[0003] The Chinese patent publication number is "CN116051442A", and its name is "Visible light and infrared image fusion method based on three-branch autoencoder network". This method first inputs the pre-processed infrared image and visible light image to be fused; then extracts the features of the input image through the three-branch autoencoder network; then decodes the fused infrared and visible light features through the decoder to obtain the fused image. The fused image obtained by this method effectively extracts the features of the visible light image and the infrared image, but the application scenario is single and the fusion efficiency is low. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] In view of the shortcomings of the prior art, the present invention provides a multi-scene dual-band image fusion method, which solves the problems of the existing infrared image and visible light image fusion method having a single application scenario, low fusion efficiency, and insufficient feature extraction of different scenes.

[0006] (II) Technical solution

[0007] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0008] A multi-scene dual-band image fusion method, device and storable medium, including three aspects;

[0009] In a first aspect, the present invention provides an image fusion method, comprising the following steps:

[0010] Step 1: Prepare the data set: prepare infrared and visible light image data sets for training and testing scene recognition and image fusion networks, and enhance the images to improve image details;

[0011] Step 2: Build a network model: The entire network is divided into a scene recognition network and an image fusion network. The fusion network includes an encoder, a fusion module, and a decoder. The scene recognition network includes a convolutional neural network and a graph convolutional network.

[0012] Step 3, training the network model: training the scene recognition network model, inputting the data set preprocessed in step 1 into the scene recognition network model constructed in step 2 for training, obtaining the scene training weight, and then inputting the data set into the fusion network model to obtain the fusion training weight;

[0013] Step 4, select a suitable loss function and determine the optimal evaluation index of this method: the scene recognition network model selects a suitable loss function to minimize the loss between the output and the manual true label value, and the fusion network model selects a suitable loss function to minimize the loss between the output fusion image and the input image. Set the training loss threshold, and continuously iterate and optimize the model until the number of training times reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can be considered to have been pre-trained, save the model parameters, select the image in the data set and input it into the model to obtain the fusion image, and use the optimal evaluation index of the fusion image effect to measure the accuracy and performance of the model;

[0014] Step 5, determine the model: solidify the network model parameters and determine the final network model. For example, when performing the infrared image and visible light image fusion task, the multimodal image can be directly input into the trained end-to-end network model to obtain the final fused image;

[0015] Furthermore, in step 1, the scene recognition part is trained using the Places2 dataset, and the infrared image and visible light image fusion part uses the TNO dataset.

[0016] Furthermore, in step 2, the convolutional neural network in the scene recognition network is a pre-trained ResNet-50 model.

[0017] Furthermore, in step 2, the encoder of the image fusion network is composed of a convolution block, a multi-scene feature extraction module, and an attention feature extraction module, and the decoder network is composed of multiple convolution blocks.

[0018] Furthermore, in step 2, the multi-scene feature extraction module uses different strategies to extract features of images of different scenes, and extracts features of different scenes by adjusting the number of convolution layers and the size of the convolution kernel.

[0019] Furthermore, in step 2, the activation functions of the first to eighth convolution blocks use R-type functions; the sizes of the convolution kernels in all convolution blocks are unified to n×n.

[0020] Furthermore, in step 2, the attention feature extraction module is composed of channel attention and spatial attention.

[0021] Furthermore, in step 3, the scene recognition data set preprocessing requires processing the images into the same size.

[0022] Furthermore, in step 4, the scene recognition network loss function uses cross entropy loss, and the infrared image and visible light image fusion network uses perceptual loss and texture loss.

[0023] The cross entropy loss is used to measure the difference between the output of the network and the true label, minimize the cross entropy loss, make the network's prediction as close to the true label as possible, and improve the accuracy of classification.

[0024] The perceptual loss is used to measure the degree of similarity between the feature representation of the generated image and the feature representation of the real image, and minimize the perceptual loss, thereby ensuring that the generated image is perceptually similar to the real image.

[0025] The texture loss can enable the fused image to retain rich texture details, thereby visually ensuring that the generated image is similar to the real image and making the image look more natural.

[0026] In a second aspect, the present invention provides an image fusion device, comprising: an acquisition module, a preprocessing module, and a fusion module.

[0027] The acquisition module is used to acquire infrared images and visible light images.

[0028] The preprocessing module is used to perform denoising, enhancement and grayscale operations on the infrared image and visible light image collected by the acquisition module.

[0029] The fusion module is used to perform an image fusion operation on the visible light image and the infrared image processed by the preprocessing module to obtain a fused image.

[0030] In a third aspect, the present invention provides a computer storable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the image fusion method as described in any one of the first aspects.

[0031] (III) Beneficial effects

[0032] Compared with the prior art, the present invention provides a multi-scene dual-band image fusion method, which has the following beneficial effects:

[0033] The present invention designs an infrared image and visible light image fusion network embedded with scene recognition, which can input different scenes for recognition and then perform image fusion, thereby improving the efficiency of image fusion and solving the problem of the single applicable scene of the existing infrared image and visible light image fusion. Different image scenes use the same neural network, which increases the versatility of the network.

[0034] The present invention, by designing a multi-scene feature extraction module, can perform different feature extraction according to different scenes, and then fuse the image features, so as to perform more sufficient feature extraction for different scenes and better capture the details in different scenes.

[0035] The present invention designs a loss function adjusted according to the scenario and uses different loss functions for training in different scenarios, so that the trained model can be more suitable for the corresponding scenario and further reduce the impact of problems caused by different scenarios on the fusion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a multi-scene dual-band image fusion method of the present invention;

[0037] Figure 2 This is a schematic diagram of the structure of a multi-scene dual-band image fusion device of the present invention;

[0038] Figure 3 This is an overall working principle diagram of a multi-scene dual-band image fusion method of the present invention;

[0039] Figure 4 A network structure diagram for scene recognition of the present invention;

[0040] Figure 5 A graph of a convolutional neural network in a scene recognition network of the present invention;

[0041] Figure 6 This is a network structure diagram of the present invention that integrates infrared image and visible light image;

[0042] Figure 7 This is a schematic diagram of the attention feature extraction module of the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] Example

[0045] like Figure 1 As shown, a multi-scene dual-band image fusion method proposed in one embodiment of the present invention includes the following steps:

[0046] Step 1: Prepare the data set. Prepare the data set for training the scene recognition network and the image fusion network. The scene recognition network is trained using the Places365-Standard in the Places2 data set. A total of 16,000 images of three classified scenes are selected for training. The image fusion network is trained using the TNO data set.

[0047] Step 2: Build a network model. Figure 4 This is the schematic diagram of the scene recognition network. The scene recognition network consists of a convolutional neural network and a graph convolutional neural network. The convolutional neural network is a pre-trained ResNet 50 network that can obtain the global features of the image. It adaptively selects several feature points on the global feature map and regards them as vertices of the graph structure in the form of feature vectors. The points are input into the graph neural network, and then the results are converted into weights. The global features are calibrated and the global feature vectors are multiplied element by element by the weights to obtain enhanced global features. The enhanced global features are then input into the fully connected layer and the Softmax layer to obtain the scene score. Figure 3 This is the schematic diagram of the image fusion network, which includes an encoder and a decoder. The encoder consists of convolution block 1, convolution block 2, convolution block 3, convolution block 4, multi-scene feature extraction module and attention feature extraction module. The convolution kernel size of convolution blocks 1 to 4 is 3×3, and the step size and padding are both set to 1. All activation functions use R-type functions. The size of the convolution kernel in the multi-scene feature extraction module can be changed according to different scenes. The depth of the convolution layer is also related to the scene category. The schematic diagram of the attention feature extraction module is shown in the figure. The visible light features and infrared features extracted from convolution blocks 1 to 4 are spliced ​​respectively, and then the attention features are extracted. The extracted fusion features are involved in the image fusion. The decoder consists of convolution block 5, convolution block 6, convolution block 7, and convolution block 8. The convolution kernel size is 3×3, the step size and padding are both 1, and the activation function is an R-type function. The R-type function is defined as follows.

[0048]

[0049] First, the multi-scene feature extraction module extracts different features for different scenes, and then convolution blocks one to four and the attention feature extraction module fully extract the image features. After the extracted features are spliced, convolution blocks five to eight perform feature reconstruction to obtain the final fused image.

[0050] Step 3: Train the network model. Use the prepared data set to train the scene recognition network, and then train the scene recognition network and the image fusion network.

[0051] Step 4: Select a suitable loss function and determine the optimal evaluation index of this method. The scene recognition network model selects a suitable loss function to minimize the loss between the output and the manual true label value. The fusion network model selects a suitable loss function to minimize the loss between the output fusion image and the input image. Set the training loss threshold, and continuously iterate and optimize the model until the number of training times reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can be considered to have been pre-trained and the model parameters can be saved.

[0052] Cross entropy loss is used in scene recognition networks to measure the difference between the network output and the true label, minimize the cross entropy loss, make the network's prediction as close to the true label as possible, and improve the accuracy of classification. The specific formula is as follows:

[0053] L=-z logσ(y)-(1-z)log(1-σ(y)) (2)

[0054] where z represents the scene label of the input image, y refers to the output of the scene-aware network, and σ refers to the Softmax function that normalizes the scene probability to [0, 1].

[0055] The image fusion network uses perceptual loss and texture loss. The goal of perceptual loss is to make the feature representation of the generated image as close as possible to the feature representation of the real image, so as to ensure that the generated image is perceptually similar to the real image. per The specific formula is as follows:

[0056]

[0057] in, is the output image, F(y) i is the input image, ||·||2 is the L2 norm, N is the total number of pixels in the image, and y is the value of the pixel.

[0058] Texture loss can make the fused image retain rich texture details, thus visually ensuring that the generated image is similar to the real image and making the image look more natural. tex The specific formula is as follows:

[0059]

[0060] In summary, the total loss can be expressed as follows:

[0061] L total =L per +λLtex (5)

[0062] Where λ is a parameter used to adjust the two loss weights.

[0063] The evaluation indexes include information entropy (EN), standard deviation (SD), spatial frequency (SF), average gradient (AG), and mutual information (MI). Information entropy is an index used to measure the amount of image information. The larger the entropy, the more information the image contains. Infrared image fusion usually involves the fusion of image information from multiple sensors. Therefore, information entropy is also an important evaluation index. Its expression is as follows:

[0064]

[0065] Among them, p1 represents I f How often the median value 1 is used.

[0066] Standard deviation is a statistical concept used to quantify I fused The standard deviation can be used to measure the brightness change of the image. Generally, the smaller the standard deviation, the smaller the brightness change of the image, which means that the overall brightness uniformity of the image is better. On the contrary, it means that the brightness change of the image is large and the overall brightness is not uniform. The specific formula is as follows:

[0067]

[0068] in, Indicates I fused The average pixel value.

[0069] Spatial Frequency Measurement I fused The texture details in the spatial frequency index are measured by calculating the gradient difference between the fused image and the original image to measure the detail richness of the fused image. The specific formula is as follows:

[0070]

[0071] Among them, RF fused and CF fused denote the row frequency and column frequency respectively,

[0072]

[0073]

[0074] The average gradient can be obtained by calculating the gradient average of the image grayscale value. The gradient refers to the rate of change of the image at a certain position, which can usually be calculated using operators such as Sobel and Prewitt. The larger the gradient value, the more drastic the change of the image at that position, and the more likely it is that there is detail information. In infrared image fusion, the average gradient can be used to evaluate the clarity of the image and the preservation of detail information. If the clarity of the fused image is high and the detail information is well preserved, the average gradient value will be correspondingly high. Generally, the higher the average gradient value, the better the clarity of the image and the better the detail information is preserved. The specific formula is,

[0075]

[0076]

[0077] Mutual information measures the information contained in image X from image Y and is used to measure the similarity between two images. The specific formula is:

[0078]

[0079] During network training, the learning rate is set to 0.0001, the batch size is set to 8, and a total of 10,000 iterations are performed. The Adam optimizer is used to update the network parameters, and the exponential decay rate and eps value are set to (0.9, 0.009) and 1e-08, respectively. The weight value of the loss function is set to be scene-dependent, and different weights, such as 0.1 or 1.1, are used for different scenes.

[0080] Step 5, determine the model. After the network training is completed, all parameters in the network need to be saved. Then, the registered infrared image and visible light image are input into the network to obtain the fused image.

[0081] Among them, convolution, activation function, splicing operation, and ResNet 50 implementation are algorithms well known to those skilled in the art, and the specific processes and methods can be found in corresponding textbooks or technical literature.

[0082] The present invention uses a multi-scene dual-band image fusion method to achieve high-quality fusion effects of different scenes, reducing the impact of different scenes on the fusion quality of infrared images and visible light images. By calculating the relevant indicators of the image obtained by the existing method, the feasibility and superiority of the method are further verified. The comparison of the relevant indicators of the existing technology and the method proposed in the present invention is shown in Table 1:

[0083] Table 1

[0084]

[0085] It can be seen from the table that the method proposed in the present invention has higher image contrast, edge strength, spatial frequency, information entropy, average gradient and standard deviation. These indicators also further illustrate that the method proposed in the present invention has better fusion image quality.

[0086] like Figure 2 As shown, a multi-scene dual-band image fusion device proposed in one embodiment of the present invention includes an acquisition module, a preprocessing module, and a fusion module.

[0087] The acquisition module is used to acquire infrared images and visible light images.

[0088] The preprocessing module is used to perform denoising, enhancement and grayscale operations on the infrared image and visible light image collected by the acquisition module.

[0089] The fusion module is used to perform an image fusion operation on the visible light image and the infrared image processed by the preprocessing module to obtain a fused image.

[0090] The present invention provides a computer storable medium, wherein the computer storable medium stores an instruction set, a code set, and a program, and the instruction set, code set, and program can implement each step in the aforementioned method when executed by a computer processor; optionally, the computer storable medium includes: a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, and other media that can store computer program codes.

[0091] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-scene dual-band image fusion method, characterized in that: The method comprises the following steps: Step 1: Prepare the data set and obtain infrared images and visible light images; Step 2: Build a network model, including a scene recognition network and an image fusion network; The step 2 includes a scene recognition network and an image fusion network; wherein the scene recognition network includes a pre-trained convolutional neural network, a graph neural network, a fully connected layer and a Softmax layer; the convolutional neural network is used to extract global features of the image, the graph neural network is used to enhance global features, and the fully connected layer and the Softmax layer are used to classify features to obtain scene scores; The image fusion network includes a decoder and an encoder; the encoder includes convolution block 1, convolution block 2, convolution block 3, convolution block 4, a multi-scene feature extraction module and an attention feature extraction module, and the decoder includes convolution block 5, convolution block 6, convolution block 7 and convolution block 8; the encoder is used to extract features from infrared images and visible light images to obtain infrared features and visible light features, and the decoder is used to reconstruct the spliced ​​infrared features and visible light features to obtain a fused image; The convolution block 1, the convolution block 2, the convolution block 3, the convolution block 4, the convolution block 5, the convolution block 6, the convolution block 7, and the convolution block 8 all include convolution layers, the convolution kernel size is 3×3, the step size and the padding are 1, and the activation function is an R-type activation function; The multi-scene feature extraction module performs different feature extractions according to different image scenes based on an algorithm; The size of the convolution kernel in the multi-scene feature extraction module changes according to the scene, and the depth of the convolution layer is also related to the scene category; The attention feature extraction module adds the infrared image and visible light image features extracted by convolution blocks 1 to 4, obtains fusion features through the attention network, and participates in the feature fusion of the image; Step 3: Train the network model, use the prepared data set to train the scene recognition network, and then train the scene recognition network and image fusion network; Step 4, select a suitable loss function and determine the optimal evaluation index of this method. The scene recognition network model selects a suitable loss function to minimize the loss between the output and the manual true label value. The fusion network model selects a suitable loss function to minimize the loss between the output fusion image and the input image. Set the training loss threshold, and continuously iterate and optimize the model until the number of training times reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can be considered to have been pre-trained, save the model parameters, select the image in the data set and input it into the model to obtain the fusion image, and use the optimal evaluation index of the fusion image effect to measure the accuracy and performance of the model; Step 5: Determine the model, solidify the network model parameters, determine the final network model, and when performing the infrared image and visible light image fusion task, directly input the multimodal image into the trained end-to-end network model to obtain the final fused image.

2. The multi-scene dual-band image fusion method according to claim 1, characterized in that: In step 1, the Places2 dataset is used for training the scene recognition network, image data of different scenes are used to annotate the scene of each image, and the TNO dataset is used for fusion of infrared images and visible light images.

3. The multi-scene dual-band image fusion method according to claim 1, characterized in that: The step 4 selects a suitable loss function, specifically, the scene recognition network uses cross entropy loss; the image fusion network uses perceptual loss and texture loss.

4. The multi-scene dual-band image fusion method according to claim 3, characterized in that: The cross entropy loss is used to measure the difference between the network output and the true label, minimize the cross entropy loss, make the network prediction as close to the true label as possible, and improve the accuracy of classification; The goal of the perceptual loss is to make the feature representation of the generated image as close as possible to the feature representation of the real image, thereby ensuring that the generated image is perceptually similar to the real image; The texture loss enables the fused image to retain rich texture details, thereby visually ensuring that the generated image is similar to the real image and making the image look more natural.

5. The multi-scene dual-band image fusion method according to claim 1, characterized in that: The evaluation index of step 4 is: Including,information entropy (EN), standard deviation (SD), spatial frequency (SF), average gradient (AG), and mutual information (MI) as indicators for evaluating the image fusion effect.

6. A multi-scene dual-band image fusion device, comprising an acquisition module, a preprocessing module, and a fusion module, characterized in that: The acquisition module is used to acquire infrared images and visible light images; The preprocessing module is used to perform denoising, enhancement and grayscale operations on the infrared image and visible light image collected by the acquisition module; The fusion module is used to perform an image fusion operation on the visible light image and the infrared image processed by the preprocessing module to obtain a fused image. The multi-scene dual-band image fusion device implements the image fusion method as described in any one of claims 1 to 5 above.

7. A computer storable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image fusion method as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method and device and computer storage medium

    CN114022742A

  • Unpaired infrared image colorization method based on comparative learning

    CN116503502A