Infrared light and visible light image fusion method, device and system, and storage medium
Through the combination of image quality evaluation network and multimodal image fusion network, the degradation problem of infrared light and visible light images is automatically judged and processed, and the high-quality fusion of infrared light and visible light images is achieved, solving the problem of poor fusion effect caused by ignoring the degradation problem in the prior art.
Patent Information
- Application Number
- CN202510115104.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
The existing infrared light and visible light image fusion method ignores the degradation problem of source images, resulting in the inability to obtain the most ideal fusion results intelligently and adaptably.
An infrared light and visible light image fusion method is adopted to determine the degradation type through image quality evaluation network, acquire the degraded text description features, and input them into the multimodal image fusion network for feature learning and fusion. This method uses adaptive optimization to adjust the target loss function of the fusion network to realize the fusion of multimodal images of any quality input.
There is no need to manually enter the prompt text, automatically determine the type of degradation, reduce labor costs, and avoid the consistency of human subjective judgments. This method can flexibly handle multiple degradation types, improving the quality and generalization capabilities of fusion results.
Smart Images

Figure CN119941530A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image processing, and in particular relates to a method and device, a system, and a storage medium for fusing infrared light and visible light images. Background Art
[0002] Existing infrared and visible light image fusion methods mostly focus on processing directly based on image features, while ignoring the complexity of degradation factors and their impact on the fusion effect. For example, problems such as low contrast, blur, and noise can significantly reduce the quality of the fusion result. In recent years, the CLIP model has performed well in matching text and image features, but it has not been fully applied to guide image fusion to solve the problems caused by source image degradation.
[0003] In recent years, the latest related research has used semantic text guidance for degradation perception and interactive image fusion to start preliminary exploration of the above problems. For example, Text-IF allows users to input text to achieve text-guided interactive degraded image fusion. However, it does not handle different degradation combinations where both source images are degraded well, so the resulting network is difficult to generalize and adapt to the needs of multimodal image fusion in various degradation situations. At the same time, although the degradation quality of single-modal images is considered in related methods, manual input of guidance text is still required during the fusion process, and true full-process intelligent interpretation and processing cannot be achieved. In addition, when considering the combination of relevant fusion losses and degradation problem-related losses during network training, the weights of each loss are fixed, and it is impossible to adaptively constrain the network adjustment direction according to the characteristics of the input data. Therefore, there is an urgent need for an image fusion technology combined with text guidance to achieve adaptive multimodal image fusion processing for various degradation types. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method and device, system and storage medium for fusion of infrared light and visible light images, so as to solve the problem that the existing image fusion ignores the defect of degradation of the source image, resulting in the inability to obtain the most ideal fusion result through intelligent adaptation.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] A method for fusing infrared light and visible light images, comprising:
[0007] Step S1, obtaining a visible light-infrared light image;
[0008] Step S2: input the visible light-infrared light image into the image quality assessment network to identify the degradation type and obtain the degradation text description feature;
[0009] Step S3, inputting the degraded text description features and the visible light-infrared light image into a multimodal image fusion network for feature learning and fusion;
[0010] Step S4, calculating multiple loss items based on the fusion result to obtain the fusion network target loss function;
[0011] Step S5: adaptively optimize and adjust the fusion network target loss function to obtain a finally trained multimodal fusion network system and realize the fusion of multimodal images with any quality input.
[0012] Preferably, in step S2, CLIP-IQA is used to perform image quality assessment to discriminate multiple image degradation combination types, and corresponding degradation text description features are read according to the degradation types.
[0013] Preferably, in step S3, the degraded text description features and the visible light-infrared light image are input into a multimodal image fusion network, and the features of spatial and deep information of the visible light image and the infrared light image are extracted through TransformerBlock and cross-attention mechanism.
[0014] Preferably, in step S4, the multiple loss items include: frequency domain loss, consistency loss, text loss, color loss and maximum loss.
[0015] The present invention also provides an infrared light and visible light image fusion device, comprising:
[0016] A first processing module, used for acquiring a visible light-infrared light image;
[0017] The second processing module is used to input the visible light-infrared light image into the image quality assessment network to identify the degradation type and obtain the degradation text description feature;
[0018] The third processing module is used to input the degraded text description features and the visible light-infrared light image into the multimodal image fusion network for feature learning and fusion;
[0019] The fourth processing module is used to calculate multiple loss items based on the fusion result to obtain the fusion network target loss function;
[0020] The fifth processing module is used to adaptively optimize and adjust the target loss function of the fusion network to obtain the final trained multimodal fusion network system and realize the fusion of multimodal images with any quality input.
[0021] Preferably, the second processing module is used to use CLIP-IQA to perform image quality assessment to achieve image multi-degradation combination type discrimination, and read corresponding degradation text description features according to the degradation type.
[0022] Preferably, the third processing module is used to input the degraded text description features and the visible light-infrared light image into the multimodal image fusion network, and extract the features of spatial and deep information of the visible light image and the infrared light image through TransformerBlock and cross-attention mechanism.
[0023] Preferably, the multiple loss items include: frequency domain loss, consistency loss, text loss, color loss and maximum loss.
[0024] The present invention also provides an infrared light and visible light image fusion system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the infrared light and visible light image fusion method when executed by the processor.
[0025] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the infrared light and visible light image fusion method when running.
[0026] The present invention does not require manual input of prompt text, and the judgment of degradation type is handed over to the network according to the image quality evaluation results, which reduces labor costs while avoiding the consistency problem of human subjective judgment. In terms of degradation processing types, the combination of source image degradation types is more flexible. Compared with Text-IF which only considers 4 different situations of single image degradation, the present invention considers 9 situations more comprehensively. In fact, it is not limited to these nine types. Based on the technical strategy of the present invention, more multimodal image degradation problem evaluation and removal fusion can be integrated. In addition, in addition to considering the spatial domain information considered by most networks, the present invention additionally considers the frequency domain information of the image, and better copes with the impact of source image degradation through the frequency domain loss that can better reflect the degradation effect, and in the process of constructing the loss objective function, adaptive adjustments are made to the degradation characteristics of the input data, and the constraint direction is continuously adjusted according to the degradation type during the training process, which further improves the generalization of the fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0028] Figure 1 This is a flow chart of the infrared light and visible light image fusion method according to an embodiment of the present invention;
[0029] Figure 2 Schematic diagram of various degradation type data sets;
[0030] Figure 3 Output a result plot for the model. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Embodiment 1:
[0034] like Figure 1 As shown, an embodiment of the present invention provides a method for fusing infrared light and visible light images, comprising:
[0035] Step 1: Select a pair of visible light and infrared light images from the random image types. The image type selection of the image pair includes two types: affected by degradation and not affected by degradation. The degradation types of the visible light image and the infrared light image are as follows: Figure 2 As shown, there are 9 combinations of degradation types that can be randomly generated in the end.
[0036] Step 2: Send the selected image pairs to the image quality assessment network, and use the CLIP-IQA technology to obtain the scores of 'Quality', 'Brightness', 'Sharpness', 'Noisiness', 'Colorfulness', and 'Contrast' of the corresponding images, and then judge the degradation type corresponding to the two images based on the scores. Subsequently, read the corresponding degradation text description statement according to the degradation type, and splice the degradation type description texts of the two modalities as the final fusion guidance text, which is sent to the fusion network together with the source image.
[0037] Step 3: Load the CLIP model, select ViT-B / 32 for the image encoder, use tokenize to perform feature mapping on the guidance text obtained in step 2, and pass the mapped guidance text and the source image into the fusion network.
[0038] Step 4: In the fusion network, four levels of Transformer (encoder_level1 to encoder_level4) are constructed in the encoder, and each level contains a specified number of TransformerBlocks. As the number of levels increases, the dimension of the feature map is expanded by multiplying by a power of 2 to capture higher-level features. Two-way cross-attention calculations are performed on the features of the two different modalities layer by layer to fully bridge and associate the semantic information of the feature space of the two modalities, and the result is residually connected with the input at the output to retain the input feature information and enhance gradient propagation.
[0039] Step 5: Perform feature fusion on the features of each level, normalize the fusion results first, and then convolve to generate Query, Key and Value. Use Query and Key to calculate the attention weight to represent the dependency relationship between each position and other positions. Use the attention weight to weight the sum of the Value to generate the output features after spatial fusion. Finally, the residual connection output features and input features are used to obtain the final output.
[0040] Step 6: Input the fusion result after the spatial attention mechanism calculation together with the text features into the dynamic feature adjustment module based on text embedding, and guide the affine transformation of the input feature map through the text information. Specifically, use MLP to map the text embedding, generate adjustment parameters gamma and beta that match the number of channels of the input feature map, and reshape the input feature map according to gamma and beta in the manner of x = (1 + gamma) * x + beta. Subsequently, the feature maps obtained from each layer are input into the decoder, and the features are decoded and processed layer by layer through the stacked Transformer blocks to extract high-level semantic information and decode to restore the fused image.
[0041] Step 7: During the training process, the groundtruth images of the two modalities and the fused image obtained in step 6 are input when calculating the loss function. The corresponding frequency domain loss, consistency loss, text loss, color loss and maximum loss are calculated based on the input. The frequency domain loss is specifically calculated by comparing the difference between the Fourier transform (FFT) of the original image and the generated image, and the difference is calculated at three different sizes: original size, 256*256 and 128*128. By calculating the frequency domain difference at multiple different scales, the model can capture the frequency characteristics at different sizes, thereby improving the quality of the generated image. Ultimately, it ensures that the generated images are not only similar at the pixel level, but also consistent in the frequency domain, further improving the quality of the generated image.
[0042] Step 8: During the training process, record the frequency domain loss, consistency loss, text loss, color loss and maximum loss obtained each time, and update the loss weight every 4 epochs. Specifically, first set the initial loss component weight to [0.2, 0.2, 0.2, 0.2, 0.2] to facilitate the initial total loss calculation. Next, for each loss component, first calculate the rate of change of the loss, then calculate the average loss value of each loss component, and finally, call the softmax method to calculate the weight of each loss component.
[0043] Step 9: After multiple rounds of training, the final fusion model is obtained, which can input two degraded visible light and infrared light images to obtain a fusion result that is not affected by degradation.
[0044] As an implementation method of the present invention, the weights of each loss item in the loss function during the multimodal image fusion network training process are not fixed, and a strategy for adaptively adjusting the weights of each loss item is proposed. Specifically, the values of each loss item in 5 rounds are first collected, and then the importance of each current loss item is calculated, and the weight is adjusted according to the importance. Calculated, where s ik represents the rate of change of the kth loss component at the i-th iteration. It is usually expressed as s ik =f k (x i )-f k (x i -1). That is, the difference between the current and previous losses, β is an adjustable hyperparameter used to determine the weight distribution strategy
[0045] The present invention implements image quality assessment by using a multimodal image quality degradation type judgment network to obtain each image degradation type and the corresponding degradation type text description. The identified degradation type text will be further projected by the CLIP model to the same embedding space as the image features, and input into the image fusion network together with the source image. In the image fusion network, the text semantic encoder and semantic interaction fusion decoder of TransformerBlock and cross-attention technology will generate the corresponding fused image while completing the correction of image degradation problems. At the same time, the fusion network uses a combination loss function including maximum loss, consistency loss, text loss and frequency loss to deal with various degradation problems existing in the source image. In addition, the loss training of the image fusion network of this system can adaptively adjust the weights of each loss item, so that the training can adaptively adjust the loss according to the different degradation types of the data image. Therefore, the multimodal image fusion system of the present invention has extremely strong robustness and generalization ability, and can fuse and process any source images with different degradation types and give high-quality fusion results. The specific fusion result examples are as follows. Figure 3 shown.
[0046] Embodiment 2:
[0047] The embodiment of the present invention further provides an infrared light and visible light image fusion device, comprising:
[0048] A first processing module, used for acquiring a visible light-infrared light image;
[0049] The second processing module is used to input the visible light-infrared light image into the image quality assessment network to identify the degradation type and obtain the degradation text description feature;
[0050] The third processing module is used to input the degraded text description features and the visible light-infrared light image into the multimodal image fusion network for feature learning and fusion;
[0051] The fourth processing module is used to calculate multiple loss items based on the fusion result to obtain the fusion network target loss function;
[0052] The fifth processing module is used to adaptively optimize and adjust the target loss function of the fusion network to obtain the final trained multimodal fusion network system and realize the fusion of multimodal images with any quality input.
[0053] As an implementation mode of the embodiment of the present invention, the second processing module is used to use CLIP-IQA to perform image quality assessment to achieve image multi-degradation combination type discrimination, and read the corresponding degradation text description features according to the degradation type.
[0054] As an implementation mode of an embodiment of the present invention, the third processing module is used to input the degraded text description features and the visible light-infrared light image into the multimodal image fusion network, and extract the features of spatial and deep information of the visible light image and the infrared light image through the TransformerBlock and cross-attention mechanism.
[0055] As an implementation manner of the embodiment of the present invention, the multiple loss items include: frequency domain loss, consistency loss, text loss, color loss and maximum loss.
[0056] Embodiment 3:
[0057] An embodiment of the present invention further provides an infrared light and visible light image fusion system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes an infrared light and visible light image fusion method when executed by the processor.
[0058] Embodiment 4:
[0059] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the infrared light and visible light image fusion method when running.
[0060] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for fusing infrared light and visible light images, characterized in that: include: Step S1, obtaining a visible light-infrared light image; Step S2, inputting the visible light-infrared light image into the image quality assessment network to identify the degradation type and obtain the degradation text description feature; Step S3, inputting the degraded text description features and the visible light-infrared light image into a multimodal image fusion network for feature learning and fusion; Step S4, calculating multiple loss items based on the fusion result to obtain the fusion network target loss function; Step S5: adaptively optimize and adjust the fusion network target loss function to obtain a finally trained multimodal fusion network system and realize the fusion of multimodal images with any quality input.
2. The infrared light and visible light image fusion method according to claim 1, characterized in that: In step S2, CLIP-IQA is used to perform image quality assessment to distinguish multiple image degradation combination types, and the corresponding degradation text description features are read according to the degradation type.
3. The infrared light and visible light image fusion method according to claim 2, characterized in that: In step S3, the degraded text description features and the visible light-infrared light image are input into the multimodal image fusion network, and the features of spatial and deep information of the visible light image and the infrared light image are extracted through the TransformerBlock and cross-attention mechanism.
4. The infrared light and visible light image fusion method according to claim 3, characterized in that: In step S4, the multiple loss items include: frequency domain loss, consistency loss, text loss, color loss and maximum loss.
5. An infrared light and visible light image fusion device, characterized in that: include: A first processing module, used for acquiring a visible light-infrared light image; The second processing module is used to input the visible light-infrared light image into the image quality assessment network to identify the degradation type and obtain the degradation text description feature; The third processing module is used to input the degraded text description features and the visible light-infrared light image into the multimodal image fusion network for feature learning and fusion; The fourth processing module is used to calculate multiple loss items based on the fusion result to obtain the fusion network target loss function; The fifth processing module is used to adaptively optimize and adjust the target loss function of the fusion network to obtain the final trained multimodal fusion network system and realize the fusion of multimodal images with any quality input.
6. The infrared light and visible light image fusion device according to claim 5, characterized in that: The second processing module is used to use CLIP-IQA to perform image quality assessment to achieve image multi-degradation combination type discrimination, and read the corresponding degradation text description features according to the degradation type.
7. The infrared light and visible light image fusion device according to claim 2, characterized in that: The third processing module is used to input the degraded text description features and visible light-infrared light images into the multimodal image fusion network, and extract the features of spatial and deep information of visible light images and infrared light images through TransformerBlock and cross-attention mechanism.
8. The infrared light and visible light image fusion device according to claim 7, characterized in that: The multiple loss items include: frequency domain loss, consistency loss, text loss, color loss and maximum loss.
9. An infrared light and visible light image fusion system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the infrared light and visible light image fusion method as described in any one of claims 1 to 4 is executed.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when running, executes the infrared light and visible light image fusion method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Deep learning multitask magnetic resonance heart segmentation and quantification method
CN117115112A
Image restoration method and device based on multi-modal large language model, and medium
CN118691511A