Image fusion method and device based on dynamic relativity enhancement

By introducing dynamic relativity enhancement and cross-modal enhancement modules in image fusion, the combination of image fusion and image recovery tasks is solved, image quality and robustness are improved, and the repair ability of multiple degradation types is achieved.

CN120088146AActive Publication Date: 2025-06-03TIANJIN UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510093976.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-03
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art lacks effective mutual combination in image fusion and image recovery tasks, resulting in a degradation of image quality under the influence of noise and blur.

Method used

Using an image fusion method based on dynamic relativity enhancement, the image fusion and image repair tasks are integrated with relative dominance by introducing natural multimodal complementarity and prompt-based cross-modal enhancement modules, and image quality is enhanced by relative dominance.

Benefits of technology

The robustness and performance of image fusion are improved, especially the image quality in the respective modal disadvantage areas, and the perception and repair of unknown degradation types are realized, which enhances the practical application capabilities of image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088146A_ABST
    Figure CN120088146A_ABST
Patent Text Reader

Abstract

The invention discloses an image fusion method and device based on dynamic relativity enhancement, and the method comprises the steps: inputting the features of a multi-source image into a fusion network, and generating a fusion image; calculating fusion loss, executing back propagation, training the fusion basic network, and obtaining fusion basic network parameters; freezing fusion basic network parameters, and inserting a prompt-based cross-modal enhancement module into the fusion basic network to obtain an image fusion network based on dynamic relativity enhancement; inputting the degraded visible light image into an image fusion network based on dynamic relativity enhancement to obtain features of the visible light image and the infrared image, and inputting the features into a cross-modal enhancement module based on prompt; after obtaining a fused image, respectively carrying out loss calculation on the fused image and a clear image which is not degraded, and carrying out back propagation; and deploying the trained image fusion network based on dynamic relativity enhancement, and completing degradation repair and image fusion tasks by the image fusion network based on dynamic relativity enhancement. The device comprises a processor and a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image fusion, and in particular, to an image fusion method and device based on dynamic relativity enhancement. Background Art

[0002] Image fusion aims to integrate multi-source image information acquired by different sensors to generate a comprehensive single image. Multimodal (visible light and infrared) image fusion has been widely applied in fields such as autonomous driving, unmanned aerial vehicles, and forest fire monitoring. Images captured by different imaging sensors often have different characteristics. Infrared images can adapt to different lighting conditions but usually lack detailed textures; while visible light images can capture rich colors and details in sufficient light, but often experience quality degradation in complex environments. Image fusion takes advantage of the complementary information between modalities, retaining the advantages of both visible light and infrared images and minimizing their limitations as much as possible. However, in real-world environments, due to sensor failures or external interference, visible light and infrared images are often affected by noise and blur (such as Gaussian noise and JPEG noise), resulting in a decrease in image quality.

[0003] To enhance low-quality images, some existing methods have achieved remarkable results by learning the distribution of the dataset. Most existing methods regard image fusion and image restoration as two independent tasks. Recently, some studies have attempted to handle these two tasks simultaneously in one framework because they have similar requirements in extracting effective information. From the perspective of multimodal complementarity, image fusion can extract the dominant information of different modalities, thereby further improving the effect of image restoration. Therefore, in real-world scenarios, it is necessary to fuse these two tasks to capture their internal consistency. Summary of the Invention

[0004] The present invention provides an image fusion method and device based on dynamic relativity enhancement. The present invention provides a dynamic relative enhancement image fusion framework, aiming to jointly improve the performance of image fusion and enhancement. The present invention introduces the natural multimodal complementarity of image fusion to enhance the quality of cross-modal images, especially in the disadvantaged regions of their respective modalities. The present invention designs a simple and easy-to-integrate module called the hint-based cross-modal enhancement module. This module promotes cross-modal enhancement in relatively weak quality regions by capturing the relative advantages of each modality, thereby generating enhanced multimodal features. Based on the mutual guidance between modalities, the present invention can simultaneously handle the degradation problem and the image fusion problem, as described in detail below:

[0005] In a first aspect, a method for image fusion based on dynamic relativity enhancement, the method includes the following steps:

[0006] Input the visible light image and the infrared image into a two-stream backbone network; input the features of the multi-source images into a fusion network to generate a fused image; calculate the fusion loss and perform backpropagation to train the fusion basic network and obtain the fusion basic network parameters;

[0007] Freeze the fusion basic network parameters, insert a prompt-based cross-modal enhancement module into the fusion basic network to obtain an image fusion network based on dynamic relativity enhancement; input the degraded visible light image into the image fusion network based on dynamic relativity enhancement to obtain the features of the visible light image and the infrared image, and input the features into the prompt-based cross-modal enhancement module;

[0008] After obtaining the fused image, calculate the loss with the undegraded clear image respectively and perform backpropagation; deploy the trained image fusion network based on dynamic relativity enhancement, and the image fusion network based on dynamic relativity enhancement completes the degradation repair and image fusion tasks.

[0009] Among them, the prompt-based cross-modal enhancement module includes two components: a relative dominance component and a cross-modal enhancement component.

[0010] Among them, the relative dominance component is calculated by integrating the features of multiple modalities and inputting them into a gating network;

[0011] Given a visible light feature F v and an infrared image feature F I , its relative dominance RD i is calculated by a relative dominance gate G:

[0012] RD i =G([F i ; F v )

[0013] Among them, G(x)=σ(Conv(x)), σ is an activation function. Given the relative dominance RD m , obtain the modal prompt information p m containing the relative dominance prompt = F m ·RD m .

[0014] Among them, the cross-modal enhancement component is: the prompt information p m is passed to two convolutional layers and a Sigmoid layer, and the output will be used as an attention map The attention map represents the dominant information of another modality, guides the reconstruction process and is used as prompt information to enhance the performance of the inferior modality.

[0015] Among them, the method further includes: perception and repair of unknown degradation types;

[0016] Generate a self-enhancing feature: SEF based on the current feature i = Cov(Re(∑ c z·σ(∑ H ∑ W (F i ))), where z is a learnable parameter, σ is an activation function, and Re represents a bilinear upsampling operation; use as a cross-modal prompt to assist in feature enhancement and repair, and obtain relatively enhanced features where SA is a self-attention module and FFN is a linear layer.

[0017] In a second aspect, an apparatus for image fusion based on dynamic relativity enhancement, the apparatus includes: a processor and a memory, and program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the apparatus to execute the method described in any one of the first aspect.

[0018] In a third aspect, a computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in any one of the first aspect.

[0019] The beneficial effects of the technical solution provided by the present invention are as follows:

[0020] 1. The present invention utilizes the principle of complementarity between modalities to design an image fusion method based on dynamic relativity enhancement. This method integrates image fusion and image restoration into one network, uses the relative dominance generated by image fusion to assist image restoration, and enhances the robustness of image fusion through image restoration;

[0021] 2. The present invention designs a prompt-based cross-modal enhancement module for modality feature enhancement and repair, which can be plug-and-play; this module captures the relatively dominant regions of each modality to specifically enhance the defective regions in other modalities; the module can be easily and flexibly integrated into a general fusion model, thus effectively improving the fusion performance of degraded images;

[0022] 3. The present invention enables the model to perceive and repair unknown degradation types through learnable parameters, so that one model can perform repairs for multiple degradation types, thereby enhancing the practical application ability of image fusion;

[0023] 4. The present invention designs an image-scalable inference method. For the mechanism of Transformer, during the inference process, any image size can be input into the network for inference, increasing the usability and applicable scenarios of this method in practical applications. Description of the Drawings

[0024] Figure 1 It is a flowchart of the image fusion method based on dynamic relativity enhancement proposed by the present invention;

[0025] Figure 2 It is a schematic diagram of the cross-modal enhancement module of the present invention;

[0026] Figure 3 It is a schematic diagram of the image fusion network structure based on dynamic relativity enhancement proposed by the present invention. Detailed Description of the Invention

[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes in detail the embodiments of the present invention.

[0028] To solve the technical problems existing in the background art, an embodiment of the present invention proposes an image fusion network based on dynamic relativity enhancement. This network enhances and repairs the weak areas of the image through the dynamic perception of the relative dominance between different modalities during the image fusion process, thereby further improving the performance of the model. This network can perform image fusion and image repair through one model. It uses image fusion to guide image repair and improves the robustness of image fusion through image repair. The network structure effectively solves the technical problems existing in the background art, can better combine these two tasks, and thus improves the performance and robustness of image fusion.

[0029] As Figure 1 shown, an image fusion network based on dynamic relativity enhancement consists of two parts, and the functions of the content of each part are as follows:

[0030] In the basic network for image fusion, the embodiment of the present invention includes a two-stream backbone network and a fusion network. The two-stream backbone network consists of multiple Transformers, which include: an encoder and a decoder. The encoder extracts features from multi-source images, and the decoder restores the features.

[0031] Furthermore, the feature extraction of multi-source images can be expressed as Given a visible light image and an infrared image The features for each modality can be expressed by the following formula:

[0032]

[0033] where represents the corresponding feature of the given modality, represents the decoder, Let \(E\) denote the encoder, \(i\) denote the infrared modality, \(v\) denote the visible light modality, \(m\) denote one of the set of infrared and visible light modalities, and \(F\) l denote the features obtained by the encoder, \(F\) h denote the features obtained by the decoder, and \(H\) and \(W\) denote the height and width of the image.

[0034] The goal of the embodiments of the present invention is to fuse the complementary information of two original images to generate a fused image \(I\) f , therefore, the fusion network can be expressed as where, \([;]\) represents the concatenation operation, denote the fusion network, \(F\) i denote the features of the infrared modality, \(F\) v denote the features of the visible light modality.

[0035] Furthermore, the hint-based cross-modal enhancement includes two components: the relative dominance component and the cross-modal enhancement component. Among them, the relative dominance of each modality can be calculated by integrating the features of multiple modalities and inputting them into the gating network. In practical applications, taking infrared as an example, given a visible light feature \(F\) v and infrared image feature \(F\) I , its relative dominance \(RD\) i can be calculated by the relative dominance gate \(G\): \(RD\) i = G([F\) i ; F\) v , where \(G(x)=\sigma(Conv(x))\), and \(\sigma\) is the activation function. Given the relative dominance \(RD\) m , further, the modality hint information \(p\) containing the hint with relative dominance can be obtained m = F\) m ·RD\) m . Where \(\cdot\) is the dot product operation.

[0036] Furthermore, the cross-modal image enhancement is as Figure 2 shown. The hint information \(p\) m is passed to two convolutional layers and a Sigmoid layer. Its output will be used as the attention map This attention map represents the dominant information of another modality and points out the insufficient part that needs to be enhanced. It can guide the reconstruction process and be used as hint information to enhance the performance of the inferior modality.

[0037] Furthermore, as Figure 2 shown, the specific process of perceiving and repairing unknown degradation types is: according to the current features, generate a self-enhanced feature: \(SEF\) i = Cov(Re(\(\sum\) c z·\(\sigma(\sum\) H \(\sum\)W (F i ))))), where z is a learnable parameter, σ is an activation function, and Re represents a bilinear upsampling operation. Taking as a cross-modal prompt to assist in feature enhancement and repair, improving inferior features and obtaining relatively enhanced features where SA is a self-attention module and FFN is a linear layer.

[0038] Example 1

[0039] The embodiment of the present invention provides a method for image fusion based on dynamic relativity enhancement, and the method includes the following steps:

[0040] 101: The image fusion basic network is composed of a two-stream backbone network and a fusion network;

[0041] Among them, the two-stream backbone network is composed of multiple Transformers, including: an encoder and a decoder. The structures of each stream are the same, but the parameters are not shared. The fusion network is composed of multiple Transformer layers and a convolutional layer.

[0042] Specifically, in the two-stream backbone network, the encoder is responsible for feature extraction of multi-source images, and the decoder is responsible for representing and restoring the features. The fusion network is used to fuse the features extracted by the two-stream backbone network and finally generate a fused image.

[0043] 102: Input the clear visible light image and infrared image into the two-stream backbone network;

[0044] Among them, the two-stream backbone network encodes the multi-source images through the encoder to obtain encoded features, and then inputs the encoded features into the decoder to obtain decoded features.

[0045] 103: Input the features of the multi-source images into the fusion network to generate a fused image;

[0046] Among them, the features of the multi-source images are input into the fusion network after being concatenated. After passing through multiple Transformer Blocks, the features are finally reconstructed into a 3-channel fused image through the convolutional layer.

[0047] 104: After obtaining the fused image, calculate the fusion loss, perform backpropagation, and then train the fusion basic network to obtain the fusion basic network parameters;

[0048] Among them, the fusion loss function used when calculating the fusion loss includes: maximum pixel loss (Max-Pixel), maximum gradient loss (Max-Grad), structural consistency loss (SSIM), and color consistency loss (Color).

[0049] 105: Freeze the basic network parameters of the fusion, and insert a prompt-based cross-modal enhancement module into the basic fusion network to obtain an image fusion network based on dynamic relativity enhancement.

[0050] Specifically, a prompt-based cross-modal enhancement module is added to each layer of the decoder.

[0051] 106: Degrade clear visible light and infrared images in a random type to obtain degraded visible light images and degraded infrared images;

[0052] Among them, the random degradation types include: one of Gaussian noise, Poisson noise, speckle noise, JEPG compression, downsampling compression, and dynamic blur.

[0053] 107: Input the degraded visible light image into the image fusion network based on dynamic relativity enhancement to obtain the features of the visible light image and the infrared image, and input the features into the prompt-based cross-modal enhancement module for feature enhancement;

[0054] Among them, the number of network layers of the decoder is the same as the number of modules of the prompt-based cross-modal enhancement. Therefore, the network will include several steps of 105-107. Finally, the enhanced features are input into the basic fusion network.

[0055] 108: After obtaining the fused image, calculate the loss with the clear image that has not been degraded respectively, and perform backpropagation;

[0056] Specifically, the loss calculation is the same as in step 104.

[0057] 109: Deploy the trained image fusion network based on dynamic relativity enhancement. The image fusion network based on dynamic relativity enhancement can complete the degradation repair and image fusion tasks. By using learnable parameters to perceive the degradation type, the image fusion network based on dynamic relativity enhancement can repair multiple degradation types without knowing the specific degradation type. Repair and fusion can be efficiently integrated into one model, saving memory and computational overhead, and promoting the development of the image fusion field and the image restoration field.

[0058] In summary, the embodiments of the present invention gradually form a network with repair capabilities and generate fused images through two parts: the basic image fusion network and the image prompt module based on prompts; utilize the dominance of image fusion to generate a dominant prompt area to guide image enhancement; utilize image enhancement to improve the result and robustness of image fusion; utilize prompt learning to enable the model to perceive multiple types of degradation without prior knowledge of the degradation type. By connecting the two tasks of image fusion and image enhancement, one method can complete two tasks, improving the performance of the tasks and saving computational time and computational overhead.

[0059] Example 2

[0060] The solution in Example 1 will be further introduced below in combination with specific examples and calculation formulas. See the following description for details:

[0061] I. Data Preparation

[0062] The experiments of the embodiments of the present invention were carried out on two publicly available datasets, LLVIP and MSRS.

[0063] LLVIP is a visible light-infrared paired dataset for low-light vision tasks. It contains 15,488 paired data, most of which are taken in very dark scenes. All images are strictly aligned in time and space. Pedestrians in the dataset are labeled. The dataset has a high level of fineness and standardization and can be used to estimate and improve the performance of existing image fusion algorithms, low-light illumination human detection, image translation and other methods.

[0064] MSRS is a multi-spectral dataset for infrared and visible light image fusion, containing 1,444 pairs of high-quality, aligned infrared and visible light images. This dataset is derived from the MFNet dataset. After removing unaligned image pairs, it contains 715 pairs of daytime images and 729 pairs of nighttime images.

[0065] II. Image Fusion Network Structure Based on Dynamic Relativity Enhancement

[0066] The image fusion network process based on dynamic relativity enhancement is as Figure 1 shown and consists of two parts. The specific network details are as Figure 3 shown. In the basic network for image fusion, the embodiments of the present invention include: a two-stream backbone network and a fusion network. The two-stream backbone network consists of multiple Transformers, which include: an encoder and a decoder. The encoder extracts features from multi-source images, and the decoder characterizes and restores the features.

[0067] Furthermore, the feature extraction of multi-source images can be expressed as Given a visible light image and an infrared image The features for each modality can be expressed by the following formula:

[0068]

[0069] where represents the corresponding features of a given modality. The goal of the embodiments of the present invention is to fuse the complementary information of the two original images to generate a fused image I f , therefore, the fusion network can be expressed as Among them, [;] represents the concatenation operation.

[0070] Furthermore, the hint-based cross-modal enhancement includes two components: relative dominance and cross-modal enhancement. Among them, the extraction of the relative dominance of each modality is achieved by integrating the features of multiple modalities and inputting them into a gating network. Specifically, taking infrared as an example, given a visible light feature F v and an infrared image feature F I , its related dominance RD i can be calculated by the relative dominance gate G: RD i = G([F i ; F v ), where G(x) = σ(Conv(x)), and σ is the activation function. Given the related dominance RD m , further, the modality hint information p m containing the hint with related dominance can be obtained: p m = F m ·RD

[0071] Furthermore, the cross-modal image enhancement is as shown in Figure 2 . The hint information p m is passed to two convolutional layers and one Sigmoid layer. Its output will be used as the attention map This attention map represents the dominant information of another modality and points out the insufficient part that needs to be enhanced. It can guide the reconstruction process and be used as hint information to enhance the performance of the inferior modality.

[0072] Furthermore, as shown in Figure 2 , the specific process of perceiving and repairing unknown degradation types is as follows: According to the current features, a self-enhanced feature is generated: SEF i = Cov(Re(∑ c z·σ(∑ H ∑ W (F i ))), where z is a learnable parameter, σ is the activation function, and Re represents the bilinear upsampling operation. Taking as the cross-modal hint to assist in the enhancement and repair of features, the inferior features are improved, and relatively enhanced features can be obtained where SA is the self-attention module and FFN is the linear layer.

[0073] III. Evaluation Metrics and Protocols

[0074] The embodiments of the present invention use 5 metrics to verify the effectiveness of the method. Entropy (EN), gradient-based metric Qabf, Structural Similarity (SSIM), Mutual Information (MI), and Visual Information Fidelity Fusion (VIFF).

[0075] IV. Details of Model Usage

[0076] 1. Data Augmentation: Due to limited computing resources, the embodiments of the present invention adopt the methods of random flipping and shearing to improve data diversity. Specifically: First, randomly crop the visible light and infrared images into a size of 96×96, and then randomly flip the visible light and infrared images horizontally and vertically with a probability of 50%.

[0077] 2. Model Optimization: The batch size in the training process of the embodiments of the present invention is 12, and the AdamW optimization algorithm with β 1 =0.9 and β 2 =0.999 is used. The learning rate decreases from the initial 1×10 -6 to 1×10 -8 in the form of a cosine curve.

[0078] 3. Hyperparameter Setting: The number of hint-based cross-modal enhancement modules in the embodiments of the present invention is 4.

[0079] 4. Loss Setting: The loss in the embodiments of the present invention is an unsupervised loss. It is composed of the maximum pixel loss (Max-Pixel), the maximum gradient loss (Max-Grad), the structural consistency loss (SSIM), and the color consistency loss (Color). The four are weighted and summed in a ratio of 4:1:1:10 to obtain the final loss.

[0080] The embodiments of the present invention provide a dynamic relative enhancement image fusion framework, aiming to jointly improve the performance of image fusion and enhancement. The embodiments of the present invention introduce the natural multimodal complementarity of image fusion to enhance the quality of cross-modal images, especially in the disadvantaged regions of each modality. The embodiments of the present invention design a simple and easy-to-integrate module called cross-modal promotion enhancement. This module captures the relative advantages of each modality, and then promotes the cross-modal enhancement of relatively weak quality regions, thereby generating enhanced multimodal features. Based on the mutual guidance between modalities, the embodiments of the present invention can handle both the degradation problem and the image fusion problem simultaneously.

[0081] The embodiments of the present invention all have the following three key creative points:

[0082] I. Propose an image fusion method based on dynamic relative enhancement

[0083] Technical effect: This method integrates image fusion and image inpainting into one network, uses the relative dominance generated by image fusion to assist image inpainting, and enhances the robustness of image fusion through image inpainting.

[0084] II. A prompt-based cross-modal enhancement module is proposed.

[0085] Technical effect: This module is used for prompting cross-modal enhancement and has the characteristics of plug-and-play. This module captures the relatively dominant regions of each modality to specifically enhance the defective regions in other modalities. Its design enables the module to be flexibly and easily integrated into a general image fusion model, thus effectively improving the fusion performance of degraded images.

[0086] III. A blind inpainting method is proposed.

[0087] Technical effect: This method can sense and repair unknown degradation types through learnable parameters, so that a single model can perform blind inpainting for multiple degradation types, thereby enhancing the practical application ability of image fusion.

[0088] IV. An image-scalable inference method is proposed.

[0089] Technical effect: Aiming at the mechanism of Transformer, small images are used for training during the training process. During the inference process, images of any size can be input into the network for inference, increasing the usability and practical application significance of actual available scenarios.

[0090] Embodiment 3

[0091] The method proposed in the embodiment of the present invention is compared with a variety of existing technical methods on a multi-modal dataset to verify the performance of image fusion.

[0092] I. Comparison on a multi-modal dataset without degradation.

[0093] As shown in Table 1, the embodiments of the present invention are evaluated on the LLVIP and MSRS datasets through five evaluation metrics. On the LLVIP dataset, the method of the embodiments of the present invention outperforms other comparison methods in all five metrics, showing significant advantages, especially in the Qabf, SSIM, MI, and VIFF metrics. Specifically, the highest EN and MI scores indicate that the method can retain the most information, thanks to its ability to jointly enhance non-dominant regions. Higher Qabf scores reflect better alignment of local gradients and intensities between the source image and the fused image, indicating that the method can better retain valuable information from multiple modalities. In addition, higher VIFF scores indicate that the fused image retains a large amount of visual information in the source image, indicating superior fusion quality. The SSIM score measures structural similarity, indicating that the method can retain more structural details. These qualitative results demonstrate that the method achieves excellent fusion performance in the dynamic relative enhancement of image quality.

[0094] Table 1

[0095]

[0096] II. Comparison of multi-modal datasets with degradation.

[0097] For degraded data, the experimental results are shown in Table 2. The results show that even without a dedicated restoration model, the method can still achieve competitive performance. Specifically, EN, SSIM, and VIFF are close to the best values, while Qabf and MI reach the best scores. Higher MI and VIFF scores indicate that the method can effectively retain the mutual information between source images, ensuring the retention of necessary details and context information in the final output while maintaining subjective visual quality. This performance highlights the method's ability to effectively maintain the authenticity of images and enhance degraded features without relying on a dedicated restoration model.

[0098] Table 2

[0099]

[0100] Embodiment 4

[0101] An apparatus for image fusion based on dynamic relative enhancement, the apparatus includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the apparatus to execute the following method steps in Embodiment 1:

[0102] Input the visible light image and the infrared image into the dual-stream backbone network; input the features of the multi-source image into the fusion network to generate a fused image; calculate the fusion loss and perform backpropagation to train the fusion basic network to obtain the fusion basic network parameters;

[0103] Freeze the basic network parameters of the fusion, insert the prompt-based cross-modal enhancement module into the basic fusion network to obtain an image fusion network based on dynamic relativity enhancement; input the degraded visible light image into the image fusion network based on dynamic relativity enhancement to obtain the features of the visible light image and the infrared image, and input the features into the prompt-based cross-modal enhancement module;

[0104] After obtaining the fused image, calculate the loss with the clear image without degradation respectively and perform backpropagation; deploy the trained image fusion network based on dynamic relativity enhancement, and the image fusion network based on dynamic relativity enhancement completes the degradation repair and image fusion tasks.

[0105] Among them, the prompt-based cross-modal enhancement module includes two components: the relative dominance component and the cross-modal enhancement component.

[0106] Among them, the relative dominance component is calculated by integrating the features of multiple modalities and inputting them into the gated network;

[0107] Given a visible light feature F v and an infrared image feature F I , its relative dominance RD i is calculated by the relative dominance gate G:

[0108] RD i = G([F i ; F v )

[0109] Among them, G(x) = σ(Conv(x)), σ is the activation function, given the relative dominance RD m , obtain the modal prompt information p containing the relative dominance prompt m = F m ·RD m .

[0110] Among them, the cross-modal enhancement component is: the prompt information p m is passed to two convolutional layers and a Sigmoid layer, and the output will be used as the attention map The attention map represents the dominant information of another modality, guides the reconstruction process and is used as prompt information to enhance the performance of the inferior modality.

[0111] Among them, the device also includes: perception and repair of unknown degradation types;

[0112] According to the current features, generate a self-enhanced feature: SEF i = Cov(Re(∑ c z·σ(∑ H ∑W (F i ))))), where z is a learnable parameter, σ is an activation function, and Re represents a bilinear upsampling operation; taking as a cross-modal prompt to assist in feature enhancement and repair, obtaining relatively enhanced features where SA is a self-attention module and FFN is a linear layer.

[0113] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated here.

[0114] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. Specifically, in implementation, the embodiments of the present invention do not limit the execution subject and select according to the needs in actual applications.

[0115] Data signals are transmitted between the memory and the processor through a bus, and the embodiments of the present invention will not elaborate on this.

[0116] Based on the same inventive concept, the embodiments of the present invention also provide a computer-readable storage medium. The storage medium includes a stored program that controls the device where the storage medium is located to execute the method steps in the above embodiments when the program runs.

[0117] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.

[0118] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the method description in the embodiments, and the embodiments of the present invention will not be elaborated here.

[0119] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.

[0120] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium or a semiconductor medium, etc.

[0121] In the embodiments of the present invention, unless otherwise specified for the models of each device, the models of other devices are not limited, and any device that can perform the above functions may be used.

[0122] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0123] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for image fusion based on dynamic relativity enhancement, characterized in that: The method comprises the following steps: Input the visible light image and infrared image into the dual-stream backbone network; input the features of the multi-source image into the fusion network to generate a fused image; calculate the fusion loss and perform back propagation to train the fusion basic network and obtain the fusion basic network parameters; Freeze the parameters of the fusion base network, insert the hint-based cross-modal enhancement module into the fusion base network, and obtain the image fusion network based on dynamic relativity enhancement; input the degraded visible light image into the image fusion network based on dynamic relativity enhancement, obtain the features of the visible light image and the infrared image, and input the features into the hint-based cross-modal enhancement module; After obtaining the fused image, the loss is calculated with the clear image that has not been degraded, and back propagation is performed; the trained image fusion network based on dynamic relativity enhancement is deployed, and the image fusion network based on dynamic relativity enhancement completes the degradation restoration and image fusion tasks.

2. The method for image fusion based on dynamic relativity enhancement according to claim 1, characterized in that: The cue-based cross-modal enhancement module includes two components: a relative dominance component and a cross-modal enhancement component.

3. The method for image fusion based on dynamic relativity enhancement according to claim 2, characterized in that: The relative dominance component is calculated by integrating the features of multiple modalities and inputting them into a gating network; Given a visible light feature F v and infrared image feature F I , its related dominant RD i Calculated by the relative dominant gating G: RD i =G([F i ;F v ] Where G(x) = σ(Conv(x)), σ is the activation function, given the correlation dominance RD m , obtain the modal prompt information p containing relevant dominant prompts m =F m ·RD m .

4. The method for image fusion based on dynamic relativity enhancement according to claim 2, characterized in that: The cross-modal enhancement component is: prompt information p m Passed to two convolutional layers and a Sigmoid layer, the output will be used as the attention map The attention map represents the dominant information of another modality. Guide the reconstruction process and serve as hints to enhance the performance of the inferior modality.

5. The method for image fusion based on dynamic relativity enhancement according to claim 1, characterized in that: The method also includes: sensing and repairing unknown degradation types; Based on the current features, generate a self-improvement feature: SEF i =Cov(Re(∑ c z·σ(∑ H ∑ W (F i )))), where z is a learnable parameter, σ is an activation function, and Re represents a bilinear upsampling operation; As a cross-modal hint to assist in feature enhancement and repair, and obtain relatively enhanced features Among them, SA is the self-attention module and FFN is the linear layer.

6. A device for image fusion based on dynamic relativity enhancement, characterized in that: The device comprises: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method

    CN111709903A

  • Underwater image enhancement method based on context decomposition feature fusion

    CN114913083A

  • Dynamic illumination face image quality enhancement method based on multi-scale attention mechanism

    CN115880225A

  • Infrared and visible light image fusion method based on information interaction and edge guidance

    CN118333881A

  • Visible light and infrared image fusion method and device, equipment and storage medium

    CN118967462A