Visible light-infrared image restoration fusion method based on hybrid expert model
Through a two-stage method based on a hybrid expert model, complex degradation problems such as stripe noise in infrared images and haze in visible light images are solved, high-quality image restoration and fusion are achieved, and the visual quality and information transmission capability of the image are improved.
Patent Information
- Application Number
- CN202510767575.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-10
AI Technical Summary
Existing multimodal image fusion methods fail to effectively solve complex degradation problems such as streak noise in infrared images and haze in visible light images, resulting in image quality degradation and information loss. In addition, existing methods ignore the coupling problem between image restoration and fusion tasks.
A two-stage method based on a hybrid expert model is adopted. In the first stage, image restoration is performed through degradation-aware gating and multi-expert collaborative modules. In the second stage, feature fusion is performed, and high-quality fusion results are generated using fusion-aware gating and multi-expert collaborative modules.
It achieves high-quality image restoration and fusion, reduces mutual interference between tasks, improves the visual quality and information transmission capability of the image, and demonstrates robustness and generalization capabilities in different degradation scenarios.
Smart Images

Figure CN120765475A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a visible-infrared image restoration fusion method based on a hybrid expert model and belongs to the technical field of image processing. BACKGROUND
[0002] Multi-modal image fusion aims to integrate complementary information from different imaging modalities and plays a key role in rescue operations, monitoring systems, and remote sensing tasks. By combining the advantages of visible light and infrared imaging modalities, in infrared imaging, infrared focal plane arrays are key components for capturing thermal information. However, due to the use of different readout circuits by different column sensors, the bias voltage difference of these circuits often leads to stripe noise in infrared images, appearing as alternating bright and dark stripes. This degradation severely affects the visual quality and interpretability of infrared images. At the same time, visible light images often suffer from different degradation types such as haze, noise, and blur, which are very common in real-world scenarios and significantly affect the information conveyed by visible light images and the imaging quality. These complex and interwoven degradation patterns pose significant challenges to existing multi-modal image fusion methods and affect the implementation of various practical tasks.
[0003] Infrared and visible light image data in low-altitude environments have high complementarity: infrared images can penetrate complex environments such as smoke and weak light through thermal radiation imaging, enabling all-weather target detection, but have limitations such as low resolution and lack of detailed features; visible light images can provide rich texture, color, and geometric details, but are significantly affected by lighting conditions and weather. This difference in characteristics makes it difficult for a single sensor to meet the comprehensive perception needs of low-altitude scenes. Through data fusion, the synergistic advantages of both can be fully utilized: on the one hand, infrared data can compensate for the perceptual blind area of visible light in harsh environments, ensuring all-weather monitoring capability; on the other hand, visible light data can supplement fine visual features to infrared images, significantly improving target recognition accuracy. Especially in applications such as unmanned aerial vehicle inspection and disaster rescue, the fused data can not only achieve rapid positioning of large-scale heat sources but also perform high-precision target analysis, greatly improving the environmental adaptability and task completion of the system.
[0004] In the early days of digital image processing, people mainly relied on spatial and frequency domain-based methods to deal with different image degradations. Linear or nonlinear filtering is often used to smooth images to deal with different levels of noise; inverse filtering or Wiener filtering is used in the frequency domain to try to restore spectral information to solve motion blur and defocus blur problems; for haze problems, the dark channel prior is often used to estimate the atmosphere light and transmission rate, and then to restore the haze-free image, and histogram equalization is used to enhance the visual effect. These methods, although computationally efficient, usually assume that the degradation model is known and linear, and have limited adaptability to complex degradation, laying the foundation for subsequent deep learning-based algorithms.
[0005] Fusing visible light images and infrared images is a popular research direction, compared with single modal images, multi-modal images contain more information. In recent years, different deep learning-based methods have promoted image fusion to adapt to different scene tasks, such as DenseFuse, SwinFuse, CDDFuse, EMMA, etc., but these image fusion algorithms only focus on information extraction between different modalities, and do not focus on the degradation phenomenon of single modal images and the influence of different degradation on the fused images.
[0006] Current degradation image fusion algorithms are mostly two-stage methods, the first stage uses an image restoration model to realize image restoration, and the second stage uses an image fusion model to fuse the restored images. However, this kind of method ignores the coupling problem produced by the two cascaded networks, and is easy to cause serious loss of details. Some advanced methods recently try to consider the image degradation problem in the fusion process, such as AWFusion, DRMF and Text-if models, but these methods focus on limited degradation types and have insufficient generalization ability. How to combine the image restoration task with the image fusion task and propose an integrated restoration fusion model remains to be studied. SUMMARY
[0007] The technical problem to be solved by the present application is to provide a visible light-infrared image restoration fusion method based on a hybrid expert model, which divides the complex restoration fusion task into two stages, the first stage uses a degradation perception gate and a multi-expert collaboration module to guide a Transformer network to restore visible light and infrared modal images, and extracts restoration features; in the second stage, the restoration features of the two modalities are pre-fused, and a fusion perception gate and a multi-expert collaboration module are used to obtain the fusion result.
[0008] The technical problem to be solved by the present application is to provide a visible light-infrared image restoration fusion method based on a hybrid expert model, which divides the complex restoration fusion task into two stages, the first stage uses a degradation perception gate and a multi-expert collaboration module to guide a Transformer network to restore visible light and infrared modal images, and extracts restoration features; in the second stage, the restoration features of the two modalities are pre-fused, and a fusion perception gate and a multi-expert collaboration module are used to obtain the fusion result.
[0009] A visible light-infrared image restoration fusion method based on a hybrid expert model, comprising the following steps:
[0010] Step 1, obtain a visible light-infrared fusion dataset, and pre-process the fusion dataset, divide the pre-processed visible light-infrared fusion dataset into a one-stage training set for restoration task, a two-stage training set for restoration and fusion task and a test set; the visible light-infrared fusion dataset includes a visible light-infrared degraded image pair, a non-degraded image pair corresponding to the degraded image pair, and a fusion image corresponding to the degraded image pair;
[0011] Step 2, a mixed expert mechanism-based image restoration fusion model is constructed, including a one-stage image restoration network and a two-stage image fusion network; wherein the one-stage image restoration network includes first and second two branches, and the two branches are respectively used for restoring visible light and infrared degraded images, the input of the first branch is the preprocessed visible light degraded image, and the output is a visible light restored image and a restored feature, the input of the second branch is the preprocessed infrared degraded image, and the output is an infrared restored image and a restored feature; the two-stage image fusion network is used for fusing the visible light and infrared restored features, and outputs a fused image;
[0012] Step 3, the one-stage image restoration network is trained by using the one-stage training set, and a trained one-stage image restoration network is obtained; based on the trained one-stage image restoration network, the entire image restoration fusion model is trained by using the two-stage training set, and a trained image restoration fusion model is obtained;
[0013] Step 4, the visible light-infrared degraded image pair in the test set is input into the trained image restoration fusion model to obtain a fused image.
[0014] Compared with the prior art, the above technical scheme has the following technical effects:
[0015] 1. The image restoration and image fusion tasks are completed by the two-stage training strategy, which effectively distinguishes the optimization of the two tasks and greatly reduces the mutual interference between the two tasks, and realizes the goal of integrated processing.
[0016] 2. The task-aware gating module is designed, which performs degradation-aware gating in the first stage and fusion-aware gating in the second stage, dynamically adapts to different degradations, actively selects different scene information, and completes the aggregation of multi-modal features to realize high-quality fusion.
[0017] 3. Extensive experiments on multiple benchmarks show that the framework has superior performance and robustness in different degradation scenarios, verifying its effectiveness in real-world applications. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a task-aware gating and multi-expert collaborative module structure diagram based on a mixed expert mechanism;
[0019] Figure 2 is a whole structure diagram of a visible light-infrared image restoration fusion method based on a mixed expert model of the application;
[0020] Figure 3 is a flowchart of the visible light-infrared image restoration fusion method based on the mixed expert model of the application. DETAILED DESCRIPTION
[0021] Embodiments of the present application are described below in detail, examples of which are shown in the accompanying drawings. The embodiments described below by reference to the drawings are exemplary only, and are for the purpose of explanation of the present application, and are not to be construed as limiting the present application.
[0022] As shown in Figure 3 , the present application proposes a visible-infrared image restoration fusion method based on a hybrid expert model, which uses a two-stage training strategy to complete the image restoration and image fusion tasks, as follows:
[0023] Step 1, obtain a visible-infrared fusion dataset containing multiple scenes, and pre-process it, then divide the pre-processed dataset into a training set and a test set.
[0024] Step 2, construct an image restoration fusion model based on a hybrid expert mechanism, and train the model in stages to obtain a trained image restoration fusion model.
[0025] As shown in Figure 2 , the image restoration fusion model based on the hybrid expert mechanism includes a one-stage image restoration network and a two-stage image fusion network. Different modal images are first sent into the one-stage image restoration network, and expert selection is realized by using a degradation perception gating module and a multi-task collaboration module (as shown in Figure 1 ), the restoration feature is extracted through a Transformer module, and the restored image is generated through a decoder, and the restoration network parameters are optimized. Then the restoration network parameters are frozen, the restoration feature is pre-fused through trainable weights, the feature fusion is realized through a fusion perception gating and a multi-task collaboration module, and the fused image is generated.
[0026] The image restoration fusion model based on the hybrid expert mechanism in the present application uses a Transformer system with a U-Net structure as the backbone to extract the restoration feature. Different modal images are sent into the one-stage image restoration network, and expert weights E i (x) are generated by using a degradation perception gating module, and expert G i (x) selection is realized through a multi-task collaboration module, expert results y are generated, and expert balance loss L balance is generated. The restoration feature is extracted through a Transformer module, and the restored image is generated through a decoder.
[0027]
[0028] where f i is the expert selection frequency, P i is the average weight of the degradation perception gating module for the expert, and λ is a hyperparameter set artificially.
[0029] For the restoration task, the restoration result I restored with the un-degraded target image I target , the L1 loss degradation expert balance loss and gradient loss Optimize the parameters of the one-stage image restoration network:
[0030]
[0031] Wherein, sobel is a gradient operator that can extract gradient information of an image.
[0032] Fuse the visible light restoration feature F V and the infrared restoration feature F I through trainable weights α and β, and fuse the fused feature F M Use the fusion perception gate and the multi-task cooperative module to realize feature fusion, and generate a fused image through a shallow Transformer module.
[0033] F M =α·F V +β·F I
[0034] For the fusion task, the fusion result I fusion obtained by using the whole image restoration fusion model is used, the un-degraded visible light image I Vtarget and the un-degraded infrared image I Itarget , the L1 loss intensity loss fusion expert balance loss and gradient loss Optimize the model parameters:
[0035]
[0036] Step 3: input the visible light and infrared degraded images in the test set into the trained image restoration fusion model, and output the final fused image.
[0037] The present application selects correlation coefficient (Correlation Coefficient, CC), mean squared error (Mean Squared Error, MSE), peak signal-to-noise ratio (Peak Signal-to-Noise Ratio, PSNR), no-reference quality assessment based on blur and noise factors (No-reference assessment based on blur and noise factors, N abfand Multiscale structural similarity (MS_SSIM) as image fusion evaluation indicators.
[0038] Correlation coefficient is a statistical indicator to measure the strength and direction of linear relationship between two variables, with a value range of [-1, 1].
[0039] When CC = 1, it represents a complete positive correlation, when CC = -1, it represents a complete negative correlation, and when CC = 0, there is no linear correlation.
[0040]
[0041] Wherein, cov represents covariance calculation, and std represents standard deviation calculation.
[0042] Mean square error is a commonly used indicator to measure the difference between predicted values and true values, and the calculation method is the average value of the square of the prediction error. The smaller the value of MSE, the better the fitting effect of the prediction model and the true data; otherwise, the error is larger. Wherein, n is the total number of test samples.
[0043]
[0044] Peak signal-to-noise ratio compares the ratio between the maximum possible power of the original signal (such as the maximum pixel value of the image) and the mean square error of the noise or distortion signal, and the higher the PSNR value, the smaller the distortion and the closer the quality to the original signal.
[0045]
[0046] Wherein, is the maximum value of the pixels of the fused image.
[0047] The no-reference quality evaluation based on blur and noise factors evaluates the image quality by calculating the blur degree and noise level of the image.
[0048] Multiscale structural similarity analyzes the brightness, contrast and structural information of the image at different resolutions, and calculates the similarity score by comprehensive weighting, and the larger the MS_SSIM value, the smaller the distortion, which can simulate the perception of the human visual system to multi-level visual features.
[0049] The effectiveness of the method proposed in the present application is verified below on the DeMMI-RF dataset and the EMS dataset.
[0050] The model is compared with advanced image restoration and fusion algorithms on the infrared visible light data set DeMMI-RF to prove that the model has high performance and generalization ability compared with the existing two-stage restoration and fusion method and the existing single-stage method. In addition, the model is compared with advanced restoration and fusion algorithms on the public data set EMS to prove that the model has superiority in the same field.
[0051] The images are selected from the existing visible light-infrared data set DeMMI-RF, and the corresponding degradation information is added to the images according to the target task. The image training set only contains 6 single degradation types, wherein the single-mode image training set for image restoration contains a total of 8787 images, and the dual-light image training set for image restoration and image fusion contains a total of 26631 images. The test set includes 6 single degradation types and 17 mixed degradation types, a total of 9895 single degradation types and 510 mixed degradation types.
[0052] The EMS data set is proposed by Yi et al. and is an infrared visible light fusion data set including low light, overexposure, rain, fog, blur and random noise in visible images, low contrast, stripe noise and random noise in infrared images, etc. Degradation types provide rich scenes for infrared visible light image fusion containing multiple degradations.
[0053] Model parameters: the mixed expert selection is set to 11:6, and the one-stage training round is 30 rounds. The model uses the Adam optimizer, and the learning rate is 1x10 -4 , the parameters of the one-stage training are frozen in the two-stage, and the other parameters are kept unchanged.
[0054] The method TG-ECNet proposed by the application is compared with 6 traditional double network structures and 3 restoration and fusion models on 6 degradation types of low-density Gaussian noise, medium-density Gaussian noise, high-density Gaussian noise, haze, defocus blur and stripe noise. The 6 traditional double network structures use AdaIR as a pre-restoration model to obtain a restored image, and then use DenseFuse, SwinFuse, CDDFuse, SeAFusion, MGDN and EMMA for image fusion. The 3 restoration and fusion models are AWFusion, DRMF and Text-if.
[0055] The experimental results on the DeMMI-RF data set are shown in Table 1, wherein the method of the application reaches the optimal value in 5 indexes, and the PSNR exceeds the second place by 0.4514, and the MS_SSIM exceeds the second place by 0.0339.
[0056] Table 1 Performance comparison of image restoration and fusion models on DeMMI-RF data set
[0057] CC MSE PSNR <![CDATA[N abf ]]> MS_SSIM DenseFuse 0.5436 0.0636 30.2758 0.0459 0.3698 SwinFuse 0.5455 0.0644 30.2877 0.0511 0.3867 CDDFuse 0.5455 0.0632 30.2913 0.0635 0.3606 SeAFusion 0.5459 0.0649 30.2317 0.0595 0.3614 MGDN 0.5468 0.0487 30.8572 0.0623 0.3742 EMMA 0.5438 0.0610 30.3803 0.0628 0.3576 AWFusion 0.5432 0.0676 30.1597 0.0760 0.3398 Text-if 0.5464 0.0650 30.1645 0.0494 0.3591 DRMF 0.5438 0.0807 29.6901 0.0786 0.3187 TG-ECNet 0.5512 0.0397 31.3086 0.0223 0.4206
[0058] In order to enhance the generalization of the results of the embodiments of the present application, the same degraded data set of the EMS data set is used for comparison test, and the experimental results are shown in Table 2. The method of the present application achieves the optimal in the five indexes, which fully embodies the generalization ability of the model of the present application.
[0059] Table 2 Performance comparison of image restoration fusion model on EMS data set
[0060]
[0061]
[0062] Based on the same inventive concept, the embodiment of the present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the visible light-infrared image restoration fusion method based on the hybrid expert model when executing the computer program.
[0063] Based on the same inventive concept, the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the visible light-infrared image restoration fusion method based on the hybrid expert model when executed by a processor.
[0064] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0065] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The means for performing the functions specified in a flow or multiple flows and / or blocks.
[0066] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 The functions of a flow or multiple flows and / or a block or multiple blocks in conjunction with the disclosed methods can be implemented on a computer or other programmable data processing apparatus. Figure 1
[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 The functions of a flow or multiple flows and / or a block or multiple blocks in conjunction with the disclosed methods can be implemented on a computer or other programmable data processing apparatus. Figure 1
[0068] The above embodiments are merely illustrative of the technical idea of the present application and cannot limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical solutions falls within the protection scope of the present application.
Claims
1. A visible light-infrared image restoration and fusion method based on a hybrid expert model, characterized in that: The steps include: Step 1: Obtain a visible light-infrared fusion dataset and preprocess the fusion dataset. The preprocessed visible light-infrared fusion dataset is divided into a first-stage training set for the restoration task, a second-stage training set for the restoration and fusion tasks, and a test set. The visible light-infrared fusion dataset includes visible light-infrared degraded image pairs, non-degraded image pairs corresponding to the degraded image pairs, and fused images corresponding to the degraded image pairs. Step 2: Construct an image restoration and fusion model based on a hybrid expert mechanism, which includes a first-stage image restoration network and a second-stage image fusion network. The first-stage image restoration network includes the first and second branches, which are used to restore visible light and infrared degraded images respectively. The first branch takes the preprocessed visible light degraded image as input and outputs the visible light restored image and restored features. The second branch takes the preprocessed infrared degraded image as input and outputs the infrared restored image and restored features. The second-stage image fusion network is used to fuse the visible light and infrared restored features and output the fused image. Step 3: Use the first-stage training set to train the first-stage image restoration network to obtain a trained first-stage image restoration network; based on the trained first-stage image restoration network, use the second-stage training set to train the entire image restoration fusion model to obtain a trained image restoration fusion model; Step 4: Input the visible light-infrared degraded image pairs in the test set into the trained image restoration fusion model to obtain the fused image.
2. The visible light-infrared image restoration and fusion method based on the hybrid expert model according to claim 1 is characterized in that: In step 1, for the visible light-infrared degraded image pair, the degradation type of the visible light degraded image is low light, overexposure, rain, fog, blur or random noise, and the degradation type of the infrared degraded image is low contrast, streak noise or random noise.
3. The visible light-infrared image restoration and fusion method based on hybrid expert model according to claim 1 is characterized in that: In step 2, the first and second branches have the same structure, both including a degradation-aware gating module and a Transformer system; wherein the degradation-aware gating module is used to identify the degradation type of the input preprocessed visible light degraded image or infrared degraded image; The Transformer system has a U-Net structure and includes a degradation-aware encoder and a decoder. The degradation-aware encoder includes the first to fourth encoding layers, each of which is composed of a multi-expert collaborative module and a first Transformer module. The decoder includes the first to fourth decoding layers, each of which includes a second Transformer module. For the second to fourth coding layers, the input of the multi-expert collaborative module in each coding layer is the output of the first Transformer module and the output of the degradation-aware gating module in the previous coding layer. The input of the multi-expert collaborative module in the first coding layer is the output of the degradation-aware gating module and the preprocessed visible light degraded image or infrared degraded image. For the first to third decoding layers, the input of the second Transformer module of each decoding layer is the input of the second Transformer module of the next decoding layer and the output of the first Transformer module in the encoding layer corresponding to the current decoding layer. The input of the second Transformer module of the fourth decoding layer is the output of the first Transformer module of the fourth encoding layer; the second Transformer module of the first decoding layer outputs visible light restoration features or infrared restoration features.
4. The visible light-infrared image restoration and fusion method based on the hybrid expert model according to claim 3 is characterized in that: In step 2, the two-stage image fusion network includes a fusion perception gating module and a multi-expert collaboration module; the two-stage image fusion network fuses the visible light restoration features output by the first branch and the infrared restoration features output by the second branch to obtain the fused features, and then uses the fusion perception gating module to extract image information and scene features from the fused features to obtain fusion prompt components and expert gating weights; The multi-expert collaborative module selects experts to reconstruct fusion features based on expert gating weights and finally generates a fused image.
5. The visible light-infrared image restoration and fusion method based on hybrid expert model according to claim 1 is characterized in that: In step 3, when training the first-stage image restoration network, L1 loss is used. Degenerate Expert Balance Loss With gradient loss Jointly optimize the parameters of the one-stage image restoration network, that is, the loss function L used in the training of the one-stage image restoration network stage1 for: Among them, I restored Represents the restored image output by the one-stage image restoration network, I target represents the non-degraded image corresponding to the input visible light degraded image or infrared degraded image, λ is the preset hyperparameter, N is the total number of experts, and f i is the frequency of expert selection, P i is the average weight of the degradation-aware gating module on expert i, and Sobel is the gradient operator; When training the entire image restoration fusion model, L1 loss is used Fusion Expert Balanced Loss Gradient Loss and strength loss Jointly optimize the parameters of the entire image restoration fusion model, that is, the loss function used in the training of the entire image restoration fusion model for: Among them, I Vtarget represents the non-degraded visible light image, I Itarget represents the non-degraded infrared image, I fusion Represents the fused image output by the entire image restoration fusion model.
6. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the visible light-infrared image restoration and fusion method based on the hybrid expert model are implemented as described in any one of claims 1 to 5.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the visible light-infrared image restoration and fusion method based on a hybrid expert model are implemented.
Citation Information
Cited By
Cross-modal image enhancement fusion method fusing degradation identification and dynamic recovery mechanism
CN121660921A
Multi-modal feature fusion-based severe environment image enhancement and restoration method
CN122115236A
A method for image enhancement and restoration in harsh environments based on multimodal feature fusion
CN122115236B