Anti-degradation image fusion method and system based on prior knowledge and hybrid expert model

The hybrid expert model with diffusion network and semantic segmentation addresses image fusion challenges in complex environments, ensuring stable performance and improved semantic consistency for high-level vision tasks.

CN120318101APending Publication Date: 2025-07-15SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510395812.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of image degradation in complex environments, resulting in loss of image details, reduced contrast, and color distortion. The fusion network is unstable in a multi-task environment and lacks adaptability, which affects semantic consistency and the accuracy and robustness of advanced visual tasks.

Method used

An anti-degradation image fusion method based on prior knowledge and mixed expert models is adopted, and a degradation-free pseudo-label is generated through the diffusion model, combining the defog, brightness and rain degradation expert models, a degradation removal backbone network is built, and a semantic destruction segmentation module is introduced to optimize the fusion image loss function, and multimodal feature fusion and semantic segmentation are enhanced.

Benefits of technology

It realizes adaptive adjustment of various degradation types in complex environments, improves stability, enriches semantic information of the fused image, and enhances the accuracy and robustness of advanced visual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318101A_ABST
    Figure CN120318101A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-degradation image fusion method and system based on prior knowledge and a hybrid expert model. The method comprises the steps of obtaining a visible light image, an infrared image and text information in a degradation scene; constructing a degradation removal network based on a diffusion model, and generating a pseudo tag; constructing a degradation removal backbone network in combination with prior knowledge and a hybrid expert model to obtain visible light image features after degradation removal; inputting the multi-modal features into a trained fusion task head module, and outputting a fusion image; wherein the multi-modal features are combined with the text information to be input into a trained semantic deconstruction segmentation module, and a semantic segmentation result is obtained; and obtaining segmentation loss based on a semantic segmentation result, and optimizing the fused image in combination with the fused image loss. According to the method, the fused image is richer and more accurate in semantic level, more semantic information is reserved, and the precision and robustness of a subsequent advanced vision task are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and in particular to an anti-degradation image fusion method and system based on prior knowledge and a mixture of experts model. Background Art

[0002] Due to the limitations of imaging technology, a single type of sensor can only obtain partial information of a scene, while through the collaborative work of multiple sensors, different features of the scene can be comprehensively captured. Therefore, multi-modal image fusion has become an important research topic in the field of digital image processing technology. As one of the classic representatives of multi-modal image fusion technology, infrared-visible image fusion aims to simultaneously retain the thermal radiation information in the infrared image and the detailed texture and physical characteristics of objects in the visible light image, ensuring that the fusion result not only meets the intuitive needs of human visual perception but also provides accurate semantic expressions for high-level visual tasks. More importantly, this technology is of great significance for enhancing the practicality in fields such as autonomous driving and intelligent security. It not only promotes the progress of information extraction technology but also improves the data analysis and processing capabilities, thus supporting the comprehensive development of smart cities and smart life.

[0003] Autoencoder (AE), Convolutional Neural Network (CNN), and Generative Adversarial Network (GAN) technologies have become the mainstream methods in the field of multi-modal image fusion (MMIF). However, although numerous image fusion methods have been proposed currently, most methods are mainly designed and optimized based on normal scenes and are often difficult to effectively cope with image degradation problems in complex environments. First, the light change, noise interference, and other degradation factors in complex scenes will cause the visible light image to have problems such as detail loss, contrast reduction, and color distortion. This quality degradation not only seriously affects the input characteristics of the image, making it difficult for subsequent fusion to fully utilize the information of the source image, but also brings the problem of difficult label acquisition. In harsh environments, the real degradation scenes are complex and diverse, and it is difficult to obtain high-quality data as labels, thus affecting the training effect of the model. Second, the fusion network needs to complete image degradation removal and information fusion in a multi-task and multi-modal environment, but the coupling between multiple tasks easily leads to the instability of the network. Especially when dealing with multiple degradation forms, traditional methods lack adaptability to different degradation types, and the fusion network is prone to mode collapse due to interference between tasks, resulting in the model being unable to effectively generate high-quality fusion images. Finally, semantic information is easily obscured or distorted in complex degradation environments, resulting in the fusion result being unable to maintain semantic consistency. This semantic loss directly affects the accuracy and robustness of subsequent high-level visual tasks (such as object detection and object tracking). Summary of the Invention

[0004] In view of the deficiencies in the existing technology, the present invention provides an anti-degradation image fusion method and system based on prior knowledge and a mixture of experts model, aiming to solve the technical problems mentioned in the above background art.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In the first aspect, the present invention provides an anti-degradation image fusion method based on prior knowledge and a mixture of experts model, including:

[0007] Obtain visible light images and infrared images under a degradation scenario, preprocess the visible light images and infrared images, and obtain the text information of the visible light images;

[0008] Construct a degradation removal network based on a diffusion model, input the preprocessed visible light image into the degradation removal network based on the diffusion model, and output a non-degraded visible light image as a pseudo-label;

[0009] Based on the pseudo-label, construct a degradation removal backbone network by combining prior knowledge and a mixture of experts model, input the preprocessed visible light image into the degradation removal backbone network, pass through different expert models in sequence, couple the output results of different expert models, and output the features of the visible light image after degradation removal;

[0010] Extract the infrared image features, fuse the features of the visible light image after degradation removal to generate multi-modal features, input the multi-modal features into the trained fusion task head module, and output a fused image; wherein the multi-modal features are also combined with the text information and input into the trained semantic deconstruction segmentation module to output modulated fusion features and obtain a semantic segmentation result;

[0011] Obtain a segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

[0012] As a further technical solution, the degradation removal network based on the diffusion model includes two Markov chains, a forward process and a reverse process. The forward process gradually adds Gaussian noise to the preprocessed visible light image, and the reverse process gradually removes the added Gaussian noise; in the reverse process, a noise prediction network is used to predict the Gaussian noise corresponding to each pixel point.

[0013] As a further technical solution, the degradation removal backbone network constructed by combining prior knowledge and a mixture of experts model includes a defogging expert model, a brightness expert model, a rain degradation processing expert model, and a non-degraded visible light image processing expert model; and multiple expert models are coupled through a gating function.

[0014] As a further technical solution, the defogging expert model is specifically expressed as:

[0015] E fog = [x - A(1 - t(x))] / t(x); where A represents the atmospheric light value, which is generated by a fully connected network, and t(x) is the transmittance;

[0016] The brightness expert model is specifically expressed as:

[0017] E light (x) = x n + α n ·x n ·(1 - x n );

[0018] x n = x n-1 + α n-1 ·x n-1 ·(1 - x n-1 );

[0019] x n = x; where E light (x) represents the brightness expert model, x n represents the image brightness value after the nth iteration; α n represents the brightness adjustment coefficient used in the nth iteration, x n-1 represents the image brightness value after the (n - 1)th iteration, and x represents the brightness value of the original input image;

[0020] The rain degradation processing expert model and the non - degraded visible light image processing expert model are both composed of consecutive convolutional layers and activation layers.

[0021] As a further technical solution, the semantic deconstruction segmentation module includes a segmentation task head, which outputs modulated fusion features based on the input multi - modal features and text information, and then generates a semantic segmentation result, specifically expressed as:

[0022] where, represents the modulated fusion feature, MLP() is a fully connected layer, ⊙ represents the Hadamard product, F f represents the input multi - modal feature, text represents the input text information, F text represents the result obtained by encoding text using the text encoder of CLIP.

[0023] As a further technical solution, the segmentation loss is obtained based on the modulated fusion feature, specifically expressed as:

[0024] where N represents the number of pixel points in the segmented image, C represents the number of categories, y i,c represents the one - hot encoding of the true label of pixel i, pi,c The probability predicted by the semantic deconstruction and segmentation module.

[0025] As a further technical solution, the fused image loss is specifically expressed as: L f =

[0026] L col + L pix + L gard ; where L f represents the fused image loss, L col represents the color loss, L pix represents the pixel loss, and L gard represents the gradient loss;

[0027] Among them, the definition of the color loss L col is as follows:

[0028] Among them, respectively represent the cb channel and the cr channel of the fused image, and ‖‖1 represents the first norm, respectively represent the cb channel and the cr channel of the non-degraded visible light image;

[0029] The definition of the pixel loss L pix is as follows:

[0030] L pix = ||I f - max(I g , I inf )||1; where max represents taking the maximum element-wise;

[0031] The definition of the gradient loss L gard is as follows:

[0032] Among them, Δ represents the gradient value of the input image calculated using the Sobel operator.

[0033] In a second aspect, the present invention provides an anti-degradation image fusion system based on prior knowledge and a mixture of experts model, including the following modules:

[0034] An image and text acquisition module, configured to: acquire visible light images and infrared images in a degraded scene, preprocess the visible light images and infrared images, and acquire the text information of the visible light images;

[0035] A pseudo-label generation module, configured to: construct a degradation removal network based on a diffusion model, input the preprocessed visible light image into the degradation removal network based on the diffusion model, and output a non-degraded visible light image as a pseudo-label;

[0036] The degradation removal backbone network construction module is configured to: based on the pseudo labels, construct a degradation removal backbone network by combining prior knowledge and a mixture of experts model, input the preprocessed visible light image into the degradation removal backbone network, sequentially pass through different expert models, couple the output results of different expert models, and output the visible light image features after removing degradation;

[0037] The segmentation and fusion module is configured to: extract infrared image features, fuse the visible light image features after removing degradation, generate multi-modal features, input the multi-modal features into the trained fusion task head module, and output a fused image; wherein the multi-modal features are also combined with text information and input into the trained semantic deconstruction segmentation module to output modulated fusion features and obtain a semantic segmentation result;

[0038] The loss optimization module is configured to: obtain a segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

[0039] One or more technical solutions of the present invention have the following beneficial effects:

[0040] 1. The present invention constructs a degradation removal backbone network by combining prior knowledge and a mixture of experts model, defines multiple experts to handle different degradation types, and couples them through a gating function, enabling the present invention to make adaptive adjustments for different degradation types, avoiding mode collapse, being able to effectively handle multiple degradation types, and maintaining stable performance during multi-task processing.

[0041] 2. The present invention obtains images in a degradation scenario and uses a degradation removal network based on a diffusion model to generate non-degraded visible light images as high-quality pseudo labels, effectively solving the technical problem in the prior art that in a harsh environment, due to the complexity and diversity of the real degradation scenario, it is difficult to obtain high-quality data as labels, thus affecting the training effect of the model.

[0042] 3. The present invention introduces a semantic deconstruction segmentation module, outputs modulated fusion features through an additional segmentation task head, and uses multi-modal features and text information, and guides the generation of a semantic segmentation result through a segmentation loss. This makes the fused image richer and more accurate at the semantic level, retains more semantic information, and enhances the accuracy and robustness of subsequent high-level vision tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0044] Figure 1 It is a flowchart of the image fusion method provided in Embodiment 1 of the present invention;

[0045] Figure 2 It is a schematic framework diagram of the image fusion method provided in the first embodiment of the present invention;

[0046] Figure 3 It is a schematic framework diagram of generating pseudo - labels in the first embodiment of the present invention;

[0047] Figure 4 It is a schematic framework diagram of a degradation removal backbone network constructed by combining prior knowledge and a mixture - of - experts model in the first embodiment of the present invention;

[0048] Figure 5 It is a schematic framework diagram of the semantic deconstruction segmentation module in the first embodiment of the present invention; Detailed implementation mode

[0049] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0050] Embodiment 1

[0051] In this embodiment, an anti - degradation image fusion method based on prior knowledge and a mixture - of - experts model is provided, as Figure 1 and 2 shown. The specific steps of the method are as follows:

[0052] S1: Obtain visible - light images and infrared images in a degraded scenario, pre - process the visible - light images and infrared images, and obtain the text information of the visible - light images;

[0053] S2: Construct a degradation removal network based on a diffusion model, input the pre - processed visible - light image into the degradation removal network based on the diffusion model, and output a non - degraded visible - light image, which is used as a pseudo - label;

[0054] S3: Based on the pseudo - label, construct a degradation removal backbone network by combining prior knowledge and a mixture - of - experts model, input the pre - processed visible - light image into the degradation removal backbone network, pass through different expert models in sequence, couple the output results of different expert models, and output the visible - light image features after removing degradation;

[0055] S4: Extract infrared image features, fuse the visible - light image features after removing degradation to generate multi - modal features, input the multi - modal features into the trained fusion task head module, and output a fused image; wherein the multi - modal features are also combined with the text information and input into the trained semantic deconstruction segmentation module to output modulated fusion features and obtain a semantic segmentation result;

[0056] S5: Obtain the segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

[0057] In step S1, in order to fully evaluate the performance of the model under various complex degradation scenarios, in this embodiment, a dataset is collected from different scenarios. The dataset includes various complex degraded image scenarios, covering rainy days, foggy days, overexposure, low light, and other scenarios to ensure the diversity and representativeness of the dataset.

[0058] For the dataset to be used subsequently for the degradation removal network based on the diffusion model, in this embodiment, the D-HAZY and LoL datasets are mixed, and additional raindrops are added to the mixed dataset, and then it is used to train the degradation removal network based on the diffusion model. For the dataset to be used subsequently for constructing the degradation removal backbone network by combining prior knowledge and the mixture of experts model, in this embodiment, the M3FD, MSRS, and LLVIP datasets are selected for training, and the degradation removal network based on the diffusion model is used to generate pseudo-labels to guide the training of the fusion network. Specifically, the pseudo-labels generated by the diffusion model and the fused image are used together to calculate the value of the loss function, and then these loss values are passed back to the network through backpropagation to guide the training of the fusion network.

[0059] In step S1, the preprocessing is as follows: for the infrared light image, ensure that it is in the format of a single-channel grayscale image; for the visible light image, keep its original RGB three-channel structure. Subsequently, both types of images are normalized, and the pixel values are standardized to the interval [0,1] and converted to the tensor format. The text information of the visible light image is obtained by inputting the image into the LLava large model.

[0060] In step S2, due to the problem that it is difficult to obtain paired non-degraded images in the existing visible light-infrared image fusion data, in this embodiment, a degradation removal network based on the diffusion model is proposed to generate non-degraded images as pseudo-labels. Specifically, in this embodiment, the diffusion model is used to remove the degradation of the visible light image and generate a non-degraded visible light image as a pseudo-label. The pseudo-label participates in the calculation of multiple loss functions and serves as a supervision signal during the training process to guide the fusion network to recover clear visible light features from the degraded image.

[0061] As Figure 3 shown, the degradation removal network based on the diffusion model in this embodiment includes two Markov chains, a forward process and a reverse process. The forward process gradually adds Gaussian noise to the preprocessed visible light image, and the reverse process gradually removes the added Gaussian noise to generate a clear image. To accelerate the diffusion process, the forward process can be directly defined by the initial image and the noise parameters, and the sampling process can be made differentiable through reparameterization, that is, the random noise is rewritten as a combination of a deterministic variable and a Gaussian random variable.

[0062] In this embodiment, each forward process step can be defined as:

[0063] Among them, x t represents the image state at time step t, x t-1 Indicates that at time step t - The image state at 1, ψ t represents the noise coefficient at time step t, which increases with the increase of time step and is used to control the intensity of noise addition. I represents the unit matrix, which is used to define the covariance matrix of Gaussian noise. In order to accelerate the diffusion process, x t In the forward process, it can be obtained directly from x0.

[0064] Definition φ t =1-ψ t , represents the cumulative noise coefficient at time step t, represents the cumulative noise coefficient from time step 1 to time step t. The forward process can be redefined as:

[0065]

[0066] Considering that random sampling is not differentiable, the gradient descent method cannot be used directly for training. Therefore, this embodiment uses a reparameterization method to make the sampling process differentiable. Specifically, the random noise n~N(μ+σ 2 ) can be rewritten as n = μ + σ∈, where ∈ is a random variable that follows a Gaussian distribution. Therefore, the forward process can be restated as:

[0067] in, Represents the actual Gaussian noise added, which is used to compare with the predicted noise and calculate the loss.

[0068] The reverse process gradually removes noise by learning a Gaussian distribution with a mean and a fixed variance. In this embodiment, a noise prediction network is used to predict the Gaussian noise corresponding to each pixel point during the reverse process. Specifically, in this embodiment, Unet suitable for the pix2pix task is selected as the overall framework of the noise prediction network. Since there are various degradation information in visible light images, different styles of noise information need to be predicted for different de-degradation tasks. Therefore, this embodiment uses a mixture of experts model to improve the noise prediction network. Specifically, the proposed noise prediction network consists of an encoder part and a decoder part. In the encoder part, the input features are subjected to multi-scale feature extraction through 4 consecutive mixture of experts layers and downsampling layers; in the decoder part, the opposite operation is taken, and 4 consecutive mixture of experts layers and upsampling layers are used for noise prediction; skip connections are used between the corresponding layers of the encoder and the decoder to avoid potential information loss. After the training of the noise prediction network is completed, the clean image will be generated from a randomly sampled pure noise image, so as to guide the training of the fusion network in the subsequent process.

[0069] In step S3, as Figure 4 shown, the degradation removal backbone network constructed by combining prior knowledge and the mixture of experts model includes a defogging expert model, a brightness expert model, a rain degradation processing expert model, and a visible light image processing expert model without degradation; and multiple expert models are coupled through a gating function.

[0070] Specifically: First, this embodiment defines a defogging expert model for removing haze interference in images, which is specifically expressed as:

[0071] E fog = [x - A(1 - t(x))] / t(x); where x represents the visible light image feature. The above formula is extended from the atmospheric scattering model [], A represents the atmospheric light value, which is generated by a fully connected network, and t(x) is the transmittance, which is predicted from the input x by another MLP layer. The method for obtaining the visible light image feature is as follows: The preprocessed visible light image first undergoes preliminary feature extraction through four convolutional layers; subsequently, these preliminarily extracted features are further subjected to deep feature extraction through a series of convolutional layers and LeakyReLU activation functions, and finally the visible light image feature input to the defogging expert is obtained.

[0072] Further considering the influence of illumination degradation on the network, this embodiment defines a brightness expert model for adjusting the brightness of the image, which is specifically expressed as:

[0073] E light (x) = x n + α n ·x n ·(1 - xn )

[0074] x n = x n-1 + α n-1 · x n-1 · (1 - x n-1 )

[0075] x n = x; where, E light (x) represents a brightness expert model for adjusting the brightness of the input image. x n represents the image brightness value after the nth iteration. α n represents the brightness adjustment coefficient used in the nth iteration for controlling the intensity of brightness adjustment. x n-1 represents the image brightness value after the (n - 1)th iteration. x represents the brightness value of the original input image. The above formula is extended from the brightness adjustment curve model, and the brightness value of the input image is continuously adjusted in a recursive manner. α n = {α1, α2, …, α n} is predicted by a fully connected layer. In this embodiment, the value of n is set to 8.

[0076] Considering the diversity of the input image, this embodiment additionally defines an expert module E rain (x) for processing rain degradation and an expert module E norm (x) for processing non-degraded visible light images, both of which are composed of consecutive convolutional layers and activation layers.

[0077] In this embodiment, multiple expert models are coupled through a gating function:

[0078] G(x) = {g1(x), g2(x), g3(x), g4(x)}

[0079]

[0080] where g i (x) represents the activation weight for the ith expert, and Φ i (x) is a scoring function, which is implemented using a fully connected layer in this embodiment. Finally, the final output of the visible light feature extractor can be defined as:

[0081]

[0082] where Top K (·) represents selectively activating the top k experts with dominant activation weights, and k is set to 1 in this embodiment.

[0083] In step S4, as Figure 5As shown, in order to enable the image fusion network to retain as much semantic information as possible during the execution of the de - degradation task and the multi - modal information fusion task, in this embodiment, an additional semantic deconstruction segmentation module (segmentation task head) is added outside the conventional fusion task head module. Its function is to use the de - degraded and fused multi - modal features to generate high - quality semantic segmentation results, thereby forcing the fused features to have more semantic information. Specifically, in this embodiment, the SegFormer architecture is selected for expansion. Specifically, the first - layer feature extraction unit in the SegFormer model is removed, and the fused multi - modal features are used as the input of the SegFormer model.

[0084] Among them, infrared feature extraction is achieved through the infrared feature activation module, which consists of four - layer convolutional networks, and each layer is followed by a LeakyReLU activation function. Intermediate features at all levels are retained during the feature extraction process, and an expert mixture module is introduced after the second layer to enhance the feature expression ability. This hierarchical extraction method captures both the global thermal distribution information of the infrared image and retains local hot - spot details, providing a complete thermal imaging feature representation for subsequent multi - modal fusion.

[0085] The infrared features extracted by the infrared feature activation module and the visible - light features extracted by the degradation - removal backbone network constructed by combining prior knowledge and the mixture - of - experts model are first concatenated in the channel dimension to form a single feature map. Then, this feature map is processed by two - layer convolutional networks for dimensionality reduction and added to the intermediate - layer features of the two modalities to enhance the detailed information. The finally obtained multi - modal fusion features will be used as the input of the SegFormer model.

[0086] Considering that the GT labels used for training the segmentation task head are generated by large models (such as ZipCLIP, SAM), there are category information differences in the labels for different input scenarios. To avoid the negative impact of different categories on the training of the segmentation task, this embodiment uses text information to modulate the features input to SegFormer, thereby providing more semantic context. Specifically, ZipCLIP preprocesses the image according to different degradation modes of the input image by generating transformation boxes of the image and extracts different regional features of the image. Then, SAM segments the image based on these generated boxes to produce predicted semantic labels or features as the preliminary results of semantic segmentation. Subsequently, this embodiment uses the text encoder of CLIP to convert the input text information into a representation compatible with the image features. By tokenizing and encoding the input text, global feature representations are generated, and these text features are embedded into the image feature space to enhance the semantic expression of the image features. Finally, these embedded text information and the image features generated by ZipCLIP and SAM act on SegFormer together to provide additional semantic guidance, thereby improving the performance of the segmentation model.

[0087] In this embodiment, the text information is defined as text, and the multi-modal feature is F f , and the feature input to SegFormer can be defined as:

[0088]

[0089] represents the modulated fusion feature, MLP() is the fully connected layer, ⊙ represents the Hadamard product, F f represents the input multi-modal feature, text represents the input text information, F text represents the result obtained by encoding text using the text encoder of CLIP. In the above way, the additional text information can be used as additional guiding information to address the training collapse problem caused by category differences in the environment.

[0090] In step S5, in the image fusion task in extreme environments, complex environmental factors often cause serious image degradation phenomena, resulting in the loss of key details and semantic information. To effectively address these challenges, this embodiment proposes an improved loss function optimization strategy aimed at simultaneously improving the quality of the fused image from both the detail and semantic levels.

[0091] The loss of the network consists of two parts, and the loss function can be defined as:

[0092] L = L f + L seg

[0093] Where Lf denotes the fused image loss, which is used to constrain the fused task head to generate high-quality fused images. L seg denotes the segmentation loss, which is used to constrain the semantic segmentation task head to generate high-quality segmentation results.

[0094] L f is specifically represented as follows:

[0095] L f = L col + L pix + L gard ;

[0096] where L f denotes the fused image loss, L col denotes the color loss, L pix denotes the pixel loss, L gard denotes the gradient loss.

[0097] The purpose of the color loss is to ensure that the fused image has vivid colors. The color loss L col is defined as follows:

[0098]

[0099] where, respectively represent the cb channel and the cr channel of the fused image, ‖‖1 represents the first norm, respectively represent the cb channel and the cr channel of the non-degraded visible light image.

[0100] The purpose of the pixel loss is to ensure that the fused image can integrate multi-modal input information at the pixel level. The pixel loss L pix is defined as follows:

[0101] L pix = ||I f - max(I g , I inf )||1;

[0102] where, max represents taking the maximum element-wise, I f represents the fused image, I g represents the pseudo-label generated using the diffusion model, I inf represents the infrared light image.

[0103] The purpose of the gradient loss is to ensure that the fused image has excellent texture detail information. The gradient loss L gard is defined as follows:

[0104]

[0105] where, It represents calculating the gradient value of the input image using the Sobel operator.

[0106] For the segmentation loss, it is composed of the cross-entropy loss:

[0107]

[0108] where N represents the number of pixel points in the segmented image, and C represents the number of categories. y i,c represents the one-hot encoding of the true label of pixel i, and p i,c is the probability predicted by the semantic deconstruction segmentation module.

[0109] Example Two

[0110] In this example, an anti-degradation image fusion system based on prior knowledge and a mixture of experts model is provided, including the following modules:

[0111] An image and text acquisition module, configured to: acquire visible light images and infrared images in a degraded scenario, preprocess the visible light images and infrared images, and acquire the text information of the visible light images and infrared images;

[0112] A pseudo-label generation module, configured to: construct a degradation removal network based on a diffusion model, input the preprocessed visible light image into the degradation removal network based on the diffusion model, and output a non-degraded visible light image as a pseudo-label;

[0113] A degradation removal backbone network construction module, configured to: based on the pseudo-label, construct a degradation removal backbone network by combining prior knowledge and a mixture of experts model, input the preprocessed visible light image into the degradation removal backbone network, pass through different expert models in sequence, couple the output results of different expert models, and output the feature of the visible light image after removing degradation;

[0114] A segmentation and fusion module, configured to: extract the infrared image feature, fuse the feature of the visible light image after removing degradation to generate a multi-modal feature, input the multi-modal feature into the trained fusion task head module, and output a fused image; wherein the multi-modal feature is also combined with the text information and input into the trained semantic deconstruction segmentation module to output a modulated fusion feature and obtain a semantic segmentation result;

[0115] A loss optimization module, configured to: obtain a segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

[0116] For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An anti-degradation image fusion method based on prior knowledge and a mixture of experts model, characterized in that Including: Obtain visible light images and infrared images in a degraded scenario, preprocess the visible light images and infrared images, and obtain the text information of the visible light images; Construct a degradation removal network based on a diffusion model, input the preprocessed visible light image into the degradation removal network based on the diffusion model, and output a non-degraded visible light image, which is used as a pseudo-label; Based on the pseudo-label, construct a degradation removal backbone network by combining prior knowledge and a mixture of experts model. Input the preprocessed visible light image into the degradation removal backbone network, and successively pass through different expert models. Couple the output results of different expert models to output the feature of the visible light image after degradation removal; Extract the infrared image feature, fuse the feature of the visible light image after degradation removal to generate a multi-modal feature, input the multi-modal feature into the trained fusion task head module, and output a fused image; wherein the multi-modal feature is also combined with the text information and input into the trained semantic deconstruction segmentation module to output a modulated fusion feature and obtain a semantic segmentation result; Obtain a segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

2. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 1, wherein The degradation removal network based on the diffusion model includes two Markov chains, a forward process and a reverse process. The forward process gradually adds Gaussian noise to the preprocessed visible light image, and the reverse process gradually removes the added Gaussian noise; in the reverse process, a noise prediction network is used to predict the Gaussian noise corresponding to each pixel point.

3. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 1, wherein The degradation removal backbone network constructed by combining prior knowledge and a mixture of experts model includes a defogging expert model, a brightness expert model, a rain degradation processing expert model, and a non-degraded visible light image processing expert model; and multiple expert models are coupled through a gating function.

4. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 3, wherein, The defogging expert model is specifically represented as: E fog = [x - A(1 - t(x))] / t(x); where A represents the atmospheric light value, which is generated by a fully connected network, and t(x) is the transmittance; The brightness expert model is specifically represented as: E light f(x) = x n + α n · x n · (1 - x n ); x n = x n-1 + α n-1 · x n-1 · (1 - x n-1 ); x n = x; where, E light (x) represents the brightness expert model, x n represents the image brightness value after the nth iteration; α n represents the brightness adjustment coefficient used in the nth iteration, x n-1 represents the image brightness value after the (n - 1)th iteration, and x represents the brightness value of the original input image; Both the rain degradation processing expert model and the non-degraded visible light image processing expert model are composed of consecutive convolutional layers and activation layers.

5. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 1, wherein The semantic deconstruction segmentation module includes a segmentation task head. The segmentation task head outputs a modulated fusion feature based on the input multi-modal feature and text information, and then generates a semantic segmentation result, which is specifically represented as: Among them, represents the modulation fusion feature, MLP() is the fully connected layer, ⊙ represents the Hadamard product, and F f represents the input multi-modal feature, text represents the input text information, and F text is obtained by encoding text using the text encoder of CLIP.

6. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 1, wherein Obtain a segmentation loss based on the modulated fusion feature, which is specifically represented as: Among them, N represents the number of pixel points in the segmented image, C represents the number of categories, and y i,c represents the one-hot encoding of the true label of pixel i, and p i,c is the probability predicted by the semantic deconstruction segmentation module.

7. The anti-degradation image fusion method based on prior knowledge and a mixture of experts model according to claim 1, wherein The fusion image loss is specifically expressed as: L f = L col + L pix + L gard ; where L f represents the fusion image loss, L col represents the color loss, L pix represents the pixel loss, and L gard represents the gradient loss; Among them, the color loss L col is defined as follows: Among them, respectively represent the cb channel and the cr channel of the fused image, and ‖‖1 represents the first norm, respectively represent the cb channel and the cr channel of the non-degraded visible light image; Pixel loss L pix is defined as follows: L pix = ||I f - max(I g , I inf )||1; where max represents taking the element-wise maximum; Gradient loss L gard is defined as follows: Among them, denotes calculating the gradient value of the input image by using the Sobel operator.

8. An anti-degradation image fusion system based on prior knowledge and a mixture of experts model, characterized in that, Including the following modules: An image and text acquisition module, configured to: obtain visible light images and infrared images in a degraded scenario, preprocess the visible light images and infrared images, and obtain the text information of the visible light images; A pseudo-label generation module, configured to: construct a degradation removal network based on a diffusion model, input the preprocessed visible light image into the degradation removal network based on the diffusion model, and output a non-degraded visible light image, which is used as a pseudo-label; A degradation removal backbone network construction module, configured to: based on the pseudo-label, construct a degradation removal backbone network by combining prior knowledge and a mixture of experts model. Input the preprocessed visible light image into the degradation removal backbone network, and successively pass through different expert models. Couple the output results of different expert models to output the feature of the visible light image after degradation removal; The segmentation and fusion module is configured to: extract infrared image features, fuse the visible light image features after removing degradation, generate multi-modal features, input the multi-modal features into the trained fusion task head module, and output a fused image; wherein the multi-modal features are also combined with text information and input into the trained semantic deconstruction and segmentation module, output modulated fusion features, and obtain a semantic segmentation result. The loss optimization module is configured to: obtain a segmentation loss based on the semantic segmentation result and combine it with the fused image loss to optimize the fused image.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the anti-degradation image fusion method based on prior knowledge and the mixture of experts model according to any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the anti-degradation image fusion method based on prior knowledge and the mixture of experts model according to any one of claims 1-7.

Citation Information

Cited By

  • Cross-modal image enhancement fusion method fusing degradation identification and dynamic recovery mechanism

    CN121660921A