An unsupervised infrared guided visible image defogging method and device based on multispectral domain adaptation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0011]针对现有技术中真实雾天场景下域适应框架难以充分利用红外模态信息,导致合成域到真实域的去雾能力迁移受限,以及现有融合策略缺乏针对浓雾区域的有效处理机制、难以实现浓雾区域稳健去雾的问题,本发明提出了一种基于多光谱域适应的无监督红外引导可见光图像去雾方法及装置,通过构建红外辅助的多光谱去雾模型,并结合面向真实雾天场景的跨域迁移训练与区域自适应融合机制,实现真实雾天图像的有效去雾,提高浓雾区域的恢复能力和模型的跨场景泛化能力
[0076] First, the unsupervised infrared-guided visible light image dehazing method and apparatus based on multispectral domain adaptation of the present invention abandons the dependence of existing dehazing techniques on supervision by real paired clear images. It adopts an unsupervised domain adaptation training strategy from the synthetic domain to the real domain, and transfers the multispectral dehazing capability obtained by the source domain training to the real foggy scene in the target domain through a cross-domain multispectral distillation mechanism. At the same time, it avoids the modal conflict and artifact problems of infrared and visible light in the domain transfer process. Even in the absence of real paired clear image supervision, it can still achieve effective dehazing of real domain images, thereby improving the model's cross-scene generalization ability and practical application value.
Smart Images

Figure CN122550415A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, specifically to an unsupervised infrared-guided visible light image dehazing method and apparatus based on multispectral domain adaptation. Background Technology
[0002] Image dehazing is a fundamental task in computer vision, aiming to restore clear scenes from degraded images affected by haze scattering. The dehazing effect not only impacts the visual quality of the image itself but also directly affects the accuracy and stability of downstream vision tasks such as autonomous driving, remote sensing interpretation, outdoor monitoring, object detection, object recognition, and object tracking. Therefore, image dehazing technology has received widespread attention.
[0003] Existing image dehazing methods mainly include those based on prior models and those based on deep learning. Early methods typically relied on manually designed physical assumptions such as dark channel priors, color attenuation models, and fog line models to estimate transmittance or atmospheric light parameters. These methods could achieve certain results in specific scenarios, but they were sensitive to scene conditions and had limited generalization ability. With the development of deep learning, data-driven methods based on convolutional neural networks or Transformers have gradually become mainstream. These methods can directly learn the mapping relationship from foggy images to clear images, achieving better restoration performance in conventional dehazing tasks.
[0004] However, most existing deep learning dehazing methods are still based on single-modal visible light images. These methods typically rely solely on the degradation information of the foggy visible light image itself for recovery. While they may achieve some results in light or moderate fog scenes, in real-world complex fog scenes such as dense fog, heavy fog, non-uniform fog, and distant fog, the texture, edge, and structural information in the visible light image are severely obscured. This makes it difficult for the network to fully perceive and recover the obscured areas, easily leading to a series of difficult-to-repair problems such as severe residual fog, loss of detail, structural blurring, color shift, recovery distortion, and edge breakage. This severely limits their usability in real-world outdoor scenes.
[0005] To improve defogging capabilities in dense fog scenes, introducing infrared modes to provide supplementary information for visible light defogging has become a promising research direction. Infrared images, with their longer wavelengths and relatively less scattering by fog, typically maintain relatively stable target outlines and scene structure information even in foggy conditions, thus serving as an effective structural prior for visible light image defogging. Based on this, joint modeling of infrared and visible light images helps enhance the ability to recover structures and preserve details under dense fog conditions.
[0006] However, the application of existing infrared-assisted defogging technology in real foggy scenarios still faces two key technical problems, which have become key obstacles restricting the practical deployment of multispectral defogging.
[0007] First, it is difficult to obtain a dataset of strictly corresponding foggy visible light images, clear visible light images, and infrared images under real foggy conditions. This makes it difficult for existing methods to obtain sufficient and reliable supervision information in real scenes, thus limiting the model's defogging effect and generalization ability on real foggy images.
[0008] Regarding the first problem mentioned above, most existing infrared-assisted dehazing methods are still based on supervised learning frameworks, typically relying on triplets of strictly corresponding hazy visible light images, clear visible light images, and infrared images for training. However, under real foggy conditions, changes in illumination, atmospheric disturbances, camera movement, and dynamic scene changes can all lead to registration failures and missing labels, making it impossible to construct large-scale, high-quality, spatially strictly aligned triplet supervised datasets. Due to this limitation, existing methods often rely on synthetic foggy data for training, resulting in a significant domain shift between the dehazing mapping learned by the model and the real fog degradation process. This leads to problems such as incomplete dehazing, heavy residual fog, insufficient detail recovery, weak generalization ability, and poor robustness in real-world scenes. Some unsupervised domain-adaptive dehazing methods have attempted to transfer dehazing capabilities from the synthetic domain to real foggy images to alleviate the training difficulties caused by the lack of supervision from real, clear images. However, most of these methods revolve around single-modal visible light images, primarily modeling the domain shift between synthetic and real foggy visible light images. They haven't fully incorporated infrared modes for collaborative dehazing, thus failing to leverage the structural preservation advantages of infrared images in foggy conditions. Furthermore, when the unsupervised domain adaptation framework is extended to infrared-visible multispectral dehazing scenarios, the model must not only address the differences in visible light degradation distribution between the synthetic and real domains but also simultaneously contend with imaging differences between infrared and visible light modes, as well as distribution shifts between source and target domain multispectral data. Without effective constraints and distillation mechanisms for cross-domain multispectral features, modal conflicts, feature misalignments, and cross-domain artifacts can easily occur, leading to structural distortion, texture anomalies, and failure to recover dense fog regions in the dehazing results. Therefore, the stability and generalization ability of existing methods in real-world complex foggy scenarios remain limited, failing to meet practical application requirements.
[0009] Second, existing multispectral defogging methods often struggle to effectively differentiate and adaptively process different regions based on their degradation levels when fusing infrared and visible light information. This is especially true in dense fog regions where visible light information is severely degraded or even distorted. If simple splicing, direct superposition, or fixed-weight fusion methods are still used, it can easily lead to modal conflicts, weaken the structural recovery capability of dense fog regions, and thus affect the overall defogging quality.
[0010] Regarding the second issue mentioned above, existing multispectral fusion strategies still fall short in handling the modal differences between infrared and visible light. Infrared and visible light images differ in their imaging mechanisms, and the reliability and complementarity of the two modal information vary across different spatial locations and fog concentrations. Existing methods often employ simple stitching, direct weighting, or fixed fusion ratios for joint modeling, lacking dynamic, adaptive, and region-specific fusion mechanisms for different degraded regions. In dense fog regions, forcibly fusing visible light features heavily polluted by fog with infrared features directly introduces distorted information, leading to modal inconsistencies, increased artifacts, blurred edges, and structural distortions. In high-visibility regions, failing to fully preserve effective texture, color, and detail information from visible light results in texture loss, unnatural colors, and poor visual effects. Therefore, existing technologies consistently fail to achieve a balance between structural restoration in dense fog regions and texture preservation in clear regions, struggling to simultaneously ensure effective multimodal complementarity and robust structural restoration in dense fog regions, thus failing to meet the high-precision defogging requirements of real-world, complex foggy scenes. Summary of the Invention
[0011] To address the limitations of existing technologies in fully utilizing infrared modal information in real-world foggy scenarios, which restricts the transfer of defogging capabilities from the synthetic domain to the real domain, and the lack of effective processing mechanisms for dense fog regions in existing fusion strategies, making it difficult to achieve robust defogging in dense fog areas, this invention proposes an unsupervised infrared-guided visible light image defogging method and apparatus based on multispectral domain adaptation. By constructing an infrared-assisted multispectral defogging model and combining cross-domain transfer training and region adaptive fusion mechanisms for real-world foggy scenarios, effective defogging of real-world foggy images is achieved, improving the recovery capability of dense fog regions and the model's cross-scene generalization ability.
[0012] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:
[0013] In a first aspect, the present invention discloses an unsupervised infrared-guided visible light image dehazing method based on multispectral domain adaptation, the method comprising:
[0014] S1. Construct a spatially registered source domain synthetic multispectral dataset and a target domain real multispectral dataset; the source domain multispectral dataset includes a clear visible light image of the source domain, a corresponding infrared image of the source domain, and a synthetic hazy visible light image generated based on the clear visible light image of the source domain; the target domain real multispectral dataset includes a real hazy visible light image of the target domain and its corresponding infrared image.
[0015] S2. Construct a multispectral dehazing model (MFD); wherein the multispectral dehazing model (MFD) includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder VisEncoder for extracting hazy visible light features, a second encoder IREncoder for extracting infrared features, a region adaptive multispectral fusion module ReMix for performing region adaptive fusion of the hazy visible light features and infrared features, a decoder FusionDecoder for reconstructing the dehazed visible light image based on the fused features, and a dehazed image output terminal;
[0016] S3. The multispectral dehazing model is trained using a phased training strategy:
[0017] In the first stage, supervised pre-training is performed using the source domain multispectral dataset to obtain the source domain teacher multispectral dehazing model. During the pre-training process, a visibility confidence map is generated based on the gradient difference between hazy visible light features and infrared features. The visibility confidence map is then subjected to segmented weight hardening to obtain a three-segment region fusion weight. The visible light features and infrared features are then fused differentially according to the fusion weight: visible light features are hard-retained in high visibility regions, infrared features are hard-replaced in low visibility regions, and adaptive weighted fusion is performed in transition regions.
[0018] In the second stage, the target domain student multispectral dehazing model is initialized with the parameters of the source domain teacher multispectral dehazing model. The real foggy visible light image and the corresponding infrared image of the target domain are simultaneously input into the teacher model and the student model. Cross-domain multispectral distillation is used to impose a dual consistency constraint on the fusion representation and dehazing prediction results of the student model. Domain adaptation training is completed under unsupervised conditions to obtain a target domain student model adapted to real foggy scenes.
[0019] S4. Input the real foggy visible light image to be processed and the corresponding infrared image into the trained target domain student multispectral dehazing model, and output the dehazed visible light image.
[0020] Furthermore, the region adaptive multispectral fusion module ReMix includes a confidence generation unit, a segmented hardening unit, and a feature fusion unit;
[0021] The confidence generation unit receives the first encoder VisEncoder output of the first... Visible light characteristics of layered fog and the output of the second encoder IREncoder Layer infrared features The gradient difference between the visible light features and infrared features of foggy conditions is calculated using the following formula, and a visibility confidence map is generated. , represented as:
[0022]
[0023] In the formula, Represents the spatial gradient operator, express Norm, This represents the difference sensitivity control coefficient. Indicates the first Spatial location in the layer feature map;
[0024] The segmented hardening unit is based on the visibility confidence map. and the Visible light characteristics of layered fog Generate visibility confidence map , represented as:
[0025]
[0026] In the formula, This represents the activation function. Represents a convolution mapping. The balance coefficient is then used; based on the aforementioned visibility confidence map... Perform segmented weight hardening to generate a fused weight graph. , represented as:
[0027]
[0028] In the formula, and These represent the low threshold and the high threshold, respectively. Finally, the weighted graph will be merged. Expanded to the first in the channel dimension corresponding to the layer feature map One channel;
[0029] The feature fusion unit receives the fusion weight map after dimensionality expansion. , No. Visible light characteristics of layered fog and the Layer infrared features Perform segmented weight hardening processing, in The region retains visible light characteristics, in Perform infrared feature hard replacement in the region, in The transition region is weighted and fused with visible light and infrared features to obtain the first... Layer fusion features , represented as:
[0030]
[0031] in, This indicates element-wise multiplication.
[0032] Furthermore, the decoder FusionDecoder is used to process the fusion features output by the region adaptive multispectral fusion module ReMix. Reconstruction is performed to obtain the source region dehazing prediction map. The decoder FusionDecoder includes a fused feature input, a supplementary branch input, a cross-layer feature aggregation and decoding module, and an output mapping and dehazing prediction module.
[0033] The fusion feature input terminal is used to receive the multi-layer fusion features output by the ReMix regional adaptive multispectral fusion module. ;in, Indicates the first Layer fusion features, dimension 1 , The layer index is represented, with a range of [1,5].
[0034] The supplementary branch input is used to receive multi-layered hazy visible light features extracted by the first encoder, VisEncoder. This is used as a supplementary branch for the local texture of the corresponding layer, and the fusion weight map is utilized. Position-by-position suppression or retention control is applied to the visible light characteristics of the corresponding layer containing fog.
[0035] The cross-layer feature aggregation and decoding module adopts a five-layer upsampling decoding method from deep to shallow, using the multi-layer fused features as the main basis for reconstruction, and combining the decoded features of the previous layer with the fused weight map. After element-wise weighted control, the corresponding layer's hazy visible light features are subjected to cross-layer feature aggregation and reconstruction to obtain the first layer. Layer decoding features , represented as:
[0036]
[0037]
[0038] in, Indicates the deepest decoded features. Indicates the first Layer decoding features, This represents upsampling and convolution operations; This indicates a feature concatenation operation. This represents element-wise multiplication;
[0039] The output mapping and dehazing prediction module will decode the first layer of features. Perform output mapping to obtain the dehazing prediction map. , represented as:
[0040]
[0041] in, This indicates the output mapping convolution operation.
[0042] Furthermore, in step S3, in the first stage, the source domain dehazing loss function is adopted. Multispectral dehazing model for source domain teachers Supervised training was conducted to learn how to synthesize hazy visible light images from the source domain. Source domain infrared images Clear visible light image of the source domain The cross-modal dehazing mapping relationship is established, and edge details and overall structural information are preserved during the dehazing process; the source domain dehazing loss function is described. Represented as:
[0043]
[0044] in, , , They represent , , Weighting coefficients; for The loss function is expressed as:
[0045]
[0046] in, Represents the pixel number of the image. Indicates the total number of pixels in the image; The structural similarity loss function is expressed as:
[0047]
[0048] in, Represents a structural similarity function; The perceptual loss function is expressed as:
[0049]
[0050] in, The first feature extraction operator represents the... Layer feature mapping, This represents the set of feature layers involved in the calculation of perceptual loss.
[0051] Furthermore, in step S3, the second stage of training is as follows:
[0052] The source domain teacher multispectral dehazing model was trained. Parameter initialization target domain student multispectral dehazing model ;
[0053] Real foggy visible light image of the target domain and its corresponding infrared image Simultaneously input the source domain teacher multispectral dehazing model and the target domain student multispectral dehazing model The source domain teacher defogging prediction map was obtained respectively. Target Domain Student Defogging Prediction Map , No. Layered source domain teacher integration characteristics and the Layer target domain student integration characteristics :
[0054] ;
[0055] ;
[0056] Using joint loss function Multispectral dehazing model for target domain students Unsupervised domain adaptation training is performed, and the joint loss function is... Represented as:
[0057]
[0058] in, and Represent the domain consistency loss function respectively Image-text joint semantic constraint loss Weighting coefficients;
[0059] The domain consistency loss function Consistency loss due to fusion features Consistency loss with defogging prediction Composition, represented as:
[0060]
[0061] in, and They represent and Weighting coefficients;
[0062] The consistency loss of the fusion features Represented as:
[0063] ;
[0064] The defogging prediction consistency loss Represented as:
[0065] ;
[0066] in, This represents the number of mini-samples in the target domain's real foggy multispectral samples. Indicates the first One sample, This indicates that gradient calculation has stopped. express Norm;
[0067] Constructing a cross-domain multispectral distillation mechanism CdMD, using the domain consistency loss function Multispectral dehazing model for source domain teachers Multispectral dehazing model for students in the target domain The fusion features and dehazing prediction results on real foggy multispectral samples in the target domain are subject to dual consistency constraints, and a recursive teacher update strategy based on exponential moving average is adopted to improve the source domain teacher multispectral dehazing model. The parameters are updated using a time-cumulative integrated update method, which is as follows:
[0068]
[0069] in, This indicates the source domain teacher multispectral dehazing model. Parameters during the unsupervised adaptation training process in the target domain. Represents the target domain student multispectral dehazing model The parameters, This represents the recursive update coefficient.
[0070] Secondly, the present invention discloses an unsupervised infrared-guided visible light image dehazing device based on multispectral domain adaptation, the device comprising a dataset generation module, a multispectral dehazing model MFD, a model training module, and a dehazing module;
[0071] The dataset generation module is used to construct a spatially registered source domain synthetic multispectral dataset and a target domain real multispectral dataset. The source domain multispectral dataset includes a clear visible light image of the source domain, a corresponding infrared image of the source domain, and a synthetic hazy visible light image generated based on the clear visible light image of the source domain. The target domain real multispectral dataset includes a real hazy visible light image of the target domain and its corresponding infrared image.
[0072] The multispectral dehazing model MFD includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder VisEncoder for extracting hazy visible light features, a second encoder IREncoder for extracting infrared features, a region adaptive multispectral fusion module ReMix for performing region adaptive fusion of the hazy visible light features and infrared features, a decoder FusionDecoder for reconstructing the dehazed visible light image based on the fused features, and a dehazed image output terminal.
[0073] The model training module is used to train the multispectral dehazing model using a phased training strategy. In the first phase, supervised pre-training is performed using the source domain multispectral dataset to obtain a source domain teacher multispectral dehazing model. During pre-training, a visibility confidence map is generated based on the gradient difference between hazy visible light features and infrared features. The visibility confidence map is then subjected to segmented weight hardening to obtain three-segment region fusion weights. Differential fusion of visible light features and infrared features is performed according to these fusion weights: visible light features are hard-preserved in high-visibility regions, infrared features are hard-replaced in low-visibility regions, and adaptive weighted fusion is performed in transition regions. In the second phase, the target domain student multispectral dehazing model is initialized with the parameters of the source domain teacher multispectral dehazing model. Real hazy visible light images and corresponding infrared images of the target domain are simultaneously input into both the teacher and student models. Cross-domain multispectral distillation applies a dual consistency constraint to the fusion representation and dehazing prediction results of the student model. Domain adaptation training is completed under unsupervised conditions to obtain a target domain student model adapted to real foggy scenes.
[0074] The dehazing module is used to input the real foggy visible light image to be processed and the corresponding infrared image into the trained target domain student multispectral dehazing model, and output the dehazed visible light image.
[0075] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0076] First, the unsupervised infrared-guided visible light image dehazing method and apparatus based on multispectral domain adaptation of the present invention abandons the dependence of existing dehazing techniques on supervision by real paired clear images. It adopts an unsupervised domain adaptation training strategy from the synthetic domain to the real domain, and transfers the multispectral dehazing capability obtained by the source domain training to the real foggy scene in the target domain through a cross-domain multispectral distillation mechanism. At the same time, it avoids the modal conflict and artifact problems of infrared and visible light in the domain transfer process. Even in the absence of real paired clear image supervision, it can still achieve effective dehazing of real domain images, thereby improving the model's cross-scene generalization ability and practical application value.
[0077] Second, the unsupervised infrared-guided visible light image dehazing method and apparatus based on multispectral domain adaptation of the present invention abandons the processing method of uniform fusion or fixed fusion ratio in existing multimodal dehazing technology, and proposes a regional adaptive multispectral fusion mechanism based on dual-modal gradient difference and segmented weight hardening. According to local visibility, the visible light features and infrared features are fused in a regionally differentiated manner, so that the infrared features dominate the feature construction in areas with severe fog obscuration and missing visible light information. In the subsequent decoding and reconstruction process, its structural information is used to form a reconstruction representation oriented towards visible light recovery, thereby reducing modal conflict, residual fog and structural artifacts from the root, and significantly improving the dehazing integrity and visual naturalness in foggy scenes. Attached Figure Description
[0078] Figure 1 This is a schematic diagram of the overall process of the unsupervised infrared-guided visible light image dehazing method based on multispectral domain adaptation of the present invention.
[0079] Figure 2 This is a schematic diagram of the overall defogging principle designed for this invention;
[0080] Figure 3 This is a schematic diagram of the ReMix regional adaptive multispectral fusion module and its fusion and reconstruction process.
[0081] Figure 4 This is a diagram illustrating the defogging effect. Detailed Implementation
[0082] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0083] See Figure 1 This invention discloses an unsupervised infrared-guided visible light image dehazing method based on multispectral domain adaptation, the method comprising:
[0084] S1. Construct a spatially registered source domain synthetic multispectral dataset and a target domain real multispectral dataset; the source domain multispectral dataset includes a clear visible light image of the source domain, a corresponding infrared image of the source domain, and a synthetic hazy visible light image generated based on the clear visible light image of the source domain; the target domain real multispectral dataset includes a real hazy visible light image of the target domain and its corresponding infrared image.
[0085] S2. Construct a multispectral dehazing model (MFD); wherein the multispectral dehazing model (MFD) includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder VisEncoder for extracting hazy visible light features, a second encoder IREncoder for extracting infrared features, a region adaptive multispectral fusion module ReMix for performing region adaptive fusion of hazy visible light features and infrared features, a decoder FusionDecoder for reconstructing the dehazed visible light image based on the fused features, and a dehazed image output terminal;
[0086] S3. The multispectral dehazing model is trained using a phased training strategy:
[0087] In the first stage, supervised pre-training was performed using the source domain multispectral dataset to obtain the source domain teacher multispectral dehazing model. During the pre-training process, a visibility confidence map was generated based on the gradient difference between the hazy visible light features and infrared features. The visibility confidence map was then subjected to segmented weight hardening to obtain a three-segment region fusion weight. The visible light features and infrared features were then fused differentially according to the fusion weight: visible light features were hard retained in high visibility regions, infrared features were hard replaced in low visibility regions, and adaptive weighted fusion was performed in transition regions.
[0088] In the second stage, the target domain student multispectral dehazing model is initialized with the parameters of the source domain teacher multispectral dehazing model. The real foggy visible light image and the corresponding infrared image of the target domain are simultaneously input into the teacher model and the student model. Cross-domain multispectral distillation is used to impose a dual consistency constraint on the fusion representation and dehazing prediction results of the student model. Domain adaptation training is completed under unsupervised conditions to obtain a target domain student model adapted to real foggy scenes.
[0089] S4. Input the real foggy visible light image to be processed and the corresponding infrared image into the trained target domain student multispectral dehazing model, and output the dehazed visible light image.
[0090] Example
[0091] Figure 2 This is a schematic diagram illustrating the overall dehazing principle. The principles of each step of this invention are explained below with specific examples. The numerical settings such as the dataset partitioning ratio and image resolution are only for the convenience of explaining the principle and are not technical limitations.
[0092] (a) Training dataset
[0093] In this embodiment, a source domain multispectral dataset and a target domain multispectral dataset are constructed: , .in, For a clear visible light image of the source region, For the source domain corresponding infrared image, This is a synthesized hazy visible light image generated from the source domain clear visible light image. For the target domain, a real hazy visible light image. The infrared image corresponding to the real foggy visible light image of the target domain;
[0094] The dataset was randomly divided into training and validation sets in an 8:2 ratio. All images in the dataset have a resolution of 640×512. Visible light images are RGB three-channel images, and infrared images are grayscale images. To ensure the spatial consistency of different modalities during subsequent network training, during the dataset construction phase, visible light and infrared images were paired and registered, cropped, and resized.
[0095] In this embodiment, a source domain synthesized hazy visible light image is obtained. Clear visible light image from the source domain Generated using a fogging algorithm. The fogging algorithm employs an atmospheric scattering model:
[0096]
[0097] in, Represents a clear visible light image. This represents a visible light image after fogging. Represents global atmospheric light. This represents a transmittance map. By randomly setting global atmospheric light and transmittance parameters, synthetic visible light images with fog of different concentrations and spatial distributions are generated to expand the diversity of the source domain training samples.
[0098] In this embodiment, the source domain multispectral dataset is used for supervised pre-training of the source domain teacher multispectral dehazing model in the first stage; the target domain multispectral dataset is used for unsupervised domain adaptation training in the second stage.
[0099] (II) Constructing a multispectral dehazing model
[0100] In this embodiment, a multispectral dehazing model (MFD) is constructed, such as... Figure 2As shown, the multispectral dehazing model (MFD) includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder (VisEncoder), a second encoder (IREncoder), a region-adaptive multispectral fusion module (ReMix), a decoder (FusionDecoder), and a dehazed image output terminal. During forward propagation, the hazy visible light image is input to the first encoder (VisEncoder) to extract multi-scale hazy visible light features, and the infrared image is input to the second encoder (IREncoder) to extract multi-scale infrared features. Subsequently, the region-adaptive multispectral fusion module (ReMix) performs region-by-region fusion of the two modal features at each scale. Finally, the decoder (FusionDecoder) reconstructs the fused features layer by layer to output the dehazed visible light image.
[0101] In this embodiment, the input is a hazy visible light image. The resolution is 640×512×3, and the input infrared image is... Both have a resolution of 640×512×1. The first encoder, VisEncoder, and the second encoder, IREncoder, both employ a five-layer convolutional coding structure. Each layer consists of two consecutive 3×3 convolutional layers, followed by a batch normalization layer and a ReLU activation function. Adjacent layers are compressed using a 3×3 downsampling convolutional layer with a stride of 2. After five layers of coding, five layers of hazy visible light features and five layers of infrared features are obtained, denoted as... to , to The corresponding feature sizes are 640×512×64, 320×256×64, 160×128×128, 80×64×256 and 40×32×256, respectively, thus ensuring that the spatial size of the two modes is consistent at the same level, which facilitates subsequent layer-by-layer fusion.
[0102] like Figure 3 As shown, the forward propagation of the ReMix regional adaptive multispectral fusion module includes three steps: visibility confidence generation, segmented weight hardening, and fusion feature generation. For the first... Layer features, first based on the first Visible light characteristics of layered fog and the Layer infrared features Sobel gradient operators are applied to the gradient differences of each channel, and the gradients are converged along the L1 norm in the channel dimension to generate a single-channel gradient difference map. :
[0103]
[0104] in, This indicates the spatial location in the current scale feature map. The sensitivity control coefficient for difference is set to 0.25. Gradient difference plot. The size is ,in These correspond to 640×512, 320×256, 160×128, 80×64, and 40×32 respectively. Then, With the Visible light characteristics of layered fog The convolutional mapping results are jointly input into the visibility estimation unit to generate a visibility confidence map. :
[0105]
[0106] in, For balance coefficient, The Sigmoid activation function is used. A 3×3 convolutional layer with 1 output channel is used. Subsequently, the visibility confidence map is... Perform segmented weight hardening to generate a fused weight graph. Set a threshold. Upper threshold .when At that time, the location is identified as a heavy fog area, and the weights are integrated. ;when At that time, the location is identified as a high-visibility region, and the weights are merged. ;when When this location is identified as a transition region, the fusion weights are determined linearly. Finally, the weighted graph will be merged. , No. Visible light characteristics of layered fog , No. Layer infrared features Perform region adaptive fusion to obtain the first Layer fusion features :
[0107]
[0108] in, This indicates element-wise multiplication. Through this method, ReMix can adaptively adjust the fusion ratio of visible light and infrared features based on local visibility. The fused feature sizes output by each layer are 640×512×64, 320×256×64, 160×128×128, 80×64×256, and 40×32×256, respectively.
[0109] The FusionDecoder decoder employs a U-shaped layer-by-layer upsampling decoding structure. to Reconstruction is performed from deep to shallow; simultaneously, the corresponding layer's hazy visible light features extracted by the first encoder, VisEncoder, are used as input to the decoder as a local texture supplement branch, and the fused weight map is utilized. The visible light features of the corresponding layer with haze are weighted element-wise before being included in the decoding. The deepest layer's decoded features are represented as follows:
[0110]
[0111] The decoding features of the remaining layers are represented as follows:
[0112]
[0113] in, Indicates the deepest decoded features. Indicates the first Layer decoding features, This represents the upsampling and convolution operations. This represents the feature concatenation operation. Through the aforementioned layer-by-layer decoding method, the reconstruction process uses region-adaptive fusion features as the primary basis for reconstruction, while also utilizing the fusion weight map. The visible light features of the corresponding foggy layer are controlled element-wise, so that the corresponding visible light texture information in the high visibility region is retained and participates in the supplementation, the corresponding visible light texture information in the transition region participates in the supplementation according to the weight, and the corresponding visible light features in the heavy fog region are suppressed. In this way, while maintaining the structural restoration ability of the heavy fog region, the texture integrity and visual naturalness of the defogging result are improved.
[0114] Decode the first layer features Perform output mapping to obtain the dehazed visible light prediction map. , represented as:
[0115]
[0116] in, This indicates the output mapping convolution operation, and the dehazed visible light prediction map. The dimensions are 640×512×3. Through the above-mentioned layer-by-layer upsampling decoding and cross-layer feature aggregation method, the decoder FusionDecoder integrates global structural information from deep features and local texture information from shallow features during the reconstruction process, and finally outputs a clear, detailed, and dehazed visible light image.
[0117] (III) Phased Training
[0118] 1. First Phase Training
[0119] The first stage is source domain-supervised pre-training. This stage utilizes the source domain to synthesize hazy visible light images. Source domain infrared images and source domain clear visible light image Supervised training was performed on the multispectral dehazing model MFD to obtain a source-domain teacher multispectral dehazing model with basic multispectral dehazing capabilities. .
[0120] Synthesize a hazy visible light image from the source region. Source domain infrared images The first encoder, VisEncoder, and the second encoder, IREncoder, are input respectively to extract the hazy visible light features and infrared features. These features are then fused layer by layer by the region adaptive multispectral fusion module, ReMix, and finally output by the decoder, FusionDecoder, as a source domain dehazing prediction map. Source Domain Teacher Multispectral Dehazing Model Dehazing loss function Represented as:
[0121]
[0122] in, , , They represent , , The weighting coefficients; in this embodiment, , , . for The loss function is expressed as:
[0123]
[0124] in, Represents the pixel number of the image. Indicates the total number of pixels in the image; The structural similarity loss function is expressed as:
[0125]
[0126] in, Represents a structural similarity function; The perceptual loss function is expressed as:
[0127]
[0128] in, The first feature extraction operator represents the... Layer feature mapping, This represents the set of feature layers involved in the calculation of perceptual loss.
[0129] In the first training phase, the Adam optimizer was used, with a batch size of 8, an initial learning rate of 0.0001, and 20 training epochs. Cosine annealing decay was employed for the learning rate. PSNR and SSIM were calculated on the validation set, and the model weights corresponding to the highest PSNR were saved as the source domain teacher multispectral dehazing model after the first training phase.
[0130] 2. Second Phase Training
[0131] The second stage is unsupervised adaptation training of the target domain. This involves using the source domain teacher multispectral dehazing model trained in the first stage. The pre-trained parameters are used to initialize the target domain student multispectral dehazing model. During training, real-world foggy visible light images of the target domain are used. and its corresponding infrared image Simultaneously input the source domain teacher multispectral dehazing model and target domain student multispectral dehazing model The dehazing results of the source domain teacher model were obtained respectively. Dehazing results of student model in target domain , No. Layered source domain teacher integration characteristics and the Layer target domain student integration characteristics .
[0132]
[0133]
[0134] Using joint loss function Multispectral dehazing model for target domain students Unsupervised domain adaptation training is performed, and the joint loss function is expressed as:
[0135]
[0136] in, and Represent the domain consistency loss function respectively Image-text joint semantic constraint loss The weighting coefficients. Take... , .
[0137] Domain consistency loss function Consistency loss due to fusion features Consistency loss with defogging prediction Composition, represented as:
[0138]
[0139] in, and They represent and The weighting coefficient. In the specific implementation parameters, take... , .
[0140] Fusion feature consistency loss Represented as:
[0141] ;
[0142] Defogging prediction consistency loss Represented as:
[0143] ;
[0144] in, This represents the number of mini-batch samples of real foggy multispectral samples in the target domain, taken as... ; Indicates the first One sample; This indicates that gradient calculation has been stopped; express Norm. During training, the consistency loss is calculated for the fused features at five scales, and the average value is taken as the norm. The calculation result is as follows:
[0145]
[0146] By fusing feature consistency loss The target domain student multispectral dehazing model approximates the output of the source domain teacher multispectral dehazing model at the intermediate fusion representation level; consistency loss is predicted through dehazing. The constrained target domain student multispectral dehazing model approximates the output of the source domain teacher multispectral dehazing model at the final dehazing prediction level.
[0147] In the unsupervised domain adaptation training process of the target domain, a cross-domain multispectral distillation mechanism CdMD is constructed, using the domain consistency loss function. Multispectral dehazing model for source domain teachers Multispectral dehazing model for students in the target domain Consistency constraints are imposed on the fusion features and dehazing prediction results on real foggy multispectral samples in the target domain. A recursive teacher update strategy based on exponential moving average (EMA) is adopted to update the source domain teacher multispectral dehazing model. The parameters are updated cumulatively over time. This recursive teacher update strategy is expressed as:
[0148]
[0149] in, This indicates the source domain teacher multispectral dehazing model. Parameters during the unsupervised adaptation training process in the target domain. Represents the target domain student multispectral dehazing model The parameters, Represents the recursive update coefficients, taking... In this way, the changes in the teacher model parameters are smoother, providing a stable distillation reference for the student model.
[0150] Image-text joint semantic constraint loss The dehazed images used to guide the target domain student multispectral dehazing model output are semantically close to the "clear visible light image" distribution and far from the "hazy visible light image" distribution. During training, the target domain student dehazing prediction image is used... The input image encoder extracts image features, and the sets of clear semantic cue words and hazy semantic cue words are respectively input into the text encoder to extract text features; the average similarity between the student's dehazed image features and the clear semantic text features is denoted as . The average similarity between the features of the foggy semantic text and the features of the foggy semantic text is denoted as Then the semantic constraint loss of the image and text joint Represented as:
[0151]
[0152] in, The temperature coefficient is selected from a set of specific implementation parameters. The set of clear semantic cue words can be taken as "a clear image", and the set of hazy semantic cue words can be taken as "a hazy image".
[0153] During the second phase of training, the source domain teacher multispectral dehazing model was frozen. The gradient backpropagation path is only applicable to the target domain student multispectral dehazing model. Parameters are updated; the optimizer is Adam, the batch size is 8, and the initial learning rate is... The training rounds are 20, and the weight decay coefficient is... The learning rate is decayed using cosine annealing. After each training epoch, PSNR and SSIM are calculated on the validation set, and the target domain student multispectral dehazing model parameters corresponding to the highest PSNR are saved as the final domain adaptation model.
[0154] Through the above two-stage training, the first stage obtains the source domain teacher multispectral dehazing model; the second stage utilizes the cross-domain multispectral distillation mechanism CdMD to transfer the source domain multimodal dehazing capability to the target domain real foggy day multispectral data, thereby achieving effective dehazing of real foggy day scenes in the absence of real paired clear supervision.
[0155] The visible light and infrared images of a real foggy scene to be defogged are input into the trained target domain student multispectral defogging model. This generates a dehazed visible light image. Figure 4 This is a diagram illustrating the defogging effect in two scenarios.
[0156] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0157] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An unsupervised infrared-guided visible image defogging method based on multispectral domain adaptation, characterized in that, The method includes: S1. Construct a spatially registered source domain synthetic multispectral dataset and a target domain real multispectral dataset; the source domain multispectral dataset includes a clear visible light image of the source domain, a corresponding infrared image of the source domain, and a synthetic hazy visible light image generated based on the clear visible light image of the source domain; the target domain real multispectral dataset includes a real hazy visible light image of the target domain and its corresponding infrared image. S2. Construct a multispectral dehazing model (MFD); wherein the multispectral dehazing model (MFD) includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder VisEncoder for extracting hazy visible light features, a second encoder IREncoder for extracting infrared features, a region adaptive multispectral fusion module ReMix for performing region adaptive fusion of the hazy visible light features and infrared features, a decoder FusionDecoder for reconstructing the dehazed visible light image based on the fused features, and a dehazed image output terminal; S3. The multispectral dehazing model is trained using a phased training strategy: In the first stage, supervised pre-training is performed using the source domain multispectral dataset to obtain the source domain teacher multispectral dehazing model. During the pre-training process, a visibility confidence map is generated based on the gradient difference between hazy visible light features and infrared features. The visibility confidence map is then subjected to segmented weight hardening to obtain a three-segment region fusion weight. The visible light features and infrared features are then fused differentially according to the fusion weight: visible light features are hard-retained in high visibility regions, infrared features are hard-replaced in low visibility regions, and adaptive weighted fusion is performed in transition regions. In the second stage, the target domain student multispectral dehazing model is initialized with the parameters of the source domain teacher multispectral dehazing model. The real foggy visible light image and the corresponding infrared image of the target domain are simultaneously input into the teacher model and the student model. Cross-domain multispectral distillation is used to impose a dual consistency constraint on the fusion representation and dehazing prediction results of the student model. Domain adaptation training is completed under unsupervised conditions to obtain a target domain student model adapted to real foggy scenes. S4. Input the real foggy visible light image to be processed and the corresponding infrared image into the trained target domain student multispectral dehazing model, and output the dehazed visible light image.
2. The method of claim 1, wherein the method is based on multi-spectral domain adaptation for unsupervised infrared guided visible dehazing. The ReMix regional adaptive multispectral fusion module includes a confidence generation unit, a segmented hardening unit, and a feature fusion unit. The confidence generation unit receives the first encoder VisEncoder's output of the first... Visible light characteristics of layered fog and the output of the second encoder IREncoder Layer infrared features The gradient difference between the visible light features and infrared features of foggy conditions is calculated using the following formula, and a visibility confidence map is generated. , represented as: In the formula, Represents the spatial gradient operator, express Norm, This represents the difference sensitivity control coefficient. Indicates the first Spatial location in the layer feature map; The segmented hardening unit is based on the visibility confidence map. and the Visible light characteristics of layered fog Generate visibility confidence map , represented as: wherein denotes an activation function, denotes a convolutional mapping, is a balancing coefficient; based on the visibility confidence map segmentation weight hardening is performed, generating a fusion weight map denoted as: In the formula, and These represent the low threshold and the high threshold, respectively. Finally, the weighted graph will be merged. Expanded to the first in the channel dimension corresponding to the layer feature map One channel; The feature fusion unit receives the fusion weight map after dimensionality expansion. , No. Visible light characteristics of layered fog and the Layer infrared features Perform segmented weight hardening processing, in The region retains visible light characteristics, in Perform infrared feature hard replacement in the region, in The transition region is weighted and fused with visible light and infrared features to obtain the first... Layer fusion features , is represented as: wherein represents an element-wise multiplication.
3. The unsupervised infrared-guided visible image defogging method based on multispectral domain adaptation according to claim 2, characterized in that, The decoder, FusionDecoder, is used to process the fusion features output by the region-adaptive multispectral fusion module ReMix. Reconstruction is performed to obtain the source region dehazing prediction map. The decoder FusionDecoder includes a fused feature input, a supplementary branch input, a cross-layer feature aggregation and decoding module, and an output mapping and dehazing prediction module. The fusion feature input terminal is used to receive the multi-layer fusion features output by the ReMix regional adaptive multispectral fusion module. ;in, Indicates the first Layer fusion features, dimension 1 , The layer index is represented, with a range of [1,5]. The supplementary branch input is used to receive multi-layered hazy visible light features extracted by the first encoder, VisEncoder. This is used as a supplementary branch for the local texture of the corresponding layer, and the fusion weight map is utilized. Position-by-position suppression or retention control is applied to the visible light characteristics of the corresponding layer containing fog. The cross-layer feature aggregation and decoding module adopts a five-layer upsampling decoding method from deep to shallow, using the multi-layer fused features as the main basis for reconstruction, and combining the decoded features of the previous layer with the fused weight map. After element-wise weighted control, the corresponding layer's hazy visible light features are subjected to cross-layer feature aggregation and reconstruction to obtain the first layer. Layer decoding features , is represented as: in, Indicates the deepest decoded features. Indicates the first Layer decoding features, This represents upsampling and convolution operations; This indicates a feature concatenation operation. This represents element-wise multiplication; The output mapping is performed on the first layer decoded feature to obtain a defogging prediction map , which is expressed as: wherein, represents an output mapping convolution operation.
4. The method of claim 1, wherein the method is based on multi-spectral domain adaptation for unsupervised infrared guided visible dehazing. In step S3, in the first stage, the source domain dehazing loss function is adopted. Multispectral dehazing model for source domain teachers Supervised training was conducted to learn how to synthesize hazy visible light images from the source domain. Source domain infrared images Clear visible light image of the source domain The cross-modal dehazing mapping relationship is established, and edge details and overall structural information are preserved during the dehazing process; the source domain dehazing loss function is described. Represented as: in, , , They represent , , Weighting coefficients; for The loss function is expressed as: wherein, represents the pixel number of the image, represents the total pixel number of the image; is a structural similarity loss function, which is represented as: wherein, denotes a structural similarity function; is a perceptual loss function, denoted as: in, The first feature extraction operator represents the... Layer feature mapping, This represents the set of feature layers involved in the calculation of perceptual loss.
5. The unsupervised infrared-guided visible light image dehazing method based on multispectral domain adaptation according to claim 1, characterized in that, In step S3, the second stage of training is as follows: The trained source domain teacher multi-spectral defogging model is adopted The parameter initialization target domain student multi-spectral defogging model ; Real foggy visible light image of the target domain and its corresponding infrared image Simultaneously input the source domain teacher multispectral dehazing model and the target domain student multispectral dehazing model The source domain teacher defogging prediction map was obtained respectively. Target Domain Student Defogging Prediction Map , No. Layered source domain teacher integration characteristics and the Layer target domain student integration characteristics : ; ; Using joint loss function Multispectral dehazing model for target domain students Unsupervised domain adaptation training is performed, and the joint loss function is... Represented as: wherein, and respectively represent the weight coefficients of the domain consistency loss function and the joint semantic constraint loss of image and text. respectively. The domain consistency loss function By fusion feature consistency loss And defogging prediction consistency loss The domain consistency loss function is constituted and expressed as: wherein and respectively represent and weighting coefficients; The fusion feature consistency loss is represented as: ; the defogging prediction consistency loss is represented as: ; in, This represents the number of mini-samples in the target domain's real foggy multispectral samples. Indicates the first One sample, This indicates that gradient calculation has stopped. express Norm; Constructing a cross-domain multispectral distillation mechanism CdMD, using the domain consistency loss function Multispectral dehazing model for source domain teachers Multispectral dehazing model for students in the target domain The fusion features and dehazing prediction results on real foggy multispectral samples in the target domain are subject to dual consistency constraints, and a recursive teacher update strategy based on exponential moving average is adopted to improve the source domain teacher multispectral dehazing model. The parameters are updated using a time-cumulative integrated update method, which is as follows: in, This indicates the source domain teacher multispectral dehazing model. Parameters during the unsupervised adaptation training process in the target domain. Represents the target domain student multispectral dehazing model The parameters, This represents the recursive update coefficient.
6. An unsupervised infrared guided visible image defogging device based on multispectral domain adaptation, characterized in that, The device includes a dataset generation module, a multispectral dehazing model (MFD), a model training module, and a dehazing module. The dataset generation module is used to construct a spatially registered source domain synthetic multispectral dataset and a target domain real multispectral dataset. The source domain multispectral dataset includes a clear visible light image of the source domain, a corresponding infrared image of the source domain, and a synthetic hazy visible light image generated based on the clear visible light image of the source domain. The target domain real multispectral dataset includes a real hazy visible light image of the target domain and its corresponding infrared image. The multispectral dehazing model MFD includes an infrared image input terminal, a hazy visible light image input terminal, a first encoder VisEncoder for extracting hazy visible light features, a second encoder IREncoder for extracting infrared features, a region adaptive multispectral fusion module ReMix for performing region adaptive fusion of the hazy visible light features and infrared features, a decoder FusionDecoder for reconstructing the dehazed visible light image based on the fused features, and a dehazed image output terminal. The model training module is used to train the multispectral dehazing model using a phased training strategy. In the first phase, supervised pre-training is performed using the source domain multispectral dataset to obtain a source domain teacher multispectral dehazing model. During pre-training, a visibility confidence map is generated based on the gradient difference between hazy visible light features and infrared features. The visibility confidence map is then subjected to segmented weight hardening to obtain three-segment region fusion weights. Differential fusion of visible light features and infrared features is performed according to these fusion weights: visible light features are hard-preserved in high-visibility regions, infrared features are hard-replaced in low-visibility regions, and adaptive weighted fusion is performed in transition regions. In the second phase, the target domain student multispectral dehazing model is initialized with the parameters of the source domain teacher multispectral dehazing model. Real hazy visible light images and corresponding infrared images of the target domain are simultaneously input into both the teacher and student models. Cross-domain multispectral distillation applies a dual consistency constraint to the fusion representation and dehazing prediction results of the student model. Domain adaptation training is completed under unsupervised conditions to obtain a target domain student model adapted to real foggy scenes. The dehazing module is used to input the real foggy visible light image to be processed and the corresponding infrared image into the trained target domain student multispectral dehazing model, and output the dehazed visible light image.