An Adaptive Image Restoration Method Based on Joint Awareness of Degradation Type and Degree

By constructing a degradation-aware composite weather image dataset and a dynamic convolutional Transformer module, adaptive convolutional kernels and attention weights are generated, solving the problem that existing technologies fail to fully integrate the differences in degradation types and degrees, and achieving accurate adaptive restoration of composite weather degradation images.

CN120953136BActive Publication Date: 2026-01-30SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511484182.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-30
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing image restoration methods fail to fully integrate the differences in degradation type, degree, and image content information when faced with complex weather degradation, making it difficult to achieve accurate and efficient restoration. Furthermore, there is a lack of high-quality datasets specifically designed for real-world complex weather degradation.

Method used

A degradation-aware composite weather image dataset DSD is constructed. A degradation-aware scene descriptor generator generates multi-dimensional degradation-aware scene descriptors, and a degradation-aware dynamic convolution Transformer module is used to adaptively generate convolution kernels and attention weights to achieve adaptive image restoration.

Benefits of technology

It achieves accurate perception and adaptive restoration of images degraded by complex weather conditions, improves the ability to restore image details and resist interference, and significantly enhances the restoration performance under complex weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953136B_ABST
    Figure CN120953136B_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive image restoration method based on joint perception of degradation type and degree, relating to the fields of computer vision and artificial intelligence. First, a hierarchical parsing mechanism for degradation type-degree fusion is constructed, achieving multi-dimensional and accurate perception of complex weather by jointly generating type and degree descriptors. Second, a degradation adaptive restoration network based on dynamic convolution is designed, autonomously adjusting the restoration strategy according to the degradation descriptor. Its core is the integration of a degradation-aware dynamic convolution Transformer module, employing dynamic convolution kernels and a cross-attention mechanism to achieve directional control of the image restoration process by the degradation descriptor. Finally, this invention constructs a complex weather dataset (DSD) based on the RAISE database. This invention provides key technical support for fields relying on visual perception, such as autonomous driving and security monitoring, by improving image restoration capabilities under complex weather conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an adaptive image restoration method based on joint perception of degradation type and degree, belonging to the field of computer vision and multimodal learning. Background Technology

[0002] In complex weather conditions such as rain, fog, snow, and low light, image and video quality deteriorates significantly due to environmental interference. For example, fog causes color distortion, blurring, and reduced contrast, making object outlines difficult to discern; rain and snow create dynamic noise and stripe occlusion, compromising the integrity of targets; and low light conditions result in localized underexposure, increased noise, and loss of critical information. These degradation phenomena not only affect human visual perception but also pose a serious threat to intelligent systems such as autonomous driving and security monitoring that rely on high-precision visual information. Therefore, image quality restoration technologies for complex weather conditions are crucial.

[0003] Current image restoration methods are mainly divided into two categories: models focused on specific weather degradations (such as rain removal, snow removal, low-light enhancement, and fog removal) and integrated restoration models designed to handle multiple degradations. Thanks to breakthroughs in deep learning technologies such as Transformer and Stable Diffusion, restoration methods focused on specific weather degradations (such as Transformer-based FE-DARFormer and Stable Diffusion-based Diff-Dehazer) have demonstrated superior performance in single degradation scenarios. These methods typically focus on a single weather degradation, but real-world scenarios often face complex degradation problems such as "rain-fog" or "rain-snow," making integrated restoration methods more valuable for practical applications.

[0004] Existing integrated restoration methods mainly include: multi-degradation restoration methods based on a unified architecture and multi-degradation restoration methods based on cue learning. Multi-degradation restoration methods based on a unified architecture (such as AirNet and WeatherDiff) handle multiple degradations through a single network structure, exhibiting good generalization ability and efficiency. However, they generally rely on fixed-parameter models to handle all degradation types, significantly limiting their adaptability and effectiveness when facing complex and variable composite degradation weather. Multi-degradation restoration methods based on cue learning (such as PromptIR and CoCoOp) enhance the model's flexibility and control by introducing cue learning. However, existing methods often fail to fully integrate multi-dimensional information related to degradation type, degree differences, and image content, making it difficult to achieve accurate and efficient restoration when processing complex weather-degraded images.

[0005] In real-world scenarios, complex weather conditions often lead to the overlapping of multiple degradation factors (such as rain, fog, and low illumination), resulting in images exhibiting complex types and varying degrees of quality degradation, such as "light rain-moderate fog" or "low illumination-moderate snow." Effective restoration of these images places higher demands on the model's degradation perception and adaptation capabilities. However, current methods often fail to fully integrate information related to multiple types and degrees of degradation and image content, making it difficult to achieve accurate and efficient restoration when processing complex weather-degraded images. Furthermore, there is a lack of high-quality datasets specifically designed for real-world complex weather degradation. Existing representative datasets, such as CDD, while including various degradation types, lack a quantitative evaluation system for the degree of degradation, making it difficult to meet the needs of joint modeling and performance evaluation of multiple degradations. Summary of the Invention

[0006] To address the issues of parameter fixation and insufficient utilization of degradation information in image restoration, this invention proposes an adaptive image restoration method, DegRestorNet, based on joint perception of degradation type and degree. The aim is to mitigate the impact of fixed parameters and insufficient information fusion during training on the model's restoration performance under complex weather conditions.

[0007] The technical solution of the present invention includes the following steps:

[0008] (1) By establishing the Degradation Sensing Composite Weather Image Dataset (DSD), the present invention systematically simulates typical weather degradation and its combinations such as rain, snow, fog, and low illumination, generating a total of 11 degradation types and 5 degradation degree levels, realizing the simulation and coverage of multiple composite weather degradation scenarios;

[0009] (2) By using the Degradation-Aware SceneDescriptor Generator (DASDG), type descriptors and degree descriptors are generated respectively. After further fusion, a degradation-aware scene descriptor containing multi-dimensional information is generated, and a hierarchical parsing mechanism for degradation features is constructed to provide degradation priors for image restoration.

[0010] (3) Through the Degradation-Aware Dynamic Convolutional Transformer Module (DADCT), the Degradation-Aware Scene Descriptor Guided Module generates dynamic convolutional kernels and attention weights. For each degradation type and degree, the network can adaptively generate corresponding weight parameters and use them in the subsequent image restoration process.

[0011] (4) By using Dynamic Convolution-based Degradation-Adaptive Restoration Network (DARN), the restoration strategy is adaptively adjusted with the help of dynamic convolution modules, enabling the model to dynamically adapt to different types and degrees of quality degradation such as rain, fog, snow, and low light, and achieve adaptive restoration of complex weather-degraded images.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] 1. This invention introduces a scene descriptor generator for degradation perception, constructs a hierarchical parsing mechanism that fuses degradation type and degree, and generates a degradation-perceived scene descriptor, thereby achieving multi-dimensional and accurate perception of complex weather.

[0014] 2. This invention introduces a degradation-aware dynamic convolution Transformer module to adaptively generate convolution kernels and attention weights that match the type and degree of degradation, thereby improving the model's ability to restore image details and resist interference in complex degradation scenarios.

[0015] 3. This invention introduces a degradation adaptive restoration network based on dynamic convolution. By integrating a degradation-aware dynamic convolution Transformer module, it can autonomously adjust the image restoration strategy according to the input degradation descriptor, thereby effectively achieving adaptive restoration in complex weather degradation scenarios. Attached Figure Description

[0016] Figure 1: A block diagram illustrating the principle of the adaptive image restoration method based on joint perception of degradation type and degree according to the present invention;

[0017] Figure 2: A schematic diagram of the generation of the dynamic adaptive image restoration dataset according to the present invention;

[0018] Figure 3: Block diagram of the Degradation-Aware Dynamic Convolution Transformer module of the present invention;

[0019] Figure 4 This invention provides a comparison of the image restoration results under "low illumination-fog-rain" test conditions.

[0020] Figure 5 This invention provides a comparison of the image restoration results for a "low illumination-rain" test. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0022] like Figure 1 As shown, an adaptive image restoration method based on joint perception of degradation type and degree includes the following steps:

[0023] (1) A degradation-sensing composite weather image dataset (DSD) was constructed based on a physical degradation model. For each degradation image in the dataset... A structured metadata annotation file containing degradation types and quantization parameters was established, and an image-text pair dataset was constructed.

[0024] (2) The image-text dataset is used to train a scene descriptor generator for degradation perception. The training aims to maximize the semantic similarity between image features and text embedding vectors, thereby achieving cross-modal alignment between visual features and text semantics.

[0025] (3) For a given input degraded image Degradation-aware scene descriptors generated by a degradation-aware scene descriptor generator (Text information) guides the degradation-aware dynamic convolution Transformer module to dynamically generate convolution kernels and attention weights, which are then used in the subsequent image restoration process.

[0026] (4) For a given input degraded image Generate the corresponding degradation-aware scene descriptor Then, the images are input together into a degradation adaptive restoration network based on dynamic convolution for image restoration.

[0027] The detailed steps are as follows:

[0028] Step (1): A Degradation-Sensitive Composite Weather Image Dataset (DSD) was constructed based on a physical degradation model.

[0029] Figure 2 shows a schematic diagram of the dataset generation. 1383 high-resolution images were selected from the RAISE database. Mathematical modeling was performed on four typical degradation types: "rain," "snow," "low light," and "haze," simulating the coupling of multiple degradation factors in real-world scenes. This generated a dataset containing 11 composite degradation types. These 11 degradation types are: "rain," "snow," "low light," "haze," "low-rain," "low-snow," "low-haze," "haze-rain," "haze-snow," "low-haze-rain," and "low-haze-snow." During the dataset construction process, the low, medium, and high levels of the four typical degradation types were systematically quantified into five degradation levels: "clear," "0," "low," "middle," and "high." The final dataset consists of 15,213 images with a resolution of 1080×720 (13,013 images in the training set and 2,200 images in the test set).

[0030] Specifically, during the construction of the DSD dataset, a structured metadata annotation file containing degradation type and quantization parameters was established for each degraded image, accurately recording the degradation type and degree (e.g., "light fog - moderate rain"), forming a standardized description system for multi-dimensional degradation information. Therefore, when training the degradation-aware scene descriptor generator, this invention uses the degradation type and degree text from the aforementioned structured metadata as auxiliary information to guide cross-modal alignment of visual features and textual semantics. After training, degraded images can be directly input into the generator to automatically generate corresponding degradation-aware scene descriptors, which serve as input prompts for subsequent degradation restoration networks, achieving adaptive image restoration without manual annotation.

[0031] Step (2): Train a scene descriptor generator for degradation perception. By aligning visual features with text semantics across modalities, the visual degradation features of the input image are mapped to the corresponding text description, i.e., the scene descriptor.

[0032] This invention proposes a scene descriptor generator for degradation perception, comprising two core modules: a degradation type identification module and a degradation severity evaluation module, which generate type descriptors and severity descriptors, respectively. The type descriptor describes different degradation modes such as rain, fog, snow, and low light, while the severity descriptor characterizes the severity of each degradation phenomenon. The two are fused to generate a degradation-perceived scene descriptor.

[0033] Degradation Type Recognition Module: To achieve coarse-grained classification of complex weather degradation images, a cross-modal feature space capable of representing semantic relationships needs to be constructed. Therefore, this invention employs a cross-modal feature construction method combining text semantic encoding and visual feature extraction. In text semantic encoding, a GloVe pre-trained word vector model is used to extract semantic features of five basic degradation types: "clear," "low light," "rain," "snow," and "haze." These features are then mapped to a 12-dimensional text vector space using a multilayer perceptron (MLP), explicitly modeling the semantic combination relationships of these 12 modes: "clear," "rain," "snow," "low light," "haze," "low-rain," "low-snow," "low-haze," "haze-rain," "haze-snow," "low-haze-rain," and "low-haze-snow." In visual feature extraction, ResNet-18 is used to capture the global structural features of the image, and then a mapping network consisting of convolutional layers, dropout layers, and fully connected layers is used to generate visual feature vectors. By combining basic degradation types, diverse degradation patterns in real-world scenarios are covered, avoiding the limitations of modeling a single degradation type. Text vectors and visual feature vectors share a 12-dimensional feature space, providing a unified benchmark for subsequent semantic matching.

[0034] To achieve precise semantic alignment between degraded images and text descriptions, this paper references classic cross-modal retrieval methods and characterizes semantic relevance by measuring the geometric distance between visual feature vectors and text vectors. The cosine similarity between the visual feature vector and the text vector is calculated. After normalization, the text embedding corresponding to the highest probability value is selected as the matching result, which can be expressed as follows:

[0035] (1)

[0036] in and These are text vectors and visual feature vectors. It is a temperature factor. It is the number of text vectors. For cosine similarity, This represents the similarity score.

[0037] Degradation Level Assessment Module: To achieve fine-grained quantification of degradation levels, independent intensity assessment branches are constructed for each single degradation type, such as low light, fog, rain, and snow, employing a cross-modal feature modeling framework similar to that used in the degradation type identification module. In text semantic encoding, five degree description texts ("clear", "0", "low", "middle", and "high") are used as input. A 5-dimensional text vector space is generated using a GloVe pre-trained word vector model and an MLP, explicitly representing the degree semantic features of different degradation types (such as low light intensity and fog concentration level). The visual feature extraction process shares the same ResNet-18 network as the degradation type identification module. The visual feature vectors output by the mapping network and the text vectors are semantically aligned using cosine similarity. Normalization enables the classification of degradation levels.

[0038] like Figure 1 As shown, during inference, a degraded image is input, and the scene descriptor generator for degradation perception generates semantic information output by the degradation type recognition module. Degradation information output by the degradation assessment module After channel splicing, degradation-aware scene cues are generated by MLP. .

[0039] Step (3): The Degradation-Aware Dynamic Convolution Transformer module dynamically generates convolution kernels and attention weights.

[0040] The degradation-aware dynamic convolutional Transformer module achieves adaptive restoration of diverse degradation scenarios by transforming degradation descriptors into dynamic network parameters. The structure is as follows: Figure 3 As shown, it consists of two modules: a dynamic convolutional kernel hyperparameter network and a degradation feature Transformer module. The dynamic convolutional kernel hyperparameter network takes degradation-aware scene descriptors as input and constructs a mapping from semantic representation to network parameters; the degradation feature Transformer module accepts parameters generated by the hyperparameter network and drives the dynamic adjustment of convolutional kernels and attention mechanisms.

[0041] Specifically, the dynamic convolutional kernel hyperparameter network uses three fully connected layers and a Gaussian error linear unit activation function to process the degradation-aware scene descriptor. The formula for nonlinear feature extraction is as follows:

[0042] (2)

[0043] in Let Gaussian error be the activation function of the linear unit. Degradation-aware scene descriptor To make the parameter generation more stable, the final hidden state... application Functions to ensure the generation of hyperparameters It is positive and normalized, suitable for multi-core fusion, and the formula is as follows:

[0044] (3)

[0045] in Belongs to the The weights of each convolution kernel, and the kernel coefficients satisfy... Hyperparameters The dynamic convolution weights of the Transformer module for degenerate features are calculated by weighted summation. The formula is as follows:

[0046] (4)

[0047] The weight matrix , Include A predefined set of convolutional kernels, each set containing These parameters collectively constitute all the convolutional kernel parameters required for the cross-attention module, self-attention module, and feedforward network convolutional layers in DADCT. The generated dynamic convolutional weights... It is used to replace the fixed convolutional kernels in these modules, enabling the network parameters to be adaptively adjusted according to the degradation features. For example, when the input descriptor represents "moderate rain-light fog", the module will generate corresponding convolutional kernels and attention weights to achieve customized restoration of specific degradation patterns, effectively improving the model's adaptability and restoration accuracy in complex scenes.

[0048] Step (4): Use a degradation adaptive restoration network based on dynamic convolution, with degradation descriptors guiding the restoration of the degradation image.

[0049] The restoration network is a typical encoder-decoder structure network to ensure efficient modeling of both the global structure and local details of the image. The SDTB module is used as the backbone unit of the encoder, and the previously generated degradation-aware scene descriptor serves as the query input for the cross-attention mechanism. By constructing a "degradation descriptor-image feature" cross-attention mechanism, deep fusion of degradation prior semantic information and image visual features is achieved, ensuring that the features extracted by the encoder contain explicit degradation semantic information. The decoder, based on the semantically enhanced features output by the encoder, utilizes a degradation-aware dynamic convolutional Transformer module to adaptively adjust the restoration strategy. This module takes the degradation-aware scene descriptor as conditional input and dynamically generates convolutional kernel parameters and attention weight matrices that match the current degradation mode through a dynamic convolutional kernel hyperparameter network. This drives the network to perform customized image restoration for specific degradation types and degrees, achieving parameter adaptive optimization guided by degradation scene awareness.

[0050] Overall loss during model training The definition is as follows:

[0051] (5)

[0052] in , and These represent L1 loss, MS-SSIM loss, and composite degradation recovery loss, respectively. , , For hyperparameters, , , , These represent positive samples, anchor points, input negative samples, and other negative samples, respectively. In the design of the contrastive loss, the restored image serves as the anchor point, the degraded input image serves as the input negative sample, the clear image serves as the positive sample, and other negative samples are introduced to enhance the discriminative ability. and Minimize the distance between the anchor point and the positive sample; Minimize the distance between the anchor point and positive samples, and maximize the distance between the anchor point and all negative samples. This improves the model's ability to perceive degradation in the representation space.

[0053] To verify the effectiveness of the method of this invention, tests and comparisons were conducted on the DSD dataset. Two multi-degradation restoration methods based on a unified architecture and five multi-degradation restoration methods based on cue learning were selected as comparison methods, specifically:

[0054] Method 1: NAFNet proposed by Chen et al., see reference "Chen L, Chu X, Zhang X, et al. Simple baselines for image restoration / / European conference on computervision. Tel Aviv, 2022: 17-33."

[0055] Method 2: OKNet proposed by Cui et al., see reference "Cui Y, Ren W, Knoll A. Omni-kernel network for image restoration / / Proceedings of the AAAI conference on artificial intelligence. Vancouver, 2024, 38(2): 1426-1434."

[0056] Method 3: PromptIR proposed by Potlapalli et al., see reference "Potlapalli V, Zamir SW, Khan SH, et al. Promptir: Prompting for all-in-one image restoration. Advances in Neural Information Processing Systems, 2023, 36: 71275-71293."

[0057] Method 4: WGWSNet proposed by Zhu et al., reference "Zhu Y, Wang T, Fu

[0058] Method 5: RestorNet proposed by Wang et al., see reference "Wang X, Chen H, Gou H, He J, Wang Z, He X, Qing L, Sheriff RE. RestorNet: An efficient network for multiple degradation image restoration. Knowledge-Based Systems, 2023, 282:111116."

[0059] Method 6: DEMNet proposed by Yang et al., see reference "Yang Y, Wang X, Lin X, Chen H. DEMNet: A degradation difference enabled multi-stage network for multiple degradation image restoration. Knowledge-Based Systems, 2025, 318: 113426."

[0060] Method 7: OneRestore proposed by Guo et al., see reference "Guo Y, Gao Y, Lu Y, et al. Onerestore: A universal restoration framework for composite degradation / / European Conference on Computer Vision. Milan, 2024: 255-272."

[0061] As shown in Table 1, this invention uses Peak Signal-to-Noise Ratio (SSIM), Structural Similarity Index (PSNR), and Floating-Point Operations (FLOPs) as evaluation metrics. All comparison methods were retrained on the DSD training set of this invention. Compared with the other seven methods, this invention has a significant advantage in performance on the DSD dataset.

[0062] Table 1 shows the comparison results with other methods on the DSD dataset, with the best results shown in bold.

[0063]

[0064] Figure 4 and Figure 5This paper showcases the performance of different image restoration methods in handling combined "low-light-fog-rain" and "low-light-rain" scenes. Results show that PromptIR and OKNet retain significant rain streaks in both combined scenes, while DEMNet introduces some noise in the "low-light-fog-rain" scene. Compared to existing image restoration methods, the proposed DegRestorNet algorithm demonstrates state-of-the-art performance, primarily due to the design of our degradation-aware scene descriptor generator and degradation-aware dynamic convolutional Transformer module. Both simultaneously address multiple degradation factors and robustly reconstruct scene details.

[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive image restoration method based on joint perception of degradation type and degree, characterized in that Comprise the following steps: (1) By establishing a degradation-aware composite weather image dataset DSD, simulating 11 degradation types and 5 degradation levels of weather degradation, realizing the simulation and coverage of various composite weather degradation scenarios, wherein the 11 degradation types include: "rain", "fog", "snow", "low illumination", "low illumination-rain", "low illumination-snow", "low illumination-fog", "fog-rain", "fog-snow", "low illumination-fog-rain", "low illumination-fog-snow", and the 5 degradation levels include: "clean", "no related degradation", "low degradation level", "medium degradation level", "high degradation level"; (2) Through the Degradation-Aware Scene Descriptor Generator (DASDG), type descriptors and degree descriptors are extracted respectively, and a hierarchical analysis mechanism of degradation features is used to fuse and generate degradation-aware scene descriptors, providing degradation prior for image restoration; (3) Through the Degradation-Aware Dynamic Convolutional Transformer Module (DADCT), dynamic convolution kernels and attention weights are generated guided by degradation descriptors, realizing the directional regulation of degradation descriptors on the image restoration process, and applying it to the subsequent restoration task, the Degradation-Aware Dynamic Convolutional Transformer Module is composed of a dynamic convolution kernel hyperparameter network and a degradation feature Transformer module, the dynamic convolution kernel hyperparameter network performs nonlinear feature extraction on the degradation-aware scene descriptors through 3 fully connected layers and Gaussian error linear unit activation functions, and its calculation process is as follows: (1) where, is a Gaussian error linear unit activation function, is a degenerate perceptual scene descriptor To make the parameter generation more stable, the final hidden state is applied a function to ensure that the generated hyperparameters are positive and normalized, computed as follows: (2) wherein belongs to the weight of the th convolution kernel, the convolution kernel coefficient satisfies , the hyperparameter The dynamic convolution weight of the degenerate feature Transformer module is calculated by weighted summation , and the formula is as follows: (3) wherein the weight matrix , comprises a predefined kernel group, each kernel group comprising parameters, the dynamic convolution weight is used to replace the fixed convolution kernel in these modules, realizing adaptive adjustment of network parameters with degraded features; (4) Through the Dynamic Convolution-based Degradation-Adaptive Restoration Network (DARN), the adaptive restoration of complex weather degradation images is realized, specifically including: The network adopts a typical encoder-decoder structure to ensure efficient modeling of global structure and local details of the image. The Scene Descriptor-guided Transformer Block (SDTB) is used as the backbone unit of the encoder. The previously generated degradation-aware scene descriptor is used as the query input of the cross-attention mechanism. The cross-attention mechanism between the degradation descriptor and the image features is constructed to realize the deep fusion of the degradation prior semantic information and the image visual features, so that the features extracted by the encoder contain clear degradation semantic information. Based on the semantic enhanced features output by the encoder, the decoder uses the degradation-aware dynamic convolution Transformer module to complete the adaptive adjustment of the restoration strategy. The module takes the degradation-aware scene descriptor as the conditional input, dynamically generates the convolution kernel parameters and attention weight matrix matching the current degradation mode through the dynamic convolution kernel hyperparameter network, and drives the network to perform customized image restoration according to the degradation type and degree. 2.The adaptive image restoration method based on joint perception of degradation type and degree according to claim 1, characterized in that, In the step (1), the DSD dataset is constructed. 1383 high-definition images are selected from the RAISE database and superimposed with degradation. Mathematical modeling is performed on four typical degradations, i.e., "rain", "snow", "low illumination", and "fog", to simulate the coupling phenomenon of multiple degradation factors in real scenes and generate 11 degradation type datasets containing compound degradation types. During the dataset construction process, the low, medium, and high degrees of the four typical degradations are quantified systematically. The quantification is divided into five degradation degrees, i.e., "clean", "no related degradation", "low degradation degree", "medium degradation degree", and "high degradation degree". During the construction of the DSD dataset, a structured metadata annotation file containing the degradation type and quantization parameters is established for each degraded image, accurately recording the degradation type and degree. 3.The adaptive image restoration method based on joint perception of degradation type and degree according to claim 1, characterized in that, The DASDG in the step (2) comprises two core modules: a degradation type identification module and a degradation degree evaluation module, which respectively generate a type descriptor and a degree descriptor , and generate a degradation-aware scene cue by fusion.

4. The adaptive image restoration method based on joint perception of degradation type and degree according to claim 3, characterized in that, The step (2) further includes the following steps: first, training the type descriptor; A GloVe pre-trained word vector model was used to extract semantic features of five basic degradation types: "clean," "low light," "rain," "snow," and "fog." These features were then mapped to a 12-dimensional text vector space using a multilayer perceptron (MLP). For visual feature extraction, ResNet-18 was used to capture the global structural features of the image, and a mapping network consisting of convolutional layers, dropout layers, and fully connected layers was used to generate visual feature vectors. To achieve precise semantic alignment between the degraded image and the text description, the cosine similarity between the degraded image vector and the text vector was calculated. After normalization, the text embedding corresponding to the highest probability value is selected as the matching result, which can be expressed as follows: (4) wherein and is a text vector and a visual feature vector, is a temperature factor, is a number of text vectors, is a cosine similarity, is a similarity score.

5. The adaptive image restoration method based on joint perception of degradation type and degree according to claim 3, characterized in that, The step (2) further includes the following steps: training the degree descriptor; For single degradation types such as "low illumination", "fog", "rain", and "snow", independent degradation intensity evaluation branches are constructed. In the text semantic encoding, the five-level degree description texts, i.e., "clean", "no related degradation", "low degradation degree", "medium degradation degree", and "high degradation degree", are used as inputs. The GloVe pre-trained word vector model and MLP are used to generate a 5-dimensional text vector space, which explicitly represents the semantic feature of the degree of different degradation types. The process of visual feature extraction and realization of semantic precise alignment between the degraded image and the text description is consistent with the degradation type recognition module. The scene descriptor generator oriented to degradation perception inputs a degraded image during reasoning, and outputs semantic information from a degradation type identification module and degree information from a degradation degree evaluation module After channel splicing, the scene prompter with degradation perception is generated through an MLP .

6. The adaptive image restoration method based on joint perception of degradation type and degree according to claim 1, characterized in that, The degradation-aware dynamic convolution Transformer module DADCT in the step (3) is composed of a dynamic convolution kernel hyperparameter network and a degradation feature Transformer module; the degradation feature Transformer module includes a cross-attention layer, a self-attention layer and a fully connected feedforward network layer; the dynamic convolution kernel hyperparameter network takes the degradation-aware scene descriptor as input, and constructs a mapping from semantic representation to network parameters; the degradation feature Transformer module accepts the generated parameters, and takes the degradation descriptor as a cross-attention query vector Q, to drive the dynamic adjustment of the convolution kernel and the attention mechanism.

Citation Information

Patent Citations

  • Multi-degraded image restoration method based on self-adaptive prompt

    CN119515734A

  • Bad weather image restoration method based on multi-modal state space model

    CN120634891A