CLIP-Enhanced Spatial Adaptive Blind Image Restoration Method and Device
Through the CLIP-enhanced blind image recovery method, the spatial prompt generation and gated prompt interaction module are used to dynamically generate spatial dynamic prompt images, solving the problem that image recovery methods in the prior art are difficult to deal with spatial heterogeneous degradation, and achieving flexible and intelligent image recovery effects.
Patent Information
- Application Number
- CN202510632834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Existing image recovery methods are difficult to effectively deal with spatial heterogeneous degradation, resulting in excessive smoothing, loss of details or artifacts. Large-scale pre-trained models are limited in application in image restoration tasks and are difficult to adapt to local changes.
Using a blind image recovery method based on CLIP enhancement, the blind image recovery network is combined with the encoder-decoder structure, and the spatial prompt generation module and the gated prompt interaction module are used to dynamically generate spatial dynamic prompt diagrams, provide differentiated guidance information for different regions, and adjust the influence intensity of the prompt diagram through the gated network.
It realizes the flexibility and intelligence of image details to effectively handle spatial heterogeneous degradation without predicting the type and degree of degradation, and improve image quality.
Smart Images

Figure CN120147195B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method and device for spatially adaptive blind image restoration based on CLIP enhancement. Background Art
[0002] As an important carrier of visual information, images play a crucial role in today's information age and have become one of the main ways of people's daily communication and information exchange. However, in actual application scenarios, due to device limitations, environmental factors, and shooting conditions, digital images often suffer from various quality degradation problems during the acquisition process. For example, motion blur caused by device jitter or object movement, noise interference caused by sensor noise or harsh environments, and rain and snow occlusion caused by natural weather factors. These problems not only affect the visual quality and aesthetics of the images but also reduce the accurate transmission of information.
[0003] Image restoration technology aims to recover high-quality clear images from degraded images, which is crucial for enhancing the visual experience and improving the performance of downstream computer vision tasks, such as object detection, face recognition, medical image analysis, etc. Existing image restoration methods are mainly divided into two categories: prior knowledge-based modeling methods and deep learning-based direct restoration methods. Prior knowledge-based modeling methods process the basic hint module by estimating the degradation kernel or establishing a physical model to generate multiple basic hint vectors, but usually require strong assumption conditions such as global consistency, which are difficult to meet in complex real-world scenarios. Deep learning-based direct restoration methods directly learn the mapping relationship from degraded images to clear images using deep learning. Although they can handle more complex scenarios, since real-world images often have multiple degradations such as blur and noise at the same time, and this method is usually designed for a single degradation or simply superimposed for processing, it is difficult to effectively decouple and specifically repair mixed degradations.
[0004] To cover various degradation scenarios, a common approach is to train and store model copies for each degradation type and different severities respectively. This strategy results in a large model library, complex management, and huge consumption of computing and storage resources, making it difficult to be effectively deployed on resource-constrained devices such as mobile terminals and embedded systems.
[0005] Moreover, image degradation is often spatially non-uniform, that is, it has spatial heterogeneity. For example, some regions of the image may be severely blurred while others are relatively clear; or, noise and rain streaks only appear in some regions. Existing methods mostly adopt global processing strategies or have limited local receptive fields, making it difficult to perform refined adaptive processing according to the specific content and degradation characteristics of different regions of the image, which easily leads to over-smoothing, detail loss, or artifact generation.
[0006] In addition, although large-scale pre-trained models, especially Contrastive Language-Image Pretraining (CLIP) models, have demonstrated powerful capabilities in understanding image content and semantics and have been preliminarily used in image restoration tasks, the existing application methods still have limitations. Some works only use the CLIP model for rough degradation classification or extract global features as assistance, simply taking the features extracted by the CLIP model as the fixed prior input of the image restoration network, which is difficult to adapt to local changes. Summary of the Invention
[0007] The present invention provides a method and device for CLIP-enhanced spatially adaptive blind image restoration to solve the defects existing in the prior art.
[0008] The present invention provides a method for CLIP-enhanced spatially adaptive blind image restoration, including:
[0009] Obtain the image to be restored;
[0010] Based on the blind image restoration network, obtain the restored image corresponding to the image to be restored;
[0011] Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers, and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. The plurality of encoder layers are skip-connected to the decoder module in the plurality of decoder layers;
[0012] The encoder layer is used to extract the image encoding features of the image to be restored;
[0013] The decoder module is used to fuse the input features to obtain the fused features;
[0014] The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module, and a gated prompt interaction module;
[0015] The basic prompt module is used to generate a plurality of basic prompt vectors;
[0016] The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded based on the CLIP model for a specified image, and the specified image is obtained by sampling the image to be restored to the resolution of the input features;
[0017] The gated prompt interaction module is used to fuse the fused features and the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer;
[0018] The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0019] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the basic prompt module is further used for:
[0020] Embedding the CLIP text features into the multiple basic prompt vectors;
[0021] Wherein, the CLIP text features are obtained by encoding a text prompt template based on the CLIP model.
[0022] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, there are multiple text prompt templates, and the CLIP text features are obtained by fusing the encoding results of the multiple text prompt templates;
[0023] The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0024] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the spatialized prompt generation module is specifically used for:
[0025] Based on an adaptive pooling layer, modifying the spatial resolution of the CLIP image features, and based on a convolution operation, modifying the number of channels of the CLIP image features to obtain a first feature matching the input features;
[0026] Concatenating the first feature with the input features to obtain a first concatenated feature;
[0027] Performing a convolution operation on the first concatenated feature and, based on an activation function, obtaining a spatial attention weight map;
[0028] Based on the spatial attention weight map, fusing the multiple basic prompt vectors to generate the spatial dynamic prompt map.
[0029] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the gated prompt interaction module is specifically used for:
[0030] Concatenating the fused features with the spatial dynamic prompt map to obtain a second concatenated feature;
[0031] Based on a Transformer block, transforming the second concatenated feature to obtain a transformed feature;
[0032] Based on a gating network, a gated graph matching the size of the transformed features is generated, and based on the gated graph, the transformed features and the fused features are fused element-wise weighted to obtain the output features.
[0033] According to a spatial adaptive blind image restoration method based on CLIP enhancement provided by the present invention, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image;
[0034] The training sample set includes first degraded images of various degradation types.
[0035] According to a spatial adaptive blind image restoration method based on CLIP enhancement provided by the present invention, the blind image restoration network is trained based on the following steps:
[0036] An image data set is obtained, and the image data set includes second degraded images of various degradation types and second clear images corresponding to the second degraded images;
[0037] The second degraded image is input into the initial restoration network to obtain the output image of the initial restoration network;
[0038] Based on the second clear image and the output image, the content loss and the edge loss are calculated respectively;
[0039] Based on the content loss and the edge loss, the initial restoration network is iteratively trained to obtain the blind image restoration network.
[0040] The present invention also provides a spatial adaptive blind image restoration device based on CLIP enhancement, including:
[0041] An image acquisition module for acquiring an image to be restored;
[0042] An image restoration module for obtaining a restored image corresponding to the image to be restored based on the blind image restoration network;
[0043] Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer includes a decoder module and a prompt enhancement module, and a plurality of the encoder layers are jump-connected to the decoder module in a plurality of the decoder layers;
[0044] The encoder layer is used to extract the image coding features of the image to be restored;
[0045] The decoder module is used to fuse the input features to obtain fused features;
[0046] The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module, and a gated prompt interaction module;
[0047] The basic prompt module is used to generate a plurality of basic prompt vectors;
[0048] The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and the CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are obtained by encoding a specified image based on the CLIP model, and the specified image is obtained by sampling the image to be restored to the resolution of the input features;
[0049] The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer;
[0050] The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for spatially adaptive blind image restoration based on CLIP enhancement as described in any one of the above is implemented.
[0052] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for spatially adaptive blind image restoration based on CLIP enhancement as described in any one of the above is implemented.
[0053] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for spatially adaptive blind image restoration based on CLIP enhancement as described in any one of the above is implemented.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The spatial adaptive blind image restoration method and device based on CLIP enhancement provided by the present invention utilize a blind image restoration network. Without the need to know the specific degradation type and degree of the image to be restored in advance, they can effectively utilize the powerful text-image understanding and association capabilities of the CLIP model, introduce its rich semantic and visual prior knowledge into the general blind image restoration task, and thus effectively restore the image to be restored that has suffered from single or mixed degradations such as blur, noise, rain, snow, haze, and low light. The blind image restoration network combines an encoder-decoder structure and integrates a prompt enhancement module in the decoder. Through the spatialized prompt generation module, it can dynamically generate a spatial dynamic prompt map according to the local features of the image to be restored, providing differentiated and targeted guiding information for different regions of the image to be restored, thereby effectively dealing with spatially heterogeneous degradations and better retaining local details. At the same time, the gated prompt interaction module introduces a gating mechanism through a gating network, which can dynamically adjust the influence intensity of the spatial dynamic prompt map on the fused features, and further achieve more flexible and intelligent feature fusion and restoration control. Description of the Drawings
[0056] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 is one of the schematic flowcharts of the spatial adaptive blind image restoration method based on CLIP enhancement provided by the present invention;
[0058] Figure 2 is the schematic structural diagram of the blind image restoration network provided by the present invention;
[0059] Figure 3 is the schematic structural diagram of the spatialized prompt generation module provided by the present invention;
[0060] Figure 4 is the schematic structural diagram of the spatial adaptive blind image restoration device based on CLIP enhancement provided by the present invention;
[0061] Figure 5 is the schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0062] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Apparently, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0063] Figure 1 It is a schematic flowchart of a spatial adaptive blind image restoration method based on CLIP enhancement provided in an embodiment of the present invention. As Figure 1 shown, the method includes:
[0064] S1. Obtain the image to be restored;
[0065] S2. Based on the blind image restoration network, obtain the restored image corresponding to the image to be restored;
[0066] Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers, and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. A plurality of the encoder layers are skip-connected to the decoder modules in the plurality of decoder layers;
[0067] The encoder layer is used to extract the image coding features of the image to be restored;
[0068] The decoder module is used to fuse the input features to obtain fused features;
[0069] The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module, and a gated prompt interaction module;
[0070] The basic prompt module is used to generate a plurality of basic prompt vectors;
[0071] The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded from a specified image based on the CLIP model, and the specified image is obtained by sampling the image to be restored to the resolution of the input features;
[0072] The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map to obtain the output features of the decoder layer;
[0073] The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0074] Specifically, in the embodiments of the present invention, the provided method for spatially adaptive blind image restoration based on CLIP enhancement has an execution subject which is a device for spatially adaptive blind image restoration based on CLIP enhancement. This device can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.
[0075] First, step S1 is executed to obtain the image to be restored. The image to be restored refers to a low-quality image with degradation defects that needs to be restored to a clear image. The image to be restored may have motion blur caused by device jitter or object movement, noise interference caused by sensor noise or a harsh environment, annealing defects such as rain, snow, or occlusion caused by natural weather factors, etc., and no specific limitation is made here.
[0076] Thereafter, preprocessing can also be performed on the image to be restored, such as operations like cropping and denoising the image to be restored, so that the size of the image to be restored is adjusted to H×W×3, where H is the height of the image to be restored, W is the width of the image to be restored, and 3 is the number of channels.
[0077] Then, step S2 is executed. The image to be restored is input into the blind image restoration network, and one forward pass is performed, and the restored image corresponding to the image to be restored output by the blind image restoration network can be obtained. This restored image is a clear image of the image to be restored. The whole process is blind, that is, no information about the type or degree of degradation of the image to be restored needs to be provided, and the blind image restoration network automatically adapts to the type or degree of degradation of the image to be restored.
[0078] As Figure 2 shown, this blind image restoration network can adopt a U-Net architecture, which can include an encoder, a decoder, and an output layer. The encoder includes multiple encoder layers connected in sequence, the decoder includes multiple decoder layers connected in sequence, and the last encoder layer in the encoder is connected to the first decoder layer in the decoder through a bottleneck layer, and the last decoder layer in the decoder is connected to the output layer. Figure 2 Here, only the example where the encoder includes three encoder layers and the decoder includes three decoder layers is used for illustration.
[0079] Each decoder layer includes a decoder module and a prompt enhancement module. The decoder module is connected to the prompt enhancement module. The decoder module in each decoder layer is connected to the prompt enhancement module in the previous decoder layer, and the prompt enhancement module in each decoder layer is connected to the decoder module in the next decoder layer.
[0080] The decoder modules of each encoder layer in the encoder are skip-connected to those of each decoder layer in the decoder to directly transmit information.
[0081] Each encoder layer in the encoder and each decoder layer in the decoder use a unique Transformer block design, which can effectively perform local self-attention calculations within a window, thereby reducing the computational cost and improving the restoration effect. The self-attention of each window captures the dependencies within the local window by calculating the linear projections of queries (Q), keys (K), and values (V) and performing a weighted sum. By adding relative position bias values within each window, the sensitivity of the self-attention within the window to relative positions is enhanced, enabling features to align more precisely with local position information.
[0082] The encoder is used to extract image encoding features of the image to be restored layer by layer. Each encoder layer is used to extract image encoding features of different scales of the image to be restored. Each encoder layer may include a feature extraction module and a downsampling module. The feature extraction module contains several Transformer blocks and a feed-forward neural network. Through each Transformer block, the input features are windowed and self-attention is calculated. Through the feed-forward neural network, the feature extraction module can handle local and global dependencies simultaneously, providing fine-grained feature information for image restoration. The downsampling module downsamples the output features obtained by the feature extraction module through a depth convolutional layer with a convolution kernel size of 4×4 and a stride of 2, reducing the resolution and increasing the number of feature channels. Here, the depth convolutional layer can better extract features and is particularly suitable for detail reconstruction in image restoration tasks.
[0083] The decoder includes multiple decoder layers symmetric to the encoder, except that it also fuses the image encoding features from the output of the corresponding encoder layer through skip connections. The main task of the decoder is to gradually restore the spatial resolution of the features output by the encoder, reduce the number of channels, and reconstruct image details. Here, through skip connections, not only can the hierarchical feature aggregation advantage of the U-Net architecture be retained, but also each Transformer block can be fused to enhance the modeling ability of the blind image restoration network.
[0084] The decoder module in each decoder layer includes an upsampling layer and a fusion layer connected in sequence. The upsampling layer upsamples the output features of the previous decoder layer at the beginning of the decoder layer through a transposed convolutional kernel of size 2×2 to magnify the resolution of the output features of the previous decoder layer and reduce the number of feature channels. The upsampled features and the image encoding features output by the corresponding encoder layer are fused through the fusion layer for skip connection fusion operations to combine low-level features (i.e., the input features of the decoder layer) and high-level features (i.e., the upsampled features).
[0085] Here, the output features of the previous decoder layer and the image encoding features corresponding to the output of the encoder layer are both used as the input features of the current decoder layer, that is, the input features of the decoder module in the current decoder layer. The transposed convolutional kernel can achieve depth convolution to better extract features, which is particularly suitable for detail reconstruction in image restoration tasks.
[0086] The hint enhancement module in each decoder layer is used to dynamically generate spatial hints and perform gated interactions. The hint enhancement module can include a basic hint module, a spatialized hint generation module (SPGM), and a gated hint interaction module (GPIM) that work together.
[0087] The basic hint module is used to generate multiple basic hint vectors. Each basic hint vector is a set of learnable parameters globally shared in the blind image restoration network. The number of basic hint vectors can be N, and each basic hint vector can be a D-dimensional vector. Then, the set of basic hint vectors composed of each basic hint vector can be expressed as , is the first basic hint vector, is the Nth basic hint vector. The purpose of each basic hint vector is to learn to capture general representations related to different degradation patterns, image structures, repair primitives, or abstract contexts during training, providing basic elements for the generation of dynamic hints.
[0088] The spatialized hint generation module is used to fuse multiple basic hint vectors based on the input features of the decoder layer and the CLIP image features to generate a spatial dynamic hint map to adapt to image local features and heterogeneous degradation.
[0089] The CLIP image features can be obtained by encoding a specified image through the image encoder of the CLIP model. The specified image can be obtained by sampling the image to be restored to the resolution of the input features of the decoder layer. The CLIP model can be pre-trained on large-scale datasets such as ImageNet and COCO, learning general image-text alignment capabilities. Inputting the specified image into the image encoder of the CLIP model can obtain the CLIP image features encoded by the image encoder of the CLIP model for the specified image. Here, in order to obtain the dense embedding of the specified image, the last pooling layer of the image encoder is removed, and then the CLIP image features are generated by inputting the specified image. The CLIP image features are an image embedding matrix, where each element represents the semantic information of the specified image at that position.
[0090] The gated prompt interaction module can fuse the fused features output by the fusion layer of the decoder module with the spatially dynamic prompt map generated by the spatialized prompt generation module based on the gated network to obtain the output features of the decoder layer. Here, the gated network can include a convolutional layer and an activation layer, and the activation layer uses the sigmoid function to implement the activation operation. Through the gated network, a gating mechanism can be introduced to dynamically adjust the influence intensity of the spatially dynamic prompt map on the fused features, thereby achieving more intelligent and more robust feature fusion.
[0091] The output layer of the blind image restoration network can obtain the restored image according to the output features of the last decoder layer and the image to be restored. Here, the output layer can include an output convolutional layer and a superimposing layer. The output convolutional layer can map the output features of the last decoder layer to the three-channel red (R), green (G), and blue (B) image space to obtain the residual image of the restored image. The superimposing layer can superimpose the residual image on the image to be restored to obtain the restored image. Here, the output convolutional layer can also be a depth convolutional layer, which can better extract features and is particularly suitable for detail reconstruction in image restoration tasks.
[0092] The activation function in the U-Net architecture adopted in the blind image restoration network is the GeLU function, which is a non-linear activation function that determines whether to activate neurons through a probability distribution. The definition of the GeLU function is as follows:
[0093] ;
[0094] where GELU(x) is the dependent variable of the GeLU function, and x is the tensor input to the GeLU function. is the cumulative distribution function of the standard normal distribution. This GeLU function can enhance the non-linear ability of the blind image restoration network.
[0095] In the embodiment of the present invention, the spatially adaptive blind image restoration method based on CLIP enhancement first obtains the image to be restored; then uses a blind image restoration network to obtain the restored image corresponding to the image to be restored. This method uses a blind image restoration network, which can effectively utilize the powerful text-image understanding and association capabilities of the CLIP model without prior knowledge of the specific degradation type and degree of the image to be restored, and introduce its rich semantic and visual prior knowledge into the general blind image restoration task, so as to effectively restore the image to be restored suffering from single or mixed degradation such as blur, noise, rain, snow, haze, low light, etc. The blind image restoration network combines an encoder-decoder structure and integrates a prompt enhancement module in the decoder. Through the spatialized prompt generation module, it can dynamically generate a spatial dynamic prompt map according to the local features of the image to be restored, providing differentiated and targeted guiding information for different regions of the image to be restored, so as to effectively process spatially heterogeneous degradation and better retain local details. At the same time, the gated prompt interaction module introduces a gating mechanism through a gating network, which can dynamically adjust the influence intensity of the spatial dynamic prompt map on the fused features, and then achieve more flexible and intelligent feature fusion and restoration control.
[0096] Based on the above embodiment, the basic prompt module is further configured to:
[0097] Embed the CLIP text features into the multiple basic prompt vectors;
[0098] Wherein, the CLIP text features are obtained by encoding a text prompt template based on the CLIP model.
[0099] Specifically, after generating each basic prompt vector, the basic prompt module can also embed the CLIP text features into multiple basic prompt vectors. The CLIP text features can be obtained by inputting the text prompt template into the text encoder of the CLIP model, and encoded and output by the text encoder. Here, the text prompt template can be set differently according to different types of the image to be restored. For example, when the type of the image to be restored is noise reduction / rain removal / blur removal, the corresponding text prompt templates can be set as "it is a part of noise / rainy / blurry image" respectively. The text prompt template can include one or more.
[0100] Here, in order to obtain the dense embedding of the text prompt template, the last pooling layer of the image encoder is removed, and then the CLIP text features are generated by inputting the text prompt template. The CLIP text features are a text embedding matrix, where each element represents the semantic information of the text prompt template at that position.
[0101] In the embodiments of the present invention, the introduction of CLIP text features can provide clear semantic guidance.
[0102] Based on the above embodiments, there are multiple text prompt templates, and the CLIP text features are obtained by fusing the encoding results of multiple text prompt templates;
[0103] The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0104] Specifically, when there are multiple text prompt templates, such as ["This is a blurry region", "This part is affected by motion blur", "Blur exists in this area"], the ambiguity that may exist in a single text description can be reduced.
[0105] Each text prompt template can obtain an encoding result when input into the text encoder of the CLIP model. By fusing the encoding results of each text prompt template, the CLIP text features can be obtained.
[0106] That is:
[0107] ;
[0108] Where, is the CLIP text feature, is the encoding result of the i-th text prompt template, and n is the number of text prompt templates.
[0109] The degradation degree of each position in the image to be restored can be determined by the similarity between the CLIP image features and the CLIP text features. Here, the similarity can be the cosine similarity, which can effectively measure the direction consistency. The higher the similarity, the higher the degradation degree of the corresponding position. Finally, through The function normalizes the similarity to the probability space, which is expressed as follows:
[0110] ;
[0111] Where, is the cosine similarity, is the CLIP image feature.
[0112] Based on the above embodiments, the spatialized prompt generation module is specifically used for:
[0113] Modify the spatial resolution of the CLIP image features based on the adaptive pooling layer, and modify the number of channels of the CLIP image features based on the convolution operation to obtain the first feature that matches the input features;
[0114] Concatenate the first feature with the input features to obtain the first concatenated feature;
[0115] Perform a convolution operation on the first concatenated feature and obtain the spatial attention weight map based on the activation function;
[0116] Fuse the multiple basic prompt vectors based on the spatial attention weight map to generate the spatial dynamic prompt map.
[0117] Specifically, as Figure 3 shown, the spatialized prompt generation module may include a concatenation layer, a first convolutional layer, and a fusion unit. The concatenation layer includes an adaptive pooling layer, a second convolutional layer, and a concatenation unit. The adaptive pooling layer is used to modify the spatial resolution of the CLIP image features. The second convolutional layer can be a 1×1 convolutional layer, which is used to modify the number of channels of the CLIP image features to obtain the first feature that matches the input features of the decoder module.
[0118] After that, the first feature is concatenated with the input features of the decoder module through the concatenation unit to obtain the first concatenated feature: ;
[0119] wherein, is the first concatenated feature, is the input feature of the decoder module, is the first feature.
[0120] After that, the first convolutional layer performs a convolution operation on the first concatenated feature and obtains the spatial attention weight map based on the activation function. Here, the activation function can be the softmax function. The spatial attention weight map obtained through the activation function provides a normalized attention distribution to each basic prompt vector at each spatial position.
[0121] Subsequently, the fusion unit uses the weight vectors at each spatial position in the spatial attention weight map to fuse each basic prompt vector to generate the spatial dynamic prompt map. Here, the fusion method can be weighted summation, and the obtained spatial dynamic prompt map contains different prompt information adaptively generated according to local features at different spatial positions.
[0122] Based on the above embodiments, the gated prompt interaction module is specifically used for:
[0123] Concatenate the fusion feature with the spatial dynamic prompt map to obtain the second concatenated feature;
[0124] Based on the Transformer block, transform the second concatenated feature to obtain a transformed feature;
[0125] Based on the gating network, generate a gating map that matches the size of the transformed feature, and based on the gating map, perform element-wise weighted fusion on the transformed feature and the fused feature to obtain the output feature.
[0126] Specifically, the gated prompt interaction module can be responsible for integrating the spatial dynamic prompt map generated by the spatialized prompt generation module into the fused feature obtained by the decoder module in a flexible and controlled manner . Its working process includes two key steps:
[0127] First, concatenate the spatial dynamic prompt map and the fused feature to obtain a second concatenated feature, and input the second concatenated feature into the Transformer block to obtain a transformed feature.
[0128] Second, the gated prompt interaction module uses the gating network to generate a gating map G that matches the size of the transformed feature.
[0129] Finally, use the gating map G to perform element-wise weighted fusion on the transformed feature and the fused feature to obtain the output feature:
[0130] ;
[0131] wherein, is the output feature, is the transformed feature, is the fused feature.
[0132] The gating mechanism adopted by the gated prompt interaction module allows the blind image restoration network to adaptively decide at each spatial position and channel whether to adopt more of the transformed features enhanced by the prompt guidance or retain the fused feature, thereby achieving more intelligent and robust feature fusion.
[0133] Based on the above embodiments, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image;
[0134] The training sample set includes first degraded images of multiple degradation types.
[0135] Specifically, to avoid the inability of directly using the pre-trained CLIP model to accurately perceive the degradation degree of the image to be restored, the CLIP model can be specifically fine-tuned to make it more suitable for the modeling and degradation estimation of degraded images.
[0136] To enhance the understanding ability of the CLIP model for low-quality images, especially for various underlying degradation features, a training sample set is constructed, which includes a first degraded image with various degradation types (such as different types of blur, different levels and distributions of noise, rain, snow, haze, low light, etc.) and its corresponding high-quality first clear image.
[0137] Meanwhile, the training sample set can also include degradation description texts associated with the first degraded image (such as "motion blur area", "Gaussian noise points", "clear edges", "rain streak occlusion", etc.). Text labels can be semi-automatically generated through templated statements such as "This is a clear image with sharp details." or "This image contains motion blur in some areas." and verified manually to ensure semantic consistency.
[0138] Using the training sample set, the LoRA (Low-Rank Adaptation) technique is used to perform a contrastive learning task to fine-tune the CLIP model. The goal of the fine-tuning is to maximize the similarity between the embedding vectors of the matching first clear image and the clear description text, while minimizing the similarity between the embedding vectors of the mismatching first clear image and the degradation description text, thereby strengthening the CLIP model's perception ability for the features related to the contrastive learning task, making the embedding vectors of the first clear image match the corresponding clear description text, and at the same time matching the embedding vectors of the first blurred image with their corresponding blurred description texts. These two are regarded as positive sample pairs.
[0139] Similarly, negative sample pairs are constructed by establishing the corresponding relationships between the embedding vectors of the first clear image and the blurred description text, and between the embedding vectors of the first blurred image and the clear description text.
[0140] Then, the NT-Xent loss function in contrastive learning is used to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs, thereby strengthening the CLIP model's modeling ability in the contrastive learning task:
[0141] ;
[0142] where is the NT-Xent loss, is a control parameter used to control the sensitivity of the comparison. is the embedding vector of the first clear image. is the embedding vector of the clear description text. is the embedding vector of the k-th description text, where K is the number of description texts in the training sample set.
[0143] Since the CLIP model is a large-scale pre-trained model, the LoRA technique is used to perform lightweight fine-tuning on it, enabling it to better adapt to the feature modeling of blurred images while avoiding the computational overhead caused by large-scale parameter updates. Specifically, by freezing the original weights of the CLIP model, low-rank matrices are inserted at specific layers. Then, the trainable low-rank adapter is used to learn the information specific to the deblurring task, thus avoiding fine-tuning the entire CLIP model and reducing the consumption of computational resources. After training, the weights of the fine-tuned CLIP model are saved and used for subsequent image restoration tasks.
[0144] Based on the above embodiments, the blind image restoration network is trained based on the following steps:
[0145] Obtain an image dataset, which includes second degraded images of various degradation types and the corresponding second clear images of the second degraded images.
[0146] Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network.
[0147] Based on the second clear image and the output image, calculate the content loss and the edge loss respectively.
[0148] Based on the content loss and the edge loss, perform iterative training on the initial restoration network to obtain the blind image restoration network.
[0149] Specifically, in the process of training the initial restoration network to obtain the blind image restoration network, first, an image dataset can be obtained. This image dataset includes real second degraded images of various degradation types (such as different types of blur, different levels of noise, rain, snow, haze, low light, etc.) from datasets such as Rain100L, SOTS, CBSD68, GoPro, and LOL, as well as the corresponding second clear images of the second degraded images. In addition, the image dataset can also include synthetic second degraded images and second clear images to mix the synthetic images with the real images and improve the generalization of the blind image restoration network.
[0150] After that, it is also necessary to preprocess the second degraded images in the image dataset and then divide them into a training set and a validation set. Among them, the second clear images and the corresponding second degraded images in the training set appear in pairs as a set of data.
[0151] Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network. Calculate the content loss and the edge loss respectively using the second clear image and the output image.
[0152] The content loss can be calculated by the following content loss function:
[0153] ;
[0154] where is the content loss, is the second clear image, is the output image. By calculating the L2 norm difference between the two and adding a small constant to stabilize the value.
[0155] The edge loss can be calculated by the following edge loss function:
[0156] ;
[0157] where is the edge loss, is the second clear image, is the output image. is the fast Fourier transform. By calculating the L2 norm difference between the two in the frequency domain and adding a small constant to stabilize the value.
[0158] After that, the total training loss can be obtained by weighted summation of the content loss and the edge loss: ; where .
[0159] Using the total training loss, calculate the error and generate the gradient, and update the model parameters of the initial restoration network by using the gradient backpropagation and gradient descent methods, and perform iterative training on the initial restoration network to make the initial restoration network converge to achieve the ideal clear image output effect.
[0160] As Figure 4 shown, based on the above embodiments, an apparatus for spatially adaptive blind image restoration based on CLIP enhancement is provided in an embodiment of the present invention, including:
[0161] An image acquisition module 41 for acquiring an image to be restored;
[0162] An image restoration module 42 for obtaining a restored image corresponding to the image to be restored based on the blind image restoration network;
[0163] Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers, and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. A plurality of the encoder layers are skip-connected to the decoder modules in the plurality of decoder layers;
[0164] The encoder layer is used to extract the image encoding features of the image to be restored;
[0165] The decoder module is used to fuse the input features to obtain fused features;
[0166] The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module, and a gated prompt interaction module;
[0167] The basic prompt module is used to generate a plurality of basic prompt vectors;
[0168] The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded based on the CLIP model for a specified image, and the specified image is obtained by sampling the image to be restored to the resolution of the input features;
[0169] The gated prompt interaction module is used to fuse the fused features and the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer;
[0170] The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0171] On the basis of the above embodiments, the basic prompt module is further used for:
[0172] Embedding the CLIP text features into the plurality of basic prompt vectors;
[0173] Among them, the CLIP text features are encoded based on the CLIP model for a text prompt template.
[0174] On the basis of the above embodiments, there are a plurality of the text prompt templates, and the CLIP text features are obtained by fusing the encoding results of the plurality of text prompt templates;
[0175] The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0176] On the basis of the above embodiments, the spatialized prompt generation module is specifically used for:
[0177] Modify the spatial resolution of the CLIP image features based on the adaptive pooling layer, and modify the number of channels of the CLIP image features based on the convolutional operation to obtain the first feature that matches the input feature;
[0178] Concatenate the first feature and the input feature to obtain the first concatenated feature;
[0179] Perform a convolutional operation on the first concatenated feature and obtain the spatial attention weight map based on the activation function;
[0180] Fuse the multiple basic prompt vectors based on the spatial attention weight map to generate the spatial dynamic prompt map.
[0181] Based on the above embodiments, the gated prompt interaction module is specifically configured to:
[0182] Concatenate the fused feature and the spatial dynamic prompt map to obtain the second concatenated feature;
[0183] Transform the second concatenated feature based on the Transformer block to obtain the transformed feature;
[0184] Generate a gated map that matches the size of the transformed feature based on the gated network, and perform element-wise weighted fusion on the transformed feature and the fused feature based on the gated map to obtain the output feature.
[0185] Based on the above embodiments, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image;
[0186] The training sample set includes the first degraded images of multiple degradation types.
[0187] Based on the above embodiments, a training module is further included, which is used to:
[0188] Obtain an image data set, where the image data set includes second degraded images of multiple degradation types and the second clear images corresponding to the second degraded images;
[0189] Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network;
[0190] Calculate the content loss and the edge loss respectively based on the second clear image and the output image;
[0191] Based on the content loss and the edge loss, iteratively train the initial restoration network to obtain the blind image restoration network.
[0192] Specifically, in the embodiments of the present invention, the functions of the modules in the CLIP-enhanced spatially adaptive blind image restoration device provided are in one-to-one correspondence with the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.
[0193] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the CLIP-enhanced spatially adaptive blind image restoration method provided in the above embodiments.
[0194] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disks, or optical disks, etc., which can store program codes.
[0195] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the CLIP-enhanced spatially adaptive blind image restoration method provided in the above embodiments.
[0196] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the spatially adaptive blind image restoration method enhanced by CLIP provided in the above embodiments.
[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0198] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A spatially adaptive blind image restoration method enhanced based on CLIP, characterized in that Including: Obtain the image to be restored; Based on the blind image restoration network, obtain the restored image corresponding to the image to be restored; Wherein, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer includes a decoder module and a prompt enhancement module, and the plurality of encoder layers are skip-connected to the decoder modules in the plurality of decoder layers; The encoder layer is used to extract the image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and the CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded based on the CLIP model for a specified image, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored; The spatialized prompt generation module is specifically used for: Based on an adaptive pooling layer, modify the spatial resolution of the CLIP image features, and based on a convolutional operation, modify the number of channels of the CLIP image features to obtain a first feature matching the input features; Concatenate the first feature with the input features to obtain a first concatenated feature; Perform a convolutional operation on the first concatenated feature and, based on an activation function, obtain a spatial attention weight map; Based on the spatial attention weight map, fuse the plurality of basic prompt vectors to generate the spatial dynamic prompt map.
2. The method for spatially adaptive blind image restoration based on CLIP enhancement according to claim 1, wherein The basic prompt module is further used for: Embed the CLIP text features into the plurality of basic prompt vectors; Wherein, the CLIP text features are encoded based on the CLIP model for a text prompt template.
3. The method for spatially adaptive blind image restoration enhanced based on CLIP according to claim 2, wherein, There are a plurality of the text prompt templates, and the CLIP text features are obtained by fusing the encoding results of the plurality of text prompt templates; The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
4. The method for spatially adaptive blind image restoration enhanced based on CLIP according to claim 1, characterized in that, The gated prompt interaction module is specifically used for: Concatenate the fused features with the spatial dynamic prompt map to obtain a second concatenated feature; Based on a Transformer block, transform the second concatenated feature to obtain a transformed feature; Based on a gated network, generate a gated map matching the size of the transformed feature, and based on the gated map, perform element-wise weighted fusion on the transformed feature and the fused features to obtain the output features.
5. The method for spatially adaptive blind image restoration enhanced based on CLIP according to any one of claims 1-4, characterized in that, The CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degradation description text associated with the first degraded image, and the clear description text associated with the first clear image; The training sample set includes first degraded images of multiple degradation types.
6. The CLIP-enhanced spatially adaptive blind image restoration method according to any one of claims 1-4, characterized in that The blind image restoration network is trained based on the following steps: Obtain an image dataset, where the image dataset includes second degraded images of multiple degradation types and the second clear images corresponding to the second degraded images; Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network; Based on the second clear image and the output image, calculate the content loss and the edge loss respectively; Based on the content loss and the edge loss, perform iterative training on the initial restoration network to obtain the blind image restoration network.
7. A spatial adaptive blind image restoration device based on CLIP enhancement, characterized in that, It includes: An image acquisition module for acquiring the image to be restored; An image restoration module for obtaining the restored image corresponding to the image to be restored based on the blind image restoration network; Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers, and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. A plurality of the encoder layers are skip-connected to the decoder modules in the plurality of decoder layers; The encoder layer is used to extract the image encoding features of the image to be restored; The decoder module is used to fuse the input features to obtain the fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module, and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and the CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded based on the CLIP model for a specified image, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features and the spatial dynamic prompt map based on the gated network to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored; The spatialized prompt generation module is specifically used for: Based on the adaptive pooling layer, modify the spatial resolution of the CLIP image features, and based on the convolutional operation, modify the number of channels of the CLIP image features to obtain the first feature matching the input features; Concatenate the first feature and the input features to obtain the first concatenated feature; Perform a convolutional operation on the first concatenated feature and, based on the activation function, obtain the spatial attention weight map; Based on the spatial attention weight map, fuse the plurality of basic prompt vectors to generate the spatial dynamic prompt map.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the CLIP-enhanced spatially adaptive blind image restoration method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the CLIP-enhanced spatially adaptive blind image restoration method according to any one of claims 1-6.
Citation Information
Patent Citations
Text feature-introduced image restoration method
CN117314778A
Text generation and information recovery method and system based on large model
CN119515702A