Space adaptive blind image restoration method and device based on CLIP enhancement
By introducing CLIP model and prompt enhancement modules into the image recovery network, dynamically generate spatial dynamic prompt diagrams and perform gated interactions, solving the problem of difficult to deal with hybrid degradation and resource-constrained device deployment in the prior art, and achieving efficient image recovery and detail retention.
Patent Information
- Application Number
- CN202510632834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Existing image recovery methods are difficult to effectively decouple and targetedly repair mixed degraded images, and are difficult to deploy on resource-constrained devices, and are difficult to perform refined adaptive processing based on the specific content and degradation characteristics of different areas of the image.
Using a spatially adaptive blind image recovery method based on CLIP enhancement, a blind image recovery network is combined with an encoder-decoder structure and prompt enhancement module to dynamically generate a spatial dynamic prompt diagram, and feature fusion and repair control are performed through the gate prompt interactive module.
It realizes effective recovery of images subject to multiple degradation without knowing the type and degree of image degradation in advance, and can be differentiated according to the characteristics of different areas of the image, improving the flexibility of retention of local details and feature fusion.
Smart Images

Figure CN120147195A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular, to a method and device for spatially adaptive blind image restoration based on CLIP enhancement. Background Art
[0002] Images, as an important carrier of visual information, play a crucial role in today's information age and have become one of the main ways of people's daily communication and information exchange. However, in actual application scenarios, due to the influence of device limitations, environmental factors, and shooting conditions, digital images often suffer from various quality degradation problems during the acquisition process. For example, motion blur caused by device jitter or object movement, noise interference caused by sensor noise or harsh environments, and rain and snow occlusion caused by natural weather factors. These problems not only affect the visual quality and aesthetics of the images but also reduce the accurate transmission of information.
[0003] Image restoration technology aims to recover high-quality clear images from degraded images, which is crucial for enhancing the visual experience and improving the performance of downstream computer vision tasks such as object detection, face recognition, medical image analysis, etc. Existing image restoration methods are mainly divided into two categories: prior knowledge-based modeling methods and deep learning-based direct restoration methods. Prior knowledge-based modeling methods process the basic hint module by estimating the degradation kernel or establishing a physical model to generate multiple basic hint vectors, but usually require strong assumption conditions such as global consistency, which are difficult to meet in complex real-world scenarios. Deep learning-based direct restoration methods directly learn the mapping relationship from degraded images to clear images using deep learning. Although they can handle more complex scenarios, since real-world images often have multiple degradations such as blur and noise simultaneously, and this method is usually designed for a single degradation or simply processed by superposition, it is difficult to effectively decouple and specifically repair mixed degradations.
[0004] To cover various degradation scenarios, a common approach is to train and store model copies for each degradation type and different severities respectively. This strategy results in a large model library, complex management, and huge consumption of computing and storage resources, making it difficult to be effectively deployed on resource-constrained devices such as mobile terminals and embedded systems.
[0005] Moreover, image degradation is often spatially non-uniform, that is, it has spatial heterogeneity. For example, some regions of the image may be severely blurred while others are relatively clear; or, noise and rain streaks only appear in some regions. Existing methods mostly adopt global processing strategies or have limited local receptive fields, making it difficult to perform refined adaptive processing according to the specific content and degradation characteristics of different regions of the image, and easily leading to over-smoothing, detail loss, or artifact generation.
[0006] In addition, although large-scale pre-trained models, especially Contrastive Language-Image Pretraining (CLIP) models, have demonstrated powerful capabilities in understanding image content and semantics and have been preliminarily used in image restoration tasks, the existing application methods still have limitations. Some works only use the CLIP model for rough degradation classification or extract global features as assistance, simply taking the features extracted by the CLIP model as the fixed prior input of the image restoration network, which is difficult to adapt to local changes. Summary of the Invention
[0007] The present invention provides a CLIP-enhanced spatially adaptive blind image restoration method and device to solve the defects existing in the prior art.
[0008] The present invention provides a CLIP-enhanced spatially adaptive blind image restoration method, including: Obtaining the image to be restored; Based on the blind image restoration network, obtaining the restored image corresponding to the image to be restored; Wherein, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. A plurality of the encoder layers are skip-connected to the decoder module in the plurality of decoder layers; The encoder layer is used to extract the image encoding features of the image to be restored; The decoder module is used to fuse the input features to obtain the fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and the CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded from a specified image based on the CLIP model, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0009] According to the CLIP-enhanced spatially adaptive blind image restoration method provided by the present invention, the basic prompt module is further used for: Embed the CLIP text features into the multiple basic prompt vectors; Among them, the CLIP text features are obtained by encoding a text prompt template based on the CLIP model.
[0010] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, there are multiple text prompt templates, and the CLIP text features are obtained by fusing the encoding results of the multiple text prompt templates; The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0011] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the spatialized prompt generation module is specifically configured to: Based on an adaptive pooling layer, modify the spatial resolution of the CLIP image features, and based on a convolution operation, modify the number of channels of the CLIP image features to obtain a first feature that matches the input features; Concatenate the first feature with the input features to obtain a first concatenated feature; Perform a convolution operation on the first concatenated feature and, based on an activation function, obtain a spatial attention weight map; Based on the spatial attention weight map, fuse the multiple basic prompt vectors to generate the spatial dynamic prompt map.
[0012] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the gated prompt interaction module is specifically configured to: Concatenate the fused features with the spatial dynamic prompt map to obtain a second concatenated feature; Based on a Transformer block, transform the second concatenated feature to obtain a transformed feature; Based on a gated network, generate a gated map that matches the size of the transformed feature, and based on the gated map, perform element-wise weighted fusion on the transformed feature and the fused features to obtain the output features.
[0013] According to a method for spatially adaptive blind image restoration enhanced by CLIP provided by the present invention, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image; The training sample set includes first degraded images of multiple degradation types.
[0014] A spatial adaptive blind image restoration method based on CLIP enhancement provided by the present invention, the blind image restoration network is trained based on the following steps: Obtain an image dataset, which includes second degraded images of various degradation types and corresponding second clear images of the second degraded images; Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network; Based on the second clear image and the output image, calculate the content loss and the edge loss respectively; Based on the content loss and the edge loss, perform iterative training on the initial restoration network to obtain the blind image restoration network.
[0015] The present invention also provides a spatial adaptive blind image restoration device based on CLIP enhancement, including: An image acquisition module for acquiring an image to be restored; An image restoration module for obtaining a restored image corresponding to the image to be restored based on the blind image restoration network; Wherein, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer includes a decoder module and a prompt enhancement module, and a plurality of the encoder layers are skip-connected to the decoder modules in the plurality of decoder layers; The encoder layer is used to extract the image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded based on the CLIP model for a specified image, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for CLIP-enhanced spatially adaptive blind image restoration as described in any one of the above.
[0017] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for CLIP-enhanced spatially adaptive blind image restoration as described in any one of the above.
[0018] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for CLIP-enhanced spatially adaptive blind image restoration as described in any one of the above.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The method and device for CLIP-enhanced spatially adaptive blind image restoration provided by the present invention utilize a blind image restoration network. Without the need to know the specific degradation type and degree of the image to be restored in advance, it can effectively utilize the powerful text-image understanding and association ability of the CLIP model, introduce its rich semantic and visual prior knowledge into the general blind image restoration task, and thus effectively restore the image to be restored suffering from single or mixed degradation such as blur, noise, rain, snow, haze, low light, etc. The blind image restoration network combines an encoder-decoder structure, and a prompt enhancement module is fused in the decoder. Through the spatialized prompt generation module, it can dynamically generate a spatial dynamic prompt map according to the local features of the image to be restored, providing differentiated and targeted guiding information for different regions of the image to be restored, so as to effectively process spatially heterogeneous degradation and better retain local details. At the same time, the gated prompt interaction module introduces a gating mechanism through a gating network, which can dynamically adjust the influence intensity of the spatial dynamic prompt map on the fused features, and further realize more flexible and intelligent feature fusion and restoration control. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings, for those of ordinary skill in the art, can also obtain other drawings without creative efforts.
[0021] Figure 1 is one of the flow schematic diagrams of the method for CLIP-enhanced spatially adaptive blind image restoration provided by the present invention; Figure 2 is the structural schematic diagram of the blind image restoration network provided by the present invention; Figure 3 It is a schematic structural diagram of the spatialized prompt generation module provided by the present invention; Figure 4 It is a schematic structural diagram of the CLIP-enhanced spatial adaptive blind image restoration device provided by the present invention; Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Figure 1 It is a schematic flowchart of a CLIP-enhanced spatial adaptive blind image restoration method provided in an embodiment of the present invention. As Figure 1 shown, the method includes: S1, obtaining the image to be restored; S2, obtaining the restored image corresponding to the image to be restored based on the blind image restoration network; Among them, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence. The decoder layer includes a decoder module and a prompt enhancement module. The plurality of encoder layers are skip-connected to the decoder modules in the plurality of decoder layers; The encoder layer is used to extract the image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain the fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the plurality of basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are encoded from a specified image based on the CLIP model, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0024] Specifically, in the embodiment of the present invention, the space adaptive blind image restoration method based on CLIP enhancement is executed by a space adaptive blind image restoration device based on CLIP enhancement. This device can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.
[0025] First, step S1 is executed to obtain the image to be restored. The image to be restored refers to a low-quality image with degradation defects that needs to be restored to a clear image. The image to be restored may have motion blur caused by device jitter or object movement, noise interference caused by sensor noise or harsh environment, annealing defects such as rain and snow occlusion caused by natural weather factors, etc., and no specific limitation is made here.
[0026] After that, preprocessing can also be performed on the image to be restored. For example, operations such as cropping and denoising are performed on the image to be restored to adjust the size of the image to be restored to H×W×3, where H is the height of the image to be restored, W is the width of the image to be restored, and 3 is the number of channels.
[0027] Then, step S2 is executed. The image to be restored is input into the blind image restoration network, and a forward pass is performed once to obtain the restored image corresponding to the image to be restored output by the blind image restoration network. This restored image is the clear image of the image to be restored. The whole process is blind, that is, no information about the degradation type or degree of the image to be restored needs to be provided, and the blind image restoration network automatically adapts to the degradation type or degree of the image to be restored.
[0028] As Figure 2 shown, the blind image restoration network can adopt a U-Net architecture, which can include an encoder, a decoder, and an output layer. The encoder includes a plurality of encoder layers connected in sequence. The decoder includes a plurality of decoder layers connected in sequence. The last encoder layer in the encoder is connected to the first decoder layer in the decoder through a bottleneck layer, and the last decoder layer in the decoder is connected to the output layer. Figure 2 In
[0029] Each decoder layer includes a decoder module and a prompt enhancement module. The decoder module is connected to the prompt enhancement module. The decoder module in each decoder layer is connected to the prompt enhancement module in the previous decoder layer, and the prompt enhancement module in each decoder layer is connected to the decoder module in the next decoder layer.
[0030] Decoder modules in each encoder layer of the encoder are skip-connected to decoder modules in each decoder layer of the decoder to directly transmit information.
[0031] Each encoder layer in the encoder and each decoder layer in the decoder use a unique Transformer block design that can effectively perform local self-attention calculations within a window, thereby reducing the computational cost and improving the restoration effect. The self-attention for each window captures dependencies within the local window by calculating the linear projections of queries (Q), keys (K), and values (V) and performing weighted summation. By adding relative position bias values within each window, the sensitivity of the self-attention within the window to relative positions is enhanced, enabling features to more precisely align with local position information.
[0032] The encoder is used to extract image encoding features of the image to be restored layer by layer, and each encoder layer is used to extract image encoding features of different scales of the image to be restored. Each encoder layer can include a feature extraction module and a downsampling module. The feature extraction module contains several Transformer blocks and a feed-forward neural network. The input features are windowed and self-attention is calculated through each Transformer block. Through the feed-forward neural network, the feature extraction module can handle local and global dependencies simultaneously, providing fine-grained feature information for image restoration. The downsampling module downsamples the output features obtained by the feature extraction module through a depth convolutional layer with a convolutional kernel size of 4×4 and a stride of 2, reducing the resolution and increasing the number of feature channels. Here, features can be better extracted through the depth convolutional layer, which is particularly suitable for detail reconstruction in image restoration tasks.
[0033] The decoder includes multiple decoder layers symmetric to the encoder, with the difference that it also fuses the image encoding features from the output of the corresponding encoder layer through skip connections. The main task of the decoder is to gradually restore the spatial resolution of the features output by the encoder, reduce the number of channels, and reconstruct image details. Here, through skip connections, not only can the hierarchical feature aggregation advantage of the U-Net architecture be retained, but also each Transformer block can be fused to enhance the modeling ability of the blind image restoration network.
[0034] The decoder module in each decoder layer includes an upsampling layer and a fusion layer connected in sequence. The upsampling layer upsamples the output features of the previous decoder layer at the beginning of the decoder layer through a transposed convolutional kernel of size 2×2 to magnify the resolution of the output features of the previous decoder layer and reduce the number of feature channels. The upsampled features and the image encoding features output by the corresponding encoder layer are fused through the fusion layer for skip connection to combine low-level features (i.e., the input features of the decoder layer) and high-level features (i.e., the upsampled features).
[0035] Here, the output features of the previous decoder layer and the image encoding features corresponding to the output of the encoder layer are both used as the input features of the current decoder layer, that is, the input features of the decoder module in the current decoder layer. The transposed convolutional kernel can achieve depth convolution to better extract features, which is particularly suitable for detail reconstruction in image restoration tasks.
[0036] The hint enhancement module in each decoder layer is used to dynamically generate spatial hints and perform gated interactions. The hint enhancement module can include a basic hint module, a spatialized hint generation module (SPGM), and a gated hint interaction module (GPIM) that work together.
[0037] The basic hint module is used to generate multiple basic hint vectors. Each basic hint vector is a set of learnable parameters globally shared in the blind image restoration network. The number of basic hint vectors can be N, and each basic hint vector can be a D-dimensional vector. Then the set of basic hint vectors can be represented as , is the first basic hint vector, is the Nth basic hint vector. The purpose of each basic hint vector is to learn to capture general representations related to different degradation patterns, image structures, repair primitives, or abstract contexts during training, providing basic elements for the generation of dynamic hints.
[0038] The spatialized hint generation module is used to fuse multiple basic hint vectors based on the input features of the decoder layer and the CLIP image features to generate a spatial dynamic hint map to adapt to image local features and heterogeneous degradation.
[0039] The CLIP image features can be obtained by encoding a specified image through the image encoder of the CLIP model. The specified image can be obtained by sampling the image to be restored to the resolution of the input features of the decoder layer. The CLIP model can be pre-trained on large-scale datasets such as ImageNet and COCO, learning general image-text alignment capabilities. Inputting the specified image into the image encoder of the CLIP model can obtain the CLIP image features after encoding the specified image output by the image encoder of the CLIP model. Here, in order to obtain the dense embedding of the specified image, the last pooling layer of the image encoder is removed, and then the CLIP image features are generated by inputting the specified image. The CLIP image features are an image embedding matrix, where each element represents the semantic information of the specified image at that position.
[0040] The gated prompt interaction module can fuse the fused features output by the fusion layer of the decoder module with the spatially dynamic prompt map generated by the spatialized prompt generation module based on the gated network to obtain the output features of the decoder layer. Here, the gated network can include a convolutional layer and an activation layer, and the activation layer uses the sigmoid function to implement the activation operation. Through the gated network, a gating mechanism can be introduced to dynamically adjust the influence intensity of the spatially dynamic prompt map on the fused features, thereby achieving more intelligent and more robust feature fusion.
[0041] The output layer of the blind image restoration network can obtain the restored image according to the output features of the last decoder layer and the image to be restored. Here, the output layer can include an output convolutional layer and a superimposing layer. The output convolutional layer can map the output features of the last decoder layer to the three-channel red (R), green (G), and blue (B) image space to obtain the residual image of the restored image. The superimposing layer can superimpose the residual image on the image to be restored to obtain the restored image. Here, the output convolutional layer can also be a depth convolutional layer, which can better extract features and is particularly suitable for detail reconstruction in image restoration tasks.
[0042] The activation function in the U-Net architecture adopted in the blind image restoration network is the GeLU function, which is a non-linear activation function that determines whether to activate neurons based on the probability distribution. The definition of the GeLU function is as follows: ; where GELU(x) is the dependent variable of the GeLU function, and x is the tensor input to the GeLU function. is the cumulative distribution function of the standard normal distribution. This GeLU function can enhance the non-linear ability of the blind image restoration network.
[0043] In the embodiment of the present invention, the spatially adaptive blind image restoration method based on CLIP enhancement first obtains the image to be restored; then uses a blind image restoration network to obtain the restored image corresponding to the image to be restored. This method uses a blind image restoration network, which can effectively utilize the powerful text-image understanding and association capabilities of the CLIP model without the need to know the specific degradation type and degree of the image to be restored in advance, and introduce its rich semantic and visual prior knowledge into the general blind image restoration task, so as to effectively restore the image to be restored that has suffered from single or mixed degradation such as blur, noise, rain, snow, haze, low light, etc. The blind image restoration network combines an encoder-decoder structure, and a prompt enhancement module is fused in the decoder. Through the spatialized prompt generation module, it can dynamically generate a spatial dynamic prompt map according to the local features of the image to be restored, providing differentiated and targeted guiding information for different regions of the image to be restored, so as to effectively process spatially heterogeneous degradation and better retain local details. At the same time, the gated prompt interaction module introduces a gating mechanism through a gating network, which can dynamically adjust the influence intensity of the spatial dynamic prompt map on the fused features, and then achieve more flexible and intelligent feature fusion and restoration control.
[0044] Based on the above embodiment, the basic prompt module is further configured to: Embed the CLIP text features into the multiple basic prompt vectors; Wherein, the CLIP text features are obtained by encoding a text prompt template based on the CLIP model.
[0045] Specifically, after generating each basic prompt vector, the basic prompt module can also embed the CLIP text features into multiple basic prompt vectors. The CLIP text features can be obtained by inputting the text prompt template into the text encoder of the CLIP model, and encoded and output by the text encoder. Here, the text prompt template can be set differently according to different types of the image to be restored. For example, if the type of the image to be restored is noise reduction / rain removal / blur removal, the corresponding text prompt templates can be set as "it is a part of noise / rainy / blurry image" respectively. The text prompt template can include one or more.
[0046] Here, in order to obtain the dense embedding of the text prompt template, the last pooling layer of the image encoder is removed, and then the CLIP text features are generated by inputting the text prompt template. The CLIP text features are a text embedding matrix, where each element represents the semantic information of the text prompt template at that position.
[0047] In the embodiment of the present invention, the introduction of CLIP text features can provide clear semantic guidance.
[0048] Based on the above embodiments, there are multiple text prompt templates, and the CLIP text features are obtained by fusing the encoding results of multiple text prompt templates; The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0049] Specifically, when there are multiple text prompt templates, such as ["This is a blurry region", "This part is affected by motion blur", "Blur exists in this area"], the ambiguity that may exist in a single text description can be reduced.
[0050] Each text prompt template input into the text encoder of the CLIP model can obtain an encoding result. By fusing the encoding results of each text prompt template, the CLIP text features can be obtained.
[0051] That is: ; Among them, is the CLIP text feature, is the encoding result of the i-th text prompt template, and n is the number of text prompt templates.
[0052] The degradation degree of each position in the image to be restored can be determined by the similarity between the CLIP image features and the CLIP text features. Here, the similarity can be the cosine similarity, which can effectively measure the direction consistency. The higher the similarity, the higher the degradation degree of the corresponding position. Finally, through The function normalizes the similarity to the probability space, which is expressed as follows: ; Among them, is the cosine similarity, is the CLIP image feature.
[0053] Based on the above embodiments, the spatialized prompt generation module is specifically used for: Based on the adaptive pooling layer, modify the spatial resolution of the CLIP image features, and based on the convolution operation, modify the number of channels of the CLIP image features to obtain the first feature matching the input features; Concatenate the first feature with the input features to obtain the first concatenated feature; Perform a convolution operation on the first concatenated feature and, based on the activation function, obtain the spatial attention weight map; Based on the spatial attention weight map, fuse the multiple basic prompt vectors to generate the spatial dynamic prompt map.
[0054] Specifically, as Figure 3 shown, the spatialized prompt generation module may include a splicing layer, a first convolutional layer, and a fusion unit. The splicing layer includes an adaptive pooling layer, a second convolutional layer, and a splicing unit. The adaptive pooling layer is used to modify the spatial resolution of the CLIP image features. The second convolutional layer can be a 1×1 convolutional layer, which is used to modify the number of channels of the CLIP image features to obtain a first feature that matches the input features of the decoder module.
[0055] After that, the first feature is spliced with the input features of the decoder module through the splicing unit to obtain a first spliced feature: ; where is the first spliced feature, is the input feature of the decoder module, the first feature.
[0056] After that, the first spliced feature is convolved through the first convolutional layer, and based on the activation function, a spatial attention weight map is obtained. Here, the activation function can be the softmax function. The spatial attention weight map obtained through the activation function provides a normalized attention distribution for each basic prompt vector at each spatial position.
[0057] Subsequently, the fusion unit uses the weight vectors at each spatial position in the spatial attention weight map to fuse each basic prompt vector to generate a spatial dynamic prompt map. Here, the fusion method can be weighted summation, and the obtained spatial dynamic prompt map contains different prompt information adaptively generated according to local features at different spatial positions.
[0058] Based on the above embodiments, the gated prompt interaction module is specifically used for: Splice the fusion feature with the spatial dynamic prompt map to obtain a second spliced feature; Based on the Transformer block, transform the second spliced feature to obtain a transformed feature; Based on the gated network, generate a gated map that matches the size of the transformed feature, and based on the gated map, perform element-wise weighted fusion on the transformed feature and the fusion feature to obtain the output feature.
[0059] Specifically, the gated prompt interaction module is responsible for integrating the spatial dynamic prompt map generated by the spatialized prompt generation module into the fusion feature obtained by the decoder module in a flexible and controlled manner . Its working process includes two key steps: First, splice the spatial dynamic prompt map and the fused feature to obtain a second spliced feature, and input the second spliced feature into the Transformer block to obtain a transformed feature.
[0060] Second, the gated prompt interaction module uses a gated network to generate a gated map G that matches the size of the transformed feature.
[0061] Finally, use the gated map G to perform element-wise weighted fusion on the transformed feature and the fused feature to obtain an output feature: ; where, is the output feature, is the transformed feature, is the fused feature.
[0062] The gated mechanism adopted by the gated prompt interaction module allows the blind image restoration network to adaptively determine whether to adopt more of the transformed features enhanced by prompt guidance or retain the fused feature at each spatial position and channel according to the spatial dynamic prompt map, so as to achieve more intelligent and robust feature fusion.
[0063] Based on the above embodiments, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image; The training sample set includes first degraded images of various degradation types.
[0064] Specifically, to avoid the pre-trained CLIP model being unable to accurately perceive the degradation degree of the image to be restored, the CLIP model can be specifically fine-tuned to make it more suitable for the modeling and degradation estimation of degraded images.
[0065] To enhance the CLIP model's understanding ability of low-quality images, especially various underlying degradation features, a training sample set including first degraded images of various degradation types (for example, different types of blur, different levels and distributions of noise, rain, snow, haze, low light, etc.) and their corresponding high-quality first clear images is constructed.
[0066] Meanwhile, the training sample set can also include degradation description texts associated with the first degraded image (e.g., "motion blur area", "Gaussian noise points", "clear edges", "rain streak occlusion", etc.). Text labels can be semi-automatically generated through templated statements such as "This is a clear image with sharp details." or "This image contains motion blur in some areas." and verified manually to ensure semantic consistency.
[0067] Using the training sample set, the LoRA (Low-Rank Adaptation) technique is used to perform a contrastive learning task to fine-tune the CLIP model. The goal of fine-tuning is to maximize the similarity between the embedding vectors of the matching first clear image and the clear description text, while minimizing the similarity between the embedding vectors of the mismatching first clear image and the degradation description text, thereby strengthening the CLIP model's perception ability of the features related to the contrastive learning task, making the embedding vectors of the first clear image match the corresponding clear description text, and at the same time matching the embedding vectors of the first blurred image with their corresponding blurred description texts. These two are regarded as positive sample pairs.
[0068] Similarly, negative sample pairs are constructed by establishing the correspondence between the embedding vectors of the first clear image and the blurred description text, and the correspondence between the embedding vectors of the first blurred image and the clear description text.
[0069] Next, the NT-Xent loss function in contrastive learning is used to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs, thereby strengthening the CLIP model's modeling ability in the contrastive learning task: ; where is the NT-Xent loss, is a control parameter used to control the sensitivity of the contrast, is the embedding vector of the first clear image, is the embedding vector of the clear description text, is the embedding vector of the k-th description text, and K is the number of description texts in the training sample set.
[0070] Since the CLIP model is a large-scale pre-trained model, the LoRA technique is used to perform lightweight fine-tuning on it, enabling it to better adapt to the feature modeling of blurred images while avoiding the computational overhead caused by large-scale parameter updates. Specifically, by freezing the original weights of the CLIP model, low-rank matrices are inserted at specific layers. Then, the trainable low-rank adapters are used to learn the information specific to the deblurring task, thus avoiding fine-tuning the entire CLIP model and reducing computational resource consumption. After training, the weights of the fine-tuned CLIP model are saved and used for subsequent image restoration tasks.
[0071] Based on the above embodiments, the blind image restoration network is trained based on the following steps: Obtain an image dataset, which includes second degraded images of various degradation types and the corresponding second clear images of the second degraded images; Input the second degraded images into the initial restoration network to obtain the output images of the initial restoration network; Based on the second clear images and the output images, calculate the content loss and the edge loss respectively; Based on the content loss and the edge loss, perform iterative training on the initial restoration network to obtain the blind image restoration network.
[0072] Specifically, in the process of training the initial restoration network to obtain the blind image restoration network, first, an image dataset can be obtained. This image dataset includes real second degraded images of various degradation types (such as different types of blur, different levels of noise, rain, snow, haze, low light, etc.) from datasets such as Rain100L, SOTS, CBSD68, GoPro, and LOL, and the corresponding second clear images of the second degraded images. In addition, the image dataset can also include synthetic second degraded images and second clear images to mix the synthetic images with the real images and improve the generalization of the blind image restoration network.
[0073] After that, it is also necessary to preprocess the second degraded images in the image dataset and then divide them into a training set and a validation set. Among them, the second clear images and the corresponding second degraded images in the training set appear in pairs as a set of data.
[0074] Input the second degraded images into the initial restoration network to obtain the output images of the initial restoration network. Using the second clear images and the output images, calculate the content loss and the edge loss respectively.
[0075] The content loss can be calculated by the following content loss function: ; Where is the content loss, is the second clear image, is the output image. By calculating the L 2 norm difference between the two, and adding a small constant to stabilize the value.
[0076] The edge loss can be calculated by the following edge loss function: ; where is the edge loss, is the second clear image, is the output image. is the fast Fourier transform. By calculating the L 2 norm difference between the two in the frequency domain, and adding a small constant to stabilize the value.
[0077] After that, the total training loss can be obtained by weighted summation of the content loss and the edge loss: ; where .
[0078] Using the total training loss, calculate the error and generate the gradient, and update the model parameters of the initial restoration network by using the gradient backpropagation and gradient descent methods, and perform iterative training on the initial restoration network to make the initial restoration network converge to achieve the ideal clear image output effect.
[0079] As Figure 4 shown, on the basis of the above embodiments, an apparatus for spatially adaptive blind image restoration based on CLIP enhancement is provided in an embodiment of the present invention, including: An image acquisition module 41 for acquiring an image to be restored; An image restoration module 42 for obtaining a restored image corresponding to the image to be restored based on a blind image restoration network; wherein, the blind image restoration network includes a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer includes a decoder module and a prompt enhancement module, and a plurality of the encoder layers are skip-connected to the decoder module in the plurality of decoder layers; The encoder layer is used to extract the image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized prompt generation module is used to fuse the multiple basic prompt vectors based on the input features and CLIP image features to generate a spatial dynamic prompt map; the CLIP image features are obtained by encoding a specified image based on the CLIP model, and the specified image is obtained by sampling the image to be restored to the resolution of the input features; The gated prompt interaction module is used to fuse the fused features with the spatial dynamic prompt map based on a gated network to obtain the output features of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
[0080] Based on the above embodiments, the basic prompt module is further used to: Embed the CLIP text features into the multiple basic prompt vectors; Wherein, the CLIP text features are obtained by encoding a text prompt template based on the CLIP model.
[0081] Based on the above embodiments, there are multiple text prompt templates, and the CLIP text features are obtained by fusing the encoding results of the multiple text prompt templates; The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image features and the CLIP text features.
[0082] Based on the above embodiments, the spatialized prompt generation module is specifically used to: Based on an adaptive pooling layer, modify the spatial resolution of the CLIP image features, and based on a convolution operation, modify the number of channels of the CLIP image features to obtain a first feature that matches the input features; Concatenate the first feature with the input features to obtain a first concatenated feature; Perform a convolution operation on the first concatenated feature and, based on an activation function, obtain a spatial attention weight map; Based on the spatial attention weight map, fuse the multiple basic prompt vectors to generate the spatial dynamic prompt map.
[0083] Based on the above embodiments, the gated prompt interaction module is specifically used to: Concatenate the fused features with the spatial dynamic prompt map to obtain a second concatenated feature; Based on a Transformer block, transform the second concatenated feature to obtain a transformed feature; Based on the gating network, a gating graph matching the size of the transformed feature is generated, and based on the gating graph, the transformed feature and the fused feature are fused element-wise to obtain the output feature.
[0084] Based on the above embodiments, the CLIP model is trained based on the first degraded image in the training sample set, the first clear image corresponding to the first degraded image, the degraded description text associated with the first degraded image, and the clear description text associated with the first clear image; The training sample set includes first degraded images of multiple degradation types.
[0085] Based on the above embodiments, a training module is further included, which is used for: Obtain an image data set, where the image data set includes second degraded images of multiple degradation types and the second clear images corresponding to the second degraded images; Input the second degraded image into the initial restoration network to obtain the output image of the initial restoration network; Based on the second clear image and the output image, calculate the content loss and the edge loss respectively; Based on the content loss and the edge loss, perform iterative training on the initial restoration network to obtain the blind image restoration network.
[0086] Specifically, in the embodiments of the present invention, the functions of the modules in the CLIP-enhanced spatially adaptive blind image restoration device are in one-to-one correspondence with the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.
[0087] Figure 5 An example of the physical structure diagram of an electronic device is shown in Figure 5 As shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the CLIP-enhanced spatially adaptive blind image restoration method provided in the above embodiments.
[0088] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0089] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for CLIP-enhanced spatially adaptive blind image restoration provided in the above-mentioned various embodiments.
[0090] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for CLIP-enhanced spatially adaptive blind image restoration provided in the above-mentioned various embodiments.
[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A spatially adaptive blind image restoration method based on CLIP enhancement, characterized in that: include: Obtaining the image to be restored; Based on a blind image restoration network, a restored image corresponding to the image to be restored is obtained; The blind image restoration network comprises a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer comprises a decoder module and a prompt enhancement module, and the plurality of encoder layers are jump-connected with the decoder modules in the plurality of decoder layers; The encoder layer is used to extract image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized hint generation module is used to fuse the multiple basic hint vectors based on the input features and CLIP image features to generate a spatial dynamic hint map; the CLIP image features are obtained by encoding a specified image based on a CLIP model, and the specified image is obtained by sampling the image to be restored to a resolution equal to that of the input features; The gated prompt interaction module is used to fuse the fusion feature with the spatial dynamic prompt map based on the gated network to obtain the output feature of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
2. The spatially adaptive blind image restoration method based on CLIP enhancement according to claim 1, characterized in that: The basic prompt module is also used for: embedding CLIP text features into the plurality of basis cue vectors; The CLIP text feature is obtained by encoding a text prompt template based on the CLIP model.
3. The spatially adaptive blind image restoration method based on CLIP enhancement according to claim 2, characterized in that: The text prompt templates include multiple ones, and the CLIP text feature is obtained by fusing the encoding results of the multiple text prompt templates; The degradation degree of each position in the image to be restored is determined based on the similarity between the CLIP image feature and the CLIP text feature.
4. The spatially adaptive blind image restoration method based on CLIP enhancement according to claim 1, characterized in that: The spatialization prompt generation module is specifically used for: Based on the adaptive pooling layer, modify the spatial resolution of the CLIP image feature, and based on the convolution operation, modify the number of channels of the CLIP image feature to obtain a first feature that matches the input feature; Concatenate the first feature with the input feature to obtain a first concatenated feature; Performing a convolution operation on the first concatenated features, and obtaining a spatial attention weight map based on an activation function; Based on the spatial attention weight map, the multiple basic prompt vectors are fused to generate the spatial dynamic prompt map.
5. The spatially adaptive blind image restoration method based on CLIP enhancement according to claim 1, characterized in that: The gate control prompt interaction module is specifically used for: Splicing the fusion feature with the spatial dynamic prompt map to obtain a second splicing feature; Based on the Transformer block, transform the second concatenated feature to obtain a transformed feature; Based on the gating network, a gating graph matching the size of the transformed feature is generated, and based on the gating graph, the transformed feature and the fused feature are weightedly fused element by element to obtain the output feature.
6. The CLIP-enhanced spatially adaptive blind image restoration method according to any one of claims 1 to 5, characterized in that: The CLIP model is trained based on a first degraded image in a training sample set, a first clear image corresponding to the first degraded image, a degradation description text associated with the first degraded image, and a clear description text associated with the first clear image; The training sample set includes first degraded images of multiple degradation types.
7. The CLIP-enhanced spatially adaptive blind image restoration method according to any one of claims 1 to 5, characterized in that: The blind image restoration network is trained based on the following steps: Acquire an image data set, wherein the image data set includes second degraded images of multiple degradation types and second clear images corresponding to the second degraded images; Inputting the second degraded image into an initial restoration network to obtain an output image of the initial restoration network; Based on the second clear image and the output image, respectively calculating a content loss and an edge loss; Based on the content loss and the edge loss, the initial restoration network is iteratively trained to obtain the blind image restoration network.
8. A spatially adaptive blind image restoration device based on CLIP enhancement, characterized in that: include: An image acquisition module, used for acquiring the image to be restored; An image restoration module, used for obtaining a restored image corresponding to the image to be restored based on a blind image restoration network; The blind image restoration network comprises a plurality of encoder layers, a plurality of decoder layers and an output layer connected in sequence, the decoder layer comprises a decoder module and a prompt enhancement module, and the plurality of encoder layers are jump-connected with the decoder modules in the plurality of decoder layers; The encoder layer is used to extract image coding features of the image to be restored; The decoder module is used to fuse the input features to obtain fused features; The prompt enhancement module includes a basic prompt module, a spatialized prompt generation module and a gated prompt interaction module; The basic prompt module is used to generate a plurality of basic prompt vectors; The spatialized hint generation module is used to fuse the multiple basic hint vectors based on the input features and CLIP image features to generate a spatial dynamic hint map; the CLIP image features are obtained by encoding a specified image based on a CLIP model, and the specified image is obtained by sampling the image to be restored to a resolution equal to that of the input features; The gated prompt interaction module is used to fuse the fusion feature with the spatial dynamic prompt map based on the gated network to obtain the output feature of the decoder layer; The output layer is used to obtain the restored image based on the output features of the last decoder layer and the image to be restored.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the spatially adaptive blind image restoration method based on CLIP enhancement as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the spatially adaptive blind image restoration method based on CLIP enhancement as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text feature-introduced image restoration method
CN117314778A
Text generation and information recovery method and system based on large model
CN119515702A
Underwater image enhancement method and system based on multi-modal degradation feature learning
CN119672509A
Image generation method and apparatus, and electronic device and storage medium
WO2025039985A1
Cited By
Image restoration method
CN121120452A
Unified image restoration method and system based on task-driven semantic adjustment and color gamut correction
CN121169758A
Lightweight general image restoration method based on frequency domain gating and space-frequency domain fusion
CN122222881A
Lightweight general image restoration method based on frequency domain gating and space-frequency domain fusion
CN122222881B