Compressed image enhancement method based on Prompt
By embedding Prompt-based Transformer module in the U-Net framework, the encoder and decoder network is built, and the problem of insufficient image recovery flexibility caused by unknown or changes in JPEG quality factors in the prior art is solved, and an adaptive image recovery effect is achieved.
Patent Information
- Application Number
- CN202510136026.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing deep learning image recovery methods rely on predefined JPEG quality factors and cannot adapt to the compression strategies that vary under different devices and conditions, resulting in insufficient flexibility and universality in practical applications.
The Prompt-based Transformer module is used as the decoder and is embedded in the U-Net framework to build an encoder and decoder network, and images of different compression levels are processed through a multi-head self-attention mechanism and a feedforward network, and image recovery is guided by learningable propt.
The adaptive recovery of image quality under unknown or changing compression factors is achieved, eliminating the parameter overhead of additional modules and improving the flexibility and effect of image recovery.
Smart Images

Figure CN120259105A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a Prompt-based compressed image enhancement method. Background Art
[0002] The rapid growth of digital images and videos has made it necessary to utilize lossy compression techniques to optimize storage and bandwidth utilization. Among popular compression codecs, JPEG is favored for its computational efficiency and simple implementation. It compresses images by dividing them into 8×8 blocks, applying the discrete cosine transform, and quantifying the coefficients. However, the compressed images have compression artifacts such as ringing, blocking, and blurring effects, thus reducing the quality of the user experience.
[0003] With the rapid development of deep learning technology, deep neural networks have demonstrated powerful capabilities in various image processing tasks, especially in the field of image restoration. In recent years, researchers have proposed a variety of deep neural network-based image restoration methods, aiming to repair and improve the quality of compressed images by learning to extract features and patterns from large amounts of data. Deep learning methods can, to a certain extent, reduce the compression artifacts of images, such as removing the ringing effect, reducing the blocking effect, and improving the clarity of images. However, most existing deep learning restoration methods still rely on predefined JPEG quality factors, which limits their flexibility and universality in practical applications. In many practical scenarios, the JPEG quality factor is often unknown or variable. For example, different devices and applications may use different quality factors for image compression, the quality factor in real-time video streams will be dynamically adjusted according to network conditions, and the compression of uploaded images on online platforms also varies according to different conditions. Therefore, restoration methods with fixed quality factors cannot cope with these variable compression strategies and complex practical conditions. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a Prompt-based compressed image enhancement method. The present invention aims to solve the problem of how to process compressed images with unknown quality factors through Prompt and effectively restore images from different compression levels.
[0005] A Prompt-based compressed image enhancement method includes the following steps:
[0006] Step 1: Dataset acquisition and preprocessing.
[0007] Step 2: Construct a Prompt-based compressed image enhancement network.
[0008] The compression image enhancement network described above includes an encoder and a decoder. The convolutional neural network ResNet is used as the encoder to extract multi-scale feature information from the compressed image, and then it is sent to the Transformer module for feature processing and enhancement. Finally, the Transformer module based on prompts (TPB) is used as the decoder to restore the size of the feature map and obtain the final enhanced image.
[0009] Step 3: Train the compression image enhancement network based on the preprocessed dataset.
[0010] Input the preprocessed training set images into the compression image enhancement network. After passing through the encoder and decoder, enhanced images are generated. Then, use the L1 loss function to calculate the loss with the GT, perform backpropagation, and optimize the weights through the selected optimizer and corresponding parameters. After training for multiple rounds, the final network model is obtained.
[0011] Step 4: Testing. Input the test images into the trained network model to obtain enhanced images, compare them with the GT images, and calculate the PSNR.
[0012] The above steps are specifically described as follows:
[0013] Furthermore, Step 1 is specifically as follows:
[0014] Based on the DF2K dataset (800 images from DIV2K and 2650 images from Flickr2K) and the LSDIR dataset, a training set is formed. Use the PIL toolkit to generate compressed training images with 7 quality factors, where the factors range from 10 to 70 with a step of 10. Data augmentation is performed through random flipping and rotation. The test images are randomly generated compressed images of high-quality images using the PIL toolkit, and the compression quality factors range from 10 to 40.
[0015] Furthermore, the specific method of Step 2 is as follows:
[0016] The compression image enhancement network described above includes an encoder and a decoder. The input image first passes through a 3x3 convolution and then through an encoder. The encoder consists of 1 ResNet module and N Transformer modules, and the number of Transformer modules is selected according to requirements. At the bottom layer of the encoder, the convolutional neural network ResNet extracts feature information from the compressed image and then sends it to the N Transformer modules for feature processing and enhancement. To effectively restore the image from different compression levels, N + 1 Transformer modules based on prompts (TPB) and 1 ResNet module are used as the decoder to restore the size of the feature map.
[0017] The structure of the described Prompt-based Transformer Block (TPB) is as follows Figure 2 shown, which consists of a multi-head self-attention mechanism network and a feed-forward network. For the input feature F ∈ R W×H×C , after passing through the Norm layer, it is element-wise multiplied with the learnable prompt, and then 1×1 convolution and 3×3 depth convolution are applied to aggregate channel context to obtain the Q, K, V matrices. Next, from Q ∈ R H×W×C , K ∈ R H×W×C , V ∈ R H ×W×C they are rearranged to obtain new matrices where h is the number of projection heads. Then, learnable prompts are introduced into the matrices:
[0018]
[0019]
[0020] where P Q , P K , P V are the prompts for the generated query matrix, key matrix, and value matrix respectively. Then, the three newly generated matrices are fed into the multi-head attention mechanism operation to calculate the attention among them:
[0021]
[0022] where is the value of the obtained attention. Then, the obtained features are fed into a feed-forward network, whose structure is as Figure 2 shown.
[0023] The features obtained by the encoder are finally restored in terms of the feature map size using N + 1 Prompt-based Transformer blocks and a ResNet block as the decoder, and finally, a 3x3 convolution is performed to obtain the final enhanced image.
[0024] Furthermore, the specific method of step 3 is as follows:
[0025] The images in the training set are input into the compressed image enhancement network. The present invention adopts a two-stage training strategy to optimize the compressed image enhancement network. Specifically, in the first stage, the compressed image enhancement network is pre-trained on 7 quality factors in the dataset. In the second stage, the model obtained from the first-stage training is fine-tuned using compressed images with quality factors randomly selected from [10, 70]. At this stage, the model pays more attention to the distortion-aware information encoding while still retaining the ability to extract content-aware information through Prompts. During the entire training process, the training image pairs are cropped into 128×128 image patches. In the first stage, the AdamW optimizer is used to optimize the model, with an initial learning rate of 2e-4. The cosine annealing scheduler is used to decay the learning rate to 1e-6, and the total number of iterations is set to 800k. In the second stage, the same strategy is adopted, the initial learning rate is modified to 1e-4, and the iteration rate is modified to 600k. The L1 loss function is used to calculate the loss with the GT.
[0026] Furthermore, in step 4, images compressed with any compression quality factor between [0, 40] are used as test images for testing. The compressed image enhancement network implicitly encodes the compressed information using Prompts, dynamically perceives the distortion of the images through Prompts, and provides guidance for image restoration, thereby adapting to test images with different compression levels of the input.
[0027] The beneficial effects of the present invention are as follows:
[0028] The present invention uses a Prompt-based Transformer module as the decoder and embeds it into the U-Net framework to process images with different compression levels. The network proposed by the present invention can guide the decoder to perceive different compression levels, can adaptively process images with different compression levels through one model, eliminating the need for any additional modules. The proposed Prompt-based perception method does not result in a significant parameter overhead. Description of the Drawings
[0029] Figure 1 It is a structural diagram of a Prompt-based compressed image enhancement network;
[0030] Figure 2 It is a structural diagram of a Prompt-based Transformer module (TPB);
[0031] Figure 3 It is a graph of test results of images with different compression levels. Embodiment
[0033] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.
[0034] A Prompt-based compressed image enhancement method, including the following steps:
[0035] Step 1: Preprocessing of the dataset.
[0036] Based on the DF2K dataset (800 images from DIV2K and 2650 images from Flickr2K) and the LSDIR dataset, a training set is formed. Use the PIL toolkit to generate compressed training images with 7 quality factors, ranging from 10 to 70 with a step of 10. Data augmentation is performed by random flipping and rotation. The test images are randomly generated compressed images of high-quality images using the PIL toolkit, with compression quality factors ranging from 10 to 40.
[0037] Step 2: Construct a Prompt-based compressed image enhancement network.
[0038] The described compressed image enhancement network consists of two parts: an encoder and a decoder. The network structure of this example is as Figure 1 shown. The input image first passes through a 3x3 convolution and then through an encoder. The encoder consists of 1 ResNet module and 2 Transformer modules. To effectively recover images from different compression levels, this example uses 3 Prompt-based Transformer modules (TPB) and 1 ResNet module as the decoder to restore the feature map size.
[0039] The structure of the described Prompt-based Transformer module (TPB) is as Figure 2 shown, and it consists of a multi-head self-attention mechanism network and a feed-forward network. For the input feature F ∈ R W×H×C , after passing through the Norm layer, it is element-wise multiplied with the learnable prompt, and then 1×1 convolution and 3×3 depth convolution are applied to aggregate channel context to obtain the Q, K, V matrices. Next, from Q ∈ R H×W×C , K ∈ R H×W×C , V ∈ R H ×W×C Rearrange to get a new matrix where h is the number of projection heads. Then, introduce the learnable prompt into the matrix:
[0040]
[0041] where, P Q , P K , P VThey are the prompts for the generated query matrix, key matrix, and value matrix respectively. Then, the three newly generated matrices are fed into the multi-head attention mechanism operation to calculate the attention among them:
[0042]
[0043] Among them are the values of the obtained attention. Then, the obtained features are fed into a feed-forward network, and its structure is as Figure 2 shown.
[0044] The features obtained by the encoder are finally restored to the feature map size using 3 prompt-based Transformer modules and a ResNet module as the decoder, and finally, a 3x3 convolution is performed to obtain the final enhanced image.
[0045] Step 3: Train the network.
[0046] Input the images in the training set into the compressed image enhancement network. The present invention adopts a two-stage training strategy to optimize the compressed image enhancement network. Specifically, in the first stage, the compressed image enhancement network is pre-trained on 7 quality factors in the dataset. In the second stage, the model trained in the first stage is fine-tuned using compressed images with randomly selected quality factors from [10, 70]. In this stage, the model pays more attention to the distortion-aware information encoding while still retaining the ability to extract content-aware information through Prompts. During the entire training process, the training image pairs are cropped into 128×128 image patches. In the first stage, the AdamW optimizer is used to optimize the model, and the initial learning rate is 2e-4. The cosine annealing scheduler is used to decay the learning rate to 1e-6, and the total number of iterations is set to 800k. In the second stage, the same strategy is adopted, the initial learning rate is modified to 1e-4, and the iteration rate is modified to 600k. The L1 loss function is used to calculate the loss with the GT.
[0047] Step 4: Test.
[0048] In actual application scenarios, the JPEG quality factor is often unknown or variable, and different devices and applications may use different quality factors for image compression. As Figure 3 shown, the test image is an image compressed by any quality factor between [0, 40]. Input this image into the trained network. The compressed image enhancement network implicitly encodes the compression information of the input picture using the prompt, dynamically perceives the distortion of the image through the prompt, and provides guidance for image restoration, so as to adapt to the test images with different compression levels of the input. Figure 3The first row shows images compressed with different quality factors, which are (30, 20, 10, 40) from left to right. They have different degrees of information loss. The second row shows the corresponding restored images. It can be observed that our model can process images compressed with different quality factors, restore texture details, and will not produce blurring.
[0049] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.
[0050] The parts not detailed in the present invention belong to the well-known technologies in the art.
Claims
1. A Prompt-based compressed image enhancement method, characterized in that, The steps are as follows: Step 1: Dataset acquisition and preprocessing; Step 2: Construct a Prompt-based compressed image enhancement network; The compressed image enhancement network includes an encoder and a decoder. The convolutional neural network ResNet is used as the encoder to extract multi-scale feature information from the compressed image, and then it is sent to the Transformer module for feature processing and enhancement. Finally, the prompt-based Transformer module is used as the decoder to restore the size of the feature map to obtain the final enhanced image; Step 3: Train the compressed image enhancement network based on the preprocessed dataset; Input the preprocessed training set images into the compressed image enhancement network, and generate enhanced images through the encoder and decoder; Then use the L1 loss function to calculate the loss with the GT, perform backpropagation, and optimize the weights through the selected optimizer and corresponding parameters; After multiple rounds of training, obtain the final network model; Step 4: Testing; Input the test image into the trained network model to obtain the enhanced image, compare it with the GT image, and calculate the PSNR.
2. The method for enhancing a compressed image based on Prompt according to claim 1, wherein Step 1 is specifically as follows: Form a training set based on the DF2K dataset and the LSDIR dataset, and use the PIL toolkit to generate compressed training images with 7 quality factors, where the factors range from 10 to 70 with a step of 10; Perform data augmentation by random flipping and rotation; The test images are randomly generated compressed images of high-quality images using the PIL toolkit, and the compression quality factors range from 10 to 40.
3. The method for enhancing a compressed image based on Prompt according to claim 1, characterized in that, The specific method of Step 2 is as follows: The compressed image enhancement network includes an encoder and a decoder. The input image first passes through a 3x3 convolution and then through an encoder. The encoder consists of 1 ResNet module and N Transformer modules, and the number of Transformer modules is selected according to requirements. At the bottom layer of the encoder, the convolutional neural network ResNet extracts feature information from the compressed image and then sends it to N Transformer modules for feature processing and enhancement. To effectively restore the image from different compression levels, N+1 prompt-based Transformer modules and 1 ResNet module are used as the decoder to restore the size of the feature map; The described prompt-based Transformer module consists of a multi-head self-attention mechanism network and a feed-forward network; for the input feature F ∈ R W×H×C , after passing through the Norm layer, it is element-wise multiplied with the learnable prompt, and then 1×1 convolution and 3×3 depth convolution are applied to aggregate channel context to obtain the Q, K, V matrices. Next, from Q ∈ R H×W×C , K ∈ R H×W×C , V ∈ R H×W×C They are rearranged to obtain new matrices where h is the number of projection heads; then, a learnable prompt is introduced into the matrix: Among them, P Q , P K , P V are the prompts of the query matrix, key matrix, and value matrix generated respectively. Then, the three newly generated matrices are fed into the multi-head attention mechanism operation to calculate the attention among them: wherein is the obtained attention value; then, the obtained features are fed into a feed-forward network; The features obtained by the encoder are finally restored by using N+1 prompt-based Transformer modules and a ResNet module as the decoder for the size of the feature map, and finally pass through a 3x3 convolution to obtain the final enhanced image.
4. The method for enhancing a compressed image based on Prompt according to claim 3, wherein, The specific method of Step 3 is as follows: Input the images in the training set into the compressed image enhancement network, and adopt a two-stage training strategy to optimize the compressed image enhancement network. Specifically, in the first stage, the compressed image enhancement network is pre-trained on 7 quality factors in the dataset; In the second stage, use the compressed images with randomly selected quality factors from [10,70] to fine-tune the model obtained in the first stage of training.
5. A method for enhancing a compressed image based on Prompt according to claim 4, characterized in that, During the entire training process, the training image pairs are cropped into 128×128 image patches; in the first stage, the AdamW optimizer is used to optimize the model, and the initial learning rate is 2e-4; The cosine annealing scheduler is used to decay the learning rate to 1e-6, and the total number of iterations is set to 800k; in the second stage, the same strategy is adopted, the initial learning rate is modified to 1e-4, and the iteration rate is modified to 600k, and the L1 loss function is used to calculate the loss with the GT.
6. A method for enhancing compressed images based on Prompt according to claim 4 or 5, characterized in that In step 4, the images compressed by any compression quality factor between [0, 40] are used as test images for testing. The compressed image enhancement network implicitly encodes the compression information using the hint, dynamically perceives the distortion of the image through the hint, and provides guidance for image restoration, so as to adapt to the test images of different compression levels of the input.