JPEG image deblocking method based on lookup table

Through the lookup table-based method, spatial convolution decoupling and point-by-point convolution decoupling are used to design adaptive index range adjustment factors, and a lightweight neural network is built, which solves the problem of large amount and poor quality of image deblocking, and achieves efficient image deblocking effect.

CN120355602APending Publication Date: 2025-07-22NORTHWEST ELECTROMECHANICAL ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510502233.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has a large amount of calculation and huge amount of parameters in image deblocking, making it difficult to apply to hardware platforms and edge devices, and the image quality recovered by traditional methods is poor.

Method used

Using a lookup table-based method, multiple small LUTs are generated through spatial convolution decoupling, combined with point-by-point convolution of parameter sharing, adaptive index range adjustment factors are designed, and the calculation amount is reduced through full sampling to build a lightweight neural network model.

Benefits of technology

It effectively solves the problem of LUT storage index explosion, expands the receptive field and channel count, reduces the calculation amount, improves image quality, and is suitable for hardware platforms and edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355602A_ABST
    Figure CN120355602A_ABST
Patent Text Reader

Abstract

The invention discloses a JPEG (Joint Photographic Experts Group) image deblocking method based on a lookup table, which comprises the following steps of: constructing DW and PW convolution decoupling units by referring to the idea of convolution decoupling, and splitting a large LUT (Look Up Table) of each layer of convolution into a plurality of small LUTs; designing a self-adaptive index range adjustment factor, and adding the self-adaptive index range adjustment factor as a weight into network training; building a lightweight neural network model by combining a mask layer, a ReLU activation layer and a DW and PW convolution decoupling unit structure as basic modules; calculating an output value of the trained neural network model in the adjusted index range through enumerated input, and mapping the output value to an LUT after uniform sampling; fine tuning is performed on the LUT in the training table look-up process; and carrying out full sampling on the LUT again, directly querying a corresponding deblocking image output value according to the size and the shape of the receptive field and the input index value of the original JPEG image, and carrying out arrangement to obtain a final deblocking image. The method improves the problems of high calculation amount and insufficient performance in the prior art, and meets the requirements of practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing and relates to a JPEG image deblocking method based on a lookup table. Background Art

[0002] Image Deblocking technology is mainly used to improve the blocking artifacts generated during the image compression process. It is widely applied in the fields of static image and video compression, especially in scenarios such as real-time video transmission, video conferencing, and high-definition television, where it has important practical value. In digital image and video coding, compression algorithms (such as JPEG, MPEG, etc.) are often used to reduce the file size and improve the storage and transmission efficiency. However, these compression algorithms process the image by dividing it into small blocks, usually resulting in obvious boundaries between the blocks, leading to a decrease in image quality and forming a visual defect called "blocking artifacts". This blocking artifact usually appears as obvious artifacts at the image edges, especially more significantly at high compression ratios.

[0003] Deblocking technology aims to reduce or eliminate these blocking artifacts through different algorithms to improve the image quality. Traditional deblocking methods mainly perform spatial domain processing, such as smoothing the block edges in the image, interpolation, or using filters to remove the sharp changes at the boundaries, such as based on Shape-Adaptive Discrete Cosine Transform (SA-DCT), but the restored image quality is poor. In recent years, with the development of deep learning, deblocking methods based on Convolutional Neural Networks (CNNs) have gradually emerged. Yu et al. proposed a deep convolutional network for compression artifact reduction (ARCNN), which can automatically identify and remove blocking artifacts by training the network, providing better deblocking effects and higher visual quality than traditional methods. However, its computational complexity and the number of parameters are huge, making it difficult to be applied to hardware platforms and edge devices.

[0004] As a storage array, the Lookup Table (LUT) has emerged in the field of image super-resolution (SR) as an alternative to traditional deep learning methods. Its look-up mapping operation replaces a large amount of calculations on the GPU, facilitating application deployment and having a faster speed. However, there is a problem of storage exponential explosion, and the application effects in other image processing directions still need to be studied. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention proposes a JPEG image deblocking method based on a lookup table. First, spatial convolution (DW) is used to decouple and generate multiple LUTs to expand the receptive field and extract spatial domain features. Then, pointwise convolution (PW) with shared parameters is used to generate multiple LUTs to achieve channel convolution decoupling. Next, the size of the complete LUT is reduced by adjusting the adaptive index range. Finally, the computational complexity of the inference process is reduced through full-sampling LUT.

[0006] A JPEG image deblocking method based on a lookup table, which comprises the following steps:

[0007] Step 1: Draw on the idea of convolutional decoupling, build DW and PW convolutional decoupling units, and split the large LUT of each layer of convolution into multiple small LUTs;

[0008] Step 2: Design an adaptive index range adjustment factor and add it as a weight to network training;

[0009] Step 3: Combine the mask layer, ReLU activation layer with the DW and PW convolutional decoupling unit structures in Step 1 as basic modules to build a lightweight neural network model;

[0010] Step 4: Calculate the output value through the enumerated input of the trained neural network model within the adjusted index range and map it to the uniformly sampled LUT;

[0011] Step 5: Fine-tune the LUT during the training look-up table process;

[0012] Step 6: Resample the LUT again and directly query the corresponding deblocked image output value according to the size, shape of the receptive field and the input index value of the original JPEG image, and arrange them to obtain the final deblocked image.

[0013] Preferably, the said Step 1 contains the following steps:

[0014] Step 1.1: The DW convolutional decoupling strategy of the present invention first uses point convolution to minimize the storage of a single LUT, and then uses multiple point convolutions to expand the receptive field size to improve network performance. When the receptive field is 2×2, 4 point convolutions are used to replace the 2×2 standard convolution, and 4 one-dimensional LUTs of the same size are generated through network model mapping, and the average value of these 4 LUTs is taken and the final output value is deduced. Let the input of the DW convolutional decoupling unit be x and the output be F DW , and the specific calculation process is as follows:

[0015]

[0016] In the formula, K w and K h represent the length and width of the convolution, i and j represent the coordinate offsets of the DW convolution, and w is the weight of the convolution.

[0017] Step 2.2: Set the model weights with parameter sharing according to the output channel number of the previous layer DW convolutional decoupling unit, decouple the high-dimensional LUT between the channel dimensions into multiple one-dimensional LUTs, and obtain channel features at a very small storage cost. Finally, take the average to fit the feature information of each channel to improve network performance. Let the input of the PW convolutional decoupling be x and the output be FPW , the specific calculation process is as follows:

[0018]

[0019] In the formula, K c represents the number of channels, i represents the channel position, and w is the weight of the convolution.

[0020] Preferably, in order to solve the problem that the index range is not fully utilized, the present invention proposes an adaptive index range adjustment strategy, designs an adaptive index range adjustment factor, and adaptively updates the quantization range of each channel index value according to the gradient, and optimizes the balance relationship between image accuracy and LUT storage through CNN. This strategy can reduce the index range of [-128, 127] to [-128×α, 127×α], where α∈(0, 1), and at the same time, the storage of LUT also decreases with this factor. Let F idx be the adjusted feature data, and the specific quantization process is as follows:

[0021] F idx = round(F*α)

[0022] In the formula, F represents the feature data before adjustment, and α represents the adjustment factor.

[0023] The corresponding storage size S of each LUT LUT is calculated as follows:

[0024] S LUT = (max(F idx )) - min(F idx )) × R

[0025] In the formula, R represents the size of the receptive field.

[0026] Preferably, step 3 uses the basic units and training strategies of step 1 and step 2, combines the Relu activation function and standard convolution to build a lightweight convolutional neural network for JPEG image deblocking based on LUT. The 48×48 image blocks cropped from the DIV2K dataset are used as inputs, and the mean square error function (MSE) is used as the loss function. After training by the above convolutional neural network, the deblocked image is obtained.

[0027] Preferably, step 4 first enumerates all index cases of LUT according to the size and shape of the receptive field and considering the adjusted range of the index, then uniformly samples LUT (Sampled-LUT) with a fixed interval W, uses the sampled result as the input, and at the same time loads the weights of the CNN model trained in the previous step, and stores it as LUT after inference calculation.

[0028] Preferably, in step 5, the LUT fine-tuning method is adopted. Referring to the pre-trained strategy, the LUT itself is used as a trainable parameter for training, and the error introduced by quantization of the intermediate layer is restored through a neural network, effectively improving the performance of the image without changing the LUT storage and inference calculation amount. At the same time, through the vectorization method, the highly repetitive serial calculation is changed to parallel vector operation, and multiple LUTs of each layer are combined together. A new indexing mechanism is designed based on the receptive field of DW convolution and the channel dimension of PW convolution, which speeds up the model fine-tuning and inference speed under this decoupled architecture.

[0029] Preferably, in step 6, after analysis, the decoupling of PW convolution brings a certain amount of calculation for multi-channel interpolation. At the same time, considering that steps 1 and 2 have a significant effect on reducing the storage of LUT, a part of the storage can be sacrificed to reduce the sampling interval to reduce the interpolation calculation amount. Therefore, the present invention designs a full-sampling method to omit the interpolation calculation. Interpolate the output result obtained by Sampled-LUT, and store the calculation result as a new full-sampling table according to the one-to-one correspondence with the input index. That is to say, the result of interpolating each input index is stored in the LUT in advance, thus omitting the subsequent interpolation calculation. At this time, the inference result of the LUT should also be exactly the same as the interpolation calculation result. Subsequent experiments also prove that full-sampling will not cause a decrease in image metrics.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] First, the present invention first migrates the application of LUT in the field of image super-resolution to the field of image deblocking, designs DW convolution and PW convolution by referring to the idea of convolution decomposition, and converts the large LUT before decoupling into multiple small LUTs, effectively solving the problem of exponential explosion of LUT storage. At the same time, it expands the equivalent receptive field and the number of channels, and improves the accuracy of the image at the cost of linear growth of the minimum storage.

[0032] Secondly, the present invention proposes an adaptive index range adjustment strategy, introduces an adjustment factor α in the model training stage, and maps the complete index range to a smaller area to reduce the storage of the complete LUT.

[0033] Finally, based on adjusting the index range, the present invention designs an efficient full-sampling method, solves the problem of gradient clearing during model fine-tuning, omits the interpolation process brought by sampling, and further reduces the calculation amount of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flowchart of a JPEG image deblocking method based on a lookup table according to the present invention;

[0035] Figure 2It is a schematic diagram of decoupling DW and PW convolutions to generate multiple LUTs in the present invention;

[0036] Figure 3 It is a schematic diagram of the adaptive index range adjustment strategy in the present invention;

[0037] Figure 4 It is a schematic diagram of the efficient full-sampling LUT in the present invention. Detailed implementation manners

[0038] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0039] As Figures 1 to 4 shown, a JPEG image deblocking method based on a lookup table includes the following steps:

[0040] S1.1: The DW convolution decoupling strategy of the present invention first uses point convolution to minimize the storage of a single LUT, and then uses multiple point convolutions to expand the receptive field size to improve network performance. As Figure 2 shown, when the receptive field is 2×2, 4 point convolutions are used to replace the 2×2 standard convolution, and 4 one-dimensional LUTs of the same size are generated through network model mapping, and the average value of these 4 LUTs is taken and the final output value is deduced. Let the input of the DW convolution decoupling unit be x and the output be F DW , and the specific calculation process is as follows:

[0041]

[0042] In the formula, K w and K h represent the length and width of the convolution, i and j represent the coordinate offsets of the DW convolution, and w is the weight of the convolution.

[0043] S1.2: As Figure 2 shown, set the model weights with parameter sharing according to the output channel number of the previous layer DW convolution decoupling unit, decouple the high-dimensional LUT between channel dimensions into multiple one-dimensional LUTs, and obtain channel features at a very small storage cost. Finally, take the average to fit the feature information of each channel to improve network performance. Let the input of the PW convolution decoupling be x and the output be F PW , and the specific calculation process is as follows:

[0044]

[0045] In the formula, K c represents the number of channels, i represents the channel position, and w is the weight of the convolution.

[0046] S2: In the LUT-based image super-resolution method, to further reduce the storage of the LUT, equidistant sampling is usually performed on the index, and then different interpolation methods are used to restore the lost precision. However, to improve the image precision, after adopting the DW and PW convolution decoupling units, although the multiple LUTs brought by the large receptive field and multiple channels are the smallest in dimension, it is found through the histogram statistics of each LUT that the entire index range is not fully utilized, and there is still a certain reduction space, which also indicates that the storage of each LUT can be further reduced.

[0047] To solve this problem, the present invention proposes an adaptive index range adjustment strategy, designs an adaptive index range adjustment factor, and adaptively updates the quantization range of each channel index value according to the gradient, and optimizes the balance relationship between the image precision and the LUT storage through the CNN. As Figure 3 shown, this strategy can reduce the index range of [-128, 127] to [-128×α, 127×α], where α∈(0, 1), and the storage of the LUT also decreases with this factor. Let F idx be the adjusted feature data, and the specific quantization process is as follows:

[0048] F idx = round(F * α)

[0049] In the formula, F represents the feature data before adjustment, and α represents the adjustment factor.

[0050] The storage size S of each corresponding LUT LUT is calculated as follows:

[0051] S LUT = (max(F idx ) - min(F idx )) × R

[0052] In the formula, R represents the size of the receptive field.

[0053] S3: Use the basic units and training strategies of steps 1 and 2, combine the Relu activation function and the standard convolution to build a lightweight convolution neural network for JPEG image deblocking based on the LUT. Crop 48×48 image patches from the DIV2K dataset as the input, and use the mean square error function (MSE) as the loss function. After training by the above convolution neural network, the deblocked image is obtained.

[0054] S4: First enumerate all index cases of the LUT according to the size and shape of the receptive field and considering the adjusted range of the index, then use uniform sampling of the LUT (Sampled-LUT) with a fixed interval W, use the sampled result as the input, and at the same time load the weights of the CNN model trained in the previous step, and store it as the LUT after inference calculation.

[0055] S5: The method of fine-tuning the LUT is adopted. Referring to the pre-trained strategy, the LUT itself is used as a trainable parameter for training, and the error introduced by quantization in the intermediate layer is recovered through the neural network, effectively improving the image performance without changing the storage and inference calculation amount of the LUT. At the same time, through the vectorization method, the highly repetitive serial calculations are changed to parallel vector operations, and multiple LUTs of each layer are combined together. A new indexing mechanism is designed based on the receptive field of DW convolution and the channel dimension of PW convolution, accelerating the model fine-tuning and inference speed under this decoupled architecture.

[0056] S6: The decoupling of PW convolution brings a certain amount of calculation for multi-channel interpolation. At the same time, considering that steps 1 and 2 have an obvious effect on reducing the storage of the LUT, a part of the storage can be sacrificed to reduce the sampling interval to reduce the interpolation calculation amount. Therefore, the present invention designs a full-sampling method to discard the interpolation calculation. As Figure 4 shown, the output result obtained by Sampled-LUT is interpolated and calculated, and the calculation result is stored as a new full-sampling table according to the one-to-one correspondence relationship with the input index. That is to say, the result of interpolating and calculating each input index is stored in the LUT in advance, thus omitting the subsequent interpolation calculation. At this time, the inference result of the LUT should also be exactly the same as the interpolation calculation result, and subsequent experiments also prove that full-sampling will not cause a decrease in image metrics.

[0057] The effects of the present invention can be further illustrated by the following verification experiments on image performance metrics, storage size, and calculation amount.

[0058] Experimental environment:

[0059] Intel(R) Core(TM) i7-14700KF CPU + NVIDIA GeForce RTX 3090; the development tool is Python; the development framework is Pytorch, the batchsize is set to 16, the loss function is MSE, and the total number of training iterations is set to 200000. The peak signal-to-noise ratio PSNR and structural similarity SSIM are used as evaluation metrics for image quality. The SA-DCT, SRLUT, and MuLUT methods are used to perform JEPG image deblocking on the test set. The method proposed by the present invention is used to perform deblocking on the test set, where the receptive field of the first layer is 5×5 and the output channels of the intermediate layer are all 32.

[0060] Experimental data set:

[0061] The Classic5 and LIVE1 test sets are used as the test sets, and the DIV2K training set is used.

[0062] Experimental results and analysis:

[0063] Experiment 1 was conducted on the Classic5 dataset, including the present invention and other classic JPEG image deblocking methods, to obtain PSNR and SSIM values. The results show that, benefiting from the powerful fitting ability of the neural network, ARCNN achieved the best metrics with PSNR and SSIM of 29.04 and 0.8111; while the PSNR and SSIM of the method of the present invention were 29.01 and 0.8102, with a difference of less than 0.1 from the former, which was almost negligible in terms of subjective visual effects; and it was much higher than the results of SA-DCT, SRLUT, and MuLUT, with the PSNR being 0.24, 0.57, and 0.14 higher respectively, and the SSIM being 0.031, 0.127, and 0.01 higher respectively. Overall, the present invention demonstrated obvious advantages in the objective performance metrics of images.

[0064] Experiment 2 was conducted on the LIVE1 dataset, including the present invention and other classic JPEG image deblocking methods, to obtain PSNR and SSIM values. The results show that, benefiting from the powerful fitting ability of the neural network, ARCNN achieved the best metrics with PSNR and SSIM of 29.13 and 0.8232; while the PSNR and SSIM of the method of the present invention were 29.04 and 0.8217, with a difference of less than 0.1 from the former, which was almost negligible in terms of subjective visual effects; and it was much higher than the results of SA-DCT, SRLUT, and MuLUT, with the PSNR being 0.24, 0.57, and 0.14 higher respectively, and the SSIM being 0.124, 0.205, and 0.043 higher respectively. Overall, the present invention demonstrated obvious advantages in the objective performance metrics of images.

[0065] Experiment 3 was to process a visible light image with a resolution of 640×360, including the present invention and other classic JPEG image deblocking methods, to obtain the storage capacity and computational amount of the LUT. The results show that the present invention performed best with a storage capacity and computational amount of 74.4KB and 28.3M, indicating the smallest storage and computational resources required; the computational amount of ARCNN reached an astonishing 41.3G, proving the drawback of the huge computational amount of deep learning methods, and the storage capacity also reached 415KB. Generally speaking, it was difficult to be applied and deployed on a hardware platform; SRLUT also performed relatively well, with a storage capacity of 81KB and a computational amount of 30.9M, but Experiments 1 and 2 showed that its performance was poor. Overall, the present invention demonstrated obvious advantages in terms of storage capacity and computational amount.

[0066] In summary, the comprehensive experimental results on the general dataset have shown the superiority of the method proposed in the present invention compared with some similar super-resolution methods in terms of image performance, computational amount, and storage.

Claims

1. A JPEG image deblocking method based on a lookup table, characterized in that: It includes the following steps: Step 1: Drawing on the idea of convolutional decoupling, build DW and PW convolutional decoupling units, and split the large LUT of each layer of convolution into multiple small LUTs; Step 2: Design an adaptive index range adjustment factor and add it as a weight to network training; Step 3, Combine the mask layer, ReLU activation layer with the DW and PW convolutional decoupling unit structures in Step 1 as basic modules to build a lightweight neural network model; Step 4, Calculate the output value through the enumerated input of the trained neural network model within the adjusted index range and map it to the uniformly sampled LUT; Step 5, Train the look-up table process to fine-tune the LUT; Step 6, Resample the LUT again and directly query the corresponding deblocked image output value according to the size, shape of the receptive field and the input index value of the original JPEG image, and arrange them to obtain the final deblocked image; In Step 1, drawing on the idea of convolutional decoupling, build DW and PW convolutional decoupling units, and split the large LUT of each layer of convolution into multiple small LUTs; It contains the following steps: Step 1.1: DW convolution decoupling strategy. First, point convolution is used to minimize the storage of a single LUT, and then multiple point convolutions are used to expand the receptive field size to improve network performance. When the receptive field is 2×2, 4 point convolutions are used to replace the 2×2 standard convolution. After being mapped by the network model, 4 one-dimensional LUTs of the same size are generated, and the average value of these 4 LUTs is taken to infer the final output value. Let the input of the DW convolution decoupling unit be x and the output be F DW , and the specific calculation process is as follows: where K w and K h represent the length and width of the convolution, i and j represent the coordinate offsets of the DW convolution, and w is the weight of the convolution; Step 2.2: Set the model weights with parameter sharing according to the number of output channels of the previous layer DW convolution decoupling unit, decouple the high-dimensional LUT between channel dimensions into multiple one-dimensional LUTs, obtain channel features at a very small storage cost, and finally take the average to fit the feature information of each channel to improve the network performance; Let the input of the PW convolution decoupling be x and the output be F PW , and the specific calculation process is as follows: where K c represents the number of channels, i represents the channel position, and w is the weight of the convolution; In Step 2, to solve the problem that the index range is not fully utilized, an adaptive index range adjustment strategy is proposed, an adaptive index range adjustment factor is designed, and the quantization range of each channel index value is updated adaptively according to the gradient, and the balance between image accuracy and LUT storage is optimized through CNN; This strategy can narrow the index range of [-128, 127] to [-128×α, 127×α], where α∈(0, 1), and at the same time, the storage of the LUT also decreases with this factor; Let F idx be the adjusted feature data, and the specific quantization process is as follows: F idx = round(F * α) In the formula, F represents the feature data before adjustment, and α represents the adjustment factor; For each corresponding LUT storage size S LUT Calculate as follows: S LUT = (max(F idx )) - min(F idx )) × R In the formula, R represents the size of the receptive field; In Step 3, use the basic units and training strategies of Step 1 and Step 2, combine the Relu activation function and standard convolution to build a lightweight convolutional neural network for JPEG image deblocking based on LUT; Crop 48×48 image patches from the DIV2K dataset as input, and use the mean square error function (MSE) as the loss function. After training with the above convolutional neural network, obtain the deblocked image; In Step 4, first enumerate all index cases of the LUT according to the size and shape of the receptive field and considering the adjusted index range, then use uniform sampling of the LUT (Sampled-LUT) with a fixed interval W, use the sampled result as input, and load the weights of the CNN model trained in the previous step at the same time, and store it as the LUT after inference calculation; In Step 5, adopt the method of LUT fine-tuning, train the LUT itself as a trainable parameter with reference to the pre-training strategy, and recover the error introduced by quantization in the middle layer through the neural network, effectively improving the performance of the image without changing the LUT storage and inference calculation amount; At the same time, through vectorization, change the highly repetitive serial calculation to parallel vector operation, combine multiple LUTs of each layer together, and design a new index mechanism based on the receptive field of DW convolution and the channel dimension of PW convolution, which speeds up the model fine-tuning and inference speed under this decoupled architecture; As analyzed in step 6, PW convolution decoupling brings a certain amount of computation for multi-channel interpolation. At the same time, considering that steps 1 and 2 have a significant effect on reducing the storage of the LUT, sacrificing a part of the storage to narrow the sampling interval to reduce the interpolation computation amount. Therefore, the present invention designs a full-sampling method to eliminate the interpolation computation; interpolate the output result obtained from the Sampled-LUT, and store the computation result as a new full-sampling table according to the one-to-one correspondence with the input index. That is to say, the results of interpolating each input index are stored in the LUT in advance, thus omitting the subsequent interpolation computation. At this time, the inference result of the LUT is also completely consistent with the interpolation computation result. Subsequent experiments prove that full sampling will not cause a decrease in image metrics.