Efficient compressed image artifact removal method based on multi-feature fusion
Through the multi-feature fusion method, combining shallow and multi-scale feature extraction, channel and pixel position attention, the problem of insufficient network training complexity and flexibility in the existing methods is solved, and the removal of efficient compressed image artifacts and visual quality improvement is achieved.
Patent Information
- Application Number
- CN202510455530.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-22
AI Technical Summary
The existing deep learning-based compressed image artifact removal methods require training multiple specific networks for images of different compression qualities, and cannot effectively utilize compression quality information, resulting in insufficient processing flexibility and efficiency.
The multi-feature fusion method is adopted, including shallow feature extraction module, multi-scale feature extraction module, image recovery module and loss function. The image recovery network is optimized through recursive residual units, channel attention and pixel position attention modules, combined with the Adam optimizer for training.
It improves the visual quality of compressed images, enhances the model's ability to characterize complex artifacts, reduces the computational complexity, is suitable for deployment on mobile or edge devices, and improves training efficiency and model stability.
Smart Images

Figure CN120355587A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an efficient method for removing artifacts from compressed images based on multi-feature fusion. Background Art
[0002] Limited by transmission bandwidth and storage capacity, images or videos captured by cameras need to be compressed in practical applications. Currently, lossy compression methods represented by JPEG compression have been widely used in various aspects of image processing. However, due to the loss of high-frequency information of the image during the quantization stage, the compressed image often contains artifacts such as block effects, ringing effects, and blurring. These compression artifacts not only seriously affect the visual quality but also cause a decline in performance in subsequent computer vision tasks.
[0003] To solve the problem of artifacts generated in compressed images, existing solutions are mainly divided into methods based on traditional image processing and methods based on deep learning. Traditional image processing methods mostly rely on filtering algorithms to reduce the impact of artifacts by smoothing the image. They can achieve good results when processing some simple images, but in the case of complex image degradation, their effects are often limited. In addition, traditional methods usually rely on fixed rules and parameters, lacking sufficient adaptability and flexibility, and it is difficult to effectively cope with the diversity of different types of compression image artifacts. Therefore, traditional image processing methods have certain limitations in the task of removing compression image artifacts. With the rapid development of deep neural networks, methods for removing artifacts from compressed images based on deep learning have gradually become the mainstream. Deep learning methods can effectively handle complex image degradation problems through their powerful feature extraction ability and non-linear mapping ability, and show more superior performance than traditional methods in removing compression artifacts. Compared with traditional methods, deep learning methods can automatically learn the latent features in the image, and thus more effectively remove the artifacts in the compressed image and improve the image clarity.
[0004] Although significant achievements have been made in methods for removing artifacts from compressed images based on deep learning, there are still some problems. Many methods need to train multiple specific networks for images with different compression qualities, which limits the practicality of the methods. Although there are also some methods that use a single model to process images with different compression qualities, these methods often ignore the effective use of compression quality information and cannot accurately reflect the degree of image degradation. Therefore, there is an urgent need to propose a new method that can effectively remove the artifacts in the compressed image, improve the visual effect, and enhance the flexibility and efficiency of processing. Summary of the Invention
[0005] Aiming at the above technical problems existing in the existing deep learning-based methods for removing artifacts from compressed images, the present invention provides an efficient method for removing artifacts from compressed images based on multi-feature fusion for removing compression artifacts.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] An efficient compression image artifact removal method based on multi-feature fusion, comprising the following steps:
[0008] S1. Data preparation: Download the publicly available high-definition datasets Flickr2K and DIV2K, and generate the required training dataset I by compressing with the MATLAB JPEG encoder L ;
[0009] S2. Construct a shallow feature extraction module, which is mainly composed of 2 convolutional kernels of size 3x3, and convolutional layers with a stride of 1 and a padding of 1;
[0010] S3. Construct a multi-scale feature extraction module, which consists of a convolutional layer and a recursive residual unit RRU; among them, each recursive residual unit is composed of multiple residual blocks, and each residual block is composed of two convolutional layers;
[0011] S4. Construct an image restoration module, which consists of a feature fusion module FAM and a convolutional layer; among them, FAM is composed of a channel attention CAM module and a pixel position attention PAM module;
[0012] S5. Construct a loss function L;
[0013] S6. Move the constructed network model to the GPU for training, use the Adam optimizer to optimize the image restoration network, calculate the loss function and then backpropagate it into the network for optimization, stop training when the loss function converges, and save the model.pth weight file for subsequent test use after training is completed.
[0014] The method for constructing the shallow feature extraction module in S2 is as follows:
[0015] S2.1. Receive the compressed image I L ∈R C×H×W , usually a color image or a grayscale image, the number of channels C of the color image is 3, and the number of channels C of the grayscale image is 1;
[0016] S2.2. The input image passes through the convolutional layer to extract the initial low-level features, and finally obtains the shallow features F0.
[0017] The method for constructing the multi-scale feature extraction module in S3 is as follows:
[0018] S3.1. Receive the compressed image I L ∈R C×H×W, after passing through the convolutional layer, it enters multiple Recursive Residual Units (RRUs) for multi-scale feature extraction;
[0019] S3.2, After being processed by multiple Recursive Residual Units and passing through the final convolutional layer, the final multi-scale feature F1 is output.
[0020] The working process of each Recursive Residual Unit (RRU) in S3.1: The input feature is first activated by ReLU. The activated feature passes through the first convolutional layer to obtain an intermediate feature. The intermediate feature is then activated by ReLU and enters the second convolutional layer. The output of the second convolutional layer is connected with the input through residual connection, and finally the output of the current Recursive Residual Unit is obtained.
[0021] The method for constructing the image restoration module in S4 is as follows:
[0022] S4.1, First, concatenate the received shallow feature F0 and the multi-scale feature F1 to obtain feature F;
[0023] S4.2, In the Channel Attention Module (CAM): First, feature F passes through the average pooling layer and the max pooling layer in parallel, and then is restored to the original number of channels through the convolutional layer. The output of the channel attention is obtained through the Sigmoid activation function and multiplied by the original feature, and finally the channel attention feature F is obtained. CAM ;
[0024] S4.3, In the Pixel Attention Module (PAM): First, reshape feature F into a two-dimensional matrix, and then calculate the correlation between different pixel positions through matrix multiplication and transpose operations to obtain the correlation matrix ω. The reshaped feature is then multiplied by the correlation matrix ω to obtain the channel attention feature F. PAM ;
[0025] S4.4, For the obtained feature F CAM and F PAM , first sum them to obtain the input feature F in , and then feature F in passes through the global average pooling layer and the convolutional layer in parallel to obtain the global feature G(F in ) and the local feature L(F in );
[0026] S4.5, Input the global feature G(F in ) and the local feature L(F in ) into the feature fusion module to obtain the fused feature F out = F in ·σ(G(F in ) + L(F in )), and then F outAfter passing through another convolutional layer, the final artifact-removed image I is obtained. H ; where σ is the Sigmoid function.
[0027] The formula for obtaining the channel attention feature F in S4.2 is as follows: CAM The formula for obtaining the channel attention feature F in S4.2 is as follows:
[0028] F CAM = σ(f(f(Avg(F)) + f(Max(F)))) · F
[0029] where σ is the Sigmoid function, f(·) represents the convolution operation, Avg(·) represents average pooling, and Max(·) represents max pooling.
[0030] The formula for obtaining the channel attention feature F in S4.3 is as follows: PAM The formula for obtaining the channel attention feature F in S4.3 is as follows:
[0031]
[0032] where δ(·) represents the Softmax function, T represents the transpose operation, represents matrix multiplication.
[0033] The formula for obtaining the global feature G(F in ) and the local feature L(F in ) in S4.4 is as follows:
[0034] F in = F CAM + F PAM
[0035] G(F in ) = (f(f RELU (f(f GAP (F in ))))
[0036] L(F in ) = f(f RELU (F in ))
[0037] where f RELU (·) represents the RELU activation function, and f GAP (·) represents global average pooling.
[0038] The method for constructing the loss function L in S5 is as follows:
[0039]
[0040] where N is the number of training samples, is the artifact-removed image of the i-th sample, is the original true image of the i-th sample, and ||·||1 represents the L1 norm.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] 1. The shallow feature extraction module of the present invention adopts a double 3×3 convolutional layer structure, which can effectively extract low-level features while retaining local details of the image (such as edges and textures), avoiding the problem of detail loss caused by an overly deep network. The multi-scale feature extraction module of the present invention captures artifact features at different scales (such as blocking effect and ringing effect) through the cascaded design of recursive residual units (RRUs), combining residual connections and multi-scale feature reuse, enhancing the model's representation ability for complex artifacts. In addition, the recursive structure reduces the computational complexity through parameter sharing and improves the model efficiency.
[0043] 2. The channel attention module (CAM) of the present invention dynamically adjusts the feature importance of different channels through parallel average pooling and maximum pooling operations, combined with channel dimension weight learning, suppressing the interference of noise-related channels and enhancing the feature expression of key channels. The pixel position attention module (PAM) of the present invention captures long-range dependencies through the calculation of the pixel correlation matrix, strengthening the ability to distinguish between artifact regions and normal regions. This module can accurately locate the boundary of the blocking effect and effectively eliminate the spatial discontinuity of the artifact through feature reshaping and matrix multiplication operations.
[0044] 3. The recursive residual unit (RRU) of the present invention adopts a residual skip connection, alleviating the problem of gradient disappearance and accelerating model convergence. At the same time, the recursive structure reduces the number of parameters, reducing the computational resource consumption while ensuring performance, and is suitable for deployment on mobile or edge devices. The loss function of the present invention uses L1 norm constraint, which can better retain high-frequency details (such as sharp edges) compared with the traditional L2 norm, reducing the problem of image over-smoothing. Combined with the adaptive learning rate adjustment of the Adam optimizer, the training efficiency and model stability are further improved. Description of the Drawings
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and those of ordinary skill in the art can also obtain other implementation drawings according to the provided drawings without creative efforts.
[0046] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0047] Figure 1 is the overall framework diagram of the network of the present invention;
[0048] Figure 2 is the schematic diagram of the multi-scale feature extraction module of the present invention;
[0049] Figure 3 is the schematic diagram of the channel attention module of the present invention;
[0050] Figure 4 is the schematic diagram of the pixel position attention module of the present invention;
[0051] Figure 5 is the schematic diagram of the feature fusion module of the present invention. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0053] The following will further describe in detail the specific implementation manners of the present invention in combination with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0054] This embodiment provides an efficient compression image artifact removal method based on multi-feature fusion, as Figure 1 shown, including the following steps:
[0055] Step 1, data preparation: Obtain high-quality public datasets Flickr2K and DIV2K, and preprocess the images using standard JPEG compression in MATLAB to obtain preprocessed images I L ∈R C×H×W , where C represents the number of channels of the image, and H and W represent the height and width of the image respectively;
[0056] Step 2, construct a shallow feature extraction module:
[0057] The shallow feature extraction module is mainly composed of two convolutional kernels of size 3x3 and convolutional layers with a stride of 1 and a padding of 1;
[0058] Step 2.1, receive the compressed image I L ∈R C×H×W , usually a color image or a grayscale image. The number of channels C of a color image is 3, and the number of channels C of a grayscale image is 1;
[0059] Step 2.2, the input image passes through the convolutional layer to extract the initial low-level features, and finally the shallow features F0 are obtained, F0 = H SF (I L ), where H SF (·) is the shallow feature extraction module;
[0060] Step 3, as Figure 2 shown, construct a multi-scale feature extraction module:
[0061] The multi-scale feature extraction module consists of a convolutional layer and a recursive residual unit RRU. Among them, each recursive residual unit is composed of multiple residual blocks, and each residual block is composed of two convolutional layers;
[0062] Step 3.1, receive the compressed image I L ∈R C×H×W , after passing through the convolutional layer, enter multiple recursive residual units (RRUs) for multi-scale feature extraction. There are two convolutional layers inside each recursive residual unit, and each layer is first activated by ReLU after convolution;
[0063] Step 3.2, after being processed by multiple recursive residual units, and output the final multi-scale features F1 through the last convolutional layer, F1 = H MMF (I L ), where H MMF (·) is the multi-scale feature extraction module;
[0064] Step 4, construct an image restoration module:
[0065] The image restoration module consists of a feature fusion module (FAM) and a convolutional layer. Among them, FAM is composed of a channel attention (CAM) module and a pixel position attention (PAM) module;
[0066] Step 4.1, first concatenate the received shallow features F0 and multi-scale features F1 to obtain the feature F;
[0067] Step 4.2, as Figure 3As shown, in the channel attention (CAM) module: First, the feature F passes through the average pooling layer and the max pooling layer in parallel, and then is restored to the original number of channels through the convolutional layer. The output of the channel attention is obtained through the Sigmoid activation function and multiplied by the original feature to finally obtain the channel attention feature F CAM ;
[0068] F CAM = σ(f(f(Avg(F)) + f(Max(F)))) · F (1)
[0069] In Equation (1), σ is the Sigmoid function, f(·) represents the convolutional operation, Avg(·) represents the average pooling, and Max(·) represents the max pooling.
[0070] Step 4.3, as Figure 4 shown, in the pixel position attention (PAM) module: First, the feature F is reshaped into a two-dimensional matrix, and then the correlation between different pixel positions is calculated through matrix multiplication and transpose operation to obtain the correlation matrix ω. The reshaped feature is then multiplied by the correlation matrix ω to finally obtain the channel attention feature F PAM ;
[0071]
[0072] In Equation (2), δ(·) represents the Softmax function, T represents the transpose operation, represents the matrix multiplication.
[0073] Step 4.4, for the obtained feature F CAM and F PAM , first perform summation to obtain the feature F in , and then the feature F in passes through the global average pooling layer and the convolutional layer in parallel to obtain the global feature G(F in ) and the local feature L(F in );
[0074] F in = F CAM + F PAM (4)
[0075] G(F in ) = (f(f RELU (f(f GAP (F in )))) (5)
[0076] L(F in ) = f(f RELU (F in )) (6)
[0077] In Equation (5), f RELU (·) represents the RELU activation function, and f GAP (·) represents global average pooling.
[0078] Step 4.5, as Figure 5 shown, input the global feature G(F in ) and the local feature L(F in ) into the feature fusion module to obtain the fused feature F out = F in ·σ(G(F in ) + L(F in ))), and then F out passes through a convolutional layer to obtain the final artifact-removed image I H . Among them, σ is the Sigmoid function.
[0079] Step 5, construct the loss function L;
[0080]
[0081] In Equation (7), N is the number of training samples, is the artifact-removed image of the i-th sample, is the original real image of the i-th sample, and ||·||1 represents the L1 norm.
[0082] Step 6, training of the network: Move the constructed network model to the GPU for training. Use the Adam optimizer to optimize the image restoration network, calculate the loss function and then backpropagate it into the network for optimization, and stop training when the loss function converges. After training is completed, save the model.pth weight file for subsequent testing.
[0083] Step 7, testing of the network:
[0084] Step 7.1, select Classic5 (5 images) and LIVE (29 images) for algorithm performance testing. The specific method is as follows: First, compress the selected test data set to different degrees, then use different network models to remove artifacts from the above compressed images, and finally calculate the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) between the restored image and the real image to perform network performance testing.
[0085] Step 7.2, select images that have been compressed multiple times in reality, create a data set Real (10 images) for algorithm performance testing. Use different network models to remove artifacts from the above Real data set, and finally calculate the Natural Image Quality Evaluator (NIQE) and Blind / Referenceless Image Spatial Quality Evaluator (BRISQUE) scores to perform network performance testing.
[0086] The above only describes in detail the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.
Claims
1. An efficient method for removing artifacts from compressed images based on multi-feature fusion, characterized in that, It includes the following steps: S1. Data Preparation: Download the publicly available high-definition datasets Flickr2K and DIV2K, and generate the required training dataset I by compressing them with the MATLAB JPEG encoder L ; S2. Construct a shallow feature extraction module, which mainly consists of two convolutional layers with 3x3 convolutional kernels, a stride of 1, and a padding of 1; S3. Construct a multi-scale feature extraction module, which consists of a convolutional layer and a recursive residual unit (RRU); each recursive residual unit consists of multiple residual blocks, and each residual block is composed of two convolutional layers; S4. Construct an image restoration module, which consists of a feature fusion module (FAM) and a convolutional layer; the FAM is composed of a channel attention module (CAM) and a pixel location attention module (PAM); S5. Construct a loss function L; S6. Move the constructed network model to the GPU for training, use the Adam optimizer to optimize the image restoration network, calculate the loss function and then backpropagate it into the network for optimization. Stop training when the loss function converges. After training is completed, save the model.pth weight file for subsequent test use.
2. An efficient compression image artifact removal method based on multi-feature fusion according to claim 1, characterized in that, The method for constructing the shallow feature extraction module in S2 is as follows: S2.
1. Receive the compressed image I L ∈R C×H×W , usually a color image or a grayscale image, with the number of channels C of the color image being 3 and the number of channels C of the grayscale image being 1; S2.
2. The input image passes through the convolutional layer to extract the initial low-level features, and finally obtain the shallow feature F0.
3. An efficient compression image artifact removal method based on multi-feature fusion according to claim 1, characterized in that, The method for constructing the multi-scale feature extraction module in S3 is as follows: S3.
1. Receive the compressed image I L ∈R C×H×W , and after passing through the convolutional layer, enter multiple recursive residual units (RRUs) for multi-scale feature extraction; S3.
2. After being processed by multiple recursive residual units, the final multi-scale feature F1 is output through the last convolutional layer.
4. An efficient compressed image artifact removal method based on multi-feature fusion according to claim 3, characterized in that, The working process of each recursive residual unit (RRU) in S3.1: The input features are first activated by ReLU. The activated features pass through the first convolutional layer to obtain intermediate features. The intermediate features are then activated by ReLU again and enter the second convolutional layer. The output of the second convolutional layer is connected to the input residually, and finally the output of the current recursive residual unit is obtained.
5. An efficient compression image artifact removal method based on multi-feature fusion according to claim 1, characterized in that The method for constructing the image restoration module in S4 is as follows: S4.
1. Concatenate the received shallow feature F0 and multi-scale feature F1 first to obtain the feature F; S4.
2. In the channel attention CAM module: First, the feature F passes through the average pooling layer and the max pooling layer in parallel, and then is restored to the original number of channels through the convolutional layer; the output of the channel attention is obtained through the Sigmoid activation function and multiplied by the original feature to finally obtain the channel attention feature F CAM ; S4.
3. In the Pixel Position Attention (PAM) module: First, reshape the feature F into a two-dimensional matrix, then calculate the correlation between different pixel positions through matrix multiplication and transpose operations to obtain the correlation matrix ω; subsequently, multiply the reshaped feature by the correlation matrix ω to obtain the channel attention feature F PAM ; S4.
4. For the obtained feature F CAM and F PAM , first perform summation to obtain the input feature F in . Then, the feature F in is passed through the global average pooling layer and the convolutional layer in parallel to obtain the global feature G(F in ) and the local feature L(F in ); S4.
5. Input the global feature G(F in ) and the local feature L(F in ) into the feature fusion module to obtain the fused feature F out = F in ·σ(G(F in ) + L(F in ))), and then F out passes through a convolutional layer to obtain the final artifact-removed image I H ; where σ is the Sigmoid function.
6. An efficient compression image artifact removal method based on multi-feature fusion according to claim 5, characterized in that, The channel attention feature F obtained in S4.2 CAM has the following formula: F CAM = σ(f(f(Avg(F)) + f(Max(F)))) · F where σ is the Sigmoid function, f(·) represents the convolution operation, Avg(·) represents the average pooling, and Max(·) represents the max pooling.
7. An efficient compressed image artifact removal method based on multi-feature fusion according to claim 5, characterized in that The channel attention feature F obtained in S4.3 PAM has the following formula: Among them, δ(·) represents the Softmax function, T represents the transpose operation, represents matrix multiplication.
8. An efficient compression image artifact removal method based on multi-feature fusion according to claim 5, characterized in that, The formulas for obtaining the global feature G(F in ) and the local feature L(F in ) in S4.4 are as follows: F in = F CAM + F PAM G(F in ) = (f(f RELU (f(f GAP (F in )))) L(F in ) = f(f RELU (F in )) Among them, f RELU (·) represents the RELU activation function, and f GAP (·) represents global average pooling.
9. An efficient compression image artifact removal method based on multi-feature fusion according to claim 1, characterized in that The method for constructing the loss function L in S5 is as follows: where N is the number of training samples, is the artifact-removed image of the i-th sample, is the original ground-truth image of the i-th sample, and ||·||1 represents the L1 norm.
Citation Information
Cited By
Railway external environment anomaly detection method and system based on global feature correlation
CN122048946A