JPEG (Joint Photographic Experts Group) image restoration method based on discrete cosine transform and Mama network

By introducing discrete cosine transform and Mamba network into the JPEG image recovery method, the DCTMamba network is constructed, and the problems of the existing technology in global information processing and computing complexity are solved, and higher image recovery fidelity and efficiency are achieved.

CN120070268AActive Publication Date: 2025-05-30UNIV OF SCI & TECH OF CHINA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510179138.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-30
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing JPEG image recovery methods perform poorly when processing tasks that require global information. The transformer-based method has high computational complexity, resulting in resource consumption and processing speed bottlenecks. The Mamba network is not suitable for non-causal image recovery tasks.

Method used

Using JPEG image recovery method based on discrete cosine transform and Mamba network, the DCTMamba network is constructed, combining shallow feature extraction, encoder, decoder and reconstruction modules, and using discrete cosine transform and Mamba's long sequence processing capabilities to establish a causal relationship and improve the fidelity of image recovery.

Benefits of technology

The fidelity of JPEG image recovery is improved, and the problem of inconsistent image recovery quality is solved. By reducing computational complexity and improving resource utilization, processing speed and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070268A_ABST
    Figure CN120070268A_ABST
Patent Text Reader

Abstract

The invention provides a JPEG (Joint Photographic Experts Group) image restoration method based on DCT (Discrete Cosine Transform) and a Mama network, which comprises the following steps of: constructing a DCT (Discrete Cosine Transform) network architecture, combining the long sequence modeling advantage of Mama with the characteristic that high-frequency components are abandoned after JPEG compression, and introducing DCT into the network so as to establish a causal relationship sequence conforming to Mama autoregression scanning characteristics; the rationality of Mama for image restoration is improved, and a scale adaptive normalization method is proposed for discrete cosine transform frequency distribution differences caused by different image sizes to ensure consistent restoration of images of different sizes. According to the method, the scanning sequence conforming to the Mama autoregression characteristic can be established, and pictures of different sizes can be flexibly processed, so that the fidelity of JPEG image restoration can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and specifically relates to a JPEG image restoration method based on discrete cosine transform and Mamba network. Background Art

[0002] JPEG is a widely used lossy compression method. Initially, an image is divided into 8×8 pixel blocks, and then the discrete cosine transform (DCT) is used to transform these blocks from the spatial domain to the frequency domain. This transformation prioritizes low-frequency components and significantly reduces high-frequency details during the quantization process. These details usually contain the subtle features of the image. Therefore, this compression method achieves a significant reduction in file size, minimizing storage and bandwidth requirements. The quality factor (QF) in JPEG determines how much high-frequency information is discarded. A lower QF value results in stronger compression and more obvious image distortion. Although JPEG compression is effective in reducing file size, the artifacts it introduces degrade the visual quality and hinder computer vision tasks such as image recognition and object detection.

[0003] With the rapid development of deep learning, the field of JPEG restoration has undergone a transformation, and neural network-based methods have gradually replaced traditional model-based methods. These neural network methods have shown excellent results. Convolutional neural networks (CNNs) use their non-linear mapping ability to restore degraded images to their original state. In addition, technologies like SwinIR utilize the long-sequence modeling potential of transformers, thus significantly expanding the receptive field. Recently, Mamba is at the forefront of long-sequence modeling and is particularly known for its powerful performance. Mamba is now used in image restoration projects to effectively utilize global information, setting a new standard in this field.

[0004] However, these advanced methods face unique challenges: (1) CNN-based methods are limited by their inherent local reduction bias, resulting in poor performance when dealing with image restoration tasks that require global information; (2) Transformer-based methods face quadratic computational complexity, making it difficult to process very long sequence inputs, thus leading to bottlenecks in resource consumption and processing speed; (3) Although Mamba is suitable for causal autoregressive tasks, it is not suitable for non-causal image restoration tasks, limiting its effectiveness in such applications. Summary of the Invention

[0005] The present invention aims to solve the above-mentioned deficiencies of the prior art and proposes a JPEG image restoration method based on discrete cosine transform and Mamba network, in order to establish a scanning sequence that conforms to the autoregressive characteristics of Mamba and can flexibly process pictures of different sizes, thereby improving the fidelity of JPEG image restoration.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A JPEG image restoration method based on discrete cosine transform and Mamba network according to the present invention is characterized in that it is carried out according to the following steps:

[0008] Step 1: Obtain a batch of JPEG images and their corresponding clean images , where represents the th JPEG image, represents the th clean image, B represents the number of images in a batch, represents the number of channels of the image, and respectively represent the height and width of the image;

[0009] Crop an image block at the same position from and its corresponding , and record them as the th JPEG image block and the th clean image block , where N represents the height and width of the image block;

[0010] Step 2: Construct a DCT-Mamba network, including: a shallow feature extraction module, an encoder module, a decoder module, and a reconstruction module;

[0011] Step 3: The shallow feature extraction module is composed of a convolutional layer with a convolutional kernel size of in series with a GELU activation function, and processes to obtain the th shallow feature ;

[0012] Step 4: The encoder module is composed of Q multi-scale feature reduction modules in series, and processes to obtain the th multi-scale encoded feature set ={ }, where represents the qth multi-scale encoded feature;

[0013] Step 5: The decoder module is composed of Q multi-scale feature expansion modules in series, and processes to obtain the ith multi-scale decoded feature set ={ ,... }, where Represents the q-th multi-scale decoded feature;

[0014] Step 6: The reconstruction module consists of R convolutional layers with a kernel size of and a stride of , and a padding parameter of / / 2, which are connected in series. After processing , the -th restored JPEG image block is obtained, where / / represents floor division;

[0015] Step 7: Calculate the total loss function , which is used for backpropagation training of the DCTMamba network, and the Adam optimizer is used to optimize the parameters of the DCTMamba network. When the maximum number of training times is reached, the training is stopped, and thus the optimal JPEG image restoration model is obtained for realizing the restoration of JPEG images;

[0016] Another feature of the JPEG image restoration method based on discrete cosine transform and Mamba network according to the present invention is that each multi-scale feature reduction module in step 4 consists of a first multi-level unit and a downsampling module connected in series;

[0017] Step 4.1: The first multi-level unit consists of a local module and a coarse-to-fine state space module connected in parallel.

[0018] When q = 1, the first multi-level unit processes to obtain the q-th encoded state space feature ;

[0019] Step 4.2: The downsampling module consists of a convolutional layer and a GELU activation function;

[0020] When q = 1, is input into the downsampling module of the q-th multi-scale feature reduction module for processing, so that the size of is reduced to times the original size, and the number of channels is expanded to a times the original, thereby obtaining the q-th multi-scale encoded feature , where a represents the scaling factor;

[0021] Step 4.3: When q = 2, 3,..., Q, the (q - 1)-th multi-scale encoded feature is input into the q-th multi-scale feature reduction module for processing, so that the Q-th multi-scale encoded feature is output by the Q-th multi-scale feature reduction module.

[0022] Further, step 4.1 includes:

[0023] Step 4.1.1: The local module is successively composed of Y dense residual modules, and each dense residual module includes Z convolutional layers;

[0024] When q = 1, is input into the q-th multi-scale feature reduction module, and after being processed by the Y dense residual modules of the local module of the first multi-level unit, the q-th encoded local feature is obtained ;

[0025] Step 4.1.2: The coarse-to-fine state space module consists of two parallel paths and a linear layer with a dilation factor of 1 / Z. Among them, the first path is composed of a scale modulation module and a scale adaptation module in series, and the second path is composed of a linear layer with a dilation factor of Z and a SiLU activation function in series; the scale modulation module is composed of a linear layer with a dilation factor of Z, a depth convolutional layer, and a SiLU activation function in series; the scale adaptation module is composed of a discrete cosine transform layer, a normalization layer, and an inverse discrete cosine transform layer;

[0026] Step 4.1.2.1: When q = 1, is input into the q-th multi-scale feature reduction module, and after being processed by the scale modulation module of the first path of the state space module, the q-th encoded scale modulation feature is obtained ;

[0027] is then input into the discrete cosine transform layer of the scale adaptation module of the first path for processing, and the q-th discrete cosine feature is obtained , and the normalization layer uses Equation (1) to obtain the q-th normalized discrete cosine feature :

[0028] (1)

[0029] In Equation (1), m and n respectively represent the row and column position numbers of the eigenvalues in represents the position coordinates of the eigenvalues in

[0030] After is processed by Mamba autoregression from top to bottom and from left to right, it is input into the inverse discrete cosine transform layer and outputs the q-th encoder scale adaptation feature ;

[0031] Step 4.1.2.2: After being processed in the second path, the q-th coded scale dilation feature is obtained ;

[0032] Step 4.1.2.3: After performing the Hadamard product operation on and and then processing it through a linear layer with a dilation factor of 1 / Z, the q-th coded state space feature is obtained .

[0033] Furthermore, each multi-scale feature expansion module in step 5 is composed of a second multi-level unit and a downsampling module connected in series;

[0034] Step 5.1: The structure of the second multi-level unit is the same as that of the first multi-level unit;

[0035] When q = 1, is input into the second multi-level unit of the q-th multi-scale feature expansion module and processed according to the process of step 4.1.1 - step 4.1.2.3 to obtain the q-th decoded state space feature ;

[0036] Step 5.2: The upsampling module is composed of a transposed convolutional layer and a GELU activation function;

[0037] When q = Q, is input into the upsampling module of the q-th multi-scale feature expansion module for processing, so that the size of is expanded to a times the original, and the number of channels is reduced to times the original, and then added to to obtain the q-th multi-scale decoded feature ;

[0038] Step 5.3: When q = Q - 1, Q - 2,..., 1, the (q - 1)-th multi-scale decoded feature is input into the q-th multi-scale feature expansion module for processing, so that the first multi-scale decoded feature is output from the first multi-scale feature expansion module

[0039] Furthermore, the total loss function in step 7 is established as follows:

[0040] Step 7.1: Construct the absolute value loss function using Equation (2) :

[0041] (2)

[0042] Step 7.2: Construct the discrete cosine loss function using Equation (3) :

[0043] (3)

[0044] In formula (3), DCT represents the discrete cosine transform operation;

[0045] Step 7.3, construct the structural similarity loss using formula (4) :

[0046] (4)

[0047] In formula (4), SSIM represents the structural similarity calculation function;

[0048] Step 7.4, construct the total loss function using formula (5) :

[0049] (5)

[0050] In formula (5), and are two hyperparameters.

[0051] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the JPEG image restoration method, and the processor is configured to execute the program stored in the memory.

[0052] A computer-readable storage medium according to the present invention, characterized in that a computer program stored on the computer-readable storage medium executes the steps of the JPEG image restoration method when run by a processor.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] 1. The present invention introduces the DCTMamba framework, which combines the JPEG compression characteristics with the advantages of Mamba for long sequence processing. By integrating the discrete cosine transform into Mamba, scanning from low frequency to high frequency is achieved, and the causal relationship between scanning sequences is established, thereby enhancing the fidelity of image restoration.

[0055] 2. To address the challenges brought about by the change in the density of frequency components due to different image sizes, the present invention introduces a normalization process for discrete cosine transform coefficients. This normalization balances the contributions of various frequency components and improves the consistency of the image restoration quality for different sizes of images. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 are the low-frequency and high-frequency components after the DCT transformation of the image;

[0057] Figure 2 It is a comparison chart before and after normalization;

[0058] Figure 3 It is the overall framework diagram of the network;

[0059] Figure 4 It is the quantitative performance index diagram of the present invention;

[0060] Figure 5 It is the visual effect diagram of the comparison between the present invention and other methods on Classic5;

[0061] Figure 6 It is the visual effect diagram of the comparison between the present invention and other methods on Twitter. Detailed implementation manners

[0062] In this embodiment, a JPEG image restoration method based on DCT and Mamba aims to solve the challenges in JPEG image restoration, constructs an image restoration framework combining JPEG compression characteristics and the advantages of Mamba, and includes the following steps:

[0063] Step 1: Obtain a batch of JPEG images and their corresponding clean images , where represents the th JPEG image, represents the th clean image, B represents the number of images in a batch, represents the number of channels of the image, and respectively represent the height and width of the image;

[0064] From and their corresponding crop an image patch at the same position, which is correspondingly recorded as the th JPEG image patch and the th clean image patch , where N represents the height and width of the image patch; In specific implementation, B = 8, N = 256, C = 3, and and are randomly rotated, flipped, etc.

[0065] Step 2: Construct a DCTMamba network, including: a shallow feature extraction module, an encoder module, a decoder module, and a reconstruction module;

[0066] Step 3: The shallow feature extraction module consists of a convolutional kernel with a size of is composed of a convolutional layer and a GELU activation function in series, and processes to obtain the th shallow feature ; In a specific implementation, has a size of 3, and the number of channels of the shallow feature is 28.

[0067] Step 4: The encoder module is composed of Q multi-scale feature reduction modules in series, and processes to obtain the th multi-scale encoded feature set ={ }, where represents the qth multi-scale encoded feature; among them, each multi-scale feature reduction module is composed of a first multi-level unit and a downsampling module in series; in a specific implementation, Q is taken as 3.

[0068] Step 4.1: The first multi-level unit is composed of a local module and a coarse-to-fine state space module in parallel;

[0069] Step 4.1.1: The local module is successively composed of Y dense residual modules, and each dense residual module includes Z convolutional layers; in a specific implementation, Y is taken as 3 and Z is taken as 4;

[0070] When q = 1, is input into the qth multi-scale feature reduction module and processed by the Y dense residual modules of the local module of the first multi-level unit to obtain the qth encoded local feature .

[0071] Step 4.1.2: The coarse-to-fine state space module is composed of two parallel paths and a linear layer with a dilation factor of 1 / Z, where the first path is composed of a scale modulation module and a scale adaptation module in series, and the second path is composed of a linear layer with a dilation factor of Z and a SiLU activation function in series; the scale modulation module is composed of a linear layer with a dilation factor of Z, a depth convolution layer, and a SiLU activation function in series; the scale adaptation module is composed of a discrete cosine transform layer, a normalization layer, and an inverse discrete cosine transform layer; in a specific implementation, Z is taken as 4.

[0072] Step 4.1.2.1: When q = 1, is input into the qth multi-scale feature reduction module and processed by the scale modulation module of the first path of the state space module to obtain the qth encoded scale modulation feature ;

[0073] Then is input into the discrete cosine transform layer of the scale adaptation module of the first path for processing,Figure 1 The high and low frequency components after the image DCT transformation are shown, and the q-th discrete cosine feature is obtained The normalization layer uses Equation (1) to obtain the q-th normalized discrete cosine feature :

[0074] (1)

[0075] In Equation (1), m and n respectively represent the position sequence numbers of the rows and columns of the eigenvalues in represents the position coordinates of the eigenvalues in Figure 2 The distribution of DCT coefficients before and after normalization is shown;

[0076] After performing Mamba autoregressive scanning processing on from top to bottom and from left to right, it is input into the inverse discrete cosine transform layer, and the q-th coding scale adaptive feature is output .

[0077] Step 4.1.2.2: Input into the second path for processing, and the q-th coding scale dilation feature is obtained;

[0078] Step 4.1.2.3: After performing the Hadamard product operation on and , and then processing it through a linear layer with a dilation factor of 1 / Z, the q-th coding state space feature is obtained; In specific implementation, Z takes 4.

[0079] Step 4.2: The downsampling module consists of a convolutional layer and a GELU activation function;

[0080] When q = 1, input into the downsampling module of the q-th multi-scale feature reduction module for processing, so that the size of is reduced to times the original size, and the number of channels is expanded to a times the original, thereby obtaining the q-th multi-scale coding feature , where a represents the scaling scale; In specific implementation, a takes 2.

[0081] Step 4.3: When q = 2, 3,..., Q, input the (q - 1)-th multi-scale coding feature into the q-th multi-scale feature reduction module for processing, so that the Q-th multi-scale coding feature is output by the Q-th multi-scale feature reduction module; In specific implementation, the number of channels of each multi-scale coding feature is 28, 56, and 112 respectively.

[0082] Step 5: The decoder module is composed of Q multi-scale feature expansion modules connected in series, and processes to obtain the i-th multi-scale decoded feature set ={ ,... }, where represents the q-th multi-scale decoded feature; each multi-scale feature expansion module is composed of a second multi-level unit and a downsampling module connected in series.

[0083] Step 5.1: The structure of the second multi-level unit is the same as that of the first multi-level unit;

[0084] When q = 1, is input into the second multi-level unit of the q-th multi-scale feature expansion module and processed according to the process of steps 4.1.1 - 4.1.2.3 to obtain the q-th decoded state space feature .

[0085] Step 5.2: The upsampling module is composed of a transposed convolutional layer and a GELU activation function;

[0086] When q = Q, is input into the upsampling module of the q-th multi-scale feature expansion module for processing, so that is enlarged to a times the original size and the number of channels is reduced to times the original, and then added to to obtain the q-th multi-scale decoded feature ; in a specific implementation, a is taken as 2, and the q-th multi-scale decoded feature has the same dimension size as the q-th multi-scale encoded feature.

[0087] Step 5.3: When q = Q - 1, Q - 2,..., 1, the (q - 1)-th multi-scale decoded feature is input into the q-th multi-scale feature expansion module for processing, so that the first multi-scale decoded feature is output from the first multi-scale feature expansion module.

[0088] Step 6: The reconstruction module is composed of R convolutional layers with a convolutional kernel size of , a stride of , and a padding parameter of / / 2 connected in series, and processes to obtain the -th restored JPEG image block , where / / represents rounding down; Figure 3 shows the overall framework diagram of the present invention.

[0089] Step 7. Calculate the total loss function for backpropagation training of the DCTMamba network, and use the Adam optimizer to optimize the parameters of the DCTMamba network. When the maximum number of training times is reached, stop the training to obtain the optimal JPEG image restoration model for realizing the restoration of JPEG images. In this example, the learning rate is set to 2e-4, and the cosine annealing algorithm is used for learning rate decay, and the decay period is set to 1000.

[0090] Step 7.1. Construct the absolute value loss function using Equation (2) :

[0091] (2)

[0092] Step 7.2. Construct the discrete cosine loss function using Equation (3) :

[0093] (3)

[0094] In Equation (3), DCT represents the discrete cosine transform operation;

[0095] Step 7.3. Construct the structural similarity loss using Equation (4) :

[0096] (4)

[0097] In Equation (4), SSIM represents the structural similarity calculation function;

[0098] Step 7.4. Construct the total loss function using Equation (5) :

[0099] (5)

[0100] In Equation (5), and are two hyperparameters. In this example, and are taken as 0.1 and 10 respectively.

[0101] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0102] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

[0103] The present invention verifies the effectiveness of DCTMamba in JPEG image restoration from both quantitative and qualitative aspects. Experiments are conducted on multiple JPEG compression datasets to evaluate the quality of the restored images. Performance evaluation is carried out using metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), and comparison is made with the prior art to confirm the superior performance of DCTMamba in restoring rough structures and details. Specifically, the present invention quantitatively compares it with previous JPEG restoration methods such as ARCNN (Artifacts Reduction Convolutional Neural Network), DnCNN (Denoising Convolutional Neural Network), DCSC (Deep Convolutional Sparse Coding), MWCNN (Multi-level Wavelet Convolutional Neural Networks), RNAN (Residual Non-local Attention Networks for Image Restoration), QGAC (Quantization Guided JPEG Artifact Correction), FBCNN (Towards Flexible Blind JPEG Artifacts Removal) and Zhao et al. (Comprehensive and Delicate: An Efficient Transformer for Image Restoration) on the LIVE1, Classic5 and BSDS500 datasets, as Figure 4 shown, in the quantitative analysis, the best-performing metrics are shown in bold. The results show that the present invention has achieved comprehensive excellent performance under four compression qualities of the three datasets. Figure 5 and Figure 6 show the comparison of the present invention with other methods in terms of visual effects. The test images are from the Classic5 (simulated) dataset and the Twitter (real) dataset respectively. It can be seen that the present method performs better in restoring the texture details of the images, and the overall visual effect is also better than other methods. The present invention confirms the effectiveness of DCTMamba in restoring JPEG images through extensive tests on multiple datasets. The results show that DCTMamba outperforms the current state-of-the-art methods.

Claims

1. A JPEG image restoration method based on discrete cosine transform and Mamba network, characterized in that: The steps are as follows: Step 1: Get a batch of JPEG images and its corresponding clean image ,in, Indicates JPEG images, Indicates clean images, B represents the number of images in a batch, Indicates the number of channels of the image, and Respectively represent the height and width of the image; from and its corresponding Cut an image block at the same position on each of the above images, and record it as JPEG image blocks and Clean image patches , where N represents the height and width of the image block; Step 2: Construct the DCTMamba network, including: shallow feature extraction module, encoder module, decoder module and reconstruction module; Step 3: The shallow feature extraction module consists of a convolution kernel size of The convolutional layer and a GELU activation function are connected in series, and Process it and get Shallow Features ; Step 4: The encoder module is composed of Q multi-scale feature reduction modules in series, and Process it and get Multi-scale encoding feature set ={ },in, represents the qth multi-scale encoding feature; Step 5: The decoder module is composed of Q multi-scale feature expansion modules in series, and Processing is performed to obtain the i-th multi-scale decoding feature set ={ ,... },in, represents the qth multi-scale decoding feature; Step 6: The reconstruction module consists of R convolution kernels with a size of , the step length is , filling parameters are / / 2 convolutional layers are connected in series and After processing, we get Restored JPEG image blocks , where / / means rounding down; Step 7: Calculate the total loss function , used to perform back-propagation training on the DCTMamba network, and use the Adam optimizer to optimize the parameters of the DCTMamba network. When the maximum number of training times is reached, the training is stopped to obtain the optimal JPEG image restoration model for realizing JPEG image restoration.

2. A JPEG image restoration method based on discrete cosine transform and Mamba network according to claim 1, characterized in that: In step 4, each multi-scale feature reduction module is composed of a first multi-level unit and a downsampling module in series; Step 4.1: The first multi-level unit is composed of a local module and a coarse-to-fine state space module in parallel. When q=1, the first multi-level unit pair Processing is performed to obtain the qth encoding state space feature ; Step 4.2: The downsampling module consists of a convolutional layer and a GELU activation function; When q=1, Input into the downsampling module of the qth multi-scale feature reduction module for processing, so that The size is reduced to the original times, the number of channels is expanded to a times of the original, thus obtaining the qth multi-scale coding feature , where a represents the scaling scale; Step 4.3: When q=2,3,…,Q, the q-1th multi-scale encoding feature The input is processed in the qth multi-scale feature reduction module, so that the Qth multi-scale feature reduction module outputs the Qth multi-scale encoding feature .

3. The JPEG image restoration method based on discrete cosine transform and Mamba network according to claim 2, characterized in that: The step 4.1 comprises: Step 4.1.1: The local module is composed of Y dense residual modules in sequence, and each dense residual module includes Z convolutional layers; When q=1, Input into the qth multi-scale feature reduction module, and processed by the Y dense residual modules of the local module of the first multi-level unit to obtain the qth encoded local feature ; Step 4.1.2: The coarse-to-fine state space module is composed of two parallel paths and a linear layer with a dilation factor of 1 / Z, wherein the first path is composed of a scale modulation module and a scale adaptation module in series, and the second path is composed of a linear layer with a dilation factor of Z and a SiLU activation function in series; the scale modulation module is composed of a linear layer with a dilation factor of Z, a deep convolution layer, and a SiLU activation function in series; the scale adaptation module is composed of a discrete cosine transform layer, a normalization layer, and an inverse discrete cosine transform layer; Step 4.1.2.1: When q=1, Input into the qth multi-scale feature reduction module and processed by the scale modulation module of the first path of the state space module to obtain the qth coded scale modulation feature ; Will Then it is input into the discrete cosine transform layer of the scale adaptive module of the first path for processing to obtain the qth discrete cosine feature The normalization layer uses formula (1) to obtain the qth normalized discrete cosine feature : (1) In formula (1), m and n represent The position numbers of the rows and columns of the eigenvalues ​​in ; express The position coordinates of the eigenvalues ​​in ; right After the Mamba autoregression from top to bottom and from left to right is processed, it is input into the inverse discrete cosine transform layer and the qth encoder scale adaptive feature is output. ; Step 4.1.2.2: After being input into the second path for processing, the qth encoding scale expansion feature is obtained ; Step 4.1.2.3: and After the Hadamard product operation, it is processed through a linear layer with a dilation factor of 1 / Z to obtain the qth encoded state space feature .

4. The JPEG image restoration method based on discrete cosine transform and Mamba network according to claim 3 is characterized in that: Each multi-scale feature expansion module in step 5 is composed of a second multi-level unit and a downsampling module in series; Step 5.1: The second multi-level unit has the same structure as the first multi-level unit; When q=1, Input into the second multi-level unit of the qth multi-scale feature expansion module and process according to the process of steps 4.1.1-step 4.1.2.3 to obtain the qth decoding state space feature ; Step 5.2: The upsampling module consists of a transposed convolutional layer and a GELU activation function; When q=Q, Input into the upsampling module of the qth multi-scale feature expansion module for processing, so that The size of is expanded to a times of the original, and the number of channels is reduced to the original After times, Add together to get the qth multi-scale decoding feature ; Step 5.3: When q=Q-1,Q-2,…,1, the q-1th multi-scale decoding feature The input is processed in the qth multi-scale feature expansion module, so that the first multi-scale feature expansion module outputs the first multi-scale decoding feature .

5. The JPEG image restoration method based on discrete cosine transform and Mamba network according to claim 4, characterized in that: The total loss function in step 7 It is established in the following steps: Step 7.1: Use formula (2) to construct the absolute value loss function : (2) Step 7.2: Use formula (3) to construct the discrete cosine loss function : (3) In formula (3), DCT represents discrete cosine transform operation; Step 7.3: Use formula (4) to construct the structural similarity loss : (4) In formula (4), SSIM represents the structural similarity calculation function; Step 7.4: Use formula (5) to construct the total loss function : (5) In formula (5), and are 2 hyperparameters.

6. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the JPEG image restoration method described in any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the JPEG image restoration method described in any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • JPEG image compression artifact elimination algorithm based on cascade residual coding and decoding network

    CN112509094A

  • JPEG (Joint Photographic Experts Group) image artifact removal method based on comparative learning and application thereof

    CN115829858A

  • Remote sensing image change detection method, system and equipment based on double-domain learning

    CN118379626A

  • Deep hash image retrieval method based on frequency domain decoupling and visual Mamba

    CN118820508A

  • Text image tampering detection method and device

    CN119027788A