Construction and application method of image reconstruction model, model, equipment and medium

By constructing a deep learning-based image reconstruction model, using the frequency domain coefficients and quantization tables of JPEG images for training, the problem of artifacts in JPEG image compression is solved, and high-quality image reconstruction and visual effect improvement are achieved.

CN120298233AActive Publication Date: 2025-07-11SHENZHEN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510744286.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-11
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing JPEG image compression technology will introduce significant artifacts under low quality factors, resulting in a decline in visual quality and cannot meet the high application needs such as medical image analysis, security monitoring and satellite remote sensing.

Method used

The image reconstruction model is constructed, and multiple iterative training is carried out using deep learning methods based on the frequency domain coefficients and quantization table of JPEG images. Through the learning offset guidance module and the domain transformation module, the quantization error and chromaticity distortion are gradually reduced, and the details and color consistency of the compressed image are restored.

Benefits of technology

It significantly improves the visual quality of JPEG images, reduces artifacts, enhances the overall quality and detail recovery ability of the image, and adapts to compressed images with different quality factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298233A_ABST
    Figure CN120298233A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, and provides an image reconstruction model construction and application method, a model, equipment and a medium. According to the scheme, a frequency domain coefficient and a quantization table of a JPEG (Joint Photographic Experts Group) image serve as a data set to train an initial image reconstruction model; by inputting priori knowledge (namely a quantization table used during image compression) of a compression process, an initial image reconstruction model is guided to gradually reduce a distortion phenomenon caused by image compression so as to obtain an image reconstruction model, and the image reconstruction model obtained based on the scheme of the invention can better enhance the visual quality of a JPEG (Joint Photographic Experts Group) image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a method, model, device, and medium for constructing and applying an image reconstruction model. Background Art

[0002] In the field of modern digital image processing, the rapid growth of image data has brought huge challenges to storage, transmission and processing. In order to cope with this problem, image compression technology came into being. Its core purpose is to reduce data redundancy while retaining visual information as much as possible, thereby reducing storage space and transmission bandwidth requirements. JPEG, as a widely used lossy compression standard, significantly reduces file size by quantizing and encoding to remove high-frequency details that the human eye is not sensitive to, while maintaining a high subjective visual quality. This efficient compression method makes JPEG the mainstream format for image storage and transmission, especially in resource-constrained scenarios (such as network transmission, mobile devices, etc.), its advantages are more prominent.

[0003] At the same time, in the image compression process, especially lossy compression methods such as JPEG, the compression ratio is usually represented by the quality factor (QF). At low QF, significant artifacts are usually introduced, resulting in a decrease in visual quality, affecting the visual effect of the image and the performance of subsequent processing tasks. However, in many practical applications (such as medical image analysis, security monitoring, satellite remote sensing, etc.), the image quality requirements are high, and directly using compressed images may not meet the actual application needs.

[0004] Therefore, there is an urgent need for a solution that can effectively remove artifacts that appear in compressed images. This solution can restore lost details in compressed images, suppress distortion, and improve the overall quality of the image, thereby achieving better visual experience and higher application value based on efficient compression. Summary of the invention

[0005] The present application provides a method, model, device, and medium for constructing an image reconstruction model to solve the problem of artifacts appearing in existing compressed JPEG images.

[0006] A first aspect of the present application provides a method for constructing an image reconstruction model, where the image reconstruction model is used to reconstruct a JPEG image. The method comprises: Acquire a JPEG image and decode it to obtain frequency domain coefficients and a quantization table of the JPEG image; Build an initial image reconstruction model based on deep learning methods; The frequency domain coefficients and the quantization table constitute a data set, and the data set is used to perform multiple iterations of training on the initial image reconstruction model until convergence, so as to obtain the image reconstruction model.

[0007] In some embodiments of the present application, the initial image reconstruction model takes the frequency domain coefficients and quantization table as inputs and outputs the reconstructed image corresponding to the JPEG image. The initial image reconstruction model includes: A learnable offset guiding module, configured to calculate a first corrected frequency domain coefficient feature based on the frequency domain coefficients and the quantization table; A domain transformation module, configured to transform the first corrected frequency domain coefficient feature into the pixel domain to obtain a first pixel feature of the JPEG image, and calculate the reconstructed image corresponding to the JPEG image based on the first pixel feature.

[0008] In some embodiments of the present application, the learnable offset guiding module includes: A quantization offset sub-module, configured to calculate a quantization offset feature based on the frequency domain coefficients and the quantization table; A frequency domain correction sub-module, configured to calculate a first corrected frequency domain coefficient feature based on the quantization offset feature and the frequency domain coefficients.

[0009] In some embodiments of the present application, the learnable offset guiding module further includes a dequantization sub-module, and: The dequantization sub-module is configured to calculate a dequantized frequency domain coefficient feature based on the first corrected frequency domain coefficient feature and the quantization table; The domain transformation module is further configured to transform the dequantized frequency domain coefficient feature into the pixel domain to obtain a second pixel feature of the JPEG image, and calculate the reconstructed image corresponding to the JPEG image based on the second pixel feature.

[0010] In some embodiments of the present application, the image reconstruction model further includes: An encoder-decoder module, configured to calculate a quantization guiding feature of the JPEG image based on the quantization table, the first pixel feature or the second pixel feature, and calculate the reconstructed image corresponding to the JPEG image based on the quantization guiding feature.

[0011] In some embodiments of the present application, the encoder-decoder module includes: An encoder sub-module, configured to calculate a target intermediate feature of the JPEG image based on the quantization table, the first pixel feature or the second pixel feature; A decoder sub-module, configured to calculate a quantization guiding feature of the JPEG image based on the target intermediate feature and the quantization table, and calculate the reconstructed image corresponding to the JPEG image based on the quantization guiding feature.

[0012] The second aspect of the present application provides an image reconstruction model, and the image reconstruction model includes: A learnable offset guiding module, configured to calculate a first corrected frequency domain coefficient feature based on the frequency domain coefficients and the quantization table of the obtained JPEG image; A domain transformation module is used to transform the corrected first feature of the frequency domain coefficients into the pixel domain to obtain the first pixel feature of the JPEG image, and calculate the reconstructed image corresponding to the JPEG image based on the first pixel feature.

[0013] The third aspect of the present application provides a method for applying an image reconstruction model, and the method includes: Obtain a JPEG image and decode it to obtain the frequency domain coefficients and quantization table of the JPEG image; Use the image reconstruction model as described in the above embodiments to calculate the reconstructed image corresponding to the JPEG image based on the frequency domain coefficients and the quantization table.

[0014] The fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of the first aspect and the third aspect in the above embodiments.

[0015] The fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method described in any one of the first aspect and the third aspect in the above embodiments.

[0016] The present application has the following beneficial effects: The present application uses the frequency domain coefficients and quantization table of the PEG image as a data set to train an initial image reconstruction model. By inputting the prior knowledge of the compression process (i.e., the quantization table used when compressing the image), the initial image reconstruction model is guided to gradually reduce the distortion phenomenon caused by image compression, so as to obtain an image reconstruction model. The image reconstruction model obtained based on the solution of the present application can better enhance the visual quality of the JPEG image. Description of the Drawings

[0017] The drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.

[0018] Figure 1 It is a schematic diagram of an example of the process of compressing an image into the JPEG format provided by the present application; Figure 2 It is a schematic diagram of the process of an embodiment of the method for constructing an image reconstruction model provided by the present application; Figure 3 It is a schematic diagram of the framework of the first embodiment of the image reconstruction model provided by the present application; Figure 4 It is a schematic diagram of the framework of the second embodiment of the image reconstruction model provided by the present application; Figure 5It is a schematic framework diagram of an embodiment of the learnable offset guidance module provided by this application; Figure 6 It is a schematic framework diagram of an embodiment of the encoding quantization table guidance unit provided by this application; Figure 7 It is a schematic framework diagram of an embodiment of the decoding quantization table guidance unit provided by this application; Figure 8 It is a schematic framework diagram of the third embodiment of the image reconstruction model provided by this application; Figure 9 It is a schematic framework diagram of an embodiment of the electronic device provided by this application; Figure 10 It is a schematic framework diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners

[0019] The solutions of the embodiments of this application will be described in detail below with reference to the accompanying drawings of the specification.

[0020] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand this application.

[0021] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0022] The inventors have found through research that in order to effectively remove artifacts in JPEG images, relevant researchers have proposed model-based methods. These methods are mainly based on filter design and only solve limited blocking and ringing artifacts. In recent years, inspired by the success of deep neural networks in image classification, researchers have begun to explore using deep neural networks to remove artifacts in JPEG images. Some methods for removing JPEG artifacts guide the network model to eliminate artifacts by dynamically predicting quality factors or quality factor maps, etc. However, these methods do not fully utilize the compression prior knowledge of JPEG, that is, directly inputting a JPEG image into the network model. The model only passively fits the input data and learns the mapping from the JPEG image to the reconstructed image. This method also does not consider the root cause of image distortion (i.e., quantization errors and chromatic distortion will occur in the compression method of the JPEG format), resulting in incorrect details and unnatural colors being restored in the image.

[0023] Therefore, the inventors provide a new solution for removing JPEG artifacts, also known as a blind JPEG artifact removal solution based on inverse JPEG compression. In the solution of this application, the prior knowledge of the compression process (i.e., the quantization table used when compressing the image) is input to guide the initial image reconstruction model to gradually reduce the image distortion phenomenon (i.e., this application considers the root cause of image distortion), thereby better enhancing the details and color consistency of the JPEG image, that is, better enhancing the visual quality of JPEG.

[0024] To better understand that the solution for removing JPEG artifacts proposed in this application considers the root cause of image distortion, the following is combined with Figure 1 to illustrate the specific process of compressing an image into the JPEG format.

[0025] As Figure 1 shown, it is the process of compressing an image into a JPEG format image, and the process is as follows: (1) Color space conversion: The real image is first converted from the RGB color space to the YCbCr color space. Among them, the Y component represents the luminance component, and Cb and Cr represent the blue component and the red component respectively; (2) Downsampling: In the YCbCr color space, the chrominance components (Cb and Cr) are usually downsampled to reduce the amount of data. This step introduces chrominance distortion. (3) Discrete cosine transform (DCT): Apply the discrete cosine transform (DCT) to each 8x8 pixel block. The DCT converts the pixel values into frequency domain coefficients (i.e., the frequency domain coefficients mentioned in this application), so that the energy is concentrated in the low-frequency part, and the high-frequency part contains less energy. (4) Quantization: A quantization table is used to quantize the frequency domain coefficients. High-frequency coefficients are usually discarded or greatly reduced, while low-frequency coefficients retain more details. This step introduces quantization errors, but the human eye is less sensitive to these high-frequency details. (5) Nearest neighbor rounding: The quantized frequency domain coefficients are further simplified by nearest neighbor rounding to reduce the amount of data. (6) Bitstream encoding: Use entropy encoding (such as Huffman encoding or arithmetic encoding) to encode the quantized frequency domain coefficients to generate the final JPEG file.

[0026] Therefore, the technical solution of this application is Figure 1 the inverse compression process of JPEG format compression for images in, that is, this application decodes the JPEG compressed image to obtain its corresponding frequency domain coefficients and quantization table, and inputs the quantization table and other compression prior knowledge into the initial image reconstruction model to simulate the JPEG inverse compression process to compensate for quantization errors and / or chrominance distortion, so that the model can reconstruct high-quality images.

[0027] The following will describe this application in detail with reference to the accompanying drawings and specific embodiments.

[0028] According to an embodiment of this application, as Figure 2 shown, this application provides a method for constructing an image reconstruction model. This method obtains the image reconstruction model by performing the following steps S1 - S3: S1. Obtain a JPEG image and decode it to obtain the frequency domain coefficients and quantization table of the JPEG image; S2. Construct an initial image reconstruction model based on the deep learning method; S3. Construct a data set from the frequency domain coefficients and quantization table, and use the data set to perform multiple iterative trainings on the initial image reconstruction model until convergence to obtain the image reconstruction model. Each step will be described in detail below.

[0029] I. Step S1 The inventors have found through research that relevant researchers train specific grid models for each quality factor, lacking the flexibility to learn a single grid model for different JPEG quality factors. Moreover, since the quality factors are usually unknown in practical applications, this greatly limits the practicality of the network model.

[0030] To this end, according to an embodiment of the present application, in step S1, the obtained JPEG images are compressed using different quality factors respectively, and the JPEG images corresponding to different quality factors are decoded respectively to obtain the frequency domain coefficients and quantization tables of the JPEG images corresponding to different quality factors.

[0031] As can be seen from the above embodiments, the present application constructs a data set using JPEG images corresponding to different quality factors to guide the initial image enhancement model to gradually reduce the distortion phenomenon caused by image compression, so that the obtained image reconstruction model can perform image enhancement on images compressed by any quality factor, improving the flexibility and practicality of the trained image reconstruction model; at the same time, the generalization of the image reconstruction model is also improved, that is, when training with JPEG images corresponding to quality factors that did not appear during training, the model can also perform high-quality reconstruction on JPEG images compressed based on that quality factor.

[0032] II. Step S2 According to an embodiment of the present application, as Figure 3 shown, the initial image reconstruction model takes the frequency domain coefficients and quantization table as input and the reconstructed image corresponding to the JPEG image as output, and the initial image reconstruction model includes: a learnable offset guiding module for calculating a first corrected frequency domain coefficient feature according to the frequency domain coefficients and quantization table; a domain transformation module for transforming the first corrected frequency domain coefficient feature to the pixel domain to obtain the first pixel feature of the JPEG image, and calculating the reconstructed image corresponding to the JPEG image based on the first pixel feature.

[0033] As can be seen from the above description, the learnable offset guiding module of the above embodiment of the present application can correct the quantization error of the frequency domain coefficients introduced during image compression according to the quantization prior (quantization table), effectively restoring the frequency domain information lost during the image compression process to generate high-quality reconstructed images.

[0034] In order to more clearly understand how the learnable offset guiding module corrects the frequency domain coefficients of the compressed JPEG image and how the domain transformation module realizes the final image reconstruction, the structures and functions of the learnable offset guiding module and the domain transformation module will be introduced in detail below.

[0035] (I) Learnable offset guiding module Among them, according to an embodiment of the present application, the frequency domain coefficients are divided into luminance coefficients and chrominance coefficients, and still referring to Figure 3 shown, the initial image reconstruction model includes two learnable offset guiding modules, one for calculating a first corrected luminance coefficient feature according to the luminance coefficients and quantization table, and the other for calculating a first corrected chrominance coefficient feature according to the chrominance coefficients and quantization table.

[0036] As can be seen from the above description, in the above embodiments of the present application, two independent learnable offset guiding modules respectively take the luminance coefficient and the chrominance coefficient as inputs, and use the quantization table prior to correct the quantization errors of these coefficients (luminance coefficient and chrominance coefficient), effectively restoring the luminance information and chrominance information lost during the image compression process to generate high-quality reconstructed images.

[0037] Among them, according to an embodiment of the present application, the learnable offset guiding module includes: a quantization offset sub-module for calculating a quantization offset feature according to the frequency domain coefficient and the quantization table; a frequency domain correction sub-module for calculating a first corrected frequency domain coefficient feature according to the quantization offset feature and the frequency domain coefficient.

[0038] As can be seen from the above description, the frequency domain correction sub-module of the above embodiments of the present application calculates a first corrected frequency domain coefficient feature by combining the quantization offset feature and the frequency domain coefficient. This method can effectively compensate for the quantization errors introduced during the compression process, thereby being able to restore the lost detail information of the image and improve the image reconstruction quality (that is, the learnable offset guiding module aims to learn the quantization offset to reduce the quantization errors caused by rounding operations and correct the frequency domain coefficients based on the quantization offset to generate high-quality reconstructed images).

[0039] The following will describe Figure 3 , Figure 4 and Figure 5 the sub-modules of the learnable offset guiding module and the functions they perform.

[0040] 1. Rearrangement sub-module According to an embodiment of the present application, as Figure 4 shown, the learnable offset guiding module further includes: a first rearrangement sub-module for rearranging the input frequency domain coefficients and transmitting the rearranged frequency domain coefficients to the quantization offset sub-module for subsequent processing; a second rearrangement sub-module for rearranging the first corrected frequency domain coefficient feature output by the frequency domain correction sub-module or the dequantized frequency domain coefficient feature output by the dequantization sub-module and transmitting the rearranged result to the domain transformation module for processing.

[0041] For example, rearranging the frequency domain coefficient X to obtain , that is, rearranging the frequency domain coefficients to , so that the rearranged frequency domain coefficients have an inherent frequency correlation in the channel dimension while maintaining the spatial positions between different 8×8 blocks.

[0042] As can be seen from the above description, in the above embodiments of the present application, by rearranging the frequency domain coefficients, the spatial positions between different 8×8 blocks can be kept unchanged, and at the same time, they are arranged in the frequency order in the channel dimension. This arrangement enables the model to more conveniently capture the correlations in the frequency dimension, so as to better adapt to the characteristics of the image reconstruction model, and significantly improve the quality of the reconstructed images of the compressed images by the model.

[0043] 2. Quantization offset sub-module According to an embodiment of the present application, the quantization offset sub-module includes: a parameter pair calculation unit for calculating a first parameter and a second parameter according to a quantization table; and a quantization offset calculation unit for calculating a quantization offset feature according to the frequency domain coefficient, the first parameter, and the second parameter.

[0044] Among them, as Figure 5 shown, the parameter pair calculation unit is composed of a multi-layer perceptron (MLP), a Tanh activation function, and a Sigmiod activation function. In the embodiments of the present application, the quantization table is flattened into a 128-dimensional vector, and a four-layer multi-layer perceptron (MLP) is used to learn the mapping from this vector to the first parameter and the second parameter. The quantization offset calculation unit is composed of a convolution block (the convolution block is composed of two convolution layers (1×1 convolution layers) with a PReLU activation function and a Tanh activation function).

[0045] Among them, according to an embodiment of the present application, still referring to Figure 5 shown, the parameter pair calculation unit is configured to obtain the quantization offset feature in the following manner:

[0046] Among them, represents the quantization offset feature, is the first parameter, represents the intermediate feature of the frequency domain coefficient, is the second parameter.

[0047]

[0048] Among them, represents a convolution layer, represents another convolution layer, is the PReLU activation function, represents the rearranged frequency domain coefficient.

[0049] As can be seen from the above description, and the first parameter and the second parameter are used for the affine transformation of the intermediate feature so that the image reconstruction model can adaptively refine the feature map to generate high-quality reconstructed images.

[0050] 3. Frequency domain correction sub-module According to an embodiment of the present application, still referring to Figure 5 as shown, the frequency domain correction sub-module is configured to obtain the first feature of the corrected frequency domain coefficient in the following manner:

[0051] where represents the first feature of the corrected frequency domain coefficient, is the scaling factor, is the Tanh activation function. Among them, in the embodiment of the present application, is set to 0.5 to constrain the range of quantization offset.

[0052] As can be seen from the above description, in the above embodiment of the present application, by introducing a scaling factor to constrain the range of quantization offset (also known as the quantization range), the quantization error of the frequency domain coefficient introduced during image compression can be corrected according to the rounding prior (scaling factor) and quantization prior (quantization table), effectively restoring the frequency domain information lost during the image compression process to generate a high-quality reconstructed image.

[0053] It should be noted that since the range of quantization error during the image compression process is between [-0.5, 0.5), therefore, the present application presets a scaling factor to constrain the range of quantization offset. 0.5 is only one of the embodiments during the experiment of the present application. The specific setting of the scaling factor can be adjusted according to the specific value of the obtained quantization offset, and no specific limitation is made here.

[0054] As can be seen from the above description, in the above embodiment of the present application, by setting a scaling factor to constrain the range of quantization offset, the quantization error generated during the compression process can be effectively controlled, ensuring that the first feature of the corrected frequency domain coefficient is more accurate and stable. By appropriately adjusting the scaling factor, the quantization offset can be limited within a reasonable range, thereby reducing unnecessary noise introduction and information loss, and improving the quality of image reconstruction or restoration. This method not only helps to optimize the performance of the model, but also improves the processing efficiency, making the frequency domain correction sub-module have stronger adaptability and higher accuracy in different application scenarios.

[0055] 4. Dequantization sub-module According to an embodiment of the present application, still referring to Figure 5 as shown, the learnable offset guidance module further includes a dequantization sub-module, and: the dequantization sub-module is used to calculate the dequantized frequency domain coefficient feature according to the first feature of the corrected frequency domain coefficient and the quantization table; the domain transformation module is further used to transform the dequantized frequency domain coefficient feature into the pixel domain to obtain the second pixel feature of the JPEG image, and calculate the reconstructed image corresponding to the JPEG image based on the second pixel feature.

[0056] Among them, according to an embodiment of the present application, still referring to Figure 5 As shown, the dequantization sub-module is configured to obtain the dequantized frequency-domain coefficient features in the following manner:

[0057] Wherein, represents the dequantized frequency-domain coefficient features, represents the quantization table.

[0058] As can be seen from the above description, the dequantization sub-module of the above embodiment of the present application calculates the dequantized frequency-domain coefficient features based on the corrected first feature of the frequency-domain coefficients and the quantization table. This process effectively reverses the quantization operation applied during the compression process and restores more accurate frequency-domain information, thereby significantly reducing the error and distortion caused by quantization and improving the quality of the reconstructed image.

[0059] (2) Domain transformation module According to an embodiment of the present application, the domain transformation module is configured to: use the inverse discrete cosine transform to transform the corrected first feature of the frequency-domain coefficients or the dequantized frequency-domain coefficient features into the pixel domain to obtain the first pixel feature or the second pixel feature of the JPEG image, and calculate the reconstructed image corresponding to the JPEG image based on the first pixel feature or the second pixel feature.

[0060] For example, use the inverse discrete cosine transform (IDCT) to reconstruct these corrected or dequantized frequency-domain coefficients (luminance coefficients and chrominance coefficients) into a YCbCr image. Since the chrominance image CbCr is downsampled, bilinear interpolation is used to upsample the chrominance image to make its size the same as that of the luminance image Y, and then YCbCr is converted into an RGB image.

[0061] As can be seen from the above description, in the above embodiment of the present application, the domain transformation module uses the inverse discrete cosine transform (IDCT) to transform the corrected or dequantized frequency-domain coefficients (including luminance coefficients and chrominance coefficients) into the pixel domain, generates the pixel features of the JPEG image, and further calculates the reconstructed image of the JPEG image. Among them, for the downsampling characteristic of the chrominance image CbCr, bilinear interpolation is used for upsampling to make its size the same as that of the luminance image Y, and then YCbCr is converted into an RGB image. This design not only efficiently restores the frequency-domain information of the image, but also effectively retains the chrominance details through the interpolation operation, significantly improving the quality and visual expressiveness of the reconstructed image. Especially in the high compression ratio scenario, it can still maintain a high image fidelity and has wide application value.

[0062] (3) Codec module Since the 8×8 block strategy is adopted during the compression process of the JPEG format, there are significant differences in chrominance distortion between blocks.

[0063] According to an embodiment of the present application, still referring to Figure 3 and Figure 4 shown, the image reconstruction model further includes: an encoder-decoder module, configured to calculate a quantization-guided feature of a JPEG image according to a quantization table, a first pixel feature, or a second pixel feature, and calculate a reconstructed image corresponding to the JPEG image based on the quantization-guided feature.

[0064] As can be seen from the above description, the above embodiment of the present application compensates for the error of the chrominance channel in combination with the quantization table, restores more chrominance details, and thus reduces the chrominance distortion problem caused by the downsampling operation during the image compression process.

[0065] Among them, according to an embodiment of the present application, the encoder-decoder module includes: an encoder sub-module, configured to calculate a target intermediate feature of a JPEG image according to a quantization table, a first pixel feature, or a second pixel feature; a decoder sub-module, configured to calculate a quantization-guided feature of the JPEG image according to the target intermediate feature and the quantization table, and calculate a reconstructed image corresponding to the JPEG image based on the quantization-guided feature.

[0066] Among them, according to an embodiment of the present application, the encoder-decoder module includes encoders and decoders of different scales, and the scales of the encoders and decoders correspond one by one.

[0067] As can be seen from the above description, the multi-scale encoder-decoder structure of the above embodiment of the present application provides stronger context learning ability under the same number of parameters, can more accurately restore color details, reduce chrominance distortion, make the color performance of the reconstructed image more accurate and natural, and thus significantly improve the overall visual quality. At the same time, by processing information at multiple resolution levels, this architecture can flexibly handle images with various compression degrees and achieve efficient and high-quality chrominance compensation.

[0068] Next, in combination with Figure 3 , Figure 6 , Figure 7 and Figure 8 the structures of the encoder sub-module and the decoder sub-module will be described in detail.

[0069] 1. Encoder sub-module According to an embodiment of the present application, still referring to Figure 3As shown, the encoder sub-module includes encoders of three scales. Among them, the first-scale encoder is used to calculate the first encoded feature of the JPEG image according to the quantization table, the first pixel feature, or the second pixel feature; the second-scale encoder is used to calculate the second encoded feature of the JPEG image according to the quantization table and the first encoded feature; the third-scale encoder is used to calculate the target intermediate feature of the JPEG image according to the quantization table and the second encoded feature.

[0070] Among them, according to an embodiment of the present application, still referring to Figure 3 As shown, the encoders of the three scales all include a quantization table guiding unit and a residual unit.

[0071] Among them, according to an embodiment of the present application, the encoder of the first scale includes: a first encoding quantization table guiding unit, which is used to calculate the first intermediate feature of the JPEG image according to the quantization table, the first pixel feature, or the second pixel feature; a first encoding residual unit, which is used to calculate the first encoded feature of the JPEG image according to the first intermediate feature; the encoder of the second scale includes: a second encoding quantization table guiding unit, which is used to calculate the second intermediate feature of the JPEG image according to the quantization table and the first encoded feature; a second encoding residual unit, which is used to calculate the second encoded feature of the JPEG image according to the second intermediate feature; the encoder of the third scale includes: a third encoding quantization table guiding unit, which is used to calculate the third intermediate feature of the JPEG image according to the quantization table and the second encoded feature; a third encoding residual unit, which is used to calculate the target intermediate feature of the JPEG image according to the third intermediate feature.

[0072] As Figure 6 shown, it is the structure of the encoding quantization table guiding unit in multiple-scale encoders. Among them, the encoding quantization table guiding unit is composed of a PReLU activation function, two convolutional layers, a multi-layer perceptron (MLP), a Tanh activation function, and a Sigmiod activation function.

[0073] According to an embodiment of the present application, still referring to Figure 6 shown, the encoding quantization table guiding unit in the encoder sub-module is configured to calculate its corresponding output feature in the following manner:

[0074]

[0075] Among them, represents the output feature of the th encoding quantization table guiding unit, represents the input feature of the th encoding quantization table guiding unit, represents the The second convolutional layer in an encoded quantization table guiding unit is a PReLU activation function. denotes the first convolutional layer in the denotes first parameter calculated by the denotes second parameter calculated by the It should be noted that

[0076] As can be seen from the above description, in the above embodiments of the present application, by introducing a multi-scale encoder and a quantization table guiding unit, the pixel features of the JPEG image are gradually encoded at different scales, realizing refined extraction and optimization of the features. Among them, the first quantization table guiding unit, the second quantization table guiding unit, and the third quantization table guiding unit respectively combine the quantization table and the previous layer features to dynamically adjust the image features, effectively reducing the influence of quantization error on the coding performance; at the same time, the residual unit further enhances the feature expression ability by introducing a residual learning mechanism, improving the coding accuracy and robustness. This solution can significantly improve the coding efficiency and quality of the features, enabling the model to generate high-quality reconstructed images.

[0077] 2. Decoder sub-module According to an embodiment of the present application, still referring to Figure 3 shown in

[0078] wherein, according to an embodiment of the present application, still referring to Figure 3 shown in

[0079] Among them, according to an embodiment of the present application, the decoder of the first scale includes: a first decoding residual unit, which calculates the first residual feature of the JPEG image according to the target intermediate feature; a first decoding quantization table guiding unit, which is used to calculate the first decoding feature of the JPEG image according to the first residual feature and the quantization table; the decoder of the second scale includes: a second decoding residual unit, which is used to calculate the second residual feature of the JPEG image according to the first decoding feature; a second decoding quantization table guiding unit, which is used to calculate the second decoding feature of the JPEG image according to the second residual feature and the quantization table; the decoder of the third scale includes: a third decoding residual unit, which calculates the third residual feature of the JPEG image according to the second decoding feature; a third decoding quantization table guiding unit, which is used to calculate the quantization guiding feature of the JPEG image according to the third residual feature and the quantization table, and calculates the reconstructed image corresponding to the JPEG image based on the quantization guiding feature.

[0080] As can be seen from the above embodiments, in the above embodiments of the present application, by introducing a multi-scale decoder and a decoding quantization table guiding unit, the features output by the encoder are gradually decoded and reconstructed at different scales, significantly improving the quality and accuracy of image reconstruction. Among them, the decoding residual unit effectively enhances the expression ability of the features by introducing a residual learning mechanism; the decoding quantization table guiding unit combines the quantization table and the input features to dynamically adjust the image features, reducing the influence of chromatic distortion and quantization error caused during the image compression process on the reconstructed image.

[0081] As Figure 7 shown, it is the structure of the quantization table guiding unit in multiple scale decoders. Among them, the decoding quantization table guiding unit is composed of a PReLU activation function, two convolutional layers, a multi-layer perceptron (MLP), a Tanh activation function, and a Sigmiod activation function.

[0082] Among them, according to an embodiment of the present application, still referring to Figure 7 shown, the decoding quantization table guiding unit in the decoder sub-module is configured to calculate its corresponding output feature in the following manner:

[0083]

[0084] Among them, represents the output feature of the th decoding quantization table guiding unit, represents the input feature of the th decoding quantization table guiding unit, represents the second convolutional layer in the th decoding quantization table guiding unit, is the PReLU activation function, represents the first convolutional layer in the first decoding quantization table guiding unit, represents the first parameter calculated by the first decoding quantization table guiding unit, represents the second parameter calculated by the first decoding quantization table guiding unit. It should be noted that j corresponds to the above-mentioned first decoding quantization table guiding unit, second decoding quantization table guiding unit, and third decoding quantization table guiding unit.

[0085] As can be seen from the above description, the decoding quantization table guiding unit in the above embodiments of the present application adopts structures such as the PReLU activation function, convolutional layer, multi-layer perceptron (MLP), Tanh activation function, and Sigmoid activation function, further optimizing the feature extraction and parameter calculation processes, and ensuring the detail restoration and overall quality of the reconstructed image. This solution can significantly improve the reconstruction effect of JPEG images.

[0086] Among them, according to an embodiment of the present application, a four-layer multi-layer perceptron (MLP) is designed to learn the transformation parameter pairs (i.e., the first parameter and the second parameter) at three scales. The first three layers of the MLP are responsible for generating the shared embedding representation of the quantization table, while the last layer focuses on learning the specific parameter pairs at each scale.

[0087] As can be seen from the above embodiments, the above embodiments of the present application achieve efficient learning of the transformation parameter pairs (the first parameter and the second parameter) at three scales by designing a four-layer multi-layer perceptron (MLP). The first three layers of the MLP effectively capture the global features of the quantization table by generating the shared embedding representation, reducing the computational complexity; while the last layer focuses on learning the specific parameter pairs at each scale, ensuring the pertinence and adaptability of the feature transformation at different scales. This hierarchical design not only improves the efficiency and accuracy of parameter learning, but also enhances the decoding ability of the multi-scale decoder for image features, thus significantly improving the quality and detail restoration degree of image reconstruction.

[0088] Among them, according to an embodiment of the present application, each scale encoder includes 4 encoding residual units, and each scale decoder includes 4 decoding residual units. They are used to process the input features in sequence. It should be noted that the input and output of each residual unit are processed according to the conventional feature extraction method in deep learning, and the input and output of each residual unit will not be described here.

[0089] Among them, according to an embodiment of the present application, the residual units in the encoder and decoder are both composed of two PReLU activation functions with a 3×3 convolutional layer.

[0090] As can be seen from the above description, in the above embodiments of the present application, 4 residual units are introduced in the encoders and decoders at each scale, realizing the deep processing and optimization of the input features. Through residual connection and non-linear transformation, the residual units effectively alleviate the problem of gradient disappearance during the training of the image reconstruction model, enhance the feature expression ability, and at the same time can capture richer detail information. The introduction of this multi-level residual structure significantly improves the feature extraction and reconstruction capabilities of the encoders and decoders, ensuring the accurate encoding of image features and the high-quality reconstruction of images.

[0091] In summary, the encoder sub-module of the above embodiments of the present application includes encoders at 3 scales, and the decoder sub-module includes decoders at 3 scales. Under the same number of parameters, the encoder-decoder pairs at shallow scales (1 scale, 2 scales) have poor context learning ability and large computational overhead; while the encoder-decoder at deep scales (3 scales) can not only reduce the computational amount, but also enhance the artifact removal ability of the model for images with different compression degrees. At the same time, in order to effectively reduce the chromatic distortion phenomenon during the image compression process in the above embodiments, the present application uses the prior knowledge of quantization tables (chromaticity and luminance quantization tables) to perform channel-level affine transformation on the features at different scales of the encoder-decoder network. Among them, a quantization table guiding unit is used to guide the encoder-decoder network to reduce the chromatic distortion in the initially reconstructed RGB image, which is caused by the downsampling operation of chromaticity during the image compression process. The quantization table guiding unit is embedded at both ends of each scale of the encoder and decoder networks, aiming to use the prior information of the quantization tables (such as luminance quantization table and chromaticity quantization table) to guide the networks at different scales to emphasize the features crucial for chromaticity recovery; and a channel attention mechanism is introduced, enabling both the encoder and decoder to efficiently learn the common embedded representations of the quantization tables and use these representations to guide the image reconstruction model to adaptively enhance beneficial features, thereby improving the generalization performance of the image reconstruction model under different degrees of chromatic distortion.

[0092] In addition, in order to make the feature information flow better between different scales, pixel shuffling and inverse pixel shuffling are respectively applied in the model for upsampling and downsampling to reduce the loss of spatial details.

[0093] Therefore, according to an embodiment of the present application, a first sampling unit and a second sampling unit are respectively included between the encoders at 3 scales and between the decoders at 3 scales. Among them, the first sampling unit is used to perform upsampling processing on the input features of the encoder and / or decoder, and transmit the processing result to the encoding quantization table guiding unit and / or the decoding residual unit for processing; the second sampling unit is used to perform downsampling processing on the output features of the encoder and / or decoder, and transmit the processing result to the encoder and / or decoder at the next scale.

[0094] Among them, according to an embodiment of the present application, the first sampling unit and the second sampling unit are respectively configured to perform upsampling and downsampling operations by using the Pixel Shuffle and Inverse Pixel Shuffle methods, respectively.

[0095] As can be seen from the above description, the pixel shuffle in the above embodiment of the present application magnifies the resolution of the feature map by rearranging the pixel positions, while the inverse pixel shuffle reduces the resolution through the opposite operation. This method can effectively reduce the loss of spatial details in the feature map during the scaling process, thereby retaining more image detail information, making the image reconstruction model more efficient and accurate in processing multi-scale features, and improving the overall performance.

[0096] 3. Residual sub-module According to an embodiment of the present application, as Figure 8 shown, the encoding and decoding module further includes a residual sub-module for performing residual processing on the output of the encoder sub-module and outputting the processing result to the decoder sub-module for processing.

[0097] Among them, according to an embodiment of the present application, the residual sub-module is composed of two PReLU activation functions with 3×3 convolutional layers.

[0098] As can be seen from the above description, by introducing the residual sub-module in the encoding and decoding module in the above embodiment of the present application, the output of the encoder sub-module can be subjected to residual processing, thereby effectively capturing and retaining the high-frequency detail information in the input features. The residual sub-module further enhances the feature expression ability by passing the processing result to the decoder sub-module, reduces information loss, and at the same time alleviates the gradient vanishing problem in the deep network. This design significantly improves the feature extraction and reconstruction capabilities of the encoding and decoding module, ensures that the decoder can more accurately restore image details, and ultimately improves the overall performance and quality of the image processing task.

[0099] III. Step S3 According to an embodiment of the present application, step S3 includes: during each iteration training process, updating the parameters of the image reconstruction model by using a preset total loss function.

[0100] Among them, according to an embodiment of the present application, the total loss function is:

[0101] Among them, represents the total loss function, represents the mean absolute error loss function, represents the fast Fourier transform loss function, represents a hyperparameter. Among them, = 0.1. It should be noted that this value is an example value during the experiment of this application, and the specific value can be adjusted according to the specific experimental effect, and no specific limitation is made here.

[0102] Among them, according to an embodiment of the present application,

[0103]

[0104] Among them, is the number of JPEG images, represents the th real image before the JPEG image is uncompressed, represents the th reconstructed image of the JPEG image after passing through the image reconstruction model, is the fast Fourier transform function.

[0105] As can be seen from the above description, in the above embodiment of the present application, the total loss function is constructed by combining the mean absolute error loss and the fast Fourier transform loss to train the image reconstruction model, which can effectively improve the quality of the reconstructed image and the ability to restore details. Among them, the mean absolute error loss directly measures the distance between the reconstructed image and the real image in the pixel domain, ensuring the accuracy of the reconstructed image in the overall structure and low-frequency information; while the fast Fourier transform loss calculates the frequency domain distance between the two by converting the image to the frequency domain, and specifically optimizes the high-frequency information lost during the image compression process (that is, guides the model to restore high-frequency details through the frequency domain loss), so as to guide the image reconstruction model to restore the details and textures of the image. This dual-loss design not only takes into account the overall fidelity of the image, but also significantly enhances the reconstruction effect of high-frequency details, enabling the model to generate visually more realistic and delicate images even in complex compression scenarios.

[0106] Furthermore, based on the method for constructing the image reconstruction model in the above embodiment, according to an embodiment of the present application, the present application proposes an image reconstruction model, which includes: a learnable offset guidance module for calculating the first feature of the corrected frequency domain coefficient according to the frequency domain coefficient and quantization table of the obtained JPEG image; a domain transformation module for converting the first feature of the corrected frequency domain coefficient to the pixel domain to obtain the first pixel feature of the JPEG image, and calculating the reconstructed image corresponding to the JPEG image based on the first pixel feature. For the further functions of the constructed image reconstruction model, please refer to the description in the above embodiment of the construction method, and no repeated description is made here.

[0107] In addition, according to an embodiment of the present application, the present application provides a JPEG decoder, in which the image reconstruction model of the above embodiment is configured.

[0108] As can be seen from the above description, the JPEG decoder proposed in the above embodiment of the present application integrates an image reconstruction model. This decoder can effectively restore the details and textures of the image and generate a visually more realistic and delicate reconstructed image. This design not only improves the limitations of traditional JPEG decoders in compressed image processing but also provides users with a higher-fidelity image decoding experience, having broad application value and practicality.

[0109] In summary, the above embodiment of the present application gradually guides the image reconstruction model to reduce quantization error and chromatic distortion through compression prior knowledge, thereby enhancing the visual quality of the image.

[0110] To verify the effectiveness of the image reconstruction model construction scheme proposed in the above embodiment of the present application, the inventors conducted the following experiments: I. Dataset Description In this experiment, the test sets of LIVE1, BSDS500, and the ICB dataset are used to evaluate the performance of IJCN (i.e., the image reconstruction model proposed in the present application) in the color JPEG image restoration task. For the grayscale JPEG image restoration task, the Y channel is used as the grayscale image to train IJCN, and only the luminance quantization table is used to guide the network, and the performance of IJCN is evaluated on the Classic5 and LIVE1 datasets.

[0111] II. Evaluation Index Description Three evaluation metrics are adopted, namely PSNR (dB), SSIM, and PSNR-B (dB). The higher their values, the better the image restoration effect. Among them, PSNR-B is an index specifically used to evaluate the blocking effect in the image.

[0112] (1) PSNR (Peak Signal-to-Noise Ratio): PSNR is based on the mean square error (MSE) and measures the pixel-level error between the reconstructed image and the real image, with the unit of decibel (dB). Advantages: Simple to calculate and clear physical meaning (directly reflecting pixel error). Disadvantages: Weak correlation with human visual perception (HVS), sensitive to global brightness changes but insensitive to local structures (such as textures, edges).

[0113] (2) SSIM (Structural Similarity Index): SSIM evaluates image similarity from three dimensions: brightness, contrast, and structure, ranging from [-1, 1] (1 indicates complete identity). Advantages: More consistent with human visual perception, sensitive to the preservation of structural information (such as edges and textures). Comprehensively evaluates brightness, contrast, and structural differences. Disadvantages: The computational complexity is higher than PSNR, and the sensitivity to local distortion (such as blocking effects) is limited.

[0114] (3) PSNR-B (Block Effect Peak Signal-to-Noise Ratio): PSNR-B is an improved version of PSNR, which is specifically used to evaluate block effects. Block effects are common in images compressed or processed by JPEG, and appear as discontinuous boundaries between adjacent blocks. Advantages: It directly quantifies the severity of block effects, making up for the defects of PSNR and SSIM that are insensitive to block effects. It provides targeted guidance for the optimization of compression algorithms (such as JPEG). Disadvantages: It is only applicable to scenes with block effects (such as compressed images) and ignores distortion in non-border areas.

[0115] 3. Experimental results

[0116] Table 1 above is a comparison of experimental results for artifact removal of color JPEG images, where the PSNR (dB), SSIM and PSNR-B (dB) values ​​corresponding to the quality factors of existing methods such as JPEG, QGAC, FBCNN, CRL, and DAGN are set to 10, 20, 30, and 40 respectively, are not as high as those obtained by the solution proposed in this application. The higher the values ​​of these indicators, the better the image restoration effect.

[0117]

[0118] Table 2 above is a comparison of experimental results for grayscale JPEG image artifact removal, where the corresponding PSNR (dB), SSIM and PSNR-B (dB) values ​​when the quality factors of existing methods such as JPEG, QGAC, FBCNN, CRL, and DAGN are set to 10, 20, 30, and 40 respectively (the higher the quality factor, the better the compressed image quality) are not as high as those obtained by the solution proposed in this application. The higher the values ​​of these indicators, the better the image restoration effect.

[0119] As described above, the inventor compared IJCN (i.e., the image reconstruction model proposed in this application) with several state-of-the-art methods, including QGAC, FBCNN, CRL, and DAGN. The evaluation focused on the cases where the quality factors were 10, 20, 30, and 40. As shown in Tables 1 and 2, IJCN outperformed the existing methods in the color JPEG image restoration task and achieved performance comparable to that of the state-of-the-art methods in the grayscale JPEG image restoration task, demonstrating its robustness and generalization ability in enhancing JPEG images.

[0120] In summary, compared with the existing JPEG artifact removal methods, the above embodiments of this application take into account the root cause of image distortion during the image compression process, that is, inputting the prior knowledge of compression into the image reconstruction model, enabling the image reconstruction model to gradually reduce the distortion caused by irreversible operations during the compression process, thereby obtaining high-quality reconstructed images.

[0121] Based on the inventive concept of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the above embodiments are implemented. The following will be described in detail in conjunction with Figure 9 for detailed description.

[0122] As Figure 9 shown, it shows the electronic device 100 of this application, which may specifically include a processor 110 and a memory 120. The memory 120 is coupled to the processor 110.

[0123] The processor 110 is used to control the operation of the electronic device. The processor 110 may also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 110 may be an integrated circuit chip with signal processing capabilities. The processor 110 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 110 may also be any conventional processor, etc.

[0124] The memory 120 is used to store computer programs, which can be RAM, ROM, or other types of storage terminals. Specifically, the memory 120 may include one or more computer-readable storage media, which can be non-transitory or transitory. The memory 120 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage terminals and flash storage terminals. In some embodiments, the non-transitory computer-readable storage media in the memory 120 is used to store at least one program code.

[0125] The processor 110 is used to execute the computer programs stored in the memory 120 to implement the methods described in the method embodiments of this application.

[0126] In some embodiments, the electronic device may further include: a peripheral terminal interface 130 and at least one peripheral terminal. The processor 110, the memory 120, and the peripheral terminal interface 130 may be connected through a bus or signal lines. Each peripheral terminal may be connected to the peripheral terminal interface 130 through a bus, signal lines, or a circuit board. Specifically, the peripheral terminal includes at least one of a radio frequency circuit 140, a display screen 150, an audio circuit 160, and a power supply 170.

[0127] The peripheral terminal interface 130 can be used to connect at least one peripheral terminal related to I / O (Input / Output) to the processor 110 and the memory 120. In some embodiments, the processor 110, the memory 120, and the peripheral terminal interface 130 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 110, the memory 120, and the peripheral terminal interface 130 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0128] The radio frequency circuit 140 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 140 communicates with the communication network and other Internet of Things devices through electromagnetic signals, and the radio frequency circuit 140 is the communication circuit of the electronic device. The radio frequency circuit 140 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 140 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 140 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 140 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.

[0129] The display screen 150 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 150 is a touch display screen, the display screen 150 also has the ability to collect touch signals on or above the surface of the display screen 150. The touch signal can be input to the processor 110 as a control signal for processing. At this time, the display screen 150 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 150, which is provided on the front panel of the electronic device; in other embodiments, there may be at least two display screens 150, which are respectively provided on different surfaces of the electronic device or are in a foldable design; in other embodiments, the display screen 150 may be a flexible display screen, which is provided on a curved surface or a foldable surface of the electronic device. Even, the display screen 150 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 150 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0130] The audio circuit 160 may include a microphone and a speaker. The microphone is used to collect sound waves of the operator and the environment, and convert the sound waves into electrical signals for input to the processor 110 for processing, or for input to the radio frequency circuit 140 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 110 or the radio frequency circuit 140 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 160 may further include a headphone jack.

[0131] The power supply 170 is used to supply power to each component in the electronic device. The power supply 170 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 170 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0132] For a detailed description of the functions and execution processes of each functional module or component in the electronic device embodiment of the present application, reference may be made to the description in the respective method embodiments of the present application above, and details will not be repeated here.

[0133] In several embodiments provided in the present application, it should be understood that the disclosed electronic device and method can be implemented in other ways. For example, the electronic device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some data may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in electrical, mechanical or other forms.

[0134] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0136] Based on the inventive concept of the above embodiments, the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor, performs the steps of the method described in any of the above embodiments. The following is combined with Figure 10 to illustrate the execution process of the above embodiments in the computer-readable storage medium.

[0137] As Figure 10 shown, it shows the computer-readable storage medium of the present application. If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the computer-readable storage medium 200. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions / computer programs to enable an Internet of Things device (which can be a personal computer, a server, or a network terminal, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, as well as electronic terminals such as computers, mobile phones, laptop computers, tablet computers, cameras, etc. having the above storage media.

[0138] The description of the execution process of the program data in the computer-readable storage medium can be referred to the description in the above method embodiments of the present application, and will not be repeated here.

[0139] The above are only embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

[0140] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

Claims

1. A method for constructing an image reconstruction model, the image reconstruction model being used to reconstruct a JPEG image, characterized in that, The method includes: Obtaining a JPEG image and decoding it to obtain the frequency domain coefficients and quantization table of the JPEG image; Constructing an initial image reconstruction model based on a deep learning method; Forming a data set with the frequency domain coefficients and the quantization table, and using the data set to perform multiple iterative trainings on the initial image reconstruction model until convergence to obtain the image reconstruction model.

2. The method for constructing an image reconstruction model according to claim 1, wherein The initial image reconstruction model takes the frequency domain coefficients and the quantization table as inputs, and the reconstructed image corresponding to the JPEG image as the output, and the initial image reconstruction model includes: A learnable offset guidance module for calculating a first corrected frequency domain coefficient feature according to the frequency domain coefficients and the quantization table; A domain transformation module for transforming the first corrected frequency domain coefficient feature into the pixel domain to obtain a first pixel feature of the JPEG image, and calculating the reconstructed image corresponding to the JPEG image based on the first pixel feature.

3. The method for constructing an image reconstruction model according to claim 2, wherein The learnable offset guidance module includes: A quantization offset sub-module for calculating a quantization offset feature according to the frequency domain coefficients and the quantization table; A frequency domain correction sub-module for calculating the first corrected frequency domain coefficient feature according to the quantization offset feature and the frequency domain coefficients.

4. The method for constructing an image reconstruction model according to claim 2, wherein The learnable offset guidance module further includes a dequantization sub-module, and: The dequantization sub-module is used for calculating a dequantized frequency domain coefficient feature according to the first corrected frequency domain coefficient feature and the quantization table; The domain transformation module is further used for transforming the dequantized frequency domain coefficient feature into the pixel domain to obtain a second pixel feature of the JPEG image, and calculating the reconstructed image corresponding to the JPEG image based on the second pixel feature.

5. The method for constructing an image reconstruction model according to claim 4, wherein The image reconstruction model further includes: An encoder-decoder module for calculating a quantization guidance feature of the JPEG image according to the quantization table, the first pixel feature or the second pixel feature, and calculating the reconstructed image corresponding to the JPEG image based on the quantization guidance feature.

6. The method for constructing an image reconstruction model according to claim 5, wherein The encoder-decoder module includes: An encoder sub-module for calculating a target intermediate feature of the JPEG image according to the quantization table, the first pixel feature or the second pixel feature; A decoder sub-module for calculating the quantization guidance feature of the JPEG image according to the target intermediate feature and the quantization table, and calculating the reconstructed image corresponding to the JPEG image based on the quantization guidance feature.

7. An image reconstruction model obtained based on the construction method according to any one of claims 1-6, characterized in that, The image reconstruction model includes: A learnable offset guidance module for calculating a first corrected frequency domain coefficient feature according to the frequency domain coefficients and the quantization table of the obtained JPEG image; A domain transformation module for transforming the first corrected frequency domain coefficient feature into the pixel domain to obtain a first pixel feature of the JPEG image, and calculating the reconstructed image corresponding to the JPEG image based on the first pixel feature.

8. An application method of an image reconstruction model, characterized in that, The method includes: Obtaining a JPEG image and decoding it to obtain the frequency domain coefficients and quantization table of the JPEG image; Based on the frequency domain coefficients and the quantization table, use the image reconstruction model described in claim 7 to calculate the reconstructed image corresponding to the JPEG image.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of claims 1-6, 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1-6, 8.

Citation Information

Patent Citations

  • Compressed image restoration method and device, equipment and storage medium

    CN112927146A

  • JPEG image lossless compression and decompression method, system and device

    CN113810693A

  • Image reconstruction method and device, image coding and decoding methods and devices, reconstruction model training method and device and related equipment

    CN114004743A

  • Image processing method and device based on Transform model

    CN114067009A

  • Reversible steganography method and system based on frequency domain image information hiding and electronic equipment

    CN118587077A