Construction and application method and system of image reconstruction system, equipment and medium
By constructing an image reconstruction system with a two-stage task decoupling architecture, using the quality-aware prior learning model and a hierarchical guide feedback model, the adaptability problem of JPEG artifact removal method in complex image content and compression scenarios is solved, and high-quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202510775056.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing JPEG artifact removal method is difficult to adapt to complex and changeable image content and compression scenarios, resulting in the compressed JPEG image losing details and distortion cannot be effectively suppressed, and the reconstruction image quality is not high.
A two-stage task decoupling architecture is built, including a quality-aware prior learning model and a hierarchical guided feedback model, and multi-level iterative residual feature extraction and feature alignment and fusion are carried out respectively to capture the quality changes between regions and local artifact distribution characteristics, and realize artifact removal and detail recovery.
The reconstruction quality of JPEG images is significantly improved, local detail recovery and global consistency are enhanced, and the visual effect of the image is optimized.
Smart Images

Figure CN120298534A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a construction and application method, model, device, and medium of an image reconstruction system. Background Art
[0002] In the field of modern digital image processing, the rapid growth of image data has brought great challenges to storage, transmission and processing. In order to cope with this problem, image compression technology has emerged. Its core purpose is to reduce data redundancy while retaining visual information as much as possible, thereby reducing storage space and transmission bandwidth requirements. JPEG, as a widely used lossy compression standard, significantly reduces file size by removing high-frequency details that are not sensitive to the human eye through quantization and encoding, while maintaining high subjective visual quality. This efficient compression method makes JPEG the mainstream format for image storage and transmission, especially in resource-constrained scenarios (such as network transmission, mobile devices, etc.), its advantages are more prominent. However, in the image compression process, especially lossy compression methods such as JPEG, its compression ratio is usually represented by the quality factor (QF). At low QF, significant artifacts are usually introduced, resulting in a decrease in visual quality, affecting the visual effect of the image and the performance of subsequent processing tasks. However, in many practical applications (such as medical image analysis, security monitoring, satellite remote sensing, etc.), the requirements for image quality are high, and directly using compressed images may not meet the actual application needs.
[0003] Therefore, JPEG artifact removal technology came into being, which aims to reduce the distortion caused by compression and restore the visual quality of the image. Clear JPEG images are crucial for downstream low-level visual tasks that use them as input, including contrast enhancement, super-resolution, edge detection, etc. Traditional JPEG artifact removal methods mainly use manually designed filters to filter artifacts, or formulate artifact removal as a maximum a posteriori problem to introduce a priori-based method. However, traditional methods rely on manually designed prior knowledge and have limited restoration effects, so they are often difficult to adapt to complex and changing image content and compression scenarios.
[0004] In summary, there is an urgent need for a solution that can effectively remove artifacts that appear in compressed images, which can restore lost details in compressed images, suppress distortion, and improve the overall quality of the image. Summary of the invention
[0005] The present application provides a method, system, device, and medium for constructing and applying an image reconstruction system to solve the problem that the existing JPEG artifact removal solution is difficult to adapt to complex and changeable image content and compression scenarios, resulting in the inability to restore lost details in the compressed JPEG image and suppress distortion, thereby making the quality of the reconstructed JPEG image low.
[0006] The first aspect of the present application provides a method for constructing an image reconstruction system, where the image reconstruction system is used to reconstruct JPEG images. The construction method of the present application includes: Obtain a JPEG image and construct a data set based on it; Construct an initial image reconstruction system, where the initial image reconstruction system includes: a quality-aware prior learning model and a hierarchical guided feedback model, and: The quality-aware prior learning model is configured to: perform multi-level iterative residual feature extraction operations on the JPEG image to obtain hierarchical compressed prior residual features of the JPEG image; The hierarchical guided feedback model is configured to: perform feature extraction operations on the JPEG image to obtain compressed features of the JPEG image, and perform alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculate a reconstructed image of the JPEG image based on the results of the operations; Use the data set to perform multiple iterative trainings on the initial image reconstruction system until convergence to obtain the image reconstruction system.
[0007] In some embodiments of the present application, the hierarchical compressed prior residual features include end-layer compressed prior residual features, and the step of constructing the initial image reconstruction system includes: The quality-aware prior learning model is further configured to: perform a mapping operation on the end-layer compressed prior residual features to obtain a quality factor prediction value of the JPEG image; The hierarchical guided feedback model is configured to: calculate a reconstructed image of the JPEG image according to the quality factor prediction value and the local calibration features, where the local calibration features are the results of the fusion operation.
[0008] In some embodiments of the present application, the step of constructing the initial image reconstruction system further includes: The quality-aware prior learning model is further configured to: perform residual feature extraction operations on the end-layer compressed prior residual features, and perform a mapping operation on the results of the residual feature extraction operations to obtain a quality factor prediction value of the JPEG image.
[0009] In some embodiments of the present application, the hierarchical guided feedback model includes: a feature extraction sub-module, an encoder sub-module, and a decoder sub-module, and the step of constructing the initial image reconstruction system further includes: The feature extraction sub-module is configured to: perform feature extraction operations on the JPEG image to obtain compressed features of the JPEG image; The encoder sub-module is configured to: perform alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features to obtain the local calibrated features of the JPEG image, and calculate the global calibrated encoded features of the JPEG image based on the predicted quality factor value and the local calibrated features; The decoder sub-module is configured to: calculate the reconstructed image of the JPEG image based on the global calibrated encoded features.
[0010] In some embodiments of the present application, the steps of constructing the initial image reconstruction system further include: The decoder sub-module is further configured to: calculate the global calibrated decoded features of the JPEG image based on the local calibrated features or the global calibrated encoded features and the predicted quality factor value, and calculate the reconstructed image of the JPEG image based on the global calibrated decoded features.
[0011] In some embodiments of the present application, the steps of iteratively training the initial image reconstruction system using a dataset until convergence include: Iteratively training the quality-aware prior learning model using a preset first loss function until convergence; Construct the initial image reconstruction system by combining the converged quality-aware prior learning model with the hierarchical guidance feedback model, and then iteratively train the initial image reconstruction system using a preset second loss function until convergence. Among them, during each iterative training process, the parameters of the quality-aware prior learning model are frozen and participated in the training.
[0012] The second aspect of the present application provides an image reconstruction system constructed by using the construction method of the image reconstruction system in any of the above embodiments. The image reconstruction system includes: A quality-aware prior learning model for performing multi-level iterative residual feature extraction operations on the acquired JPEG image to obtain the hierarchical compressed prior residual features of the JPEG image; A hierarchical guidance feedback model for performing feature extraction operations on the JPEG image to obtain the compressed features of the JPEG image, and performing alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculating the reconstructed image of the JPEG image based on the results of the operations.
[0013] The third aspect of the present application provides an application method of an image reconstruction system. The application method of the present application includes: Obtain a JPEG image; Use the quality-aware prior learning model in the image reconstruction system as described in the above embodiments to perform multi-level iterative residual feature extraction operations on the JPEG image to obtain the hierarchical compressed prior residual features of the JPEG image; Using the hierarchical guided feedback model in the image reconstruction system as described in the above embodiments, perform feature extraction operations on the JPEG image to obtain the compressed features of the JPEG image, and perform alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculate the reconstructed image of the JPEG image based on the results of the operations.
[0014] The fourth aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in any one of the first aspect and the third aspect in the above embodiments.
[0015] The fifth aspect of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method described in any one of the first aspect and the third aspect in the above embodiments.
[0016] This application has the following beneficial effects: In the above embodiments of this application, by constructing a two-stage task decoupling architecture (that is, constructing an initial image reconstruction system composed of a quality-aware prior learning model in the first stage and a hierarchical guided feedback model in the second stage) to respectively implement the extraction of compressed prior features of JPEG images and image reconstruction (that is, artifact removal). Among them, the architecture of the first stage (that is, the quality-aware prior learning model) is used to perform multi-level iterative residual feature extraction operations on the JPEG image to obtain fine-grained hierarchical compressed prior residual features to capture the quality changes between regions; the architecture of the second stage (that is, the hierarchical guided feedback model) is used to perform alignment and fusion processing based on the compressed features of the JPEG image and the hierarchical compressed prior residual features, and can capture the distribution characteristics of local artifacts to suppress image artifacts and restore image details, thereby improving the quality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with this application and are used together with the specification to illustrate the technical solutions of this application.
[0018] Figure 1 It is a schematic flowchart of an embodiment of the method for constructing an image reconstruction system provided by this application; Figure 2 It is a schematic diagram of the framework of multiple embodiments of the image reconstruction system provided by this application; Figure 3 It is a schematic diagram of the framework of an embodiment of the quality-aware prior learning model provided by this application; Figure 4 It is a schematic diagram of the framework of an embodiment of the local alignment calibration unit provided by this application; Figure 5 It is a schematic diagram of the framework of an embodiment of the global boot calibration encoding unit provided by the present application; Figure 6 It is a schematic diagram of the framework of an embodiment of the global boot calibration decoding unit provided by the present application; Figure 7 It is a schematic diagram of the framework of an embodiment of the electronic device provided by the present application; Figure 8 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium provided by the present application. Detailed implementation manners
[0019] Next, in conjunction with the accompanying drawings of the specification, the solutions of the embodiments of the present application will be described in detail.
[0020] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0021] As used herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" herein means two or more than two. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0022] The inventors have found through research that the progress of deep learning has enabled JPEG artifact removal to evolve from methods based on hand-designed priors to methods based on deep learning. Relevant researchers have achieved significant improvements in artifact removal by constructing deep neural networks to learn the mapping relationship between compressed images and original images from large-scale paired datasets. However, this learning of the mapping relationship between compressed images and original images mainly relies on the data fitting ability of neural networks and fails to utilize the inherent prior knowledge of JPEG to guide the reconstruction process, still hindering the development of artifact removal models. The latest JPEG artifact removal methods mainly use the prior information of the quality factor to guide the restoration of the reconstructed image. However, as a global indicator, the quality factor can only roughly reflect the overall compression intensity of the image and cannot accurately describe the distribution characteristics of local artifacts. Due to the significant heterogeneity of image content, there are obvious differences in the artifact distribution between texture-rich regions and smooth regions. Simply relying on the global quality factor is difficult to achieve local adaptive artifact removal, thus restricting the overall restoration quality of the reconstructed image. In addition, most existing methods adopt a single-stage network architecture, coupling prior extraction and artifact removal in the same network model for image reconstruction, ignoring the internal conflict between the two, that is, accurate estimation of the degradation prior requires extracting compression distortion information, while artifact removal requires suppressing these distortion information to restore image details. This conflict in learning objectives makes it difficult for the single-stage network architecture to optimize the two tasks simultaneously, limiting the further improvement of the artifact removal effect and image quality.
[0023] To solve the above problems, this application proposes a new JPEG artifact removal scheme. By constructing a two-stage task decoupling architecture to separately achieve the extraction of compression prior features of JPEG images and image reconstruction (i.e., artifact removal). Among them, the architecture of the first stage (i.e., the quality-aware prior learning model) is used to perform multi-level iterative residual feature extraction operations on JPEG images to obtain fine-grained hierarchical compression prior residual features to capture the quality changes between regions; the architecture of the second stage (i.e., the hierarchical guided feedback model) is used to perform alignment and fusion processing based on the compression features of JPEG images and hierarchical compression prior residual features, which can capture the distribution characteristics of local artifacts to achieve suppressing image artifacts and restoring image details, thereby improving the quality of the reconstructed image.
[0024] The following describes this application in conjunction with the accompanying drawings and specific embodiments.
[0025] According to an embodiment of this application, as Figure 1 shown, this application proposes a method for constructing an image reconstruction system, and this method includes: S1, obtaining a JPEG image and constructing a dataset based on it; S2, constructing an initial image reconstruction system, where, as Figure 2As shown in the figure, the initial image reconstruction system includes: a quality-aware prior learning model and a hierarchical guided feedback model, and: The quality-aware prior learning model is configured to: perform a multi-level iterative residual feature extraction operation on the JPEG image to obtain the hierarchical compression prior residual features of the JPEG image; The hierarchical guided feedback model is configured to: perform a feature extraction operation on the JPEG image to obtain the compression features of the JPEG image, and perform an alignment and fusion operation on the hierarchical compression prior residual features and the compression features, and calculate the reconstructed image of the JPEG image based on the result of the operation; S3. Use the dataset to perform multiple iterative trainings on the initial image reconstruction system until convergence to obtain the image reconstruction system. Each model will be described in detail below.
[0026] I. Quality-Aware Prior Learning Model (1) The hierarchical compression prior residual features include the outputs of all residual layers According to an embodiment of the present application, the hierarchical compression prior residual features include the final layer compression prior residual features, and the steps of constructing the initial image reconstruction system include: The quality-aware prior learning model is further configured to: perform a mapping operation on the final layer compression prior residual features to obtain the quality factor prediction value of the JPEG image; The hierarchical guided feedback model is configured to: calculate the reconstructed image of the JPEG image according to the quality factor prediction value and the first local calibration feature, where the first local calibration feature is the result of the fusion operation.
[0027] From the above description, it can be seen that in the above embodiment of the present application, the quality-aware prior learning model learns the global quality factor prior (i.e., the quality factor prediction value) and the local hierarchical compression feature prior (i.e., the hierarchical compression prior residual features) from the input JPEG image. These prior information can not only reflect the global quality change, but also accurately describe the distribution characteristics of local compression artifacts, providing important guidance for subsequent image reconstruction to improve the quality of the reconstructed image.
[0028] (2) The hierarchical compression prior residual features include the outputs of some residual layers According to an embodiment of the present application, the steps of constructing the initial image reconstruction system further include: The quality-aware prior learning model is further configured to: perform a residual feature extraction operation on the final layer compression prior residual features, and perform a mapping operation on the result of the residual feature extraction operation to obtain the quality factor prediction value of the JPEG image.
[0029] From the above description, it can be seen that in the above embodiment of the present application, by further performing a residual feature extraction operation on the final layer compression prior residual features, the expression ability of the features can be further enhanced, so as to obtain a more accurate quality factor prediction value of the JPEG image.
[0030] Among them, according to an embodiment of the present application, the quality-aware prior learning model is an improved ResNet34. As Figure 2 and 3 shown, the improved ResNet34 includes a convolutional layer, 4 residual layers, a pooling layer, and a multi-layer perceptron (MLP). The process of obtaining the hierarchical compression prior residual features of the JPEG image is as follows: First, the input JPEG image is preliminarily processed through a 1×1 convolutional layer to expand the number of channels to 64; Subsequently, four residual layers are introduced, each of which is stacked by multiple residual blocks, and each residual block contains two 3×3 convolutional layers with a ReLU activation function in the middle to extract and deepen the image features layer by layer. The output residual features of the first three residual layers are used as the hierarchical compression prior residual features (i.e., when the output of the third residual layer is the final layer compression prior residual feature) to guide the feature calibration in the second stage. In order to effectively capture residual features of different scales, the output channel numbers of each residual layer are set to 64, 128, 256, and 512 respectively, gradually enhancing the feature representation ability. At the same time, after each residual layer, a convolutional layer with a stride of 2 is applied for downsampling to reduce the resolution of the residual feature map and increase the receptive field, thereby capturing more extensive context information. Still referring to Figure 3 shown, the number of residual blocks in each residual layer is 3, 4, 6, and 3 in sequence to balance the model complexity and performance. After completing the feature extraction of all residual layers, the size of the residual feature map is compressed to 1×1 through adaptive average pooling for subsequent quality factor prediction. It should be noted that constructing the quality-aware prior learning model based on ResNet34 in this application is only one embodiment of this application. With the development of deep network models, other residual networks with better performance can also be used to construct the quality-aware prior learning model in this application, which is not specifically limited here.
[0031] The inventors have found through research that using MLP to map the hierarchical compression prior residual features to the QF label (i.e., the true value of the quality factor), this method can effectively process local and complex artifacts in JPEG images. For this reason, according to an embodiment of the present application, a multi-layer perceptron (MLP) is introduced in this application. Still referring to Figure 2 and 3 shown, the MLP consists of three fully connected layers, and ReLU activation functions are connected after the first two layers to enhance the non-linear expression ability of the model. The MLP takes a 512-dimensional feature vector as input and uses a hidden layer with 512 neurons, which can effectively capture local artifacts while retaining global features. Finally, a Sigmoid activation function is applied at the output end of the MLP to generate the quality factor prediction value of the JPEG image , and its value is restricted within the range of [0,1]. The specific formula is as follows:
[0032] Among them, represents the Sigmoid activation function, and M represents the multi-layer perceptron, and N refers to the network layer before the multi-layer perceptron (i.e., the convolutional layer, four residual layers, and pooling layer), while is the input JPEG image (which can be a color three-channel or grayscale single-channel). In order to obtain a QF prediction value that better meets the actual application requirements, the output value is multiplied by 100 and rounded down to finally obtain an integer QF prediction value.
[0033] As can be seen from the above description, the above embodiments of the present application construct a quality-aware prior learning model based on a deep residual network to extract a composite prior composed of a global quality factor and local hierarchical compression features. These prior information can not only reflect global quality changes but also accurately describe the distribution characteristics of local compression artifacts, providing important guidance for subsequent image reconstruction to improve the quality of the reconstructed image.
[0034] II. Hierarchical guided feedback model (1) Introducing the quality factor prediction value at the encoder end According to an embodiment of the present application, still referring to Figure 2 as shown, the hierarchical guided feedback model includes: a feature extraction sub-module, an encoder sub-module, and a decoder sub-module. And the steps of constructing the initial image reconstruction system further include: the feature extraction sub-module is configured to: perform a feature extraction operation on the JPEG image to obtain the compression features of the JPEG image; the encoder sub-module is configured to: perform an alignment and fusion operation on the hierarchical compression prior residual features and the compression features to obtain the local calibration features of the JPEG image, and calculate the global calibration coding features of the JPEG image according to the quality factor prediction value and the local calibration features; the decoder sub-module is configured to: calculate the reconstructed image of the JPEG image according to the global calibration coding features.
[0035] As can be seen from the above embodiments, the above embodiments of the present application use the hierarchical compressed prior residual features and quality factor prediction values extracted by the quality perception prior learning model in the first stage to guide the feature calibration and image reconstruction process of the hierarchical guidance feedback model in the second stage. Among them, the hierarchical compressed prior residual features are used to align and refine local features, suppress artifacts and restore details; while the quality factor prediction values are used to guide global information integration to ensure color consistency and structural integrity. Through this hierarchical guidance strategy, the present application maintains the global consistency of the image while restoring details, significantly improving the artifact removal effect. Among them, introducing the quality factor prediction value into the encoder sub-module can effectively guide the integration of global information to ensure color consistency and structural integrity. Through the quality factor prediction value, the encoder sub-module can better capture the overall features of the JPEG image, suppress artifacts and enhance detail restoration, so as to provide higher-quality globally calibrated encoded features for the subsequent decoder sub-module.
[0036] Among them, according to an embodiment of the present application, the local calibration feature includes a first local calibration feature, the global calibration encoded feature includes a first global calibration encoded feature, and still referring to Figure 2 As shown, the encoder sub-module includes a first-scale local alignment calibration unit (i.e., LCM1), a first-scale global guidance calibration encoding unit (i.e., GCM-EN1), and a first-scale downsampling unit (i.e., DS1). Among them, the first-scale local alignment calibration unit is configured to perform alignment and fusion operations on the hierarchical compressed prior residual features and compressed features to obtain the first local calibration feature of the JPEG image; the first-scale global guidance calibration encoding unit is configured to calculate the first global calibration encoded feature of the JPEG image according to the quality factor prediction value and the first local calibration feature; the first-scale downsampling unit is configured to perform a downsampling operation on the first global calibration encoded feature and transmit the processing result to the decoder sub-module for processing. Still referring to Figure 2 As shown, the decoder sub-module includes a first-scale upsampling unit (i.e., US1), and the first-scale upsampling unit is configured to perform an upsampling operation on the first global calibration encoded feature after the downsampling operation output by the first-scale downsampling unit, and calculate the reconstructed image of the JPEG image based on the result after the upsampling operation. The encoder sub-module of this embodiment includes: LCM1, GCM-EN1, DS1, and the decoder sub-module includes: US1.
[0037] As can be seen from the above description, in the above embodiments of the present application, by introducing a downsampling operation in the encoder sub-module, high-level features of an image can be effectively extracted and compressed, reducing the computational complexity while retaining key information to guide subsequent reconstruction; by introducing an upsampling operation in the decoder sub-module, the detailed information of the image can be gradually restored. The combination of downsampling and upsampling in this encoder-decoder structure not only optimizes the efficiency of feature extraction and reconstruction, but also enhances the model's ability to restore local details and global consistency, thus significantly improving the quality of image reconstruction.
[0038] Among them, according to an embodiment of the present application, the local calibration feature further includes a second local calibration feature, the global calibration encoding feature further includes a second global calibration encoding feature, and still referring to Figure 2 As shown, the encoder sub-module further includes: a second-scale local alignment and calibration unit (i.e., LCM2), a second-scale global guidance and calibration encoding unit (i.e., GCM-EN2), and a second-scale downsampling unit (i.e., DS2). Among them, the second-scale local alignment and calibration unit is configured to perform alignment and fusion operations on the first global calibration encoding feature after the downsampling operation output by the first-scale downsampling unit to obtain the second local calibration feature of the JPEG image; the second-scale global guidance and calibration encoding unit is configured to calculate the second global calibration encoding feature of the JPEG image according to the quality factor prediction value and the second local calibration feature; the second-scale downsampling unit is configured to perform a downsampling operation on the second global calibration encoding feature and transmit the processing result to the decoder sub-module for processing. Still referring to Figure 2 As shown, the decoder sub-module further includes a second-scale upsampling unit (i.e., US2). The second-scale upsampling unit is configured to perform an upsampling operation on the second global calibration encoding feature after the downsampling operation output by the second-scale downsampling unit and transmit the result to the first-scale upsampling unit for processing to calculate the reconstructed image of the JPEG image. The encoder sub-module of this embodiment includes: LCM1, GCM-EN1, DS1, LCM2, GCM-EN2, DS2, and the decoder sub-module includes: US2, US1.
[0039] Among them, according to an embodiment of the present application, the local calibration feature further includes a third local calibration feature, the global calibration encoding feature further includes a third global calibration encoding feature, and still referring to Figure 2As shown, the encoder sub-module further includes: a third-scale local alignment calibration unit (i.e., LCM3) and a third-scale global guidance calibration encoding unit (i.e., GCM-EN3). Among them, the third-scale local alignment calibration unit is configured to perform alignment and fusion operations on the second global calibration encoding features after the downsampling operation output by the second-scale downsampling unit to obtain the third local calibration features of the JPEG image; the third-scale global guidance calibration encoding unit is configured to calculate the third global calibration encoding features of the JPEG image based on the quality factor prediction value and the third local calibration features and transmit them to the second-scale upsampling unit for processing. The encoder sub-module of this embodiment includes: LCM1, GCM-EN1, DS1, LCM2, GCM-EN2, DS2, LCM3, GCM-EN3, and the decoder sub-module includes: US2, US1.
[0040] Among them, according to an embodiment of the present application, there are 4 GCM-EN1, GCM-EN2, and GCM-EN3 each. The specific quantity is not limited by this embodiment and can be dynamically adjusted according to the experimental environment, experimental results, etc. during the experiment.
[0041] It should be noted that the number of scale settings of the local alignment calibration unit, global guidance calibration encoding unit, and downsampling unit in the encoder sub-module is not Figure 2 limited by the shown, and the specific scale and quantity settings can be dynamically adjusted according to the experimental results. Moreover, the number of scales and the scale sizes of the downsampling unit and the upsampling unit correspond one by one.
[0042] As can be seen from the above description, in the above embodiments of the present application, by introducing local alignment calibration units, global guidance calibration encoding units, and downsampling units of different scales in the encoder sub-module and upsampling units of different scales in the decoder sub-module, it is possible to perform fine-grained alignment and global guidance on features at multiple resolution levels respectively, effectively capture and fuse local details and global structure information of the image, and significantly improve the effects of artifact removal and detail restoration; at the same time, the introduction of multi-scale downsampling units and upsampling units realizes multi-level extraction and reconstruction of features, enhances the model's ability to express multi-scale information of the image, further optimizes the quality of the reconstructed image, and makes it perform excellently in terms of local details, global consistency, and visual naturalness.
[0043] (2) Introducing the quality factor prediction value at the decoder end According to an embodiment of the present application, the steps of constructing the initial image reconstruction system further include: the decoder sub-module is further configured to: calculate the global calibration decoding features of the JPEG image based on the local calibration features and the quality factor prediction value, and calculate the reconstructed image of the JPEG image based on the global calibration decoding features.
[0044] As described above, in the decoder sub-module of the present application, the predicted value of the quality factor is introduced, which can directly guide the reconstruction process to ensure that the color and structure of the final output image meet the expectations. By using the predicted value of the quality factor, the decoder sub-module can adjust the reconstruction strategy more flexibly and optimize the balance between local details and global consistency. This design enables the decoder sub-module to restore image details more accurately while simplifying the complexity of the encoder sub-module.
[0045] Among them, according to an embodiment of the present application, the global calibration decoding feature includes a first global calibration decoding feature, and still referring to Figure 2 As shown, the decoder sub-module further includes: a first-scale global guidance calibration decoding unit (i.e., GCM-DE1), and: the first-scale upsampling unit is further configured to perform an upsampling operation on the first locally calibrated feature after the downsampling operation output by the first-scale downsampling unit; the first-scale global guidance calibration decoding unit is configured to: calculate the first global calibration decoding feature of the JPEG image based on the upsampled first locally calibrated feature output by the first-scale upsampling unit and the predicted value of the quality factor, and calculate the reconstructed image of the JPEG image based on the first global calibration decoding feature. The encoder sub-module of this embodiment includes: LCM1, DS1, and the decoder sub-module includes: US1, GCM-DE1.
[0046] Among them, according to an embodiment of the present application, the global calibration decoding feature further includes a second global calibration decoding feature, and still referring to Figure 2 As shown, the decoder sub-module further includes: a second-scale global guidance calibration decoding unit (i.e., GCM-DE2), and: the second-scale upsampling unit is further configured to perform an upsampling operation on the second locally calibrated feature after the downsampling operation output by the second-scale downsampling unit; the second-scale global guidance calibration decoding unit is configured to: calculate the second global calibration decoding feature of the JPEG image based on the upsampled second locally calibrated feature output by the second-scale upsampling unit and the predicted value of the quality factor, and transmit it to the first-scale upsampling unit for processing. The encoder sub-module of this embodiment includes: LCM1, DS1, LCM2, DS2, and the decoder sub-module includes: US2, GCM-DE2, US1, GCM-DE1.
[0047] Among them, according to an embodiment of the present application, the global calibration decoding feature further includes a third global calibration decoding feature, and still referring to Figure 2As shown, the decoder sub-module further includes: a third-scale global guidance calibration decoding unit (i.e., GCM-DE3), and the third-scale global guidance calibration decoding unit is configured to: calculate the third global calibration decoding feature of the JPEG image based on the third local calibration feature and the predicted quality factor value, and transmit it to the second-scale upsampling unit for processing. The encoder sub-module of this embodiment includes: LCM1, DS1, LCM2, DS2, LCM3, and the decoder sub-module includes: GCM-DE3, US2, GCM-DE3, US1, GCM-DE1.
[0048] Among them, according to an embodiment of the present application, there are 4 GCM-DE3, GCM-DE2, and GCM-DE1 each. The specific quantity is not limited by this embodiment and can be dynamically adjusted according to the experimental environment, experimental results, etc. during the experiment.
[0049] It should be noted that the number of scale settings of the local alignment calibration unit and the global guidance calibration decoding unit in the decoder sub-module is not limited by Figure 2 the shown limitation. The specific scale and quantity settings can be dynamically adjusted according to the experimental results. The global guidance calibration encoding unit is not introduced in the encoder sub-module of the above embodiment, and there are only local alignment calibration units and downsampling units of different scales. Among them, the number of scales and the scale sizes of the local alignment calibration units correspond one by one to the global guidance calibration decoding units.
[0050] As can be seen from the above description, by introducing global guidance calibration decoding units of different scales in the decoder sub-module in the above embodiments of the present application, multi-scale feature information can be fully utilized, the global structural consistency of the image can be restored layer by layer, and the different-scale features can be dynamically calibrated in combination with the predicted quality factor value, thereby effectively improving the detail restoration degree and visual naturalness of the reconstructed image.
[0051] (3) Introduce the predicted quality factor value at both the encoder and decoder ends According to an embodiment of the present application, the steps of constructing the initial image reconstruction system further include: the decoder sub-module is further configured to calculate the global calibration decoding feature of the JPEG image based on the global calibration encoding feature and the predicted quality factor value, and calculate the reconstructed image of the JPEG image based on the global calibration decoding feature.
[0052] As can be seen from the above embodiments, in the above embodiments of the present application, by introducing the predicted quality factor value in both the encoder and decoder sub-modules, multi-level residual feature calibration and reconstruction optimization can be achieved. The encoder sub-module uses the predicted quality factor value to integrate global information to ensure the high quality of the initial features; the decoder sub-module further uses the predicted quality factor value to refine the reconstruction process, restoring details and maintaining global consistency. This strategy of introducing the quality factor prior significantly improves the artifact removal effect, while taking into account the restoration of local details and global consistency, thereby obtaining a higher-quality reconstructed image.
[0053] Among them, according to an embodiment of the present application, the third-scale global guidance calibration decoding unit is configured to: calculate the third global calibration decoding feature of the JPEG image based on the third global calibration encoding feature output by the third-scale global guidance calibration encoding unit and the predicted quality factor value, and transmit it to the second-scale upsampling unit for upsampling processing, and transmit the upsampling processing result to the second-scale global guidance calibration decoding unit for processing. The encoder sub-module of this embodiment includes: LCM1, GCM-EN1, DS1, LCM2, GCM-EN2, DS2, LCM3, GCM-EN3, and the decoder sub-module includes: GCM-DE3, US2, GCM-DE3, US1, GCM-DE1. The combination method of the hierarchical guidance feedback model can also be that the encoder sub-module includes: LCM1, GCM-EN1, DS1, and the decoder sub-module includes: GCM-DE3, US1, GCM-DE1, or the encoder sub-module includes: LCM1, GCM-EN1, DS1, LCM2, GCM-EN2, DS2, and the decoder sub-module includes: US2, GCM-DE3, US1, GCM-DE1. The specific combination method is not limited by this embodiment and can be dynamically adjusted according to the experimental environment, experimental results, etc. during the experiment.
[0054] It should be noted that the number of scale settings of the local alignment calibration unit, global guidance calibration encoding unit, downsampling unit, and upsampling unit in the encoder sub-module is not Figure 2 limited as shown, and the specific scale and number settings can be dynamically adjusted according to the experimental results. The number of scale settings of the local alignment calibration unit and global guidance calibration decoding unit in the decoder sub-module is not Figure 2 limited as shown, and the specific scale and number settings can be dynamically adjusted according to the experimental results. At the same time, when supported by the experimental hardware and experimental computing efficiency, DS1 and DS2 in the encoder sub-module and US2 and US1 in the decoder sub-module of the above embodiments can all be not designed.
[0055] As can be seen from the above description, in the above embodiments of the present application, global guidance calibration coding units of different scales are introduced in the encoder sub-module, which can dynamically calibrate features in combination with the predicted quality factor values, effectively extract and retain the global structure information of the image at different resolutions, and provide high-quality coding features for the subsequent decoding process. At the same time, global guidance calibration decoding units of different scales are introduced in the decoder sub-module, and the predicted quality factor values are further used to guide and optimize the features layer by layer, realizing the gradual reconstruction of features from coarse to fine, significantly improving the global consistency and detail restoration degree of the reconstructed image, and finally achieving better effects in artifact removal, detail retention, and overall visual quality.
[0056] In addition, according to an embodiment of the present application, the hierarchical guidance feedback model in the second stage is designed based on an asymmetric encoder-decoder architecture similar to UNet. Using the coarse-grained quality factor prediction value (i.e., QF prior) and the fine-grained hierarchical compression prior residual feature (i.e., hierarchical compression feature prior) extracted by the quality-aware prior learning model in the first stage, the artifact removal of JPEG images at different compression levels is guided from coarse to fine. Among them, the coarse-grained QF prior is used to guide global restoration, while the fine-grained prior is used to refine local details and correct region-specific artifacts, so as to achieve more accurate image reconstruction. The encoder sub-module integrates local alignment calibration units, global guidance calibration coding units, and downsampling units of different scales to process and enhance features of different scales. The downsampling unit includes a convolutional layer with a kernel size of 2 and a stride of 2 for performing downsampling operations. The decoder sub-module then upsamples the features processed by the global guidance calibration decoding unit through transposed convolution to gradually restore the details and structure of the compressed image. A skip connection is provided between the encoder sub-module and the decoder main module to retain the spatial detail information of different scales and ensure the integrity of feature transmission.
[0057] Among them, according to an embodiment of the present application, the local alignment calibration units of different scales are all configured to: perform attention processing on their corresponding inputs using the attention mechanism to obtain a spatial attention weight map, and then use the spatial attention weight map to adjust the compressed features processed by convolution, and perform alignment and fusion operations on the compressed features and the adjusted compressed features to obtain the corresponding local calibration features.
[0058] As can be seen from the above description, in the above embodiments of the present application, the local alignment calibration units are integrated into each scale of the encoder. By enhancing the feature extraction and context understanding capabilities through the attention mechanism, the features can be fused and aligned within the local alignment calibration units, thereby suppressing fine-grained artifacts and accurately restoring local details.
[0059] The structures of the different-scale local alignment calibration units and global guidance calibration encoding units of the encoder sub-module and the structure of the different-scale global guidance calibration decoding units of the decoder sub-module will be described in detail below.
[0060] (1) Local alignment calibration unit of the encoder sub-module According to an embodiment of the present application, the local alignment calibration units of different scales all include two independent 3×3 convolutional layers and 1×1 convolutional layer, and as Figure 4 shown, the local alignment calibration unit is configured to obtain local calibration features in the following manner:
[0061]
[0062] Among them, represents the spatial attention weight map output by the scale local alignment calibration unit, is the Sigmoid activation function, is the 1×1 convolutional layer, and are both 3×3 convolutional layers, represents the compressed feature of the JPEG image input to the scale local alignment calibration unit, represents the feature input to the scale local alignment calibration unit. Among them, if it is the first-scale local alignment calibration unit, then is the hierarchical compression prior residual feature input to the first-scale local alignment calibration unit; if it is the second-scale local alignment calibration unit, then is the feature output by the first-scale downsampling unit; if it is the third-scale local alignment calibration unit, then is the feature output by the second-scale downsampling unit. By element-wise adding and fusing the results of and to obtain a feature map, and further refining it using a 1×1 convolutional layer; then, applying the Sigmoid activation function to generate , which is used to adjust the compressed feature output after . Still referring to Figure 2 shown, the input and output features in each LCM have the same dimension. For example, in the first LCM, the dimensions of the input and output features are both , and the feature dimensions in subsequent LCMs are halved layer by layer. Through the spatial attention weight map By adjusting the compression features of JPEG images, LCM can enhance the feature representation of important regions, thereby effectively suppressing local artifacts and restoring details. Finally, a residual connection is introduced to align the adjusted compression features with the compression features input to LCM to prevent information loss and ensure the integrity of the features. This design enables LCM to capture the distribution characteristics of local artifacts at different scales, providing important support for subsequent image restoration.
[0063] (2) Global Guided Calibration Encoding Unit of the Encoder Submodule Integrate the quality factor prior into the encoder submodule and dynamically adjust the feature extraction strategy according to the image quality. As Figure 5 shown, the structures of the global guided calibration encoding units at different scales all include an encoding residual block, an encoding initializer, and an encoding generator. Among them, the encoding residual block consists of two convolutional layers and a ReLU activation function, which is used to capture high-frequency details and residual information, thereby refining the features obtained from the previous processing; the encoding initializer includes a three-layer multi-layer perceptron, which is used to map the QF prediction value to a high-dimensional feature space to obtain , enabling the model to adjust the feature extraction process according to the image quality; the encoding generator includes a fully connected layer and a Sigmoid or Tanh activation function to generate a modulation parameter pair ( ), and the encoding generator is configured to obtain in the following way:
[0064]
[0065]
[0066] Among them, is the Sigmoid activation function, is the fully connected layer, represents the quality factor prediction value in the high-dimensional feature space, is the hyperbolic tangent function. Then, apply the modulation parameter pair ( ) to transform the features processed by the encoding residual block. The transformation formula is:
[0067] Among them, is the output of the encoding residual block in the global guided calibration encoding unit at the scale, represents the first modulation parameter output by the encoding generator in the global guided calibration encoding unit at the scale, represents the Scale the second modulation parameter output by the encoding generator in the global guidance calibration encoding unit. Finally, a residual connection is introduced to prevent information loss and ensure the integrity of features. The features input to the encoding residual block in each scale global guidance calibration encoding unit are the outputs of the corresponding scale local alignment calibration unit.
[0068] As can be seen from the above description, the modulation parameter ( ) of the above embodiments of the present application is used to scale and offset features, enabling the model to adapt to images of different compression levels, thereby optimizing the output quality. This design enables the global guidance calibration encoding unit of the encoder sub-module to dynamically adjust the feature extraction strategy according to the image quality, thereby enhancing color consistency and structural integrity globally and further improving the effect of artifact removal.
[0069] (3) Global guidance calibration decoding unit of the decoder sub-module Integrate the quality factor prior into the decoder sub-module and dynamically adjust the feature extraction strategy according to the image quality. As Figure 6 shown, the structure of the global guidance calibration decoding unit includes a decoding residual block, a decoding initializer, and a decoding generator. The decoding residual block consists of two convolutional layers and a ReLU activation function, and is used to capture high-frequency details and residual information, thereby refining the features obtained from the previous processing; the decoding initializer includes a three-layer multi-layer perceptron, which is used to map the QF prediction value to a high-dimensional feature space to obtain , enabling the model to adjust the feature extraction process according to the image quality; the decoding generator includes a fully connected layer and a Sigmoid or Tanh activation function to generate a modulation parameter pair ( ), and the decoding generator is configured to obtain in the following manner:
[0070]
[0071]
[0072] Among them, is the Sigmoid activation function, is the fully connected layer, represents the quality factor prediction value in the high-dimensional feature space, is the hyperbolic tangent function. Then, apply the modulation parameter pair ( ) to transform the features processed by the decoding residual block, and the transformation formula is:
[0073] Among them, is the Scale the output of the decoding residual block in the global guidance calibration decoding unit, denote the first modulation parameter output by the decoding generator in the global guidance calibration decoding unit, denote the second modulation parameter output by the decoding generator in the global guidance calibration decoding unit. Finally, a residual connection is introduced to prevent information loss and ensure the integrity of the features.
[0074] As can be seen from the above description, the modulation parameter ( ) of the above embodiments of the present application is used to scale and offset the features, enabling the model to adapt to images of different compression levels, thereby optimizing the output quality. This design enables the global guidance calibration decoding unit of the decoder sub-module to dynamically adjust the feature extraction strategy according to the image quality, thereby enhancing color consistency and structural integrity globally and further improving the effect of artifact removal.
[0075] In addition, according to an embodiment of the present application, step S3 includes: using a preset first loss function to perform multiple iterative trainings on the quality-aware prior learning model until convergence; constructing an initial image reconstruction system by combining the converged quality-aware prior learning model with a hierarchical guidance feedback model, and then using a preset second loss function to perform multiple iterative trainings on the initial image reconstruction system until convergence, where, during each iterative training process, the parameters of the quality-aware prior learning model are frozen and then participate in the training.
[0076] As can be seen from the above description, the above embodiments of the present application perform iterative training on the quality-aware prior learning model through a preset first loss function until convergence to ensure that the model can effectively capture image quality features; subsequently, the converged model is combined with a hierarchical guidance feedback model to form an initial image reconstruction system, and iterative training is performed using a preset second loss function. At the same time, the parameters of the quality-aware prior learning model are frozen during the training process (the frozen parameters refer to the parameters of the quality-aware prior learning model that remain the parameters converged by training using the first loss function during the training process of the second loss function) to avoid the interference of its parameter update on the stability of the overall system. This method can significantly improve the performance of the image reconstruction system, ensure the high quality and high fidelity of the reconstructed image, and improve the training efficiency and model robustness at the same time.
[0077] Among them, according to an embodiment of the present application, the first loss function is:
[0078] where N is the number of JPEG images in the dataset, denote the predicted quality factor value of the m-th JPEG image, denote the true quality factor value of the m-th JPEG image.
[0079] As can be seen from the above description, the above embodiments of the present application adopt the first loss function based on the difference between the predicted value and the true value of the quality factor, which can effectively optimize the training process of the quality-aware prior learning model and ensure that the model can accurately predict the quality factor of JPEG images. The design of this loss function enables the model to better learn the image quality features during the training process, thereby improving the accuracy and robustness of the model's perception of image quality. In addition, by minimizing the difference between the predicted value and the true value, this embodiment can significantly improve the performance of the model in the image reconstruction task and provide more reliable quality factor prior information for subsequent image processing tasks.
[0080] Among them, according to an embodiment of the present application, the second loss function is:
[0081] Among them, represents the m-th JPEG image, represents the reconstructed image of the m-th JPEG image.
[0082] As can be seen from the above description, the above embodiments of the present application can effectively optimize the training process of the initial image reconstruction system by adopting the second loss function based on the difference between the original JPEG image and the reconstructed image, ensuring that the reconstructed image is highly consistent with the original image in terms of details and overall quality. The design of this loss function enables the system to better learn the fine features of image reconstruction during the training process, thereby significantly improving the visual quality and fidelity of the reconstructed image.
[0083] In order to verify the effectiveness of the above-described embodiment construction method, the inventors conducted the following experiments: First, the DIV2K and Flickr2K datasets were subjected to different degrees of JPEG compression with the help of the OpenCV library to obtain JPEG image datasets under different quality factors; then, the JPEG image datasets were used to fine-tune the ResNet34 model pre-trained on the large-scale ImageNet dataset to extract hierarchical compression prior residual features and predict the quality factors.
[0084] To demonstrate the advantages of the technical solution proposed in this application, three evaluation metrics are adopted to measure the performance of different JPEG artifact removal models. Among them, Peak Signal-to-Noise Ratio (PSNR) evaluates the image quality by calculating the pixel-level difference between the reconstructed image and the original image; Structural Similarity Index Measure (SSIM) is an image quality assessment method based on structural information, which takes into account the similarities in brightness, contrast, and structure of the image, and can better reflect the perceptual characteristics of the human visual system than PSNR; Peak Signal-to-Noise Ratio for Blockiness (PSNR-B), which is sensitive to blocking artifacts, is an improved version based on traditional PSNR, and better reflects the local distortion of the image by introducing block-level calculations. The larger the values of these three metrics, the better the performance of the model. This application has fully verified the proposed solution in terms of reconstruction quality and computational complexity.
[0085] (1) Reconstruction quality: To prove the superiority of this application in reconstruction quality, it is compared with JPEG artifact removal methods such as QGAC, FBCNN, EARN, and DAGN on the color LIVE1 and grayscale Classic5 datasets. Among them, the DAGN method is only for removing artifacts from color JPEG images. It can be observed from Table 1 (Table 1 shows the comparison experiment results on the color LIVE1 dataset) and Table 2 (Table 2 shows the comparison experiment results on the Classic5 dataset) that the performance of this application on both datasets is significantly better than the comparison methods, confirming the effectiveness and superiority of the image reconstruction system constructed in this application in removing artifacts and restoring image quality. In addition, this application further analyzes the influence of the embedding position of the QF prediction value in the global guidance calibration unit (corresponding to the global guidance calibration encoding unit and the global guidance calibration decoding unit) in the codec sub-module on the final image reconstruction result. It can be found from Table 3 (Table 3 shows the ablation study of the QF embedding position on the grayscale Classic5 dataset (PSNR / SSIM / PSNR-B)) that the collaborative embedding of QF prior information at both ends of the codec sub-module can more fully guide the network learning, thus improving the quality and consistency of image reconstruction.
[0086] Table 1
[0087] Table 2
[0088] Table 3
[0089] (2)Computational complexity: To evaluate the computational efficiency, the present application was compared with other advanced methods in terms of the number of parameters, the amount of floating-point operations, and the running time (comparison of the number of parameters, the amount of floating-point operations, and the running time on 512×512 color images). The experimental results are shown in Table 4. The experimental results indicate that the present application achieves superior performance with lower computational overhead, demonstrating an effective balance between quality and efficiency.
[0090] Table 4
[0091] As can be seen from the above experiments, compared with the conflict of the existing single-stage framework, the present application proposes a blind JPEG artifact removal method based on two-stage hierarchical quality-aware guidance (i.e., constructing an image reconstruction system with a two-stage architecture). Through the separated processing of "prior learning-guided reconstruction", the internal conflict in the task objectives is effectively alleviated. To accurately reconstruct the image, the present application introduces a coarse-to-fine guidance strategy, combining fine hierarchical compression prior residual features with rough QF priors to guide image reconstruction. In addition, to better guide image reconstruction, a local alignment calibration unit and a global guidance calibration encoder-decoder unit are proposed to effectively utilize QF prior information to calibrate features for prior guidance, solving the defect that existing JPEG artifact removal methods cannot capture local quality changes, thereby significantly improving the quality of JPEG images, enhancing local details while maintaining global consistency. In addition, compared with other advanced technologies, the method proposed in the present application can achieve competitive or even superior performance with only a small number of parameters and floating-point operations, demonstrating an effective balance between quality and efficiency; at the same time, for various JPEG compression datasets, whether grayscale images or color images, the present application can achieve leading reconstruction performance, indicating the effectiveness and superiority of the image reconstruction system constructed by the construction method proposed in the present application in accurately restoring image quality. This method also shows excellent reconstruction ability in visual effects, better retaining the details and edge information of the image while removing artifacts, making the restored image closer to the original image, indicating that this method has broad application prospects and practical value.
[0092] Based on the construction method of the above embodiments of the present application, according to an embodiment of the present application, the present application proposes an image reconstruction system. The image reconstruction system includes: a quality-aware prior learning model for performing multi-level iterative residual feature extraction operations on the acquired JPEG image to obtain hierarchical compressed prior residual features of the JPEG image; a hierarchical guidance feedback model for performing feature extraction operations on the JPEG image to obtain compressed features of the JPEG image, and performing alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculating a reconstructed image of the JPEG image based on the result of the operation. For the further functions of the constructed image reconstruction model, please refer to the description in the above-mentioned construction method embodiments, and will not be repeated here.
[0093] In addition, according to an embodiment of the present application, the present application proposes an application method of an image reconstruction system. The application method of the present application includes: acquiring a JPEG image; using the quality-aware prior learning model in the image reconstruction system as described in the above embodiments to perform multi-level iterative residual feature extraction operations on the JPEG image to obtain hierarchical compressed prior residual features of the JPEG image; using the hierarchical guidance feedback model in the image reconstruction system as described in the above embodiments to perform feature extraction operations on the JPEG image to obtain compressed features of the JPEG image, and performing alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculating a reconstructed image of the JPEG image based on the result of the operation.
[0094] Meanwhile, according to an embodiment of the present application, the present application proposes a JPEG decoder, in which the image reconstruction system of the above embodiment is configured.
[0095] As can be seen from the above description, the JPEG decoder proposed in the above embodiments of the present application integrates an image reconstruction system. This decoder can effectively restore the details and textures of the image and generate a visually more realistic and delicate reconstructed image. This design not only improves the limitations of traditional JPEG decoders in compressed image processing, but also provides users with a higher-fidelity image decoding experience, and has broad application value and practicality.
[0096] In summary, compared with the existing artifact removal solutions using a single-stage architecture, the present application constructs a two-stage task decoupled architecture (i.e., constructs an initial image reconstruction system composed of a quality-aware prior learning model in the first stage and a hierarchical guided feedback model in the second stage) to respectively extract the compression prior features of JPEG images and perform image reconstruction (i.e., artifact removal). Among them, the architecture in the first stage (i.e., the quality-aware prior learning model) is used to perform multi-level iterative residual feature extraction operations on JPEG images to obtain fine-grained hierarchical compression prior residual features to capture the quality changes between regions; the architecture in the second stage (i.e., the hierarchical guided feedback model) is used to perform alignment and fusion processing based on the compression features of JPEG images and the hierarchical compression prior residual features, and can capture the distribution characteristics of local artifacts to suppress image artifacts and restore image details, thereby improving the quality of the reconstructed image.
[0097] Based on the inventive concept of the above embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the above embodiments are implemented. The following will be described in detail in conjunction with Figure 7 for a detailed description.
[0098] As Figure 7 shown, it shows the electronic device 100 of the present application, which may specifically include a processor 110 and a memory 120. The memory 120 is coupled to the processor 110.
[0099] The processor 110 is used to control the operation of the electronic device. The processor 110 may also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 110 may be an integrated circuit chip with signal processing capabilities. The processor 110 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 110 may also be any conventional processor, etc.
[0100] The memory 120 is used to store computer programs, which may be RAM, ROM, or other types of storage terminals. Specifically, the memory 120 may include one or more computer-readable storage media, which may be non-transitory or transitory. The memory 120 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage terminals, flash storage terminals. In some embodiments, the non-transitory computer-readable storage medium in the memory 120 is used to store at least one program code.
[0101] The processor 110 is configured to execute the computer programs stored in the memory 120 to implement the methods described in the method embodiments of the present application.
[0102] In some embodiments, the electronic device may further include: a peripheral terminal interface 130 and at least one peripheral terminal. The processor 110, the memory 120, and the peripheral terminal interface 130 may be connected through a bus or signal lines. Each peripheral terminal may be connected to the peripheral terminal interface 130 through a bus, signal lines, or a circuit board. Specifically, the peripheral terminal includes at least one of a radio frequency circuit 140, a display screen 150, an audio circuit 160, and a power supply 170.
[0103] The peripheral terminal interface 130 can be used to connect at least one peripheral terminal related to I / O (Input / Output) to the processor 110 and the memory 120. In some embodiments, the processor 110, the memory 120, and the peripheral terminal interface 130 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 110, the memory 120, and the peripheral terminal interface 130 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0104] The radio frequency circuit 140 is configured to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 140 communicates with a communication network and other Internet of Things devices through electromagnetic signals, and the radio frequency circuit 140 is the communication circuit of the electronic device. The radio frequency circuit 140 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 140 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 140 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 140 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0105] The display screen 150 is used to display the UI (User Interface, the operator interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 150 is a touch display screen, the display screen 150 also has the ability to collect touch signals on or above the surface of the display screen 150. The touch signals can be input to the processor 110 as control signals for processing. At this time, the display screen 150 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 150, which is provided on the front panel of the electronic device; in other embodiments, there can be at least two display screens 150, which are respectively provided on different surfaces of the electronic device or are in a foldable design; in other embodiments, the display screen 150 can be a flexible display screen, which is provided on a curved surface or a folding surface of the electronic device. Even, the display screen 150 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 150 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0106] The audio circuit 160 may include a microphone and a speaker. The microphone is used to collect sound waves of the operator and the environment, and convert the sound waves into electrical signals and input them to the processor 110 for processing, or input them to the radio frequency circuit 140 to achieve voice communication. For the purpose of stereo collection or noise reduction, there can be multiple microphones, which are respectively provided at different parts of the electronic device. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signals from the processor 110 or the radio frequency circuit 140 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 160 can also include a headphone jack.
[0107] The power supply 170 is used to supply power to each component in the electronic device. The power supply 170 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 170 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0108] For a detailed description of the functions and execution processes of each functional module or component in the embodiment of the electronic device of the present application, reference can be made to the description in the above method embodiments of the present application, and details are not described herein again.
[0109] In several embodiments provided by the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the various embodiments of the electronic devices described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some data can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0110] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0112] Based on the inventive concept of the above embodiments, the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method described in any of the above embodiments are performed. The following combination Figure 8 illustrates the execution process of the above embodiments in the computer-readable storage medium.
[0113] Such as Figure 8As shown, it shows the computer-readable storage medium of the present application. If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the computer-readable storage medium 200. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions / computer programs for causing an Internet of Things device (which can be a personal computer, a server, or a network terminal, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, as well as electronic terminals such as computers, mobile phones, laptop computers, tablet computers, cameras, etc. having the above storage media.
[0114] The description of the execution process of the program data in the computer-readable storage medium can be referred to the description in the above method embodiments of the present application, and will not be repeated here.
[0115] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
[0116] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
Claims
1. A method for constructing an image reconstruction system, the image reconstruction system being used for reconstructing JPEG images, characterized in that, The method includes: Obtaining a JPEG image and constructing a dataset based on it; Constructing an initial image reconstruction system, where the initial image reconstruction system includes: a quality-aware prior learning model and a hierarchical guided feedback model, and: The quality-aware prior learning model is configured to: perform multi-level iterative residual feature extraction operations on the JPEG image to obtain hierarchical compressed prior residual features of the JPEG image; The hierarchical guided feedback model is configured to: perform feature extraction operations on the JPEG image to obtain compressed features of the JPEG image, and perform alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculate a reconstructed image of the JPEG image based on the results of the operations; Using the dataset to perform multiple iterative trainings on the initial image reconstruction system until convergence to obtain the image reconstruction system.
2. The method for constructing an image reconstruction system according to claim 1, characterized in that, The hierarchical compressed prior residual features include final-layer compressed prior residual features, and the step of constructing the initial image reconstruction system includes: The quality-aware prior learning model is further configured to: perform a mapping operation on the final-layer compressed prior residual features to obtain a quality factor prediction value of the JPEG image; The hierarchical guided feedback model is configured to: calculate a reconstructed image of the JPEG image according to the quality factor prediction value and local calibration features, where the local calibration features are the results of the fusion operation.
3. The method for constructing an image reconstruction system according to claim 2, wherein The step of constructing the initial image reconstruction system further includes: The quality-aware prior learning model is further configured to: perform residual feature extraction operations on the final-layer compressed prior residual features, and perform a mapping operation on the results of the residual feature extraction operations to obtain a quality factor prediction value of the JPEG image.
4. The method for constructing an image reconstruction system according to claim 2, wherein, The hierarchical guided feedback model includes: a feature extraction sub-module, an encoder sub-module, and a decoder sub-module, and the step of constructing the initial image reconstruction system further includes: The feature extraction sub-module is configured to: perform feature extraction operations on the JPEG image to obtain compressed features of the JPEG image; The encoder sub-module is configured to: perform alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features to obtain local calibration features of the JPEG image, and calculate global calibration coding features of the JPEG image according to the quality factor prediction value and the local calibration features; The decoder sub-module is configured to: calculate a reconstructed image of the JPEG image according to the global calibration coding features.
5. The method for constructing an image reconstruction system according to claim 4, characterized in that, The step of constructing the initial image reconstruction system further includes: The decoder sub-module is further configured to: calculate global calibration decoding features of the JPEG image according to the local calibration features or the global calibration coding features and the quality factor prediction value, and calculate a reconstructed image of the JPEG image based on the global calibration decoding features.
6. The method for constructing an image reconstruction system according to claim 1, wherein The step of using the dataset to perform multiple iterative trainings on the initial image reconstruction system until convergence includes: Performing multiple iterative trainings on the quality-aware prior learning model using a preset first loss function until convergence; Constructing the initial image reconstruction system with the converged quality-aware prior learning model and the hierarchical guided feedback model, and then performing multiple iterative trainings on the initial image reconstruction system using a preset second loss function until convergence. Wherein, during each iterative training process, the parameters of the quality-aware prior learning model are frozen and participate in the training.
7. An image reconstruction system constructed by a construction method of the image reconstruction system as described in claim 1, characterized in that, The image reconstruction system includes: A quality-aware prior learning model for performing multi-level iterative residual feature extraction operations on the acquired JPEG image to obtain the hierarchical compressed prior residual features of the JPEG image; A hierarchical guided feedback model for performing feature extraction operations on the JPEG image to obtain the compressed features of the JPEG image, and performing alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculating the reconstructed image of the JPEG image based on the results of the operations.
8. A method for applying an image reconstruction system, characterized in that, The method includes: Obtaining a JPEG image; Using the quality-aware prior learning model in the image reconstruction system as claimed in claim 7 to perform multi-level iterative residual feature extraction operations on the JPEG image to obtain the hierarchical compressed prior residual features of the JPEG image; Using the hierarchical guided feedback model in the image reconstruction system as claimed in claim 7 to perform feature extraction operations on the JPEG image to obtain the compressed features of the JPEG image, and performing alignment and fusion operations on the hierarchical compressed prior residual features and the compressed features, and calculating the reconstructed image of the JPEG image based on the results of the operations.
9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method as claimed in any one of claims 1-6, 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method as claimed in any one of claims 1-6, 8 are implemented.
Citation Information
Patent Citations
JPEG (Joint Photographic Experts Group) image artifact removal method based on comparative learning and application thereof
CN115829858A
Image compressed sensing joint reconstruction method and system, storage medium and electronic equipment
CN120125683A
Multi-realism image compression with a conditional generator
WO2024129940A1