Image restoration model processing method and device, program product and electronic equipment

By introducing a hybrid noise and diffusion modeling mechanism into the diffusion model, the problem of assuming Gaussian white noise in the image compression process is solved, achieving accurate restoration and detail enhancement of compressed images, which is suitable for industrial intelligent vision scenarios.

CN121053017APending Publication Date: 2025-12-02CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511197178.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve accurate restoration during image compression. Traditional diffusion models assume noise to be Gaussian white noise, which limits the quality of the restored image.

Method used

A diffusion model is adopted, in which mixed noise (including Gaussian noise and structured compression noise) is gradually added during the forward diffusion process and noise is gradually removed during the reverse diffusion process. By fusing the mixed noise with the diffusion modeling mechanism, structure-aware image restoration training is carried out.

Benefits of technology

It achieves artifact removal, detail enhancement, and structure restoration of compressed images, improving image restoration quality and making it suitable for industrial intelligent vision scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053017A_ABST
    Figure CN121053017A_ABST
Patent Text Reader

Abstract

The invention provides an image restoration model processing method and device, a program product and electronic equipment, and the method comprises the steps: obtaining an original image and a compressed image corresponding to the original image; taking the original image as a training target, taking a compressed image corresponding to the original image as training input, training the image restoration model, and obtaining a trained image restoration model under the condition that a preset training ending condition is met; wherein the image restoration model adopts a diffusion model, and in the forward diffusion process, mixed noise is gradually added to an original image to obtain a noise-added image; in the back diffusion process, denoising the compressed image step by step to obtain a denoised image; the mixed noise comprises Gaussian noise and structured compression noise. According to the invention, the problem of lack of real compression artifact supervision in the standard diffusion training process can be overcome, and artifact removal, detail enhancement and structure reduction of the compressed image are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image restoration model processing method, an image restoration model processing device, a computer program product, and an electronic device. Background Technology

[0002] Image compression, while reducing image data size, often results in information loss and structural damage, manifesting as blurred details, missing textures, and severe blockiness. Although image enhancement algorithms have attempted to remove artifacts, most are based on CNN (Convolutional Neural Networks) or Transformer architectures, and their restoration quality is limited by the model's expressive power. Traditional diffusion models, while capable of image generation and detail restoration, often simply assume noise to be Gaussian white noise, making it difficult to accurately restore compressed images.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] This disclosure provides an image restoration model processing method, an image restoration model processing device, a computer program product, and an electronic device to at least partially solve the problem of accurate restoration of compressed images.

[0005] According to a first aspect of this disclosure, an image restoration model processing method is provided, comprising: acquiring an original image and a compressed image corresponding to the original image; training an image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input; and obtaining a trained image restoration model when a preset training termination condition is met; wherein the image restoration model adopts a diffusion model, wherein during the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image; and during the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image; wherein the mixed noise includes Gaussian noise and structured compression noise.

[0006] According to a second aspect of this disclosure, an image restoration model processing apparatus is provided. The apparatus includes: an image acquisition module for acquiring an original image and a compressed image corresponding to the original image; and a model training module for training an image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input, and obtaining the trained image restoration model when a preset training termination condition is met. The image restoration model employs a diffusion model, in which, during forward diffusion, mixed noise is gradually added to the original image to obtain a noisy image; and during reverse diffusion, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0007] According to a third aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the image restoration model processing method of the first aspect and its possible implementations.

[0008] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the image restoration model processing method of the first aspect and possible implementations thereof by executing the executable instructions.

[0009] The technical solution disclosed herein has the following beneficial effects:

[0010] In the above image restoration model processing, the original image and its corresponding compressed image are acquired; the original image is used as the training target, and the corresponding compressed image is used as the training input to train the image restoration model. The trained image restoration model is obtained when the preset training termination condition is met. The image restoration model employs a diffusion model. During forward diffusion, mixed noise is gradually added to the original image to obtain a noisy image; during backward diffusion, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise. This disclosure, by integrating mixed noise with a diffusion modeling mechanism, guides the image restoration model to perform structure-aware image restoration training, overcoming the problem of lacking real compression artifact supervision in standard diffusion training, thereby achieving artifact removal, detail enhancement, and structure restoration of compressed images. Attached Figure Description

[0011] Figure 1 This diagram illustrates a flowchart of an image restoration model processing method according to this exemplary embodiment;

[0012] Figure 2 This illustration shows a flowchart of training an image restoration model in this exemplary embodiment;

[0013] Figure 3 This exemplary embodiment illustrates a flowchart for calculating the joint loss of a multi-branch image restoration model.

[0014] Figure 4 This illustration shows a stage diagram of image compression enhancement based on a diffusion model in this exemplary embodiment;

[0015] Figure 5 This diagram illustrates a structural block diagram of an image restoration model processing apparatus according to an exemplary embodiment of the present invention.

[0016] Figure 6 An electronic device for implementing the above-described image restoration model processing method is shown in this exemplary embodiment. Detailed Implementation

[0017] Exemplary embodiments of this disclosure will be described more fully below with reference to the accompanying drawings.

[0018] The accompanying drawings are schematic illustrations of this disclosure and are not necessarily drawn to scale. Some block diagrams shown in the drawings may be functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in hardware modules or integrated circuits, or in networks, processors, or microcontrollers. Implementations can be carried out in various forms and should not be construed as limited to the examples set forth herein. The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough description of embodiments of this disclosure. However, those skilled in the art will recognize that one or more specific details may be omitted when implementing the technical solutions of this disclosure, or other methods, components, apparatuses, steps, etc., may be used to replace one or more specific details.

[0019] In related technologies, existing image enhancement algorithms have attempted to remove artifacts, but most are based on CNN or Transformer architectures, and their restoration quality is limited by the model's expressive power. Although traditional diffusion models have the ability to generate images and restore details, they often simply assume noise to be Gaussian white noise, making it difficult to achieve accurate restoration of compressed images.

[0020] In view of the above problems, the exemplary embodiments of this disclosure provide an image restoration model processing method, an image restoration model processing apparatus, a computer program product, and an electronic device to achieve accurate restoration of compressed images.

[0021] In one alternative implementation, refer to Figure 1The image restoration model processing method shown includes the following steps S110 to S120:

[0022] Step S110: Obtain the original image and the compressed image corresponding to the original image;

[0023] Step S120: Using the original image as the training target and the compressed image corresponding to the original image as the training input, train the image restoration model, and obtain the trained image restoration model when the preset training termination condition is met.

[0024] The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0025] Figure 1 The method shown integrates hybrid noise with diffusion modeling to guide the image restoration model in structure-aware image restoration training. This overcomes the problem of lacking real compression artifact supervision in the standard diffusion training process, thereby achieving artifact removal, detail enhancement, and structure restoration of compressed images.

[0026] The following is about Figure 1 Each step in the process will be explained in detail.

[0027] In step S110, the original image and the corresponding compressed image are obtained.

[0028] In this context, a compressed image refers to an image obtained by lossy compression encoding of the original image, which may exhibit distortions such as blockiness, artifacts, and blurred details. It should be noted that, in this disclosure, artifacts refer to unnatural visual interference generated during the compression process, such as blocky boundaries, color abrupt changes, and blurring. Generally, the image quality of the original image is higher than that of the corresponding compressed image.

[0029] Optionally, the original image can be an uncompressed image, and the compressed image corresponding to the original image can be an image after the original image has been compressed.

[0030] Optionally, the original image can be a compressed image, and the compressed image corresponding to the original image can be an image after secondary compression of the original image.

[0031] For example, a preset image set can be used to generate image pairs under different QP (Quantizer Parameter) or compression rates for use in training the image restoration model. For example, the preset image set can be datasets such as COCO, DIV2K, Kodak, etc., and this disclosure does not specifically limit it.

[0032] In step S120, the image restoration model is trained using the original image as the training target and the compressed image corresponding to the original image as the training input. The trained image restoration model is obtained when the preset training termination condition is met.

[0033] The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0034] It's important to note that the diffusion model is a generative model based on a progressive noise perturbation and reverse restoration process. It reconstructs images or generates new images by learning the denoising process within the data distribution. The forward diffusion process is the gradual blurring of a sharp image within the diffusion model; its core is the addition of noise that causes image degradation. The reverse diffusion process is the gradual restoration of a sharp image within the diffusion model; its core is noise removal and image reconstruction.

[0035] Optionally, the image restoration model can be a score-based diffusion model or DDPM (Denoising Diffusion Probabilistic Models).

[0036] Optionally, to improve processing efficiency, the image restoration model can also use DDIM (Denoising Diffusion Implicit Models) or DPM-Solver (Diffusion Probabilistic Model Solver) to accelerate sampling.

[0037] It should be noted that in practical applications, developers can choose the specific type of diffusion model to build as needed, and this disclosure does not impose any specific restrictions on this.

[0038] Specifically, the compressed image corresponding to the original image can be used as the training input of the image restoration model, and the original image can be used as the training target of the image restoration model for model training.

[0039] For example, the process of compression degradation by progressively adding mixed noise to the original image can be modeled as: x c =H(x)+∈, where x represents the original image; x c Let H(x) represent the compressed image; H(x) can be used to characterize the structural losses generated during the compression degradation process, such as low-pass filtering, DCT (Discrete Cosine Transform) truncation, quantization, and other loss types, which are not specifically limited here; ∈ represents the artifacts and perturbations introduced by compression, which can be regarded as non-Gaussian, non-independent noise, and can be defined as: ∈ ~ N(0, W T Σ f W); where W is the image compression transformation matrix; W T Σ represents the transpose of matrix W; f This represents the residual between the original image and the compressed image in the transform domain; N represents a Gaussian distribution.

[0040] It should be noted that the above compression degradation process can be regarded as a special forward diffusion, which makes the compressed image far away from the data manifold in the signal space, so that it is suitable for anti-diffusion recovery through a diffusion model.

[0041] During the training phase, the image restoration model does not directly generate the compressed and degraded image using the compression algorithm. Instead, it simulates the "generation trajectory" of compression noise, replacing the actual compression operation with the training process. Optionally, the forward process of the image restoration model can be defined as: Where q represents the conditional probability distribution of the forward diffusion process; x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor (cumulative product) for the first t steps; t () denotes a degenerate operator; Σ t Used to characterize the residual between the original image and the noisy image at step t in the transform domain; t represents the diffusion step number; x represents t From a mean covariance is It was obtained by sampling from a Gaussian distribution.

[0042] In an optional implementation where the image compression process can be viewed as a gradual noise addition process, the above-mentioned gradual addition of mixed noise to the original image to obtain a noisy image can be achieved through the following steps: by calculating... A mixed-type noise is added to obtain a noisy image; where x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor for the first t steps; t() denotes the degradation operator, and W denotes the image compression transformation matrix; W T Σ represents the transpose of matrix W; t The residual between the original image and the noisy image at step t is used to characterize the difference in the transform domain; N represents a Gaussian distribution; t represents the number of diffusion steps; η t To conform to N(0,Σ) t ) Distribution of variables.

[0043] Optional, degeneracy operator H t () can gradually increase the high-frequency cutoff ratio or downsampling rate as the number of diffusion steps increases, so as to simulate a gradual increase in compression ratio.

[0044] Since traditional diffusion models typically treat noise as Gaussian white noise, and artifacts and blur distortions in compressed images are highly structured and non-independent, they cannot be simply modeled as Gaussian distributions. This disclosure introduces structured compression noise, which enables the diffusion model to be better applied to compressed image restoration scenarios.

[0045] In one optional implementation, the image restoration model is trained using the original image as the training target and the compressed image corresponding to the original image as the training input. The trained image restoration model is obtained when a preset training termination condition is met. Figure 2 To achieve this, follow the steps shown:

[0046] Step S210: Using the original image as the training target and the compressed image corresponding to the original image as the training input, calculate the joint loss of the multi-branch image restoration model.

[0047] Step S220: Adjust the network parameters of the image restoration model by using the joint loss of the multi-branch model, and obtain the image restoration model when the preset training termination condition is met.

[0048] Figure 2 By using the original image as the training target and the corresponding compressed image as the training input, the joint loss of the multi-branch image restoration model is calculated, which enables the image restoration model using the diffusion model to be better adapted to the compressed image restoration task.

[0049] The preset training termination conditions may include, but are not limited to, the number of training iterations meeting a preset threshold, the model accuracy meeting a preset performance threshold, etc., and this disclosure does not impose specific limitations on them.

[0050] In an optional implementation, in step S210, the original image is used as the training target, and the compressed image corresponding to the original image is used as the training input to calculate the joint loss of the multi-branch image restoration model. This can be achieved through methods such as... Figure 3 To achieve this, follow the steps shown:

[0051] Step S310: Using the original image as the training target and the compressed image corresponding to the original image as the training input, calculate the noise prediction loss, image reconstruction loss, and perception loss respectively.

[0052] Step S320: Based on the noise prediction loss, image reconstruction loss, and perception loss, obtain the joint loss of the multi-branch image restoration model.

[0053] Figure 3 In the steps shown, by fusing the loss from multiple branches of noise prediction loss, image reconstruction loss, and perception loss, the model accuracy can be improved, enabling the image restoration model using the diffusion model to better adapt to compressed image restoration tasks.

[0054] Specifically, in step S310, the original image is used as the training target, and the compressed image corresponding to the original image is used as the training input. The noise prediction loss, image reconstruction loss, and perception loss are calculated respectively.

[0055] Among them, noise prediction loss is used to measure the loss of the deviation between the predicted value of the noise added during the diffusion process and the actual noise.

[0056] Among them, image reconstruction loss is used to measure the loss of the difference between the denoised image generated or reconstructed by the model and the original image.

[0057] Among them, perceptual loss is used to measure the loss of sensory quality, and it measures the subjective difference of images based on the difference of features in the intermediate layers of deep networks.

[0058] In one optional implementation, the above-mentioned calculation of noise prediction loss, image reconstruction loss, and perceptual loss using the original image as the training target and the compressed image corresponding to the original image as the training input includes: calculating noise prediction loss based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps; calculating image reconstruction loss based on the denoised image after stepwise denoising of the original image and the compressed image; and calculating perceptual loss based on the denoised image after stepwise denoising of the original image and the compressed image.

[0059] Specifically, the noise prediction loss can be calculated based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps.

[0060] For example, it can be calculated The noise prediction loss is obtained.

[0061] Where L1 represents the noise prediction loss; ∈ represents the artifacts and perturbations introduced by compression, which can be regarded as non-Gaussian, non-independent noise, and can be defined as: ∈ ~ N(0, W T Σ f W); ∈θ x represents the noise predicted by the image restoration model; x represents the original image; x t x represents the noisy image at step t; c t represents the compressed image; t represents the number of diffusion steps.

[0062] The noise prediction loss can be used to fit the noise predicted in the diffusion process, thereby improving the model's prediction accuracy.

[0063] Specifically, the image reconstruction loss can be calculated based on the original image, the denoised image after stepwise denoising of the compressed image, and the denoised image.

[0064] For example, it can be calculated or The image reconstruction loss is obtained.

[0065] Where L2 represents the image reconstruction loss; x represents the original image; This represents the denoised image after progressive denoising of the compressed image.

[0066] Image reconstruction loss can be used to force the model to maintain consistency with the original image at the pixel level in the reconstruction of the compressed image, achieving pixel-level alignment and reducing artifacts.

[0067] Specifically, the perceptual loss can be calculated based on the original image, the compressed image, and the denoised image obtained by progressively denoising.

[0068] For example, it can be calculated

[0069] Where L3 represents the perceptual loss; φ is the intermediate feature extractor in the pre-trained VGG (Visual Geometry Group) model.

[0070] Perceptual loss can be used to measure the differences in images in the semantic perception space, and can improve detail preservation and subjective visual quality.

[0071] It should be noted that in practical applications, the various branch losses can be executed synchronously or asynchronously. This disclosure does not specify the execution steps for each loss.

[0072] Specifically, in step S320, the joint loss of the multiple branches of the image restoration model is obtained based on the noise prediction loss, image reconstruction loss, and perception loss.

[0073] For example, the noise prediction loss, image reconstruction loss, and perception loss can be weighted and summed to obtain the joint loss of the multiple branches of the image restoration model.

[0074] In one optional implementation, the joint loss of the multi-branch image restoration model obtained from the noise prediction loss, image reconstruction loss, and perception loss can be achieved through the following steps: obtaining the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss; and fusing the noise prediction loss, image reconstruction loss, and perception loss based on the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss to obtain the joint loss of the multi-branch image restoration model.

[0075] The weight parameters corresponding to noise prediction loss, image reconstruction loss, and perception loss can be used to control the loss contribution ratios of noise prediction loss, image reconstruction loss, and perception loss, respectively.

[0076] For example, L can be calculated total =λ1L1+λ2L2+λ3L3, which yields the joint loss of the multi-branch image restoration model.

[0077] Among them, L total L1 represents the joint loss; L2 represents the noise prediction loss; L3 represents the image reconstruction loss; and λ1, λ2, and λ3 represent the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss, respectively.

[0078] For example, as shown in Table 1 below, the weights of each loss can be adjusted according to different training stages to improve the accuracy of the joint loss.

[0079] Table 1

[0080] Training phase Weight settings initial stage <![CDATA[λ1=1.0,λ2=1.0,λ3=0.0]]> intermediate stage <![CDATA[λ1=1.0,λ2=1.0,λ3=0.3]]> Convergence phase <![CDATA[λ1=0.5,λ2=1.0,λ3=0.5]]>

[0081] For example, it can also be calculated The joint loss of the multi-branch image restoration model is obtained.

[0082] Among them, L total Indicates joint loss; L i Used to represent the loss of the i-th type, such as L1, L2, and L3 calculated above; σ i The weight parameter is used to represent the loss of the i-th type.

[0083] Optionally, the weight parameters corresponding to each branch loss can be learnable, allowing the model to dynamically adjust the contribution ratio of each loss term during training, thereby enhancing optimization robustness and generalization ability. For example, the weight parameters corresponding to each branch loss can be initialized at the beginning and then continuously learned and iterated. It should be noted that the branch losses may include, but are not limited to, the noise prediction loss, image reconstruction loss, and perceptual loss calculated above; no specific limitations are imposed here.

[0084] Once the image restoration model has been trained through the training phase described above, it can be used in the inference phase.

[0085] In an alternative implementation, the following steps may also be performed: inputting the target compressed image to be restored into the trained image restoration model, performing stepwise denoising processing, and obtaining the restored image of the target compressed image.

[0086] Here, the target compressed image to be restored refers to the image that has been compressed and needs to be restored. The restored image of the target compressed image refers to the image obtained after the target compressed image has been restored using an image restoration model. For example, the target compressed image can be an image in compression formats such as JPEG, HEVC, or VVC, etc., without specific limitations here.

[0087] For example, the target compressed image to be restored can be mapped to a noisy version of the image restoration model at the corresponding time step during the backdiffusion process, so that the image can be reconstructed step by step using the backdiffusion network to obtain the final restored image.

[0088] By inputting real compressed images into a trained image restoration model, image artifacts can be reduced, edge sharpness and texture restoration can be improved, and compressed images can be restored.

[0089] Optionally, the trained image restoration model can be embedded in a system with encoding and decoding as the structure of a compressed image enhancement module.

[0090] like Figure 4 As shown, a schematic diagram of the stages of image compression enhancement based on a diffusion model is provided, including a modeling stage 401, a diffusion stage 402, an optimization stage 403, and an inference stage 404.

[0091] In the modeling stage 401, the compression degradation process is regarded as a special forward diffusion, which makes the compressed image far away from the data manifold in the signal space so that the diffusion model can perform anti-diffusion recovery.

[0092] The diffusion stage 402 includes a forward diffusion process and a reverse diffusion process. In the forward diffusion process, the original image is simulated as a compressed and degraded image, that is, mixed noise is added; in the reverse diffusion process, noise is gradually removed to restore image details.

[0093] In the optimization phase 403, a multi-branch joint loss is used to optimize the model.

[0094] In the inference stage 404, the compressed image that needs to be restored is enhanced and output.

[0095] The image restoration model trained using this disclosure can remove compression artifacts, enhance texture details, and improve the subjective visual experience, thus achieving quality improvement. It does not depend on compression algorithm formats and has strong compatibility. It can promote the performance and accuracy of subsequent processing stages (such as object detection and segmentation). It can be optimized in conjunction with perception networks and is suitable for industrial intelligent vision scenarios.

[0096] Exemplary embodiments of this disclosure also provide an image restoration model processing apparatus, with reference to Figure 5 As shown, the image restoration model processing device 500 may include the following program modules:

[0097] Image acquisition module 510 is used to acquire the original image and the compressed image corresponding to the original image;

[0098] The model training module 520 is used to train the image restoration model with the original image as the training target and the compressed image corresponding to the original image as the training input, and to obtain the trained image restoration model when the preset training termination condition is met.

[0099] The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0100] In an optional implementation, based on the foregoing scheme, the image restoration model processing apparatus 500 further includes: a noise addition module, used for calculating... A mixed-type noise is added to obtain a noisy image; where x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor for the first t steps; t () denotes the degradation operator; W denotes the image compression transformation matrix; W T Σ represents the transpose of matrix W; tThe residual between the original image and the noisy image at step t is used to characterize the difference in the transform domain; N represents a Gaussian distribution; t represents the number of diffusion steps; η t To conform to N(0,Σ) t ) Distribution of variables.

[0101] In an optional implementation, based on the aforementioned scheme, the model training module 520 includes: a joint loss calculation module, used to calculate the joint loss of multiple branches of the image restoration model with the original image as the training target and the compressed image corresponding to the original image as the training input; and a network parameter optimization module, used to adjust the network parameters of the image restoration model through the joint loss of multiple branches of the image restoration model and obtain the image restoration model when the preset training termination condition is met.

[0102] In an optional implementation, based on the aforementioned scheme, the joint loss calculation module may include: a branch loss calculation module, used to calculate noise prediction loss, image reconstruction loss, and perception loss respectively, using the original image as the training target and the compressed image corresponding to the original image as the training input; and a branch loss fusion module, used to obtain the joint loss of the multiple branches of the image restoration model based on the noise prediction loss, image reconstruction loss, and perception loss.

[0103] In one optional implementation, based on the aforementioned scheme, the branch loss calculation module can be configured to: calculate noise prediction loss based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps; calculate image reconstruction loss based on the original image and the denoised image after stepwise denoising of the compressed image; and calculate perceptual loss based on the original image and the denoised image after stepwise denoising of the compressed image.

[0104] In an optional implementation, based on the aforementioned scheme, the branch loss fusion module can be configured to: obtain the weight parameters corresponding to the noise prediction loss, the weight parameters corresponding to the image reconstruction loss, and the weight parameters corresponding to the perception loss; and based on the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss, fuse the noise prediction loss, the image reconstruction loss, and the perception loss to obtain the joint loss of the multi-branch image restoration model.

[0105] In an optional implementation, based on the aforementioned scheme, the image restoration model processing device 500 further includes: a model inference module, used to input the target compressed image to be restored into the trained image restoration model, perform stepwise denoising processing, and obtain the restored image of the target compressed image.

[0106] The specific details of each part of the image restoration model processing apparatus 500 described above have been described in detail in the method section of the implementation. For any undisclosed details, please refer to the implementation of the method section, and therefore will not be repeated here.

[0107] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0108] Exemplary embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the image restoration model processing method described above.

[0109] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.

[0110] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.

[0111] Computer program code can be written in one or more programming languages. Examples of programming languages ​​include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).

[0112] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic radiation, and infrared rays. Electronic devices can convert signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code causes the electronic device to execute (more specifically, its processor) the method steps of various exemplary embodiments of this disclosure, such as the methods described above, which include the following steps:

[0113] Obtain the original image and its corresponding compressed image;

[0114] The image restoration model is trained using the original image as the training target and the compressed image corresponding to the original image as the training input. The trained image restoration model is obtained when the preset training termination condition is met.

[0115] The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0116] In an optional implementation, based on the aforementioned scheme, the stepwise addition of mixed noise to the original image to obtain a noisy image can be achieved through the following steps: by calculating... A mixed-type noise is added to obtain a noisy image; where x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor for the first t steps; t () denotes the degradation operator; W denotes the image compression transformation matrix; W T Σ represents the transpose of matrix W; t The residual between the original image and the noisy image at step t is used to characterize the difference in the transform domain; N represents a Gaussian distribution; t represents the number of diffusion steps; η t To conform to N(0,Σ) t ) Distribution of variables.

[0117] In an optional implementation, based on the aforementioned scheme, the above-mentioned training of the image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input, and obtaining the trained image restoration model when the preset training termination condition is met, can be achieved through the following steps: using the original image as the training target and the compressed image corresponding to the original image as the training input, calculating the joint loss of the multi-branch image restoration model; adjusting the network parameters of the image restoration model using the joint loss of the multi-branch image restoration model, and obtaining the image restoration model when the preset training termination condition is met.

[0118] In an optional implementation, based on the aforementioned scheme, the calculation of the joint loss of the multi-branch image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input can be achieved through the following steps: using the original image as the training target and the compressed image corresponding to the original image as the training input, calculate the noise prediction loss, image reconstruction loss, and perception loss respectively; and obtain the joint loss of the multi-branch image restoration model based on the noise prediction loss, image reconstruction loss, and perception loss.

[0119] In an optional implementation, based on the aforementioned scheme, the above-mentioned calculation of noise prediction loss, image reconstruction loss, and perceptual loss using the original image as the training target and the compressed image corresponding to the original image as the training input can be achieved through the following steps: calculating noise prediction loss based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps; calculating image reconstruction loss based on the original image and the denoised image after stepwise denoising of the compressed image; and calculating perceptual loss based on the original image and the denoised image after stepwise denoising of the compressed image.

[0120] In an optional implementation, based on the aforementioned scheme, the joint loss of the multi-branch image restoration model obtained from the noise prediction loss, image reconstruction loss, and perception loss can be achieved through the following steps: obtaining the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss; and fusing the noise prediction loss, image reconstruction loss, and perception loss based on the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss to obtain the joint loss of the multi-branch image restoration model.

[0121] In an alternative implementation, based on the aforementioned scheme, the following steps may also be performed: inputting the target compressed image to be restored into the trained image restoration model, performing stepwise denoising processing, and obtaining the restored image of the target compressed image.

[0122] In the above image restoration model processing, the hybrid noise and diffusion modeling mechanism are combined to guide the image restoration model to perform structure-aware image restoration training. This can overcome the problem of lack of real compression artifact supervision in the standard diffusion training process, and thus achieve artifact removal, detail enhancement and structure restoration of compressed images.

[0123] An exemplary embodiment of this disclosure also provides an electronic device capable of implementing the image restoration model processing method described above. The electronic device may include a processor and a memory. The memory stores executable instructions for the processor, such as program code. The processor executes the executable instructions to perform the method of this exemplary embodiment.

[0124] The following is for reference. Figure 6 The electronic device is illustrated by way of a general-purpose computing device. It should be understood that... Figure 6 The electronic device 600 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0125] like Figure 6 As shown, the electronic device 600 may include: a processor 610, a memory 620, a bus 630, an I / O (input / output) interface 640, and a network adapter 650.

[0126] Memory 620 may include volatile memory, such as RAM 621 and cache unit 622, and may also include non-volatile memory, such as ROM 623. Memory 620 may also include one or more program modules 624, such program modules 624 including, but not limited to: operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. For example, program module 624 may include the modules in the above-described device.

[0127] The processor 610 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit).

[0128] The processor 610 can be used to execute executable instructions stored in the memory 620, such as performing any one or more method steps in this exemplary embodiment.

[0129] For example, processor 610 may perform the following steps:

[0130] Obtain the original image and its corresponding compressed image;

[0131] The image restoration model is trained using the original image as the training target and the compressed image corresponding to the original image as the training input. The trained image restoration model is obtained when the preset training termination condition is met.

[0132] The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

[0133] In an optional implementation, based on the aforementioned scheme, the stepwise addition of mixed noise to the original image to obtain a noisy image can be achieved through the following steps: by calculating... A mixed-type noise is added to obtain a noisy image; where x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor for the first t steps; t () denotes the degradation operator; W denotes the image compression transformation matrix; W T Σ represents the transpose of matrix W; t The residual between the original image and the noisy image at step t is used to characterize the difference in the transform domain; N represents a Gaussian distribution; t represents the number of diffusion steps; η t To conform to N(0,Σ) t ) Distribution of variables.

[0134] In an optional implementation, based on the aforementioned scheme, the above-mentioned training of the image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input, and obtaining the trained image restoration model when the preset training termination condition is met, can be achieved through the following steps: using the original image as the training target and the compressed image corresponding to the original image as the training input, calculating the joint loss of the multi-branch image restoration model; adjusting the network parameters of the image restoration model using the joint loss of the multi-branch image restoration model, and obtaining the image restoration model when the preset training termination condition is met.

[0135] In an optional implementation, based on the aforementioned scheme, the calculation of the joint loss of the multi-branch image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input can be achieved through the following steps: using the original image as the training target and the compressed image corresponding to the original image as the training input, calculate the noise prediction loss, image reconstruction loss, and perception loss respectively; and obtain the joint loss of the multi-branch image restoration model based on the noise prediction loss, image reconstruction loss, and perception loss.

[0136] In an optional implementation, based on the aforementioned scheme, the above-mentioned calculation of noise prediction loss, image reconstruction loss, and perceptual loss using the original image as the training target and the compressed image corresponding to the original image as the training input can be achieved through the following steps: calculating noise prediction loss based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps; calculating image reconstruction loss based on the original image and the denoised image after stepwise denoising of the compressed image; and calculating perceptual loss based on the original image and the denoised image after stepwise denoising of the compressed image.

[0137] In an optional implementation, based on the aforementioned scheme, the joint loss of the multi-branch image restoration model obtained from the noise prediction loss, image reconstruction loss, and perception loss can be achieved through the following steps: obtaining the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss; and fusing the noise prediction loss, image reconstruction loss, and perception loss based on the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss to obtain the joint loss of the multi-branch image restoration model.

[0138] In an alternative implementation, based on the aforementioned scheme, the following steps may also be performed: inputting the target compressed image to be restored into the trained image restoration model, performing stepwise denoising processing, and obtaining the restored image of the target compressed image.

[0139] In the above image restoration model processing, the hybrid noise and diffusion modeling mechanism are combined to guide the image restoration model to perform structure-aware image restoration training. This can overcome the problem of lack of real compression artifact supervision in the standard diffusion training process, and thus achieve artifact removal, detail enhancement and structure restoration of compressed images.

[0140] Electronic device 600 can communicate with one or more external devices 700 (such as keyboard, mouse, external controller, etc.) through I / O interface 640.

[0141] Electronic device 600 can communicate with one or more networks via network adapter 650. For example, network adapter 650 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication. Network adapter 650 can communicate with other modules of electronic device 600 via bus 630.

[0142] although Figure 6As not shown in the diagram, other hardware and / or software modules may also be configured in the electronic device 600, including but not limited to: a display, microcode, device driver, redundant processor, external disk drive array, RAID (Redundant Arrays of Independent Disks) system, tape drive, and data backup storage system.

[0143] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0144] Those skilled in the art will understand that various aspects of this disclosure can be implemented as systems, methods, or program products. Therefore, various aspects of this disclosure can be embodied in entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuit,” “module,” or “system.” Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0145] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is defined only by the appended claims.

Claims

1. An image restoration model processing method, characterized in that, The method includes: Obtain the original image and the corresponding compressed image; Using the original image as the training target and the compressed image corresponding to the original image as the training input, the image restoration model is trained, and the trained image restoration model is obtained when the preset training termination condition is met. The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

2. The method according to claim 1, characterized in that, The step of progressively adding mixed noise to the original image to obtain a noisy image includes: Through calculation Add mixed noise to obtain a noisy image; Where x represents the original image; x t This represents the noisy image at step t; H represents the noise preservation factor for the first t steps; t () denotes the degradation operator; W denotes the image compression transformation matrix; W T Σ represents the transpose of matrix W; t The residual between the original image and the noisy image at step t is used to characterize the difference in the transform domain; N represents a Gaussian distribution; t represents the number of diffusion steps; η t To conform to N(0,Σ) t ) Distribution of variables.

3. The method according to claim 1, characterized in that, The step of training the image restoration model using the original image as the training target and the compressed image corresponding to the original image as the training input, and obtaining the trained image restoration model under the condition of satisfying a preset training termination condition, includes: Using the original image as the training target and the compressed image corresponding to the original image as the training input, calculate the joint loss of the multi-branch image restoration model; By adjusting the network parameters of the image restoration model through the joint loss of the multiple branches of the image restoration model, and obtaining the image restoration model when the preset training termination condition is met.

4. The method according to claim 3, characterized in that, The step of using the original image as the training target and the compressed image corresponding to the original image as the training input to calculate the joint loss of the multi-branch image restoration model includes: Using the original image as the training target and the compressed image corresponding to the original image as the training input, the noise prediction loss, image reconstruction loss, and perception loss are calculated respectively. The joint loss of the multi-branch image restoration model is obtained based on the noise prediction loss, the image reconstruction loss, and the perception loss.

5. The method according to claim 4, characterized in that, The step of using the original image as the training target and the compressed image corresponding to the original image as the training input to calculate noise prediction loss, image reconstruction loss, and perceptual loss respectively includes: Calculate the noise prediction loss based on the original image, the compressed image corresponding to the original image, the number of diffusion steps, and the noisy image corresponding to the number of diffusion steps; Calculate the image reconstruction loss based on the original image and the denoised image obtained by progressively denoising the compressed image; The perceptual loss is calculated based on the original image and the denoised image obtained by progressively denoising the compressed image.

6. The method according to claim 4, characterized in that, The step of obtaining the joint loss of the multi-branch image restoration model based on the noise prediction loss, the image reconstruction loss, and the perception loss includes: Obtain the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss; Based on the weight parameters corresponding to the noise prediction loss, the image reconstruction loss, and the perception loss, the noise prediction loss, the image reconstruction loss, and the perception loss are fused to obtain the joint loss of the multi-branch image restoration model.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: The target compressed image to be restored is input into the trained image restoration model, and stepwise denoising is performed to obtain the restored image of the target compressed image.

8. An image restoration model processing device, characterized in that, The device includes: The image acquisition module is used to acquire the original image and the compressed image corresponding to the original image; The model training module is used to train the image restoration model with the original image as the training target and the compressed image corresponding to the original image as the training input, and to obtain the trained image restoration model when the preset training termination condition is met. The image restoration model employs a diffusion model. During the forward diffusion process, mixed noise is gradually added to the original image to obtain a noisy image. During the reverse diffusion process, the compressed image is gradually denoised to obtain a denoised image. The mixed noise includes Gaussian noise and structured compression noise.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1 to 7 by executing the executable instructions.