Blind image quality evaluation method and device based on diffusion repair technology

Through the blind image quality evaluation method based on diffusion repair technology, distorted images are repaired and the characteristics of human visual system are simulated, which solves the problem of poor image quality evaluation performance in the prior art, and achieves efficient and accurate image quality evaluation.

CN119963491AActive Publication Date: 2025-05-09SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510002128.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-09
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing reference-free image quality evaluation methods do not perform well when dealing with complex scenarios and multiple distortion types, and lack in-depth consideration of the characteristics of human visual systems.

Method used

The blind image quality evaluation method based on diffusion repair technology is adopted to repair distorted images through a pre-trained stable diffusion model, generate high-quality pseudo-reference images, and combine it with the attention mechanism that simulates the human visual system to design a quality evaluation network for evaluation.

Benefits of technology

It realizes accurate, efficient and automated blind image quality evaluation, improves the accuracy and robustness of the evaluation, can work stably in a variety of practical application scenarios, and has a wide range of application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963491A_ABST
    Figure CN119963491A_ABST
Patent Text Reader

Abstract

The invention discloses a blind image quality evaluation method and device based on diffusion repair technology, and the method comprises the steps: obtaining a distorted image, and carrying out the blind denoising preprocessing of the distorted image; repairing the preprocessed image by using a pre-trained stable diffusion model to generate a repaired image; performing quality evaluation on the repaired image by adopting a full-reference image quality evaluation model to obtain a preliminary quality score; image block features are extracted from the restored image, a quality evaluation network is designed based on VGGNet, and a channel attention module and a space attention module are embedded; and distributing different weights for each image block, calculating the weights and scores of the image blocks, and obtaining a quality score of the whole image. The method can effectively evaluate the image quality under the condition of no reference image, has wide application prospect and practical significance, and can be widely applied to the field of image processing and evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and evaluation, and in particular to a blind image quality evaluation method and device based on diffusion restoration technology. Background Art

[0002] In the digital age, images and videos have become the main media for information dissemination. Image Quality Assessment (IQA) technology is crucial for ensuring the quality of multimedia content, optimizing image processing algorithms, and reducing data transmission and storage costs. Image quality assessment methods are mainly divided into Full-Reference Image Quality Assessment (FR-IQA) and No-Reference Image Quality Assessment (NR-IQA).

[0003] Full-reference image quality assessment methods rely on the original undistorted reference image to assess the quality of the distorted image. These methods can usually provide high accuracy, but in practical applications, reference images are often unavailable, which limits their scope of application. Therefore, no-reference image quality assessment methods, also known as blind image quality assessment methods, have attracted widespread attention because they do not rely on reference images. These methods attempt to assess the quality of the distorted image directly from the image itself, which is of great significance for scenarios such as quality monitoring of real-time video streams and automatic quality detection of online image sharing platforms.

[0004] Although no-reference image quality assessment methods have broad application prospects in theory, they still face many challenges in practical applications. Traditional no-reference image quality assessment methods are mainly based on natural scene statistics (NSS) and hand-designed image features, which often fail to achieve satisfactory performance when dealing with complex scenes and multiple types of distortion. In addition, these methods usually lack in-depth consideration of the characteristics of the human visual system, such as different sensitivities to different image regions.

[0005] In recent years, with the development of deep learning technology, image quality assessment methods based on deep neural networks have made significant progress. These methods can automatically learn complex feature representations from a large amount of image data, improving the accuracy of the assessment. However, existing deep learning-based no-reference image quality assessment methods still have some limitations, such as strong dependence on training data, limited generalization ability for specific distortion types, and the generative models used in image restoration tasks may produce poor quality restoration images. Summary of the invention

[0006] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the object of the present invention is to provide a blind image quality evaluation method, device and medium based on diffusion restoration technology.

[0007] The first technical solution adopted by the present invention is:

[0008] A blind image quality evaluation method based on diffusion restoration technology comprises the following steps:

[0009] Obtain the distorted image, perform blind denoising preprocessing on the distorted image to reduce the noise in the image and prepare the image for the next step of diffusion restoration;

[0010] The pre-processed image is inpainted using a pre-trained stable diffusion model to generate a inpainted image. This step aims to restore the image quality and simulate the reference image in the full reference image quality evaluation.

[0011] The full-reference image quality assessment model is used to evaluate the quality of the restored image and obtain a preliminary quality score;

[0012] Extract image patch features from the restored image, design a quality assessment network based on VGGNet, and embed channel attention modules and spatial attention modules to simulate the different sensitivities of the human visual system to different channels and image regions;

[0013] Assign different weights to each image block, calculate the weights and scores of the image blocks, and obtain the quality score of the overall image.

[0014] Furthermore, the Swin-Conv-UNet network is used to perform blind denoising preprocessing on the distorted image. The network can effectively remove different types of noise to improve the effect of subsequent diffusion restoration.

[0015] Furthermore, the pre-trained stable diffusion model is a Stable Diffusion model, which can generate high-quality image restoration results and effectively restore the original quality of the image.

[0016] Furthermore, the full-reference image quality assessment model is a WaDIQaM model, which can provide an accurate image quality score and provide a reference for subsequent diffusion repair.

[0017] Furthermore, the quality assessment network includes multiple convolutional layers, pooling layers, fully connected layers, and channel and spatial attention modules, wherein the channel attention module is used to simulate the sensitivity of the human visual system to different channels, and the spatial attention module is used to simulate the sensitivity of the human visual system to different image areas.

[0018] Furthermore, a dual-branch network is used to assign weights to each image block, where one branch network calculates the weight of the image block and the other branch network calculates the quality score of the image block.

[0019] Furthermore, a loss function is designed to optimize network parameters based on the difference between the predicted image quality score and the actual quality score, wherein the loss function is a mean square error loss function.

[0020] Further, the diffusion repair process includes the following steps:

[0021] A1. Use the pre-trained stable diffusion model to iteratively repair the distorted image after blind denoising preprocessing, where each iteration includes denoising and detail enhancement of the image to gradually restore the image quality;

[0022] A2. In each iteration, by comparing the difference between the repaired image and the distorted image, the parameters of the diffusion model are adjusted to optimize the repair effect;

[0023] A3, repeating steps A1 and A2 until a predetermined number of iterations is reached or the restoration effect satisfies a preset condition, and finally generating a restored image;

[0024] A4. Evaluate the quality of the restored image and calculate the preliminary quality score using the full reference image quality evaluation model;

[0025] A5. The difference between the preliminary quality score and the actual quality score is used as the loss function, and the parameters of the diffusion model are updated through the back-propagation algorithm to further improve the repair effect.

[0026] Furthermore, the quality assessment network is designed based on VGGNet, image block features are extracted from the restored image, and a channel attention module and a spatial attention module are embedded, including:

[0027] Segment the restored image into multiple image blocks;

[0028] For each image block, a quality assessment network designed based on VGGNet is used to extract the features of the image block;

[0029] In the quality assessment network, a channel attention module and a spatial attention module are embedded to simulate the different sensitivities of the human visual system to different channels and image regions respectively;

[0030] Through the channel attention module and the spatial attention module, different weights are assigned to each image block to highlight the key areas in the image;

[0031] The extracted image patch features and the assigned weights are taken as input and the quality score of each image patch is calculated through a fully connected layer.

[0032] Furthermore, assigning different weights to each image block, calculating the weights and scores of the image blocks, and obtaining the quality score of the overall image include:

[0033] For each image patch, the local quality score of the image patch is calculated according to its features and the assigned weight;

[0034] Calculate the global quality score of the entire image through the global branch network;

[0035] Through the preset fusion measurement, the local quality score and the global quality score are combined to obtain the quality score of the overall image;

[0036] In order to improve the accuracy of quality scores, a post-processing step is introduced to fine-tune the preliminary quality scores. The fine-tuning strategy can be adaptive adjustment based on image content or other advanced image processing techniques.

[0037] Finally, the quality score of the overall image is output, which reflects the overall quality of the image and can be used for various image quality evaluation tasks.

[0038] Furthermore, the blind image quality assessment method also includes a post-processing step, which fine-tunes the preliminary quality score to improve the accuracy of the final quality score.

[0039] Furthermore, the blind image quality assessment method can automatically process a large number of images and can adapt to different image qualities and distortion types.

[0040] The second technical solution adopted by the present invention is:

[0041] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement a blind image quality assessment method based on diffusion restoration technology as described above.

[0042] The third technical solution adopted by the present invention is:

[0043] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a blind image quality assessment method based on diffusion restoration technology as described above.

[0044] The fourth technical solution adopted by the present invention is:

[0045] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.

[0046] The beneficial effects of the present invention are as follows: the present invention uses a pre-trained stable diffusion model to perform high-quality restoration on distorted images, generates pseudo reference images, and combines the attention mechanism that simulates the human visual system to achieve accurate, efficient and automated blind image quality evaluation. This method not only improves the accuracy and robustness of the evaluation, but also can work stably in a variety of practical application scenarios without being restricted by specific distortion types or image content. In addition, the network structure optimization and loss function design of the present invention further improve the performance of the model, enabling it to demonstrate excellent generalization and robustness on multiple image quality assessment databases, surpassing existing reference-free image quality assessment methods, and has broad application prospects and practical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 is a flowchart of the steps of a blind image quality evaluation method based on diffusion restoration technology in an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of the framework of the DiffBIQA model in an embodiment of the present invention, including an image restoration module and an image quality assessment module;

[0050] Figure 3 is a schematic diagram of the structure of a spatial channel attention module in an embodiment of the present invention;

[0051] Figure 4 It is a schematic diagram of the relationship between the repaired image MOS, the distorted image MOS and the image distortion level and type in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0053] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0054] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0055] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0056] In view of the existing technical problems, the present invention proposes a blind image quality assessment method based on diffusion restoration technology. This method uses a pre-trained diffusion model to restore distorted images, generates high-quality pseudo reference images, and combines deep learning technology to perform image quality assessment. This method can not only generate high-quality restored images, but also simulate the characteristics of the human visual system to improve the accuracy and robustness of the assessment. In addition, the test results of this method on multiple image quality assessment databases show that its excellent performance and generalization ability on different distortion types surpass other existing no-reference image quality assessment methods.

[0057] Example 1

[0058] like Figure 1As shown, this embodiment provides a blind image quality evaluation method based on diffusion restoration technology, which can effectively evaluate image quality without a reference image, and has broad application prospects and practical significance. This method uses an advanced diffusion model to restore the distorted image, generates a pseudo reference image, and simulates the full reference image quality evaluation method for scoring to achieve non-reference image quality evaluation. This method is closer to the observation characteristics of the human visual system, can accurately identify key information in the image, and effectively solves the problem of inaccurate judgment of key areas of distorted images in previous methods. Specifically, it includes the following steps:

[0059] S1. Obtain a distorted image and perform blind denoising preprocessing on the distorted image to reduce the noise in the image and prepare the image for the next step of diffusion restoration.

[0060] As an implementation method, in step S1, the blind image denoising (BID) preprocessing method used is the Swin-Conv-UNet network, which can effectively remove different types of noise to improve the effect of subsequent diffusion repair. The BID preprocessing module is used to remove the distorted image (I distorted ) and generate a conditional image.

[0061] I BID =BID(I distorted )

[0062] The loss function is:

[0063]

[0064] Among them, I reference is the reference image, I distorted is a low-quality image, I BID is the intermediate image generated by the inpainting module.

[0065] S2. Use the pre-trained stable diffusion model to repair the preprocessed image to generate a repaired image. This step aims to restore the quality of the image and simulate the reference image in the full reference image quality evaluation.

[0066] As an implementation method, in step S2, the pre-trained stable diffusion model used is the StableDiffusion model, which can generate high-quality image restoration results and effectively restore the original quality of the image. Stable Diffusion pre-trains an autoencoder to convert the image x into a latent variable z through the encoder E and reconstruct it through the decoder D. The diffusion and denoising processes are performed in the latent space. During the diffusion process, the variance β at time t tA Gaussian noise of ∈(0,1) is added to the encoded latent variable z=E(x) to produce a noisy latent variable:

[0067]

[0068] Among them ∈~N(0,I),α t =1-β t ,and When t is large enough, the variable z in the latent space t It is almost a standard Gaussian distribution.

[0069] Network∈ θ It is learned by predicting the noise ∈ based on c (such as text prompt) at a randomly selected time step t. The optimization of the latent diffusion model is defined as follows:

[0070]

[0071] Where x and c are sampled from the dataset, z = E(x), t is sampled uniformly, and ∈ is sampled from a standard Gaussian distribution.

[0072] This model applies the ControlNet strategy and first uses the pre-trained VAE encoder to encode the restored conditional image:

[0073] c BID =E(I BID )

[0074] For the conditional network, we follow the ControlNet approach and copy a trainable pre-trained UNet encoder and intermediate block and train a U-Net encoder and intermediate block (denoted as Fcond), which receives conditional information and outputs control signals. This copying strategy provides a good weight initialization for the conditional network. Then, we condition c BID and the noise latent variable z at time t t Perform concatenation as the input of Fcond, denoted as:

[0075] z′ t =cat(z t ,c BID )

[0076] Among them, z t is the noise latent variable at time t, and cat represents the concatenation operation.

[0077] The ultimate goal is to minimize the following potential diffusion targets:

[0078]

[0079] Among them, ∈ θRepresents the denoising network, and the input is the concatenated z′ t .

[0080] Finally, after the entire image restoration module, our input I distorted Will produce a high quality version of the output I restorted I restorted As a pseudo reference image, together with I distorted , is input into the subsequent quality assessment network to obtain the predicted image quality score.

[0081] S3. Use the full reference image quality evaluation model to evaluate the quality of the restored image and obtain a preliminary quality score.

[0082] Exemplarily, in step S3, the full reference image quality assessment model used in this embodiment is the WaDIQaM model, which can provide an accurate image quality score and provide a reference for subsequent diffusion restoration. Through this step, a preliminary quality score can be obtained, which reflects the overall quality of the restored image.

[0083] S4. Extract image block features from the restored image, design a quality assessment network based on VGGNet, and embed channel attention module and spatial attention module to simulate the different sensitivities of the human visual system to different channels and image regions.

[0084] As an implementation method, in step S4, we extract image patch features from the repaired image and process them through a quality assessment network designed based on VGGNet. The original VGGnet design input size is 224x224 pixels. In order to adapt to the smaller size (32×32 pixels) input in the present invention, it is adjusted at the architectural level and three initial layers (conv3-32, conv3-32, maxpool) are added. In response to the challenge that traditional CNN cannot identify key areas during feature extraction, an attention module (CBAM) is integrated immediately after the second convolutional layer. The model extracts features through a carefully designed series of convolution and pooling layers, and applies a simple and effective feature fusion strategy to the extracted feature vectors and their difference features, followed by regression operations through two fully connected layers FC-512 and FC-1. To ensure that the output size is consistent with the input, all convolutional layers use 3x3 pixel kernels, are activated by the ReLU function, and are supplemented with zero padding. In order to reduce overfitting, the fully connected layer uses a dropout operation, and the specific ratio is set to 0.5. The network regression module of the embodiment of the present invention introduces an innovative dual-branch design, including a quality estimation branch that calculates the score y for each image patch i. i The weight estimation branch calculates the weight α for each image block i i, running in parallel with the original quality regression branch and sharing the same dimensional information. α is activated by the ReLU operation i , and add a smaller stabilizing term ∈, the modified make sure is always positive. By normalizing the weights:

[0085]

[0086] A global image quality estimate can be computed for:

[0087]

[0088] This design not only considers global quality perception, but also captures subtle changes in local quality perception in the image. In addition, the CBAM module is integrated in both branches to further refine the model's ability to identify and accurately weigh focal areas, significantly improving the accuracy and reliability of the evaluation.

[0089] S5. Assign different weights to each image block, calculate the weights and scores of the image blocks, and obtain the quality score of the overall image.

[0090] As an implementation method, in step S5, we assign different weights to each image block and calculate the weights and scores of the image blocks. This process is implemented by a two-branch network, in which one branch calculates the weights of the image blocks and the other branch calculates the quality scores of the image blocks. By integrating this information, we get the quality score of the overall image. In addition, the method also includes a loss function, which optimizes the network parameters based on the difference between the predicted image quality score and the actual quality score, wherein the loss function is a mean square error loss function.

[0091] The above method is supplemented with reference to the accompanying drawings and specific embodiments.

[0092] This embodiment provides a blind image quality assessment method based on diffusion restoration technology, which is applicable to assessing image quality including but not limited to noise, blur, compression distortion, etc., including:

[0093] Step 1: Image restoration module. The distorted image is preprocessed with blind denoising and the advanced stable diffusion generation model is used to restore the distorted image.

[0094] Step 2: Image evaluation module: The full reference image quality evaluation model is used to evaluate the quality of the restored image and obtain a preliminary quality score. At the same time, the image block features are extracted from the restored image, and the weight and score of each image block are calculated. Finally, this information is integrated to obtain the quality score of the overall image.

[0095] See also Figure 2 ,The specific implementation process mainly includes two aspects: image restoration and image evaluation:

[0096] First, a distorted image is obtained as input, and it is input into the blind denoising network SCUNet for denoising preprocessing to obtain an intermediate image. Then, the intermediate image is input into the pre-trained Stable Diffusion model, and a repaired image is generated after a diffusion and reconstruction process. In this process, the specific operations of the forward diffusion and reverse generation processes of the diffusion model are involved, such as adding and removing noise to latent variables in the latent space, as well as the setting and adjustment of related parameters. The diffusion model consists of a forward diffusion process and a reverse generation process. In the forward diffusion process, the model gradually adds noise to the data until the data is completely or partially converted into noise. In the reverse process, the model learns how to gradually recover the original data from the noisy data. In an embodiment of the present invention, Stable Diffusion pre-trains an autoencoder, converts the image x into a latent variable z through an encoder E and reconstructs it through a decoder D, and the diffusion and denoising processes are performed in the latent space.

[0097] Secondly, the image obtained by the image restoration model is used as a reference to obtain the final score of the distorted image. In order to more accurately estimate the weight of each image block, the embodiment of the present invention integrates an attention mechanism that emphasizes key features while suppressing irrelevant features, and applies a new attention module called "spatial channel attention module". This module consists of two parts: channel attention and spatial attention. It uses the information-rich features extracted by convolution operations, and uses channel attention and spatial attention mechanisms in turn, so that each branch can independently learn the "what" and "where" of important features in the channel and spatial dimensions.

[0098] For the input image features, first process them in the channel dimension. Send the input features to the maximum pooling layer and the average pooling layer at the same time, and then send the pooled features to a shared multi-layer perceptron network to generate the final channel attention feature map. Then process them in the spatial dimension. Perform average pooling and maximum pooling along the channel dimension, then connect the obtained feature maps, and apply convolution operations on the connected feature maps to generate the final spatial attention feature map. Finally, combine the spatial and channel attention in a serial manner, perform weighted processing on the input features, and obtain the feature map after attention enhancement. For the specific implementation of this step, please refer to the attached Figure 3 .

[0099] Similarly, the embodiment of the present invention also introduces a new image quality assessment network, which is improved based on the VGGnet architecture. The original VGGnet design input size is 224x224 pixels. In order to adapt to the smaller size (32×32 pixels) input in the present invention, it is adjusted at the architectural level and three initial layers (conv3-32, conv3-32, maxpool) are added. In response to the challenge that traditional CNN cannot identify key areas during feature extraction, an attention module is integrated immediately after the second convolution layer. The model extracts features through a carefully designed series of convolution and pooling layers, and applies a simple and effective feature fusion strategy to the extracted feature vectors and their difference features, followed by regression operations through two fully connected layers FC-512 and FC-1. To ensure that the output size is consistent with the input, all convolutional layers use 3x3 pixel kernels, are activated by the ReLU function, and supplemented with zero padding. In order to reduce overfitting, the fully connected layer uses a dropout operation, and the specific ratio is set to 0.5. The network regression module of the embodiment of the present invention introduces an innovative dual-branch design, including a quality estimation branch that calculates the score y for each image slice i i The weight estimation branch calculates the weight α for each image block i i , runs in parallel with the original quality regression branch and shares the same dimensionality information.

[0100] First, the feature map processed by the spatial channel attention module is input into the improved VGGnet architecture. Starting from the first convolutional layer, the feature extraction is carried out through the carefully designed convolution and pooling layers in sequence, and the attention module is used to focus on the key areas after the second convolutional layer. Then, the feature fusion operation is performed on the extracted feature vectors and their difference features to obtain the fused features. Then, the fused features are input into the two fully connected layers FC-512 and FC-1 for regression operations to obtain the scores and weights of the image blocks respectively, and finally combined to obtain the image quality score.

[0101] The model proposed in this paper uses the MSE loss function to optimize the quality assessment prediction performance of distorted images. Generally speaking, the blind image quality evaluation loss function is designed to narrow the gap between the predicted score and the true label. However, because our image restoration module performs the same restoration operation on all distorted images and uses the full reference model to assign score labels to the restored images, we hope to derive the relative relationship between the score difference and feature difference between the distorted image and the restored image. Figure 4 It shows that the score difference between the distorted image and the restored image is different according to different distortion levels and different distortion types. This also proves the rationality of designing the loss function based on measuring this difference. The formula is as follows:

[0102]

[0103] diff2=q t -qz

[0104] Loss = L MSE (diff1,diff2)

[0105] This loss function encourages the model to predict a quality score This score is consistent with the inpainted image quality score (q z ) and the true quality score (q t ) and the quality score of the restored image (q z ) between the two.

[0106] Example 2

[0107] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A blind image quality assessment method based on diffusion restoration technology is shown.

[0108] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0109] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.

[0110] Since the electronic device is an electronic device corresponding to a blind image quality assessment method based on diffusion restoration technology in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0111] Example 3

[0112] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A blind image quality assessment method based on diffusion restoration technology is shown.

[0113] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0114] Since the storage medium is a storage medium corresponding to a blind image quality assessment method based on diffusion restoration technology in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0115] Example 4

[0116] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a blind image quality assessment method based on a diffusion repair technology according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment may be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0117] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0118] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0119] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.

Claims

1. A blind image quality evaluation method based on diffusion restoration technology, characterized in that: The following steps are involved: Acquire a distorted image, and perform blind denoising preprocessing on the distorted image; Using the pre-trained stable diffusion model to repair the pre-processed image, generating a repaired image; The full-reference image quality assessment model is used to evaluate the quality of the restored image and obtain a preliminary quality score; Extract image patch features from the restored image, design a quality assessment network based on VGGNet, and embed channel attention module and spatial attention module; Assign different weights to each image block, calculate the weights and scores of the image blocks, and obtain the quality score of the overall image.

2. According to claim 1, a blind image quality assessment method based on diffusion restoration technology is characterized in that: The Swin-Conv-UNet network is used to perform blind denoising preprocessing on the distorted image. The network can effectively remove different types of noise to improve the effect of subsequent diffusion restoration.

3. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The pre-trained stable diffusion model is a Stable Diffusion model, which can generate high-quality image restoration results and effectively restore the original quality of the image.

4. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The full-reference image quality assessment model is the WaDIQaM model, which can provide accurate image quality scores and provide a reference for subsequent diffusion repair.

5. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The quality assessment network includes multiple convolutional layers, pooling layers, fully connected layers, and channel and spatial attention modules, wherein the channel attention module is used to simulate the sensitivity of the human visual system to different channels, and the spatial attention module is used to simulate the sensitivity of the human visual system to different image areas.

6. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: A dual-branch network is used to assign weights to each image patch, where one branch calculates the weight of the image patch and the other branch calculates the quality score of the image patch.

7. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The diffusion repair process includes the following steps: A1. Use the pre-trained stable diffusion model to iteratively repair the distorted image after blind denoising preprocessing, where each iteration includes denoising and detail enhancement of the image to gradually restore the image quality; A2. In each iteration, by comparing the difference between the repaired image and the distorted image, the parameters of the diffusion model are adjusted to optimize the repair effect; A3, repeating steps A1 and A2 until a predetermined number of iterations is reached or the restoration effect satisfies a preset condition, and finally generating a restored image; A4. Evaluate the quality of the restored image and calculate the preliminary quality score using the full reference image quality evaluation model; A5. The difference between the preliminary quality score and the actual quality score is used as the loss function, and the parameters of the diffusion model are updated through the back-propagation algorithm to further improve the repair effect.

8. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The quality assessment network is designed based on VGGNet, image block features are extracted from the restored image, and a channel attention module and a spatial attention module are embedded, including: Segment the restored image into multiple image blocks; For each image block, a quality assessment network designed based on VGGNet is used to extract the features of the image block; In the quality assessment network, a channel attention module and a spatial attention module are embedded to simulate the different sensitivities of the human visual system to different channels and image regions respectively; Through the channel attention module and the spatial attention module, different weights are assigned to each image block to highlight the key areas in the image; The extracted image patch features and the assigned weights are taken as input and the quality score of each image patch is calculated through a fully connected layer.

9. The blind image quality assessment method based on diffusion restoration technology according to claim 1 is characterized in that: The method of assigning different weights to each image block, calculating the weights and scores of the image blocks, and obtaining the quality score of the overall image includes: For each image patch, the local quality score of the image patch is calculated according to its features and the assigned weight; Calculate the global quality score of the entire image through the global branch network; Through the preset fusion measurement, the local quality score and the global quality score are combined to obtain the quality score of the overall image.

10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Blind image quality evaluation method and device based on attention and multiple tasks, equipment and medium

    CN118014962A

  • No-reference multi-modal medical fusion image quality evaluation method based on DwG2NPAN

    CN118261899A

  • Image-to-Image Mapping by Iterative De-Noising

    US20230103638A1

Cited By

  • Medical image quality scoring method and device for generating image screening

    CN121391759A

  • Method and apparatus for generating medical image quality scores for image screening.

    CN121391759B