No-reference image quality evaluation method based on generated noise estimation
By generating noisy images and converting them into image quality scores, the problem of insufficient accuracy of reference-free image quality evaluation in complex scenarios is solved, and more efficient image quality evaluation is achieved.
Patent Information
- Application Number
- CN202510077898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing reference-free image quality evaluation method is insufficient in complex scenarios to evaluate the accuracy of the evaluation, and feature extraction is difficult to effectively reflect the visual perception of the human eye.
Using a method based on generating noise estimation, noise images of different levels are generated by training the first neural network model, and these noise images are converted into image quality scores through the second neural network model, and the simulated image quality is deteriorated due to noise.
It improves the accuracy of image quality evaluation, can more comprehensively simulate the quality degradation process of images in complex environments, and enhances the reliability of evaluation.
Smart Images

Figure CN119992299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly to a reference-free image quality assessment method based on generated noise estimation. Background Art
[0002] Image quality assessment is a fundamental problem in the field of image processing and computer vision, and is widely used in image / video coding, super-resolution reconstruction, and image / video visual quality enhancement. This task aims to objectively evaluate the perceived quality of images by building a model that simulates the perception mechanism of the human visual system and measuring the overall quality of the image in a digital way.
[0003] No-reference image quality assessment (also called blind image quality assessment) is an important objective quality assessment method because it does not require the original reference image. This method has a wide range of applications, especially in situations where the original image cannot be obtained, such as satellite images, surveillance videos, etc. No-reference image quality assessment can effectively evaluate the quality of images and help improve image processing algorithms and systems.
[0004] A no-reference image quality assessment method based on deep learning has been proposed. This method constructs a deep neural network to learn the visual features of the image to build an image quality assessment model, or directly expresses the function of the distorted image to the image quality through end-to-end learning. However, due to the lack of the original reference image, the accuracy of the assessment may be affected. In addition, the changes in the content of different images also make the assessment difficult.
[0005] Therefore, how to improve the evaluation accuracy in complex scenes and explore new feature extraction methods to better reflect the visual perception of the human eye are issues that technical personnel in this field urgently need to solve. Summary of the invention
[0006] In view of this, the present invention provides a reference-free image quality assessment method based on generated noise estimation, which can accurately simulate the degradation of distorted images and learn image distortion, thereby improving the accuracy of assessment.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A no-reference image quality assessment method based on generative noise estimation comprises the following steps:
[0009] A first neural network model and a second neural network model are trained.
[0010] After the distorted image is input into the trained first neural network model, different levels of noise images are generated.
[0011] The noise images of different levels are input into a trained second neural network model to generate image quality scores.
[0012] Among them, the training steps of the first neural network model include: obtaining multiple distorted image samples; inputting the distorted image samples into the encoder to map the distorted image features into the latent space to form latent variables; sampling in the latent space and generating noise images of different levels through the decoder; inputting the distorted image into the diffusion model for degradation repair and generating a pseudo reference image; superimposing the pseudo reference image and the noise image, and calculating the loss function with the distorted image for parameter optimization.
[0013] Preferably, the first neural network model includes an acquisition module, a variational autoencoder and a diffusion model, the variational autoencoder and the diffusion model are respectively connected to the acquisition module, and the distorted image is acquired through the acquisition module.
[0014] Preferably, the latent vector obeys a normal distribution, and the sampling range is adjusted by setting different standard deviation parameters to generate noise images of different levels.
[0015] Preferably, the pseudo reference image and the noise image are superimposed, specifically: noise images of different levels are linearly weighted according to the noise level and superimposed with the pseudo reference image restored by the diffusion model to obtain a degraded image with noise.
[0016] Preferably, the second neural network model includes an input layer, a feature extraction module, a pooling layer and a fully connected layer connected in sequence; the input layer superimposes noise images of different levels and then inputs them into the feature extraction module.
[0017] Preferably, the feature extraction module is the feature extraction part of Vision Transformer; after feature extraction, multiple blocks are generated, which are synthesized into a global feature vector after average pooling through the pooling layer, and finally the global feature vector is mapped into a single scalar through the fully connected layer to obtain an image quality score.
[0018] Preferably, the training step of the second neural network model includes:
[0019] Calculate the quality scores of noisy images and distorted images respectively, calculate the model loss and perform parameter optimization.
[0020] A reference-free image quality assessment system based on generative noise estimation comprises a model training module, an image quality assessment module, an interface module and a storage module;
[0021] The interface module is used to receive training data or distortion data to be evaluated
[0022] The model training module is used to construct the first neural network model and the second neural network model, and perform model training according to the training data;
[0023] The image quality evaluation module is used to perform corresponding image processing on the distorted data to be evaluated according to the trained neural network model to obtain the image quality score;
[0024] The storage module is used to store the trained model parameters.
[0025] A computer-readable storage medium stores a computer program, which is used for the above-mentioned image quality evaluation method when executed by the computer.
[0026] It can be seen from the above technical solution that, compared with the prior art, the present invention discloses a reference-free image quality evaluation method based on generated noise estimation, which can learn the noise distribution from the distorted image to generate a noise image, which is used to simulate the degradation of image quality due to noise. It can more comprehensively simulate the quality degradation process of the image in various complex environments, thereby improving the final evaluation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0028] Figure 1 A schematic diagram of a no-reference image quality assessment method based on generated noise estimation provided by the present invention.
[0029] Figure 2 A schematic diagram of the structure of a no-reference image quality assessment system based on generated noise estimation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] Example 1
[0032] like Figure 1The embodiment of the present invention discloses a method for evaluating image quality without reference based on generating noise estimation, comprising the following steps:
[0033] S1: training a first neural network model and a second neural network model. The training steps of the first neural network model include:
[0034] S11: Acquire multiple distorted image samples.
[0035] S12: Input the distorted image sample into the encoder to map the distorted image features into the latent space to form a latent variable.
[0036] S13: Sampling in the latent space and generating different levels of noise images through the decoder.
[0037] S14: Input the distorted image into the diffusion model for degradation repair and generate a pseudo reference image.
[0038] S15 superimposes the pseudo reference image and the noise image to obtain a pseudo distorted image, and calculates a loss function based on the pseudo distorted image and the distorted image to perform parameter optimization.
[0039] S2: After inputting the distorted image into the trained first neural network model, different levels of noise images are generated.
[0040] S3: Input the noise images with different levels into the trained second neural network model to generate image quality scores.
[0041] In this embodiment, the diffusion model gradually adds noise to the original distorted image through a forward process until the image completely becomes a noise distribution, and then gradually removes the noise through a reverse process to generate a pseudo reference image that is as close to the original image as possible.
[0042] Through the encoding-decoding architecture, a noise image that meets the characteristics of the distorted image is generated based on the distorted image itself to simulate the interference that the image may face in different complex environments, and superimposed with the pseudo reference image generated by the diffusion model to simulate the image quality degradation. This simulation method can reflect the sensitivity of the image to different noises. Finally, the quality score is evaluated through the second neural network model.
[0043] In order to further implement the above technical solution, the distorted image samples in the training process are images under the influence of different types of distortion {I d}, the distorted images are uniformly resized to a specific pixel size, such as 224 × 224, through image processing software or library. This step ensures the consistency of subsequent processing and the validity of comparison.
[0044] Among them, the distorted image is formed from the clear reference image through degradation processing. This process can be described as:
[0045] y=x+ε
[0046] Among them, y represents the distorted image, x represents the original image, and ε represents the combined distortion factor.
[0047] In order to further implement the above technical solution, the first neural network model includes an acquisition module, a variational autoencoder and a diffusion model. The variational autoencoder and the diffusion model are respectively connected to the acquisition module, and the distorted image is acquired through the acquisition module.
[0048] Specifically, the image processing process of the variational autoencoder (VAE) can be described as:
[0049] I Ni = decoder(φ(encoder(I d )))
[0050] Among them, I Ni represents the generated noise image, decoder is the decoder, φ is the latent space, and encoder is the encoder.
[0051] Furthermore, suppose that the prior distribution of VAE, that is, the latent variable Z, obeys the standard normal distribution Z~N(0,I), where I is the unit matrix, indicating that the variance of each dimension is 1. In the generation phase, Z is randomly sampled from this distribution, and the formula is:
[0052] Z=μ+∈*σ
[0053] Where ∈ is random noise sampled from a standard normal distribution. By adjusting the standard deviation σ of the normal distribution, the sampling of the latent variable Z will deviate from the original noise-free region. For example, when σ = 0.5, the decoder generates a result close to the original image with only slight distortion. When σ = 2, the decoder generates an image with significant noise, and the image details may be significantly distorted. The sampled noise estimate data is input into the decoder, which attempts to map it back to the image space, thereby generating a series of images with different noise levels.
[0054] In order to further implement the above technical solution, the pseudo reference image and the noise image are superimposed, specifically: the noise images of different levels are linearly weighted and superimposed with the pseudo reference image restored by the diffusion model to obtain a degraded image with noise. The expression is as follows:
[0055] I dis' =I r +α*I noise
[0056] Among them, Idis' is the degraded image, I r is a pseudo reference image, I noise is the noise image after linear superposition, and α is the factor that controls the noise intensity.
[0057] In this embodiment, during linear superposition, i.e., linear weighting, the weight coefficient of each noise image is determined by its own noise level; since this embodiment controls the noise level by changing the size of the standard deviation, the corresponding standard deviation in each noise image can be normalized to obtain the weight coefficient of each noise image.
[0058] In order to further implement the above technical solution, the second neural network model includes an input layer, a feature extraction module, a pooling layer and a fully connected layer connected in sequence.
[0059] Specifically, the input layer performs linear weighted superposition of noise images of different levels and inputs them into the feature extraction module. The feature extraction module uses the feature extraction part of the Vision Transformer, that is, removes the remaining part of the classification head; the feature extraction module outputs multiple patches, which are averaged and pooled by the pooling layer to form a global feature vector. Finally, the global feature vector is mapped to a single scalar through the fully connected layer to obtain the image quality score.
[0060] In order to further implement the above technical solution and improve the network's recovery capability, the calculation I dis' The loss between the original distorted image and the reconstructed degraded image I is taken as the mean absolute error (MAE) dis' The loss function between can be described as:
[0061]
[0062] Where n represents the number of samples, represents a distorted image, represents the corresponding pseudo-distorted image.
[0063] Furthermore, the mean absolute error is selected as the loss function L for noisy image quality estimation s :
[0064]
[0065] in, Indicates the MOS value of the distorted image, Represents the prediction score of the noisy image.
[0066] Example 2
[0067] like Figure 2Based on the same inventive concept, an embodiment of the present invention discloses a reference-free image quality assessment system based on generated noise estimation. The system adopts the image quality assessment method in Example 1, including: a model training module, an image quality assessment module, an interface module and a storage module.
[0068] The interface module is used to receive training data or distortion data to be evaluated. The model training module is used to construct the first neural network model and the second neural network model, and perform model training according to the training data. The image quality evaluation module is used to perform corresponding image processing on the distortion data to be evaluated according to the trained neural network model to obtain an image quality score. The storage module is used to store the trained model parameters.
[0069] In this embodiment, the interface module includes an input interface and an output interface. The input interface provides an operation interface for users to upload images that need to be evaluated for quality. It can be a file upload button in a web interface, or a file selection dialog box in a desktop application. The output interface is used to display the predicted quality score results to the user. The scores can be displayed in the form of text boxes, progress bars, etc. on the web interface; the evaluation results can also be presented in a similar intuitive manner in the desktop application.
[0070] Example 3
[0071] Based on the same inventive concept, an embodiment of the present invention discloses a computer-readable storage medium, in which a computer program is stored, and when a computer is executed, it is used for the image quality evaluation method in Embodiment 1.
[0072] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0073] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A no-reference image quality assessment method based on generative noise estimation, characterized in that: The following steps are involved: training a first neural network model and a second neural network model; After inputting the distorted image into the trained first neural network model, different levels of noise images are generated; The noise images of different levels are input into a trained second neural network model to generate image quality scores. The training step of the first neural network model includes: obtaining a plurality of distorted image samples; Inputting the distorted image sample into an encoder to map the distorted image features into a latent space to form a latent variable; Sampling in the latent space to generate different levels of noise images through a decoder; Inputting the distorted image into a diffusion model for degradation repair and generating a pseudo reference image; The pseudo reference image and the noise image are superimposed, and a loss function is calculated with the distorted image to perform parameter optimization.
2. The method for no-reference image quality assessment based on generative noise estimation according to claim 1, characterized in that: The first neural network model includes an acquisition module, a variational autoencoder and a diffusion model. The variational autoencoder and the diffusion model are respectively connected to the acquisition module, and the distorted image is acquired through the acquisition module.
3. The method for no-reference image quality assessment based on generative noise estimation according to claim 2, characterized in that: The latent vector obeys the normal distribution, and the sampling range is adjusted by setting different standard deviation parameters to generate noise images of different levels.
4. The method for no-reference image quality assessment based on generative noise estimation according to claim 1, characterized in that: The pseudo reference image and the noise image are superimposed, specifically: noise images of different levels are linearly weighted according to the noise level and superimposed with the pseudo reference image restored by the diffusion model to obtain a degraded image with noise.
5. The method for no-reference image quality assessment based on generative noise estimation according to claim 1, characterized in that: The second neural network model includes an input layer, a feature extraction module, a pooling layer and a fully connected layer connected in sequence; the input layer superimposes noise images of different levels and then inputs them into the feature extraction module.
6. The method for no-reference image quality assessment based on generative noise estimation according to claim 5, characterized in that: The feature extraction module is the feature extraction part of the Vision Transformer; after feature extraction, multiple blocks are generated, which are synthesized into a global feature vector after average pooling through the pooling layer, and finally the global feature vector is mapped into a single scalar through the fully connected layer to obtain an image quality score.
7. The method for no-reference image quality assessment based on generative noise estimation according to claim 1, characterized in that: The training step of the second neural network model includes: Calculate the quality scores of noisy images and distorted images respectively, calculate the model loss and perform parameter optimization.
8. A no-reference image quality assessment system based on generative noise estimation, characterized in that: The image quality assessment method according to any one of claims 1 to 7 is adopted, comprising a model training module, an image quality assessment module, an interface module and a storage module; The interface module is used to receive training data or distortion data to be evaluated; The model training module is used to construct the first neural network model and the second neural network model, and perform model training according to the training data; The image quality evaluation module is used to perform corresponding image processing on the distorted data to be evaluated according to the trained neural network model to obtain the image quality score; The storage module is used to store the trained model parameters.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer is executed, it is used to implement the image quality evaluation method described in any one of claims 1-7.
Citation Information
Patent Citations
System and methods to measure noise and to generate picture quality prediction from source having no reference
CN102622734A
No-reference mixed distorted image quality evaluation method
CN106780446A
No-reference noise image quality evaluation method and system
CN106991670A
Cited By
Image noise evaluation method and device, electronic equipment and medium
CN121437377A