No-Reference Image Quality Assessment Method Based on Visual Salience and Gradient Features
Through the reference-free image quality evaluation method based on visual saliency and gradient characteristics, an image quality evaluation model is constructed, which solves the problem that image quality evaluation is not simple, fast and accurate enough in the prior art, and achieves a fast and accurate evaluation of image quality, which is consistent with human subjective evaluation.
Patent Information
- Application Number
- CN202210683617.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-16
AI Technical Summary
The prior art is difficult to achieve simple, fast and accurate reference-free image quality evaluation in image quality evaluation, and the existing quality evaluation model cannot accurately simulate subjective evaluation of human visual system.
The reference-free image quality evaluation method based on visual saliency and gradient features is adopted. The gradient features are extracted by the Sobel operator and the significance area detection method of frequency tuning is extracted, and an image quality evaluation model including a gradient feature extractor, a mass regression module, a weight mechanism module and a spatial domain feature extractor are constructed for training.
It realizes rapid and accurate evaluation of image quality, can predict the perceived quality level of distorted images more accurately, and performs well on different data sets, which is in line with human subjective evaluation.
Smart Images

Figure CN115082756B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a no-reference image quality assessment method. Background Art
[0002] In the fields of image processing, video image coding, etc., the quality of images decreases due to various reasons during the processes of image generation, storage, coding, transmission, etc. Evaluating the changes in image quality in these links, especially establishing a method that conforms to human visual perception and can automatically and real-time evaluate image quality, has important research significance and application value.
[0003] Subjective evaluation scores based on the intuitive feelings of observers, and the evaluation results are reliable. However, it requires a large number of professionals, is time-consuming and laborious, and it is difficult to achieve real-time evaluation of a large amount of image data. Objective evaluation automatically evaluates images by establishing a model, which is mainly divided into full-reference image quality assessment, partial-reference image quality assessment, and no-reference image quality assessment. The no-reference image quality assessment method evaluates the quality of images without any original image information, has strong flexibility, and is a research method with broad application prospects and high practicability in the field of image quality assessment.
[0004] In recent years, image quality assessment technology based on deep learning has become a research hotspot and achieved certain results. Deep learning has the ability of self-feature extraction and still has excellent processing ability when facing big data such as photos and videos that penetrate into people's daily lives. However, the current research methods are still restricted by some factors. For example, most methods focus on optimizing the structure and parameters of neural network models, the research on the human visual system is not deep enough, the existing quality assessment models cannot accurately simulate people's subjective evaluation, and the existing objective algorithms are still not ideal. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a no-reference image quality assessment method based on visual saliency and gradient features, which is simple in method, fast in evaluation speed, and accurate in evaluation.
[0006] The technical solution adopted to solve the above technical problem consists of the following steps:
[0007] (1) Select a data set
[0008] Divide the distorted images in the image quality assessment database into a training set, a validation set, and a test set, and the ratio of the number of training set images to the number of validation set images and test set images is 8:1:1.
[0009] (2) Image preprocessing
[0010] The Sobel operator is used to extract the gradient features of the distorted image. The Sobel operator estimates the gradient of the central pixel using the eight neighboring pixels and determines the gradient image G(x) according to the following formula:
[0011]
[0012]
[0013]
[0014] where * is the convolution operation and D(x) is the distorted image.
[0015] The frequency-tuned saliency region detection method is used to extract the saliency map of the distorted image. Gaussian filtering is performed, and the color space is converted from the RGB color space to the LAB color space. The means of the L, A, and B channel images of the converted image are taken respectively. The saliency value S(x, y) of each pixel is determined according to the following formula:
[0016] S(x, y) = ||I μ -I ω (x, y)||
[0017] where I μ is the average feature vector of the image, and I ω (x, y) is the image pixel vector value corresponding to the distorted image after Gaussian filtering; the saliency value S(x, y) of each pixel in the image is divided by the maximum saliency value to obtain the final saliency map.
[0018] Each distorted image, gradient map, and saliency map is cropped into image patches of 32×32 pixels in size as input samples.
[0019] (3) Construct an image quality evaluation model
[0020] The image quality evaluation model includes a gradient feature extractor, a quality regression module, a weight mechanism module, and a spatial domain feature extractor. The gradient feature extractor and the spatial domain feature extractor are in parallel, and after being in series with the quality regression module, they are in parallel with the weight mechanism module 3.
[0021] (4) Train the image quality evaluation model
[0022] The first-order moment estimation method and the second-order moment estimation method of the gradient of the adaptive moment estimation optimizer based on the gradient are used to adaptively adjust the learning rate of each parameter. The learning rate is set to 0.0001, the exponential decay rate of the first-order moment estimation is 0.9, and the exponential decay rate of the second-order moment estimation is 0.999.
[0023] The loss function L1 is determined according to the following formula:
[0024]
[0025] Among them, x represents the number of images, and H x is the subjective score of the x-th image, and S x is the predicted quality score of the x-th image.
[0026] The gradient map in the training set generates a feature vector Γ through the gradient feature extractor of the image quality evaluation model:
[0027] Γ = (γ1, γ2, …, γ k )
[0028] where γ k is an 8×8 matrix, and k ∈ {1, 2, …, 64}; the distorted image blocks generate 3 or 4 feature vectors through the spatial domain feature extractor.
[0029] The feature vector Γ and the 3 or 4 feature vectors generated by the spatial domain feature extractor are converted into a single-column matrix and vertically concatenated in sequence to construct a vertical feature vector Ν.
[0030] Taking the vertical feature vector Ν as the input of the quality regression module to obtain the quality score S i of the i-th image block, where i is the number of distorted image blocks.
[0031] The distorted image is segmented into 32×32 image blocks. According to the change of the significance image gray value, the significance map is divided into prominent regions, general regions, and non-prominent regions, and 3 different thresholds are set. The pixel value in the range of [0, 74] is the non-prominent region, the pixel value in the range of [75, 174] is the normal region, and the pixel value in the range of [175, 255] is the prominent region. Different weights are set for different regions, and the weights of the 3 different regions are {0.8, 1, 1.2} respectively to obtain the weight map W(x).
[0032] Determine the weight w of each image block according to the following formula i :
[0033]
[0034]
[0035] where N d is the number of pixels in the image block, n is the number of image blocks of a distorted image, and n takes the value of 2 α .
[0036] Determine the image score S of the distorted image according to the following formula:
[0037]
[0038] Train the image quality evaluation model until the loss function converges. In each epoch, input the validation set into the trained image quality evaluation model; select the model with the minimum loss in the validation set as the final model.
[0039] (5) Test the image quality evaluation model
[0040] Input the images in the test set into the trained image quality evaluation model to obtain the final image quality scores.
[0041] In step (3) of constructing the image quality evaluation model of the present invention, the gradient feature extractor is composed of 1 convolutional layer, 2 pyramid convolutional layers, and 2 max pooling layers. Each series of 1 pyramid convolutional layer and 1 max pooling layer constitutes a pyramid convolutional unit. One convolutional layer and two pyramid convolutional units are serially connected in sequence to form the gradient feature extractor. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2. The pyramid convolutional layer divides the input feature map into 3 groups, and each group uses different convolutional kernels, which are 3×3, 5×5, and 7×7 respectively.
[0042] In step (3) of constructing the image quality evaluation model of the present invention, the spatial domain feature extractor is composed of 4 or 5 convolutional units connected in sequence. Each convolutional unit is composed of 2 convolutional layers and 1 max pooling layer connected in series. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2.
[0043] In step (3) of constructing the image quality evaluation model of the present invention, the quality regression module is composed of at least 3 fully connected layers connected in sequence.
[0044] In step (4) of training the image quality evaluation model of the present invention, the distorted image patches generate 3 feature vectors through the spatial domain feature extractor as feature vector Σ 1 、feature vector Σ 2 、feature vector Σ 3 :
[0045]
[0046] Among them, is an 8×8 matrix, m ∈ {1, 2, …, 64}.
[0047]
[0048] Among them, is a 4×4 matrix, n ∈ {1, 2, …, 128}.
[0049]
[0050] Among them, are all 2×2 matrices, and b ∈ {1, 2, …, 256}.
[0051] Convert the eigenvector Γ, eigenvector Σ 1 , eigenvector Σ 2 , eigenvector Σ 3 into single-column matrices, and vertically concatenate them in sequence to construct a vertical eigenvector Ν:
[0052] Ν = concat(Γ, Σ 1 , Σ 2 , Σ 3 ). (4)
[0053] In the (4) training image quality evaluation model of the present invention, the distorted image block generates 4 eigenvectors as eigenvector Σ 1 , eigenvector Σ 2 , eigenvector Σ 3 , eigenvector Σ 4 :
[0054]
[0055] Among them, is an 8×8 matrix, and m ∈ {1, 2, …, 64}.
[0056]
[0057] Among them, is a 4×4 matrix, and n ∈ {1, 2, …, 128}.
[0058]
[0059] Among them, are all 2×2 matrices, and b ∈ {1, 2, …, 256}.
[0060]
[0061] Among them, are all 1×1 matrices, and v ∈ {1, 2, …, 512}.
[0062] Convert the eigenvector Γ, eigenvector Σ 1 , eigenvector Σ 2 , eigenvector Σ 3 , eigenvector Σ 4 into single-column matrices, and vertically concatenate them in sequence to construct a vertical eigenvector Ν:
[0063] N = concat(Γ, Σ 1 , Σ 2 , Σ 3 , Σ 4 ). (6)
[0064] The beneficial effects of the present invention are as follows:
[0065] Since the original images and distorted images of the present invention are divided into training sets, test sets, and validation sets, the Sobel operator is used to extract the gradient features of the images, and the saliency detection method is used to extract the saliency maps. Each distorted image, gradient map, and saliency map is cropped into image patches of 32×32 pixel size. The distorted image is input into the neural network model to obtain the image quality scores of the image patches; weights are assigned to each image patch according to the saliency map; the sum of the products of the quality scores and weight scores of each image patch is obtained as the overall quality score of the distorted image. The simulation comparison experiment shows that the method of the present invention can more accurately predict the perceptual quality level of the distorted image, has good quality prediction results for different data sets, and is consistent with the subjective quality evaluation. The present invention has the advantages of simple method, fast evaluation speed, accurate evaluation, etc., and can be used for the evaluation of image quality. Description of the Drawings
[0066] Figure 1 is the flowchart of Embodiment 1 of the present invention.
[0067] Figure 2 is Figure 1 the structural schematic diagram of the image quality evaluation model in Specific Implementation Method
[0068] The following further details the present invention in conjunction with the drawings and embodiments, but the present invention is not limited to the following embodiments.
[0069] Embodiment 1
[0070] Taking the selection of 3000 images from the TID2013 image quality evaluation database as an example, the no-reference image quality evaluation method based on visual saliency and gradient features in this embodiment consists of the following steps (see Figure 1 ):
[0071] (1) Select the data set
[0072] The distorted images in the image quality evaluation database are divided into training sets, validation sets, and test sets, and the ratio of the number of training set images to validation set images and test set images is 8:1:1.
[0073] (2) Image preprocessing
[0074] The Sobel operator is used to extract the gradient features of the distorted image. The Sobel operator estimates the gradient of the central pixel using the eight adjacent pixels, and determines the gradient image G(x) according to the following formula:
[0075]
[0076]
[0077]
[0078] where * is the convolution operation and D(x) is the distorted image.
[0079] For the distorted image, the frequency-tuned saliency region detection method is used to extract the saliency map, perform Gaussian filtering, and convert from the RGB color space to the LAB color space; the means of the images of the L, A, and B channels of the converted image are taken respectively; the saliency value S(x, y) of each pixel is determined according to the following formula:
[0080] S(x, y) = ||I μ -I ω (x, y)||
[0081] where I μ is the average feature vector of the image, and I ω (x, y) is the image pixel vector value corresponding to the distorted image after Gaussian filtering. The saliency value S(x, y) of each pixel in the image is divided by the maximum saliency value to obtain the final saliency map.
[0082] Each distorted image, gradient map, and saliency map is cropped into image patches of 32×32 pixels in size as input samples.
[0083] (3) Construct an image quality evaluation model
[0084] In Figure 2 , the image quality evaluation model of this embodiment includes a gradient feature extractor 1, a quality regression module 2, a weight mechanism module 3, and a spatial domain feature extractor 4. The gradient feature extractor 1 and the spatial domain feature extractor 4 are in parallel, and after being connected in series with the quality regression module 2, they are in parallel with the weight mechanism module 3 to form.
[0085] The gradient feature extractor 1 of this embodiment consists of 1 convolutional layer, 2 pyramid convolutional layers, and 2 max pooling layers. Among them, each series-connected pyramid convolutional layer and 1 max pooling layer form a pyramid convolutional unit. A convolutional layer and two pyramid convolutional units are connected in series in sequence to form a gradient feature extractor. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2. The pyramid convolutional layer divides the input feature map into 3 groups, and each group uses different convolutional kernels, which are 3×3, 5×5, and 7×7 respectively.
[0086] The spatial domain feature extractor 4 of this embodiment is composed of 4 convolutional units connected in series in sequence. Each convolutional unit is composed of 2 convolutional layers and 1 max pooling layer connected in series. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2.
[0087] The quality regression module 2 of this embodiment consists of 3 fully connected layers connected in series in sequence.
[0088] (4) Training the image quality evaluation model
[0089] Use the first-order moment estimation method and the second-order moment estimation method of the optimizer gradient based on gradient to adaptively adjust the learning rate of each parameter. Set the learning rate to 0.0001, the exponential decay rate of the first-order moment estimation to 0.9, and the exponential decay rate of the second-order moment estimation to 0.999.
[0090] Use the following formula to determine the loss function L1:
[0091]
[0092] Among them, x represents the number of images, H x is the subjective score of the x-th image, and S x is the predicted quality score of the x-th image.
[0093] The gradient map in the training set generates a feature vector Γ through the gradient feature extractor 1 of the image quality evaluation model:
[0094] Γ = (γ1, γ2, L, γ k )
[0095] where γ k is an 8×8 matrix, k ∈ {1, 2,..., 64}; the distorted image block generates 3 feature vectors through the spatial domain feature extractor 4 as the feature vector Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 :
[0096]
[0097] Among them, is an 8×8 matrix, and m ∈ {1, 2, …, 64}.
[0098]
[0099] Among them, is a 4×4 matrix, and n ∈ {1, 2, …, 128}.
[0100]
[0101] Among them, are 2×2 matrices respectively, and b ∈ {1, 2, …, 256}.
[0102] Convert the eigenvector Γ, eigenvector Σ 1 , eigenvector Σ 2 , eigenvector Σ 3 into single-column matrices, and vertically concatenate them in sequence to construct a vertical eigenvector Ν:
[0103] Ν = concat(Γ, Σ 1 , Σ 2 , Σ 3 ) (4)
[0104] Use the vertical eigenvector Ν as the input of the quality regression module 2 to obtain the quality score S i of the i-th image patch, where i is the number of distorted image patches.
[0105] Divide the distorted image into 32×32 image patches. According to the change of the saliency image gray value, divide the saliency map into prominent regions, general regions, and non-prominent regions, and set 3 different thresholds. Pixel values between [0, 74] are non-prominent regions, pixel values between [75, 174] are normal regions, and pixel values between [175, 255] are prominent regions. Set different weights for different regions. The weights of the 3 different regions are {0.8, 1, 1.2} respectively to obtain the weight map W(x).
[0106] Determine the weight w of each image patch according to the following formula i :
[0107]
[0108]
[0109] where N d is the number of pixels in the image patch, n is the number of image patches of a distorted image, and n takes the value of 2 α , α takes a value of at least 5. In this embodiment, α takes a value of 5, that is, n is 32.
[0110] Determine the image score S of the distorted image according to the following formula:
[0111]
[0112] Train the image quality evaluation model until the loss function converges. In each epoch, input the validation set into the trained image quality evaluation model; select the model with the minimum loss in the validation set as the final model.
[0113] (5) Test the image quality evaluation model
[0114] Input the images in the test set into the trained image quality evaluation model to obtain the final image quality score.
[0115] Complete the no-reference image quality evaluation method based on visual saliency and gradient features.
[0116] Embodiment 2
[0117] Taking 3000 images selected from the TID2013 image quality evaluation database as an example, the no-reference image quality evaluation method based on visual saliency and gradient features in this embodiment consists of the following steps:
[0118] (1) Select the data set
[0119] This step is the same as that in Embodiment 1.
[0120] (2) Image preprocessing
[0121] This step is the same as that in Embodiment 1.
[0122] (3) Construct the image quality evaluation model
[0123] The image quality evaluation model includes a gradient feature extractor 1, a spatial domain feature extractor 4, a quality regression module 2, and a weight mechanism module 3. The gradient feature extractor 1 is in parallel with the spatial domain feature extractor 4, and after being connected in series with the quality regression module 2, it is in parallel with the weight mechanism module 3 to form.
[0124] The structure of the gradient feature extractor 1 in this embodiment is the same as that in Embodiment 1.
[0125] The structure of the spatial domain feature extractor 4 in this embodiment is the same as that in Embodiment 1.
[0126] The quality regression module 2 in this embodiment is composed of 5 fully connected layers connected in series in turn.
[0127] (4) Train the image quality evaluation model
[0128] The first - order moment estimation method and the second - order moment estimation method of the gradient are used to adaptively adjust the learning rate of each parameter with the Adaptive Moment Estimation optimizer based on gradients. The learning rate is set to 0.0001, the exponential decay rate of the first - order moment estimation is 0.9, and the exponential decay rate of the second - order moment estimation is 0.999.
[0129] The loss function L1 is determined by the following formula:
[0130]
[0131] where x represents the number of images, H x is the subjective score of the x - th image, and S x is the predicted quality score of the x - th image.
[0132] The gradient map in the training set generates a feature vector Γ through the gradient feature extractor 1 of the image quality evaluation model:
[0133] Γ=(γ1,γ2,L,γ k )
[0134] where γ k is an 8×8 matrix, k∈{1,2,…,64}; the distorted image block generates 3 feature vectors as the feature vector Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 :
[0135]
[0136] where, is an 8×8 matrix, m∈{1,2,…,64}.
[0137]
[0138] where, is a 4×4 matrix, n∈{1,2,…,128}.
[0139]
[0140] where, are 2×2 matrices respectively, b∈{1,2,…,256}.
[0141] The feature vector Γ, the feature vector Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 are converted into single - column matrices and vertically concatenated in sequence to construct a vertical feature vector Ν:
[0142] Ν = concat(Γ,Σ 1 ,Σ2 ,Σ 3 ) (4)
[0143] Take the vertical feature vector Ν as the input of the quality regression module 2 to obtain the quality score S of the i-th image patch i , where i is the number of distorted image patches.
[0144] Divide the distorted image into 32×32 image patches. According to the change of the grayscale value of the saliency map, the saliency map is divided into a prominent region, a general region, and a non-prominent region, and three different thresholds are set. The pixel value between [0, 74] is the non-prominent region, the pixel value between [75, 174] is the normal region, and the pixel value between [175, 255] is the prominent region. Different weights are set for different regions, and the weights of the three different regions are {0.8, 1, 1.2} respectively to obtain the weight map W(x).
[0145] Determine the weight w of each image patch according to the following formula i :
[0146]
[0147]
[0148] where N d is the number of pixels in the image patch, n is the number of image patches of a distorted image, and n takes the value of 2 α , α takes a value of at least 5, and α in this embodiment takes a value of 6, that is, n is 64.
[0149] Determine the image score S of the distorted image according to the following formula
[0150]
[0151] Train the image quality evaluation model until the loss function converges. In each epoch, input the validation set into the trained image quality evaluation model; select the model with the smallest loss in the validation set as the final model.
[0152] (5) Test the image quality evaluation model
[0153] This step is the same as that in Embodiment 1.
[0154] Complete the no-reference image quality evaluation method based on visual saliency and gradient features.
[0155] Embodiment 3
[0156] Taking 3000 images selected from the TID2013 image quality evaluation database as an example, the no-reference image quality evaluation method based on visual saliency and gradient features in this embodiment consists of the following steps
[0157] (1) Select the dataset
[0158] This step is the same as that in Embodiment 1.
[0159] (2) Image preprocessing
[0160] This step is the same as that in Embodiment 1.
[0161] (3) Construct an image quality evaluation model
[0162] The image quality evaluation model includes a gradient feature extractor 1, a spatial domain feature extractor 4, a quality regression module 2, and a weight mechanism module 3. The gradient feature extractor 1 is in parallel with the spatial domain feature extractor 4, and after being in series with the quality regression module 2, it is in parallel with the weight mechanism module 3.
[0163] The structure of the gradient feature extractor 1 in this embodiment is the same as that in Embodiment 1.
[0164] The spatial domain feature extractor 4 in this embodiment is composed of 5 convolutional units connected in series in sequence. Each convolutional unit is composed of 2 convolutional layers connected in series with 1 max pooling layer. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2.
[0165] The structure of the quality regression module 2 in this embodiment is the same as that in Embodiment 1.
[0166] (4) Train the image quality evaluation model
[0167] Use the first-order moment estimation method and the second-order moment estimation method of the gradient of the adaptive moment estimation optimizer based on the gradient to adaptively adjust the learning rate of each parameter. Set the learning rate to 0.0001, the exponential decay rate of the first-order moment estimation to 0.9, and the exponential decay rate of the second-order moment estimation to 0.999.
[0168] Determine the loss function L1 with the following formula:
[0169]
[0170] where x represents the number of images, H x is the subjective score of the x-th image, and S x is the predicted quality score of the x-th image.
[0171] The gradient maps in the training set generate feature vectors Γ through the gradient feature extractor 1 of the image quality evaluation model:
[0172] Γ = (γ1, γ2, L, γ k )
[0173] where γ kis an 8×8 matrix, k ∈ {1, 2, …, 64}; the distorted image block generates 4 feature vectors through the spatial domain feature extractor 4 as the feature vector Σ 1 、feature vector Σ 2 、feature vector Σ 3 、feature vector Σ 4 :
[0174]
[0175] Among them, is an 8×8 matrix, m ∈ {1, 2, …, 64}.
[0176]
[0177] Among them, is a 4×4 matrix, n ∈ {1, 2, …, 128}.
[0178]
[0179] Among them, are 2×2 matrices respectively, b ∈ {1, 2, …, 256}.
[0180]
[0181] Among them, are 1×1 matrices respectively, v ∈ {1, 2, …, 512}.
[0182] Convert the feature vector Γ, the feature vector Σ 1 、feature vector Σ 2 、feature vector Σ 3 、feature vector Σ 4 into single-column matrices, and vertically concatenate them in sequence to construct a vertical feature vector Ν:
[0183] Ν = concat(Γ, Σ 1 , Σ 2 , Σ 3 , Σ 4 ) (6)
[0184] Take the vertical feature vector Ν as the input of the quality regression module 2 to obtain the quality score S i of the i-th image block, where i is the number of distorted image blocks.
[0185] The distorted image block is divided into image blocks of 32×32. According to the change of the grayscale value of the saliency image, the saliency map is divided into prominent regions, general regions, and non-prominent regions, and three different thresholds are set. The pixel value in the range of [0, 74] is the non-prominent region, the pixel value in the range of [75, 174] is the normal region, and the pixel value in the range of [175, 255] is the prominent region. Different weights are set for different regions, and the weights of the three different regions are {0.8, 1, 1.2} respectively, to obtain the weight map W(x).
[0186] Determine the weight w of each image block according to the following formula i :
[0187]
[0188]
[0189] where N d is the number of pixels in the image block, n is the number of image blocks of a distorted image, and n takes the value of 2 α , α takes a value of at least 5, and in this embodiment, α takes a value of 5, that is, n is 32.
[0190] Determine the image score S of the distorted image according to the following formula
[0191]
[0192] Train the image quality evaluation model until the loss function converges. In each epoch, input the validation set into the trained image quality evaluation model; select the model with the minimum loss in the validation set as the final model.
[0193] Other steps are the same as those in Embodiment 1. Complete the no-reference image quality evaluation method based on visual saliency and gradient features.
[0194] To verify the beneficial effects of the present invention, the inventors selected 3,000, 900, and 983 images from the TID2013, CSIQ, and LIVE image quality evaluation databases respectively, and compared and simulated the method of Embodiment 1 of the present invention with the peak signal-to-noise ratio (abbreviated as PSNR), structural similarity (abbreviated as SSIM), feature similarity (abbreviated as FSIMC), Deep similarity for image quality assessment (abbreviated as DeepSim), Convolutional neural networks for no-reference image quality assessment (abbreviated as CNN), Nested error map generation network for no-reference image quality assessment (abbreviated as NEMG-IQA), and End-to-end blind image quality prediction with cascaded deep neural network (abbreviated as CaHDC) methods. In the simulation experiment, the Pearson linear correlation coefficient PLCC and the rank correlation coefficient SROCC were determined according to the following formula:
[0195]
[0196]
[0197] where N represents the number of images in the test set, x β is the objective prediction score of the β-th image, and y β is the subjective score of the βi-th image, where β = (1, 2,..., N), are the average values of the subjective scores of the database images and the quality scores obtained by the objective algorithm respectively, and R xβ and R yβ represent the permutation orders of x β and y β respectively.
[0198] The experimental and calculation results are shown in Table 1.
[0199] Table 1 Simulation experiment results of the method of Embodiment 1 and the comparative experiment methods
[0200]
[0201] As can be seen from Table 1, the Pearson linear correlation coefficient and rank correlation coefficient of the image quality evaluation model tested by the method of the present invention on TID2013 are 0.872 and 0.890 respectively, and the Pearson linear correlation coefficient and rank correlation coefficient of the image quality evaluation model tested on the CSIQ database are 0.927 and 0.939 respectively, and the performance is better than that of the existing image quality evaluation methods; the Pearson linear correlation coefficient and rank correlation coefficient of the image quality evaluation model tested on the LIVE database are 0.975 and 0.972 respectively, which are better than most of the image quality evaluation methods, and only differ by 0.003 compared with the rank correlation coefficient of the NEMG-IQA method; it shows that the model has good quality evaluation results for different data sets.
[0202] In summary, by considering the visual saliency and gradient features of the image and performing multi-level feature extraction on the distorted image, the evaluation effect is consistent with the human subjective perception.
Claims
1. A reference - free image quality assessment method based on visual saliency and gradient features, characterized in that It consists of the following steps: (1) Select the dataset Divide the distorted images in the image quality evaluation database into a training set, a validation set, and a test set. The ratio of the number of training set images to the number of validation set images and test set images is 8:1:1; (2) Image preprocessing Use the Sobel operator to extract the gradient features of the distorted images. The Sobel operator estimates the gradient of the central pixel with eight adjacent pixels and determines the gradient image G(x) according to the following formula: where * is the convolution operation and D(x) is the distorted image; Use the frequency-tuned saliency region detection method to extract the saliency map for the distorted images, perform Gaussian filtering, and convert from the RGB color space to the LAB color space; take the mean values of the images in the L, A, and B channels of the converted image respectively; determine the saliency value S(x, y) of each pixel according to the following formula: S(x, y) = ||I μ - I ω (x, y)|| Among them, I μ is the average feature vector of the image, and I ω (x, y) is the image pixel vector value corresponding to the distorted image after Gaussian filtering; the significance value S(x, y) of each pixel in the image is divided by the maximum significance value to obtain the final significance map; Crop each distorted image, gradient map, and saliency map into image patches of 32×32 pixels as input samples; (3) Construct an image quality evaluation model The image quality evaluation model includes a gradient feature extractor (1), a quality regression module (2), a weight mechanism module (3), and a spatial domain feature extractor (4). The gradient feature extractor (1) is in parallel with the spatial domain feature extractor (4), and after being in series with the quality regression module (2), it is in parallel with the weight mechanism module (3) to form; The gradient feature extractor consists of a convolutional layer, a pyramid convolutional layer, and a max pooling layer. The spatial domain feature extractor consists of convolutional units in series. The quality regression module consists of fully connected layers; (4) Train the image quality evaluation model Use the first-order moment estimation method and the second-order moment estimation method of the gradient of the adaptive moment estimation optimizer based on the gradient to adaptively adjust the learning rate of each parameter. Set the learning rate to 0.0001, the exponential decay rate of the first-order moment estimation to 0.9, and the exponential decay rate of the second-order moment estimation to 0.999; Determine the loss function L1 with the following formula: Among them, x represents the number of images, and H x is the subjective score of the x-th image, and S x is the predicted quality score of the x-th image; The gradient map in the training set generates a feature vector Γ through the gradient feature extractor (1) of the image quality evaluation model; Γ = (γ1, γ2, …, γ k ) where γ k is an 8×8 matrix, k ∈ {1, 2, …, 64}; the distorted image block generates 3 or 4 feature vectors through the spatial domain feature extractor (4); Convert the feature vector Γ and the 3 or 4 feature vectors generated by the spatial domain feature extractor into a single-column matrix and vertically concatenate them in sequence to construct a vertical feature vector Ν; Using the vertical feature vector Ν as the input of the quality regression module (2) to obtain the quality score S of the i-th image patch i , where i is the number of distorted image patches; Divide the distorted image into 32×32 image patches, divide the saliency map into prominent regions, general regions, and non-prominent regions according to the change of the saliency image gray value, and set 3 different thresholds. The pixel value between [0, 74] is the non-prominent region, the pixel value between [75, 174] is the normal region, and the pixel value between [175, 255] is the prominent region. Set different weights for different regions. The weights of the 3 different regions are {0.8, 1, 1.2} to obtain the weight map W(x); Determine the weight w of each image block according to the following formula i : where N d is the number of pixels in the image block, n is the number of image blocks of a distorted image, and n takes a value of 2 α ; Determine the image score S of the distorted image with the following formula: Train the image quality evaluation model until the loss function converges. In each epoch, input the validation set into the trained image quality evaluation model; select the model with the smallest loss in the validation set as the final model; (5) Test the image quality evaluation model Input the images of the test set into the trained image quality evaluation model to obtain the final image quality score.
2. The no-reference image quality evaluation method based on visual saliency and gradient features according to claim 1, characterized in that: In the step (3) of constructing the image quality evaluation model, the gradient feature extractor (1) consists of 1 convolutional layer, 2 pyramid convolutional layers, and 2 max pooling layers. One pyramid convolutional unit is formed by connecting one pyramid convolutional layer and one max pooling layer in series. The gradient feature extractor is formed by connecting one convolutional layer and two pyramid convolutional units in series. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2. The pyramid convolutional layer divides the input feature map into 3 groups, and different convolutional kernels are used for each group, which are 3×3, 5×5, and 7×7 respectively.
3. The no-reference image quality assessment method based on visual saliency and gradient features according to claim 1, characterized in that: In the step (3) of constructing the image quality evaluation model, the spatial domain feature extractor (4) is formed by connecting 4 or 5 convolutional units in series. Each convolutional unit is formed by connecting 2 convolutional layers and 1 max pooling layer in series. The convolutional kernel pixels of the convolutional layer are 3×3, the stride of the sliding window is 1, the kernel pixels of the max pooling layer are 2×2, and the stride of the sliding window is 2.
4. The no-reference image quality assessment method based on visual saliency and gradient features according to claim 1, characterized in that: In the step (3) of constructing the image quality evaluation model, the quality regression module (2) consists of at least 3 fully connected layers connected in series.
5. The no-reference image quality assessment method based on visual saliency and gradient features according to claim 1, characterized in that In (4) training the image quality evaluation model, three feature vectors are generated by the spatial domain feature extractor (4) for the distorted image block, namely, feature vector Σ 1 , feature vector Σ 2 , and feature vector Σ 3 : Among them, is an 8×8 matrix, where m ∈ {1, 2, …, 64}; Among them, is a 4×4 matrix, where n ∈ {1, 2, …, 128}; Among them, are all 2×2 matrices, and b ∈ {1, 2, …, 256}; Convert the feature vectors Γ and Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 into single-column matrices, and vertically concatenate them in sequence to construct a vertical feature vector Ν: N = concat(Γ, Σ 1 , Σ 2 , Σ 3 ) (4).
6. The no-reference image quality assessment method based on visual saliency and gradient features according to claim 1, characterized in that In (4) training the image quality evaluation model, four feature vectors are generated by the spatial domain feature extractor (4) for the distorted image block, which are the feature vector Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 , the feature vector Σ 4 : Among them, is an 8×8 matrix, and m ∈ {1, 2, …, 64}; Among them, is a 4×4 matrix, and n ∈ {1, 2, …, 128}; wherein, are each a 2×2 matrix, and b ∈ {1, 2, …, 256}; Among them, are all 1×1 matrices, where v ∈ {1, 2, …, 512}; Convert the feature vectors Γ and Σ 1 , the feature vector Σ 2 , the feature vector Σ 3 , the feature vector Σ 4 into single-column matrices, and vertically concatenate them in sequence to construct a vertical feature vector Ν: N = concat(Γ, Σ 1 , Σ 2 , Σ 3 , Σ 4 ) (6).