Image quality assessment method based on multi-level feature distribution

By adopting a multi-level feature distribution method in image quality evaluation, using the VGG-16 network to extract features and combining generalized Gaussian distribution and deep forest model, the problem of low accuracy in image quality evaluation in the prior art is solved, and more efficient and accurate image quality evaluation is achieved.

CN118864369BActive Publication Date: 2025-05-13NINGXIA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410861552.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-05-13
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

When the existing image quality evaluation method uses no reference IQA, it fails to fully consider the degradation of different levels of features, resulting in low accuracy of image quality evaluation.

Method used

The image quality evaluation method based on multi-level feature distribution is adopted, and the underlying, middle and high-level features are extracted through the pre-trained 16-layer convolutional neural network VGG-16, and the feature map coefficients are fitted based on the generalized Gaussian distribution model. Finally, the feature vectors are mapped on the subjective score using the deep forest model.

Benefits of technology

It improves the accuracy and efficiency of image quality evaluation, reduces the load on the processor, can more effectively reflect degradation between levels, and solves the problem of model overfitting and not being suitable for large data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864369B_ABST
    Figure CN118864369B_ABST
Patent Text Reader

Abstract

The present invention provides an image quality assessment method based on multi-level feature distribution, which relates to the field of image processing technology, including: pre-training a 16-layer convolutional neural network VGG-16; wherein the first 13 layers of VGG-16 are convolutional layers, and the last 3 layers are fully connected layers; each convolutional layer is used to extract bottom-level features, middle-level features, and high-level features respectively; inputting the image to be evaluated into VGG-16 for feature extraction; determining a target convolutional layer combination; the target convolutional layer combination includes a shallow convolutional layer, a middle convolutional layer, and a high-level convolutional layer; fitting the feature map coefficients of the image features extracted by each convolutional layer in the target convolutional combination based on a generalized Gaussian distribution model to obtain an image feature vector; using a pre-constructed deep forest model to map the obtained image feature vector to a subjective score corresponding to the image, and obtain a quality assessment prediction value of the image to be evaluated. This scheme can improve the accuracy and efficiency of image quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image quality evaluation method based on multi-level feature distribution. Background Art

[0002] With the rapid development of electronic information technology, the 5G era has entered the era of the Internet of Everything. People can get real-time information through electronic tools such as mobile phones without leaving their homes. As a carrier for people to transmit information, images are the simplest and fastest visual information communication language. They are widely used in the fields of aerospace, industrial inspection, and health care. For example, in the field of aerospace, using satellite-borne cameras to obtain data on the moon, people can see the topography of the moon in all directions; in the field of industrial inspection, image detection can solve the problems of low accuracy and slow detection speed in manual inspection; in the field of health care, high-quality images can improve the accuracy and efficiency of doctors' diagnosis and provide better medical services for patients. However, various distortions will be introduced in the process of image acquisition and compression, resulting in a decrease in image quality, which will affect people's subjective feelings and information acquisition. Therefore, in order to improve the perceived quality of images, image quality assessment (IQA) is particularly important.

[0003] IQA is mainly divided into subjective evaluation and objective evaluation. Subjective evaluation relies on observers to score image quality. It has high accuracy, but is time-consuming and labor-intensive, and is easily affected by factors such as the environment and the observer's cognitive background. Objective evaluation is to design a model related to image quality to evaluate the quality of the image. According to the degree of dependence on the reference image, objective evaluation can be divided into three types of evaluation methods: Full-Reference IQA (FR-IQA), Reduced-Reference IQA (RR-IQA) and No-Reference IQA (NR-IQA). The FR-IQA method requires all the information of the reference image. By comparing the difference between the distorted image and the reference image, the distortion degree of the distorted image is analyzed to obtain the quality score of the distorted image. The evaluation results of this method are accurate and reliable; the RR-IQA method requires partial information of the reference image, and has the advantages of small amount of data transmission and strong flexibility; the NR-IQA method is also called blind image quality evaluation. This method does not need to consider the specific information of the reference image and is widely used in various fields. However, in practice, it is difficult for people to obtain specific information of reference images. Considering the limitation that FR-IQA and RR-IQA methods require image information during the evaluation process, people usually use NR-IQA method to evaluate image quality.

[0004] At present, there are two main methods for evaluating image quality using the NR-IQA method. The first method only extracts features of a single level, that is, the last layer of the deep neural network DNN, for quality evaluation. However, distortions at different levels will produce different degradations in the hierarchical features. The existing schemes do not fully consider the degradation of hierarchical features, resulting in a large gap with subjective perception, that is, the accuracy of image quality evaluation is low. The second method uses features of all levels to predict image quality. However, although this method has been improved, the extracted features of different levels are not obvious enough and cannot effectively reflect the degradation between levels. Moreover, this method requires the extraction of a large number of features, with a high dimensionality and a large processor load, which greatly reduces the prediction efficiency of the model. Summary of the invention

[0005] In view of this, in order to address the above shortcomings, it is necessary to propose an image quality evaluation method based on multi-level feature distribution to improve the accuracy and efficiency of image quality evaluation.

[0006] In a first aspect, the present invention provides an image quality assessment method based on multi-level feature distribution, comprising:

[0007] Pre-train a 16-layer convolutional neural network VGG-16; wherein the first 13 layers of the convolutional neural network VGG-16 are convolutional layers, and the last 3 layers are fully connected layers; the 13 convolutional layers include N L shallow convolutional layers, N M The middle convolutional layers and N H N high-level convolutional layers are used to extract low-level features, middle-level features, and high-level features respectively; L +N M +N H =13;

[0008] Inputting the image to be evaluated for image quality evaluation into the convolutional neural network VGG-16 for feature extraction;

[0009] Determine a target convolution layer combination; wherein the target convolution layer combination includes a shallow convolution layer, a middle convolution layer and a high convolution layer;

[0010] Fitting feature map coefficients of image features extracted by each convolution layer in the target convolution combination based on a generalized Gaussian distribution model to obtain an image feature vector;

[0011] The obtained image feature vector is mapped to a subjective score corresponding to the image using a pre-constructed deep forest model to obtain a quality evaluation prediction value of the image to be evaluated.

[0012] Preferably,

[0013] The underlying features include texture and edge features of the image;

[0014] The middle-level features include contour and shape features of the image;

[0015] The high-level features include the content or spatial structure of the image.

[0016] Preferably, the first 4 of the 13 convolutional layers are shallow convolutional layers, the middle 6 are medium convolutional layers, and the last 3 are high-level convolutional layers;

[0017] The determining of a target convolutional layer combination includes:

[0018] Select one from each of the shallow convolution layer, the middle convolution layer and the high convolution layer to form a primary convolution layer combination, and obtain A combination of primary convolutional layers;

[0019] The image features corresponding to each combination of primary convolutional layers are evaluated using a preset image quality evaluation dataset;

[0020] The primary convolutional layer with the best evaluation result among the primary convolutional layer combinations is determined as the target convolutional layer combination.

[0021] Preferably, the image quality assessment dataset includes at least one of LIVE, TID2008, TID2013 and CSIQ.

[0022] Preferably, the four shallow convolutional layers are {Conv1_1, Conv1_2, Conv2_1, Conv2_2}, the six middle convolutional layers are {Conv3_1, Conv3_2, Conv3_3, Conv4_1, Conv4_2, Conv4_3}, and the three high-level convolutional layers are {Conv5_1, Conv5_2, Conv5_3};

[0023] The target convolutional layer combination is {Conv2_1, Conv3_1, Conv5_2}.

[0024] Preferably, the performing feature map coefficient fitting on the image features extracted by each convolution layer in the target convolution combination based on the generalized Gaussian distribution model comprises: taking the parameters of the generalized Gaussian distribution model as the quality perception features, and performing feature map coefficient fitting;

[0025] Among them, the probability density function of the generalized Gaussian distribution model is:

[0026]

[0027] in,

[0028] In the formula, α represents the shape parameter of the distribution model, and β represents the scale parameter of the distribution model;

[0029] The parameters of the distribution model are α i,j and β i,j The representation is:

[0030]

[0031] The image feature vectors generated by all feature maps are:

[0032] f i,j ={f L,j ,f M,j ,f H,j}

[0033] in,

[0034] Among them, F i,j It represents the j-th feature map extracted by the i-th convolutional layer, where i=L,M,H, represents three shallow convolutional layers Conv2_1, Conv3_1, Conv5_2, a middle convolutional layer and a high-level convolutional layer, and j=128, 256, 512; GGD represents the generalized Gaussian distribution model.

[0035] Preferably, each layer of the deep forest model consists of 4 random forests, and each random forest contains D decision trees; D ≥ 50;

[0036] The method of mapping the obtained image feature vector to a subjective score corresponding to the image by using a pre-built deep forest model includes:

[0037] The image feature vector passes through 4 random forests in the first layer to obtain a 4-dimensional feature vector;

[0038] The original image feature vector is combined with the 4-dimensional feature vector as the input of the next layer;

[0039] And so on, until the quality evaluation prediction value is finally output.

[0040] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute any of the methods described in the first aspect.

[0041] In a third aspect, the present invention provides a computing device, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, any method described in the first aspect is implemented.

[0042] It can be seen from the above technical scheme that in the image quality assessment method based on multi-level feature distribution provided by this scheme, a 16-layer convolutional neural network VGG-16 is first pre-trained, the first 13 layers of the convolutional neural network are convolutional layers, the last 3 layers are fully connected layers, and the convolutional layers can be divided into shallow convolutional layers, middle convolutional layers and high-level convolutional layers, which are used to extract bottom-level features, middle-level features and high-level features respectively. Then the image to be evaluated for image quality is input into the trained convolutional neural network VGG-16 for feature extraction. Further, a target convolutional layer combination consisting of a shallow convolutional layer, a middle convolutional layer and a high-level convolutional layer is determined, and the feature map coefficients of the images extracted by each convolutional layer in the target convolutional combination are fitted based on the generalized Gaussian distribution model. Finally, the pre-constructed deep forest model is used to map the obtained image feature vector to the subjective score corresponding to the image to obtain the quality assessment prediction value. It can be seen that this scheme fully considers the degradation of features at different levels, and extracts features from the shallow convolution layer, the middle convolution layer and the high-level convolution layer, so as to cover the feature types that each convolution layer focuses on. By fusing the extracted bottom-level features, middle-level features and high-level features, the accuracy of image quality evaluation is greatly improved. Moreover, this scheme selects a target convolution layer combination from each of the shallow convolution layer, the middle convolution layer and the high-level convolution layer, and performs quality evaluation based on the image features extracted from the target convolution layer combination, which greatly reduces the amount and dimension of data processing, thereby reducing the processing load of the processor and improving the efficiency of image quality evaluation. In addition, this scheme also uses the deep forest method for prediction, which can solve problems such as model overfitting and unsuitability for large data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of an image quality assessment method based on multi-level feature distribution is provided in an embodiment of the present invention.

[0044] Figure 2 Feature maps of different convolutional layers of VGG-16.

[0045] Figure 3 Performance comparison results of the target convolutional layer combination and each convolutional layer feature.

[0046] Figure 4 The figure shows the PLCC performance comparison of the three regression methods on various standard data sets.

[0047] Figure 5 The following is a comparison chart of the SROCC performance of three regression methods on various standard data sets.

[0048] in, Figure 2 (a) is the feature map of the shallow convolutional layer. Figure 2 (b) is the feature map of the middle convolutional layer. Figure 2 (c) is the feature map of the high-level convolutional layer. DETAILED DESCRIPTION

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] like Figure 1 As shown, the present invention provides an image quality evaluation method based on multi-level feature distribution, which may include the following steps:

[0051] Step 101: pre-train a 16-layer convolutional neural network VGG-16; wherein the first 13 layers of the convolutional neural network VGG-16 are convolutional layers, and the last 3 layers are fully connected layers; the 13 convolutional layers include N L shallow convolutional layers, N M The middle convolutional layers and N H N high-level convolutional layers are used to extract low-level features, middle-level features, and high-level features respectively; L +N M +N H =13;

[0052] Step 102: inputting the image to be evaluated for image quality into the convolutional neural network VGG-16 for feature extraction;

[0053] Step 103: determining a target convolution layer combination; wherein the target convolution layer combination includes a shallow convolution layer, a middle convolution layer and a high convolution layer;

[0054] Step 104: performing feature map coefficient fitting on the image features extracted by each convolution layer in the target convolution combination based on a generalized Gaussian distribution model to obtain an image feature vector;

[0055] Step 105: Use a pre-built deep forest model to map the obtained image feature vector to a subjective score corresponding to the image, and obtain a quality evaluation prediction value of the image to be evaluated.

[0056] In this embodiment, the degradation of features at different levels is fully considered, and feature extraction is performed from the shallow convolution layer, the middle convolution layer, and the high-level convolution layer, respectively, so as to cover the feature types concerned by each convolution layer. By fusing the extracted bottom-level features, middle-level features, and high-level features, the accuracy of image quality evaluation is greatly improved. Moreover, this solution selects a target convolution layer combination from each of the shallow convolution layer, the middle convolution layer, and the high-level convolution layer, and performs quality evaluation based on the image features extracted from the target convolution layer combination, which greatly reduces the amount and dimension of data processing, thereby reducing the processing load of the processor and improving the efficiency of image quality evaluation. In addition, this solution also uses the deep forest method for prediction, which can solve problems such as model overfitting and unsuitability for large data sets.

[0057] For step 101, a 16-layer convolutional neural network VGG-16 is pre-trained;

[0058] Neuroscience research shows that the visual perception process is hierarchical, and the input visual signals are processed hierarchically from local details to global semantics. As the depth of the convolutional layer of the neural network increases from shallow to deep, the learned features also increase from low-level to high-level. The low-level features mainly focus on local details, the middle-level features mainly focus on regional patterns, and the high-level features contain global abstract information. Since the feature maps generated by VGG-16 perform well in terms of object categories and understanding images, this solution considers pre-training the VGG-16 network for feature extraction.

[0059] The VGG-16 network structure contains 16 layers of network parameters, including 13 convolutional layers and 3 fully connected layers. Considering that the fully connected layer plays the role of a classifier in the entire convolutional neural network for image classification, this solution ignores the fully connected layer. Figure 2 As shown in the figure, experiments show that the features of the first 4 layers of the VGG-16 network are based on shallow features, such as texture and edge features; the middle 6 layers propose more complex information, such as contour and shape features; the feature maps of the last 3 layers mainly retain the deep features of the image, such as content or spatial structure. Feature maps at different levels can capture different information of the image. Therefore, fusing features at different levels can capture information that may not be easily perceived by the human visual perception system, thereby improving the image quality evaluation model.

[0060] In step 102, the image to be evaluated for image quality evaluation is input into the convolutional neural network VGG-16 for feature extraction;

[0061] In this step, consider using the pre-trained convolutional neural network VGG-16 to extract features from the input image to be evaluated, and extract the image features corresponding to each convolutional layer.

[0062] For step 103, a target convolution layer combination is determined; wherein the target convolution layer combination includes a shallow convolution layer, a middle convolution layer and a high-level convolution layer;

[0063] In order to reduce the dimension and extract more comprehensive image features, this step considers selecting a convolution layer from the shallow convolution layer, the middle convolution layer and the high-level convolution layer to form a target convolution layer combination, and performs image quality evaluation based on the features extracted by the convolution layer in the target convolution layer combination. This not only fully considers the degradation of features at different levels, making the image quality evaluation more accurate, but also greatly reduces the amount of data processing and the dimension of data processing, thereby improving the efficiency of image quality prediction. Specifically, the first 4 of the 13 convolution layers can be divided into shallow convolution layers, the middle 6 can be divided into middle convolution layers, and the last 3 can be divided into high-level convolution layers. In this way, step 103 can be specifically implemented in the following way:

[0064] Select one from each of the shallow convolution layer, the middle convolution layer and the high convolution layer to form a primary convolution layer combination, and obtain A combination of primary convolutional layers;

[0065] The image features corresponding to each combination of primary convolutional layers are evaluated using a preset image quality evaluation dataset;

[0066] The primary convolutional layer with the best evaluation result among the primary convolutional layer combinations is determined as the target convolutional layer combination.

[0067] The four shallow convolutional layers can be {Conv1_1, Conv1_2, Conv2_1, Conv2_2}, the six middle convolutional layers can be {Conv3_1, Conv3_2, Conv3_3, Conv4_1, Conv4_2, Conv4_3}, and the three high-level convolutional layers can be {Conv5_1, Conv5_2, Conv5_3};

[0068] The target convolutional layer combination may be {Conv2_1, Conv3_1, Conv5_2}.

[0069] In this embodiment, the 13 convolutional layers are divided into three groups, {Conv1_1, Conv1_2, Conv2_1, Conv2_2}, {Conv3_1, Conv3_2, Conv3_3, Conv4_1, Conv4_2, Conv4_3}, {Conv5_1, Conv5_2, Conv5_3}, and one convolutional layer is selected from each of the three groups to represent the bottom, middle and high-level features of the image, respectively. Therefore, there are a total of Further numerical experiments are conducted in preset image quality evaluation datasets, such as LIVE, TID2008, TID2013, and CSIQ datasets, to select the best performance combination. Figure 3 Shown is the performance comparison of the Spearman rank correlation coefficient SROCC, Pearson linear correlation coefficient PLCC and Kendall rank correlation coefficient KROCC of the target convolution layer combination A of {Conv2_1, Conv3_1, Conv5_2} and the features of each layer. The experimental results show that the various objective indicators under the combination of {Conv2_1, Conv3_1, Conv5_2} have the best effect, and it is used as the target convolution layer combination of this scheme.

[0070] Of course, in some embodiments, in order to reduce the amount of experiments, the convolution layer with the best single-layer performance among the shallow convolution layer, the middle convolution layer and the high-level convolution layer can be selected to form a target convolution layer combination, thereby avoiding experiments on 72 combinations and reducing the number of experiments.

[0071] For step 104, feature map coefficients are fitted on the image features extracted by each convolution layer in the target convolution combination based on a generalized Gaussian distribution model to obtain an image feature vector;

[0072] Since the sizes of different feature images are inconsistent, this step considers statistical analysis of the feature map through generalized Gaussian distribution, and uses the statistical results as the quality perception features of the image. Specifically, step 104 uses the parameters of the generalized Gaussian distribution model as the quality perception features to fit the feature map coefficients;

[0073] Among them, the probability density function of the generalized Gaussian distribution model is:

[0074]

[0075] in,

[0076] In the formula, α represents the shape parameter of the distribution model, and β represents the scale parameter of the distribution model;

[0077] The parameters of the distribution model are α i,j and β i,j The representation is:

[0078]

[0079] The image feature vectors generated by all feature maps are:

[0080] f i,j ={f L,j ,f M,j ,f H,j}

[0081] in,

[0082] Among them, F i,j represents the jth feature map extracted by the i-th convolutional layer, where i = L, M, H, represents three shallow convolutional layers Conv2_1, Conv3_1, Conv5_2, a middle convolutional layer and a high-level convolutional layer, j = 128, 256, 512; GGD represents the generalized Gaussian distribution model. i,j is a eigenvector with dimension 1792.

[0083] In this embodiment, after extracting multiple layers of feature maps, in order to perform better feature representation, the GGD model is considered to be used to fit the feature map coefficients, and the parameters in the GGD model are used as quality perception characteristics, which can be used to process and analyze feature maps of three different sizes, thereby solving the problem of inconsistent sizes of different feature maps.

[0084] For step 105, the obtained image feature vector is mapped to a subjective score corresponding to the image using a pre-constructed deep forest model to obtain a quality evaluation prediction value of the image to be evaluated.

[0085] In this step, the deep forest model is considered to predict the predicted value of the output image quality evaluation. Specifically, each layer of the deep forest model consists of 4 random forests, and each random forest contains D decision trees; D ≥ 50;

[0086] The method of mapping the obtained image feature vector to a subjective score corresponding to the image by using a pre-built deep forest model includes:

[0087] The image feature vector passes through 4 random forests in the first layer to obtain a 4-dimensional feature vector;

[0088] The original image feature vector is combined with the 4-dimensional feature vector as the input of the next layer;

[0089] And so on, until the quality evaluation prediction value is finally output.

[0090] In this embodiment, deep forest is considered to be used for regression prediction, and the extracted image feature vector f i,j Mapped to the subjective score corresponding to the image. Each layer of the deep forest model in this paper can be assembled by 4 random forests, and each random forest contains 100 decision trees. The image feature vector f i,j The 4-dimensional vector is obtained through the 4 random forests in the first layer, and then the original image feature vector f i,jIt is merged with the 4-dimensional vector as the input of the next layer. And so on. Each layer receives the image feature vector generated by the previous layer and outputs the image feature vector generated by it to the next layer until the final output is the quality evaluation prediction value of the image to be evaluated.

[0091] Of course, after nonlinear mapping, we can also consider verifying the prediction accuracy, prediction monotonicity and prediction consistency of this scheme. For example, the reliability of the prediction results can be measured by objective evaluation indicators such as Spearman rank correlation coefficient SROCC, Kendall rank correlation coefficient KROCC, Pearson linear correlation coefficient PLCC and root mean square error RMSE.

[0092] The effect of this solution is further explained below with reference to specific application examples.

[0093] like Figure 4 and Figure 5 The PLCC performance and SROCC performance comparison diagram of three regression methods, deep forest DF, random forest RF and support vector regression SVR, on 5 standard data sets, where the control example and the embodiment have the same processing process except for the different regression methods, and MLF-DF represents this solution. It can be seen from the figure that on the 5 standard data sets, when using deep forest DF for prediction, the SROCC and PLCC values ​​are higher than RF and SVR, indicating that using DF for image fusion can significantly improve the performance of the model and has strong prediction performance.

[0094] In order to predict the overall prediction performance of the present invention, the evaluation results of the present invention and 13 comparison methods (NR-IQA) on CSIQ, LIVE, TID2008 and TID2013 data sets are given as shown in Table 1. The best performing method is marked with a bold box, and the second best performing method is marked with an underline. It can be seen from Table 1 that the present invention has very good performance on all data sets, and it can also be seen from the weighted average results that the overall performance of the present invention is better than all comparison methods.

[0095]

[0096] Table 2 lists the SROCC values ​​of individual distortion types on the LIVE and CSIQ datasets, with the best performing method in bold and the second best performing method underlined. As can be seen from Table 2, the present invention achieves the best performance on all distortion types. In the CSIQ dataset, for the WN and PN distortion types, the results obtained by the present invention are obviously sensitive to additive white Gaussian noise and additive color Gaussian noise, and can effectively handle image quality assessment tasks corrupted by WN and PN.

[0097]

[0098] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute a method in any one of the embodiments in the specification.

[0099] The present specification also provides a computing device, including a memory and a processor, wherein executable codes are stored in the memory, and when the processor executes the executable codes, a method in any embodiment of the present specification is implemented.

[0100] The modules or units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. The above disclosure is only the preferred embodiment of the present invention, and of course it cannot be used to limit the scope of the rights of the present invention. Those skilled in the art can understand that all or part of the processes of the above embodiment are implemented, and the equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. An image quality assessment method based on multi-level feature distribution, characterized in that: include: Pre-train a 16-layer convolutional neural network VGG-16; wherein the first 13 layers of the convolutional neural network VGG-16 are convolutional layers, and the last 3 layers are fully connected layers; the 13 convolutional layers include N L shallow convolutional layers, N M The middle convolutional layers and N H N high-level convolutional layers are used to extract low-level features, middle-level features, and high-level features respectively; L +N M +N H =13; Inputting the image to be evaluated for image quality evaluation into the convolutional neural network VGG-16 for feature extraction; Determine a target convolution layer combination; wherein the target convolution layer combination includes a shallow convolution layer, a middle convolution layer and a high convolution layer; Fitting feature map coefficients of image features extracted by each convolution layer in the target convolution combination based on a generalized Gaussian distribution model to obtain an image feature vector; The extracted image feature vector is mapped to a subjective score corresponding to the image using a pre-built deep forest model to obtain a quality evaluation prediction value of the image to be evaluated.

2. The image quality assessment method based on multi-level feature distribution according to claim 1, characterized in that: The underlying features include texture and edge features of the image; The middle-level features include contour and shape features of the image; The high-level features include the content or spatial structure of the image.

3. The image quality assessment method based on multi-level feature distribution according to claim 1, characterized in that: Among the 13 convolutional layers, the first 4 are shallow convolutional layers, the middle 6 are medium convolutional layers, and the last 3 are high-level convolutional layers; The determining of a target convolutional layer combination includes: Select one from each of the shallow convolution layer, the middle convolution layer and the high convolution layer to form a primary convolution layer combination, and obtain A combination of primary convolutional layers; The image features corresponding to each combination of primary convolutional layers are evaluated using a preset image quality evaluation dataset; The primary convolutional layer with the best evaluation result among the primary convolutional layer combinations is determined as the target convolutional layer combination.

4. The image quality assessment method based on multi-level feature distribution according to claim 3 is characterized in that: The image quality assessment dataset includes at least one of LIVE, TID2008, TID2013 and CSIQ.

5. The image quality assessment method based on multi-level feature distribution according to claim 3 is characterized in that: The 4 shallow convolutional layers are {Conv1_1, Conv1_2, Conv2_1, Conv2_2}, the 6 middle convolutional layers are {Conv3_1, Conv3_2, Conv3_3, Conv4_1, Conv4_2, Conv4_3}, and the 3 high-level convolutional layers are {Conv5_1, Conv5_2, Conv5_3}; The target convolutional layer combination is {Conv2_1, Conv3_1, Conv5_2}.

6. The image quality assessment method based on multi-level feature distribution according to claim 1, characterized in that: The step of fitting the feature map coefficients of the image features extracted by each convolution layer in the target convolution combination based on the generalized Gaussian distribution model includes: fitting the feature map coefficients using the parameters of the generalized Gaussian distribution model as quality perception features; Among them, the probability density function of the generalized Gaussian distribution model is: in, In the formula, α represents the shape parameter of the distribution model, and β represents the scale parameter of the distribution model; The parameters of the distribution model are α i,j and β i,j To express it, we have: The image feature vectors generated by all feature maps are: f i,j ={f L,j ,f M,j ,f H,j } in, Among them, F i,j It represents the j-th feature map extracted by the i-th convolutional layer, where i=L,M,H, represents three shallow convolutional layers Conv2_1, Conv3_1, Conv5_2, a middle convolutional layer and a high-level convolutional layer, and j=128, 256, 512; GGD represents the generalized Gaussian distribution model.

7. The image quality assessment method based on multi-level feature distribution according to claim 1, characterized in that: Each layer of the deep forest model consists of 4 random forests, and each random forest contains D decision trees; D ≥ 50; The method of mapping the extracted image feature vector to a subjective score corresponding to the image by using a pre-built deep forest model includes: The image feature vector passes through 4 random forests in the first layer to obtain a 4-dimensional feature vector; The original image feature vector is combined with the 4-dimensional feature vector as the input of the next layer; And so on, until the quality evaluation prediction value is finally output.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.

9. A computing device, comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Underwater image visual quality evaluation method

    CN106780434A

  • No-reference image quality evaluation method and system, electronic equipment and storage medium

    CN116485741A