Deep and shallow layer full reference image quality evaluation method based on visual perception
By extracting the depth and shallow features of the image and enhancing the feature, combining the Wasserstein distance and self-attention mechanism, a method of depth and shallow full reference image quality evaluation based on visual perception is designed, solving the problem of failure to effectively combine low- and high-level visual features in the prior art, and achieving an image quality evaluation that is more in line with human eye perception.
Patent Information
- Application Number
- CN202411980559.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing full reference image quality evaluation method fails to effectively combine the low-level visual characteristics and high-level visual characteristics of the image, and does not pay enough attention to the important areas in the image, resulting in the evaluation indicators not meeting the visual characteristics of the human eye.
A method of evaluating the quality of the deep and shallow-layer full-reference image based on visual perception was designed. By extracting five-layer feature maps, feature enhancement and quality evaluation of the shallow and deep feature maps were respectively performed. The specific steps include: obtaining paired reference images and distorted images, extracting five-layer feature maps, performing feature enhancement and quality evaluation, and finally obtaining image quality scores through deep learning model integration.
A more comprehensive and accurate evaluation of image quality is achieved, which can better reflect the human eye's perception of image quality, and provides an evaluation index that is more in line with the visual characteristics of the human eye.
Smart Images

Figure CN119919366A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for evaluating the quality of deep and shallow layer full-reference images based on visual perception. Background Art
[0002] Image quality assessment is a core issue in the field of image processing and computer vision. It plays an important role in many fields such as image / video coding, super-resolution reconstruction, and image / video visual quality improvement.
[0003] Currently, the disclosed full-reference image quality assessment method estimates the quality score of the distorted image by comparing the visual differences between the two images. This type of algorithm can not only be used to evaluate the effect of image compression coding, but also can be used as a test benchmark for image processing algorithms, or as a target for image processing task optimization. In the prior art, a disclosed full-reference image quality objective evaluation method compares the visual differences between the two images and performs deep fusion processing based on multiple visual features. It does not comprehensively consider the low-level and high-level visual features of the image and focus on the important areas in the image.
[0004] Therefore, regarding the important areas of the image, how to effectively extract low-level visual features and high-level visual features and design targeted processing methods for the visual features of different layers, and construct an evaluation index that is more in line with the visual characteristics of the human eye, is a technical problem that technical personnel in this field urgently need to solve. Summary of the invention
[0005] In view of this, the present invention designs a deep and shallow full-reference image quality evaluation method based on visual perception, aiming to solve the problem of how to effectively extract low-level visual features and high-level visual features and design targeted processing methods for visual features of different layers, and finally construct a method that is more in line with the evaluation indicators of human visual characteristics.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A method for evaluating the quality of a deep and shallow layer full-reference image based on visual perception is disclosed, comprising the following steps:
[0008] S1, obtaining a pair of reference images and distorted images of uniform size;
[0009] S2. Extract five layers of feature maps based on the reference image and the distorted image respectively, where the first three layers of feature maps are defined as shallow feature maps, and the last two layers of feature maps are defined as deep feature maps;
[0010] S3, perform feature enhancement on the first three shallow feature maps respectively;
[0011] S4, based on the first three shallow feature maps after feature enhancement, the Wasserstein distance is used to measure the difference between the shallow features of the reference image and the distorted image at the same layer and the quality scores of the three shallow features are calculated;
[0012] S5. Perform self-attention enhancement on the deep feature maps of the last two layers and calculate the image perception difference between the deep feature maps of the same layer of the reference image and the distorted image, which is used to measure the degree of degradation of the deep feature maps of the same layer; process the image perception difference between the two layers of deep features in sequence through the attention aggregation module, the splicing module and the fusion module, and use the weighted score and pixel prediction to obtain the quality score of the fused deep features;
[0013] S6. The three shallow feature quality scores and one deep feature quality score are integrated through the four learnable parameters in the deep learning model to obtain the final deep and shallow full reference image quality score.
[0014] Furthermore, in S1, the image sizes of the obtained reference image and the distorted image pair are unified into a pixel size of 224×224, and the image size unification method is to unify the image size using image processing software or library.
[0015] Furthermore, S2 includes:
[0016] S21, input the reference image and the distorted image of uniform size into the visual geometry group network model VGG-16 respectively;
[0017] S22, Visual Geometry Group Network model VGG-16 extracts the first three shallow feature maps of the reference image and the distorted image based on the texture, color and shape of the image;
[0018] S23, the visual geometry group network model VGG-16 extracts the last two layers of deep feature maps of the reference image and the distorted image based on the high-level semantic information of the image.
[0019] Furthermore, in S3, SENet is used to perform feature enhancement on the first three shallow feature maps of the reference image and the distorted image, respectively, and the process is as follows:
[0020] S31, SENet compresses the input shallow feature map along the spatial dimension to obtain the corresponding compression vector, which is expressed as:
[0021]
[0022] Among them, c represents a specific feature channel, u c is the shallow feature map input by feature channel c, F sq Indicates compression operation; z cThe compressed vector of the shallow feature map input by feature channel c has a dimension of 1*1*C, where C is a real number with a global receptive field, which ensures that the output dimension matches the number of input feature channels; H and W are the length and width of the input shallow feature map, respectively;
[0023] S32. The weights assigned to the output channels of the compressed shallow feature maps are identified through the Excitation operation, and the expression is:
[0024] s=F ex (z, W)=σ(g(z, W))=σ(W2δ(W1z))
[0025] Among them, s is the weight assigned to the output channel of the compressed shallow feature map, F ex represents the Excitation operation, z is the compression vector of the shallow feature map, W is the general term for the weight matrix assigned to the output channel of the shallow feature map that identifies the compression, where W1 and W2 are the weight matrices assigned by the first two full connections, and the weight matrices are all three-dimensional matrices; σ represents the sigmoid function. Before the sigmoid function, W1 and z are matrix multiplied to achieve a full connection; after being processed by the Relu layer, they are multiplied with W2 to complete the second full connection;
[0026] The two full connections are used to reduce the number of channels and the amount of calculation. The dimension of W1 is 1*1*C / r, r is a scaling parameter, and the dimension of W1 multiplied by z is 1*1*C / r. After the two are multiplied, they are processed by the Relu layer and their dimensions remain unchanged. Then they are multiplied with W2 to complete the second full connection. The dimension of W2 is 1*1*C*C / r, and the output dimension after the second full connection is 1*1*C.
[0027] S33, the neural network multiplies the input shallow feature map by the assigned weight of the feature channel to obtain a shallow feature map after feature enhancement;
[0028] The expression of the shallow feature map after feature enhancement is:
[0029]
[0030] in, is the shallow feature map after feature enhancement, F scale is the feature enhancement operation, u c ,s c They are the shallow feature map input through channel c and the assigned weights of feature channel c respectively.
[0031] Furthermore, in S4, the calculation process of the quality score of each shallow feature is:
[0032] S41, converting the shallow feature map of the same layer of the reference image and the distorted image into a one-dimensional feature vector and normalizing it to obtain a normalized probability distribution feature map;
[0033] S42, calculating the WSD value between the distorted image and the reference image using the Wasserstein distance based on the two normalized probability distribution feature maps;
[0034] S43: Map the calculated WSD value to an image quality evaluation score to evaluate the degradation degree of the distorted image.
[0035] Furthermore, in S41, the normalized probability distribution characteristic graph expression is:
[0036]
[0037] Among them, P is the normalized probability distribution feature map, f is the shallow feature map after feature enhancement, C is the feature enhancement channel, H and W are the length and width of the shallow feature map after feature enhancement, and f i Represents a single feature element in the shallow feature map after feature enhancement, and the value of i ranges from 1 to CxHxW;
[0038] In S42, the Wasserstein distance is used to calculate the WSD value between the shallow feature image pair of the same layer of the distorted image and the reference image, and its expression is:
[0039]
[0040] Among them, X and Y represent the normalized probability distribution feature maps of the reference image and the distorted image, respectively. They represent the cumulative probability distribution functions of X and Y respectively, which calculate the corresponding cumulative probability distribution functions based on the normalized probability distribution P of the reference image and the distorted image; WSD is the image degradation distance between the two normalized probability distribution feature maps of X and Y.
[0041] Furthermore, S5 includes:
[0042] S51, performing self-attention enhancement on the deep feature maps of the reference image and the two layers after the distorted image to obtain an enhanced deep feature map;
[0043] S52, calculating the perceived difference in image degradation between the reference image and the distorted image according to the enhanced deep feature map;
[0044] S53, inputting the perception difference of the two layers into the attention aggregation module for batch normalization processing to obtain an aggregated deep feature map;
[0045] S54, outputting the aggregated deep feature map to a splicing module for preliminary processing, discarding layer processing and connecting inputs to obtain a final spliced deep feature map;
[0046] S55, inputting the final spliced deep feature map into a fusion module for fusion operation to obtain a fused deep feature map;
[0047] S56. Based on the fused deep feature map, the final deep feature quality score is obtained using weighted scores and pixel predictions.
[0048] Further, in S51, the self-attention enhancement includes a normalization process for capturing the internal correlation information of the original deep feature map itself and performing feature enhancement on the feature map;
[0049] The enhanced deep feature map expression is:
[0050]
[0051] Among them, Q, K, and V represent the query vector, key vector, and value matrix vector in the self-attention mechanism respectively, and Softmax is a normalization function; dk is the dimension of the key vector, which is used to scale the dot product to avoid the problem of gradient vanishing when the dimension is large.
[0052] Furthermore, in S51, the deep feature map expression after self-attention enhancement of the i-th layer is:
[0053] f i =Attn(F i W q , F i W k , F i W v )+F i
[0054] Among them, f i is the deep feature map after self-attention enhancement in the i-th layer, F i is the deep feature map of the i-th layer before self-attention enhancement, and the value of i is 4 or 5; W q , W k and W v They represent the weights of query vector, key vector and value matrix vector in self-attention enhancement respectively;
[0055] In S52, the perceptual difference between the reference image and the distorted image feature after self-attention enhancement is calculated as:
[0056]
[0057] in, is the perceptual difference between the i-th layer deep feature map after feature enhancement of the reference image and the distorted image, They are the deep feature maps of the reference image and the distorted image after self-attention enhancement;
[0058] In S53, the expression of the aggregated deep feature map is:
[0059]
[0060] In S54, the deep feature map after preliminary processing is expressed as:
[0061]
[0062] Among them, BatchNorm represents batch normalization processing. is the deep feature map after aggregation, It is the deep feature map after preliminary processing;
[0063] The expression of the i-th layer depth feature map after the discard layer is:
[0064]
[0065] Among them, Dropout represents the discard layer operation, is the i-th layer deep feature map that is initially processed; is the i-th layer deep feature map after aggregation, is the depth feature map of the i-th layer after the discard layer processing;
[0066] The final concatenated depth feature map expression after input connection is:
[0067]
[0068] in, is the final concatenated depth feature map after connecting the inputs, They are the 4th and 5th layer depth feature maps after discarding the layer respectively;
[0069] In S55, the expression of the fused deep feature map is:
[0070]
[0071] Among them, F fusion is the fused deep feature map. The fusion process includes the final spliced deep feature map Perform 3X3 convolution, Relu activation function, two consecutive 3X3 convolutions and Relu function output;
[0072] In S56, the final deep feature quality score expression is:
[0073] d 45 =Conv(F fusion )⊙Sigmoid(Conv(F fusion ))
[0074] Among them, d 45 is the final deep feature quality score, F fusion is the fused deep feature map, which is input into the dual-branch prediction module, one branch is used for pixel prediction and the other branch is used for weighted score prediction; ⊙ represents the dot product.
[0075] Furthermore, in S6, the final deep and shallow full reference image quality score expression after the deep learning model is integrated is:
[0076] D=α1d1+α2d2+α3d3+α 45 d 45
[0077] Among them, α1, α2, α3, α 45 They are four learnable parameters in the deep learning model, indicating the criticality of different layers to the final quality evaluation score. d1, d2, and d3 are the quality evaluation scores of the shallow feature maps of the first three layers, respectively. 45 is the final deep feature quality score.
[0078] It can be seen from the above technical solution that compared with the prior art, the present invention has the following beneficial effects:
[0079] (1) Extract feature maps through VGG-16 and perform logical division to obtain shallow feature maps and deep feature maps, which provide a basis for subsequent targeted feature processing, allowing us to perform quality evaluation operations based on feature maps from the perspectives of shallow and deep layers;
[0080] (2) SENet is used to enhance shallow features. Its autonomous learning characteristics can accurately identify the importance of each convolutional feature channel and assign reasonable weights to different convolutional channels, allowing the neural network to focus on the feature channels related to the task and avoid the influence of irrelevant channels. This helps to highlight the feature information that is important for quality evaluation and lays a good foundation for the subsequent accurate calculation of shallow feature quality scores.
[0081] (3) Use Wasserstein distance to perceive the degradation distance of shallow feature maps of the same layer, so as to prepare for the subsequent quality score. This process considers the image quality reflected by shallow features from the perspective of HVS. For deep features, the self-attention mechanism is first used to enhance them so that they can better capture their own internal correlation information. Then, the perceptual difference between the two layers of deep feature maps after feature enhancement is calculated to clarify the degree of degradation of the distorted image. A series of operations such as attention aggregation module, splicing module, and fusion module are used in turn to fully integrate features at different levels. Finally, the quality score of deep features is obtained with the help of weighted scores and pixel prediction. Such a comprehensive processing method can more deeply, comprehensively and accurately evaluate the quality level of deep features and dig out the key information about image quality contained in deep features.
[0082] (4) The shallow feature quality scores and deep feature quality scores are integrated through learnable parameters. Under specific calculation methods and parameter adjustments, the mutual fusion and synergy of the two different levels of feature quality scores are achieved, thereby obtaining a final score that can comprehensively reflect the overall situation. This is used to comprehensively evaluate the performance level of relevant features at the overall level, making the final quality evaluation result more comprehensive and accurate, and closer to the actual quality status of the image.
[0083] (5) The deep and shallow layer full reference image quality assessment method based on visual perception of the present invention can achieve quality score assessment and prediction of distorted images that are closer to human eye perception through the steps of layer-by-layer feature extraction of deep and shallow layer features, feature enhancement, and final full reference image quality score calculation and integration, providing an effective technical means for accurately evaluating image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0085] Figure 1 A flowchart of a method for evaluating the quality of deep and shallow layer full reference images based on visual perception provided by the present invention;
[0086] Figure 2 The present invention provides a framework diagram of a deep and shallow layer full-reference image quality assessment method based on visual perception; DETAILED DESCRIPTION
[0087] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 2, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0088] The embodiment of the present invention provides a method for evaluating the quality of a deep and shallow layer full reference image based on visual perception, such as Figure 1 As shown, the following steps are included:
[0089] S1, obtaining a pair of reference images and distorted images of uniform size;
[0090] S2. Extract five layers of feature maps based on the reference image and the distorted image respectively, where the first three layers of feature maps are defined as shallow feature maps, and the last two layers of feature maps are defined as deep feature maps;
[0091] S3, perform feature enhancement on the first three shallow feature maps respectively;
[0092] S4, based on the first three shallow feature maps after feature enhancement, the Wasserstein distance is used to measure the difference between the shallow features of the reference image and the distorted image at the same layer and the quality scores of the three shallow features are calculated;
[0093] S5. Perform self-attention enhancement on the deep feature maps of the last two layers and calculate the image perception difference between the deep feature maps of the same layer of the reference image and the distorted image, which is used to measure the degree of degradation of the deep feature maps of the same layer; process the image perception difference between the two layers of deep features in sequence through the attention aggregation module, the splicing module and the fusion module, and use the weighted score and pixel prediction to obtain the quality score of the fused deep features;
[0094] S6. The three shallow feature quality scores and one deep feature quality score are integrated through the four learnable parameters in the deep learning model to obtain the final deep and shallow full reference image quality score.
[0095] In this embodiment, in S1, the image sizes of the acquired reference image and the distorted image are unified into a pixel size of 224×224, and the image sizes are unified by using image processing software or a library to unify the image sizes.
[0096] In this embodiment, S2 includes:
[0097] S21, input the reference image and the distorted image of uniform size into the visual geometry group network model VGG-16 respectively;
[0098] S22, Visual Geometry Group Network model VGG-16 extracts the first three shallow feature maps of the reference image and the distorted image based on the texture, color and shape of the image;
[0099] S23, the visual geometry group network model VGG-16 extracts the last two layers of deep feature maps of the reference image and the distorted image based on the high-level semantic information of the image.
[0100] In this embodiment, in S3, SENet is used to perform feature enhancement on the first three shallow feature maps of the reference image and the distorted image, respectively, and the process is as follows:
[0101] S31, SENet compresses the input shallow feature map along the spatial dimension to obtain the corresponding compression vector, which is expressed as:
[0102]
[0103] Among them, c represents a specific feature channel, u c is the shallow feature map input by feature channel c, F sq Indicates compression operation; z c The compressed vector of the shallow feature map input by feature channel c has a dimension of 1*1*C, where C is a real number with a global receptive field, which ensures that the output dimension matches the number of input feature channels; H and W are the length and width of the input shallow feature map, respectively;
[0104] S32. The weights assigned to the output channels of the compressed shallow feature maps are identified through the Excitation operation, and the expression is:
[0105] s=F ex (z, W)=σ(g(z, W))=σ(W2δ(W1z))
[0106] Among them, s is the weight assigned to the output channel of the compressed shallow feature map, F ex represents the Excitation operation, z is the compression vector of the shallow feature map, W is the general term for the weight matrix assigned to the output channel of the shallow feature map that identifies the compression, where W1 and W2 are the weight matrices assigned by the first two full connections, and the weight matrices are all three-dimensional matrices; σ represents the sigmoid function. Before the sigmoid function, W1 and z are matrix multiplied to achieve a full connection; after being processed by the Relu layer, they are multiplied with W2 to complete the second full connection;
[0107] The two full connections are used to reduce the number of channels and the amount of calculation. The dimension of W1 is 1*1*C / r, r is a scaling parameter, and the dimension of W1 multiplied by z is 1*1*C / r. After the two are multiplied, they are processed by the Relu layer and their dimensions remain unchanged. Then they are multiplied with W2 to complete the second full connection. The dimension of W2 is 1*1*C*C / r, and the output dimension after the second full connection is 1*1*C.
[0108] S33, the neural network multiplies the input shallow feature map by the assigned weight of the feature channel to obtain a shallow feature map after feature enhancement;
[0109] The expression of the shallow feature map after feature enhancement is:
[0110]
[0111] in, is the shallow feature map after feature enhancement, F scale is the feature enhancement operation, u c ,s c They are the shallow feature map input through channel c and the assigned weights of feature channel c respectively.
[0112] In this embodiment S4, the calculation process of each shallow feature quality score is:
[0113] S41, converting the shallow feature map of the same layer of the reference image and the distorted image into a one-dimensional feature vector and normalizing it to obtain a normalized probability distribution feature map;
[0114] S42, calculating the WSD value between the distorted image and the reference image using the Wasserstein distance based on the two normalized probability distribution feature maps;
[0115] S43: Map the calculated WSD value to an image quality evaluation score to evaluate the degradation degree of the distorted image.
[0116] 6. The method for evaluating the quality of deep and shallow layers based on visual perception according to claim 5, characterized in that, in S41, the normalized probability distribution feature map expression is:
[0117]
[0118] Among them, P is the normalized probability distribution feature map, f is the shallow feature map after feature enhancement, C is the feature enhancement channel, H and W are the length and width of the shallow feature map after feature enhancement, and f i Represents a single feature element in the shallow feature map after feature enhancement, and the value of i ranges from 1 to CxHxW;
[0119] In S42, the Wasserstein distance is used to calculate the WSD value between the shallow feature image pair of the same layer of the distorted image and the reference image, and its expression is:
[0120]
[0121] Among them, X and Y represent the normalized probability distribution feature maps of the reference image and the distorted image, respectively. They represent the cumulative probability distribution functions of X and Y respectively, which calculate the corresponding cumulative probability distribution functions based on the normalized probability distribution P of the reference image and the distorted image; WSD is the image degradation distance between the two normalized probability distribution feature maps of X and Y.
[0122] In this embodiment, S5 includes:
[0123] S51, performing self-attention enhancement on the deep feature maps of the reference image and the two layers after the distorted image to obtain an enhanced deep feature map;
[0124] S52, calculating the perceived difference in image degradation between the reference image and the distorted image according to the enhanced deep feature map;
[0125] S53, inputting the perception difference of the two layers into the attention aggregation module for batch normalization processing to obtain an aggregated deep feature map;
[0126] S54, outputting the aggregated deep feature map to a splicing module for preliminary processing, discarding layer processing and connecting inputs to obtain a final spliced deep feature map;
[0127] S55, inputting the final spliced deep feature map into a fusion module for fusion operation to obtain a fused deep feature map;
[0128] S56. Based on the fused deep feature map, the final deep feature quality score is obtained using weighted scores and pixel predictions.
[0129] In this embodiment, in S51, the self-attention enhancement includes a normalization process for capturing its own internal correlation information and performing feature enhancement on the original input deep feature map;
[0130] The enhanced deep feature map expression is:
[0131]
[0132] Among them, Q, K, and V represent the query vector, key vector, and value matrix vector in the self-attention mechanism, respectively, and Softmax is a normalization function; d k is the dimension of the key vector, used to scale the dot product to avoid the vanishing gradient problem when the dimension is large.
[0133] 9. The method for evaluating the quality of deep and shallow layers based on visual perception according to claim 7, wherein in S51, the deep feature map after self-attention enhancement of the i-th layer is expressed as:
[0134] f i =Attn(F i W q , F i W k , F i W v )+F i
[0135] Among them, f i is the deep feature map after self-attention enhancement in the i-th layer, F i is the deep feature map of the i-th layer before self-attention enhancement, and the value of i is 4 or 5; W q , W k and W v They represent the weights of query vector, key vector and value matrix vector in self-attention enhancement respectively;
[0136] In S52, the perceptual difference between the reference image and the distorted image feature after self-attention enhancement is calculated as:
[0137]
[0138] in, is the perceptual difference between the i-th layer deep feature map after feature enhancement of the reference image and the distorted image, They are the deep feature maps of the reference image and the distorted image after self-attention enhancement;
[0139] In S53, the expression of the aggregated deep feature map is:
[0140]
[0141] In S54, the deep feature map after preliminary processing is expressed as:
[0142]
[0143] Among them, BatchNorm represents batch normalization processing. is the deep feature map after aggregation, It is the deep feature map after preliminary processing;
[0144] The expression of the i-th layer depth feature map after the discard layer is:
[0145]
[0146] Among them, Dropout represents the discard layer operation, is the i-th layer deep feature map that is initially processed; is the i-th layer deep feature map after aggregation, is the depth feature map of the i-th layer after the discard layer processing;
[0147] The final concatenated depth feature map expression after input connection is:
[0148]
[0149] in, is the final concatenated depth feature map after connecting the inputs, They are the 4th and 5th layer depth feature maps after discarding the layer respectively;
[0150] In S55, the expression of the fused deep feature map is:
[0151]
[0152] Among them, F fusion is the fused deep feature map. The fusion process includes the final spliced deep feature map Perform 3X3 convolution, Relu activation function, two consecutive 3X3 convolutions and Relu function output;
[0153] In S56, the final deep feature quality score expression is:
[0154] d 45 =Conv(F fusion )⊙Sigmoid(Conv(F fusion ))
[0155] Among them, d 45 is the final deep feature quality score, F fusion is the fused deep feature map, which is input into the dual-branch prediction module, one branch is used for pixel prediction and the other branch is used for weighted score prediction; ⊙ represents the dot product.
[0156] In this embodiment S6, the final deep and shallow layer full reference image quality score expression after the deep learning model is integrated is:
[0157] D=α1d1+α2d2+α3d3+α 45 d 45
[0158] Among them, α1, α2, α3, α 45They are four learnable parameters in the deep learning model, indicating the criticality of different layers to the final quality evaluation score. d1, d2, and d3 are the quality evaluation scores of the shallow feature maps of the first three layers, respectively. 45 is the final deep feature quality score.
[0159] In this embodiment, the four learnable parameters of the deep learning model are the criticality of the shallow feature quality scores of the first three layers of feature maps and the final deep feature quality score in the full reference image quality evaluation.
[0160] In this embodiment, the four learnable parameters of the deep learning model can be set in advance as needed.
[0161] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0162] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A deep and shallow full-reference image quality assessment method based on visual perception, characterized in that: The steps include: S1, obtaining a pair of reference images and distorted images of uniform size; S2, extracting five layers of feature maps based on the reference image and the distorted image respectively, wherein the first three layers of feature maps are shallow feature maps, and the last two layers of feature maps are deep feature maps; S3, perform feature enhancement on the first three shallow feature maps respectively; S4, based on the first three shallow feature maps after feature enhancement, the Wasserstein distance is used to measure the difference between the shallow features of the reference image and the distorted image at the same layer and the quality scores of the three shallow features are calculated; S5. Perform self-attention enhancement on the deep feature maps of the last two layers and calculate the image perception difference between the deep feature maps of the same layer of the reference image and the distorted image, which is used to measure the degree of degradation of the deep feature maps of the same layer; process the image perception difference between the two layers of deep features in sequence through the attention aggregation module, the splicing module and the fusion module, and use the weighted score and pixel prediction to obtain the quality score of the fused deep features; S6. The deep learning model integrates three shallow feature quality scores and one deep feature quality score through four learnable parameters to obtain the final deep and shallow full-reference image quality score.
2. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 1, characterized in that: In S1, the image sizes of the obtained reference image and the distorted image pair are unified into a pixel size of 224×224, and the image size unification method is to unify the image size using image processing software or library.
3. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 1, characterized in that: S2 includes: S21, input the reference image and the distorted image of uniform size into the visual geometry group network model VGG-16 respectively; S22, Visual Geometry Group Network model VGG-16 extracts the first three shallow feature maps of the reference image and the distorted image based on the texture, color and shape of the image; S23, the visual geometry group network model VGG-16 extracts the last two layers of deep feature maps of the reference image and the distorted image based on the high-level semantic information of the image.
4. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 1, characterized in that: In S3, SENet is used to enhance the features of the first three shallow feature maps of the reference image and the distorted image, respectively. The process is as follows: S31, SENet compresses the input shallow feature map along the spatial dimension to obtain the corresponding compression vector, which is expressed as: Among them, c represents a specific feature channel, u c is the shallow feature map input by feature channel c, F sq Indicates compression operation; z c The compressed vector of the shallow feature map input by feature channel c has a dimension of 1*1*C, where C is a real number with a global receptive field, which ensures that the output dimension matches the number of input feature channels; H and W are the length and width of the input shallow feature map, respectively; S32. The weights assigned to the output channels of the compressed shallow feature maps are identified through the Excitation operation, and the expression is: s=F ex (z, W)=σ(g(z, W))=σ(W2δ(W1z)) Among them, s is the weight assigned to the output channel of the compressed shallow feature map, F ex Represents the Excitation operation, z is the compression vector of the shallow feature map, W is the general name for the weight matrix assigned to the output channel of the shallow feature map that identifies the compression, where W1 and W2 are the weight matrices assigned by the first two full connections, and the weight matrices are all three-dimensional matrices; σ represents the sigmoid function. Before the sigmoid function, W1 and z are matrix multiplied to achieve a full connection; after being processed by the Relu layer, they are multiplied with W2 to complete the second full connection, which is used to reduce the number of channels and the amount of calculation; The dimension of W1 is 1*1*C / r, r is a scaling parameter, and the dimension of W1 multiplied by z is 1*1*C / r; after the two are multiplied, the dimension remains unchanged after being processed by the Relu layer; after multiplying with W2, the dimension of W2 after the second full connection is 1*1*C*C / r, and the output dimension after the second full connection is 1*1*C; S33, the neural network multiplies the input shallow feature map by the assigned weight of the feature channel to obtain a shallow feature map after feature enhancement; The expression of the shallow feature map after feature enhancement is: in, is the shallow feature map after feature enhancement, F scale is the feature enhancement operation, u c ,s c They are the shallow feature map input through channel c and the assigned weights of feature channel c respectively.
5. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 1, characterized in that: In S4, the calculation process of the quality score of each shallow feature is: S41, converting the shallow feature map of the same layer of the reference image and the distorted image into a one-dimensional feature vector and normalizing it to obtain a normalized probability distribution feature map; S42, calculating the WSD value between the distorted image and the reference image using the Wasserstein distance based on the two normalized probability distribution feature maps; S43: Map the calculated WSD value to an image quality evaluation score to evaluate the degradation degree of the distorted image.
6. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 5, characterized in that: In S41, the normalized probability distribution characteristic graph expression is: Among them, P is the normalized probability distribution feature map, f is the shallow feature map after feature enhancement, C is the feature enhancement channel, H and W are the length and width of the shallow feature map after feature enhancement, and f i Represents a single feature element in the shallow feature map after feature enhancement, and the value range of i is 1 to CxHxW; In S42, the Wasserstein distance is used to calculate the WSD value between the shallow feature image pair of the same layer of the distorted image and the reference image, and its expression is: Among them, X and Y represent the normalized probability distribution feature maps of the reference image and the distorted image, respectively. They represent the cumulative probability distribution functions of X and Y respectively, which calculate the corresponding cumulative probability distribution functions based on the normalized probability distribution P of the reference image and the distorted image; WSD is the image degradation distance between the two normalized probability distribution feature maps of X and Y.
7. The method for evaluating the quality of a deep and shallow layer full-reference image based on visual perception according to claim 1, wherein S5 include: S51, performing self-attention enhancement on the deep feature maps of the reference image and the two layers after the distorted image to obtain an enhanced deep feature map; S52, calculating the perceived difference in image degradation between the reference image and the distorted image according to the enhanced deep feature map; S53, inputting the perception difference of the two layers into the attention aggregation module for batch normalization processing to obtain an aggregated deep feature map; S54, outputting the aggregated deep feature map to a splicing module for preliminary processing, discarding layer processing and connecting inputs to obtain a final spliced deep feature map; S55, inputting the final spliced deep feature map into a fusion module for fusion operation to obtain a fused deep feature map; S56. Based on the fused deep feature map, the final deep feature quality score is obtained using weighted scores and pixel predictions.
8. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 7, characterized in that: In S51, the self-attention enhancement includes a normalization process for capturing the internal correlation information of the original deep feature map itself and performing feature enhancement on the feature map; The enhanced deep feature map expression is: Among them, Q, K, and V represent the query vector, key vector, and value matrix vector in the self-attention mechanism, respectively, and Softmax is a normalization function; d k is the dimension of the key vector.
9. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 7, characterized in that: In S51, the deep feature map expression after self-attention enhancement of the i-th layer is: f i =Attn(F i W q ,F i W k ,F i W v )+F i Among them, f i is the deep feature map after self-attention enhancement in the i-th layer, F i is the deep feature map of the i-th layer before self-attention enhancement, and the value of i is 4 or 5; W q , W k and W v They represent the weights of query vector, key vector and value matrix vector in self-attention enhancement respectively; In S52, the perceptual difference between the reference image and the distorted image feature after self-attention enhancement is calculated as: in, is the perceptual difference between the i-th layer deep feature map after feature enhancement of the reference image and the distorted image, They are the deep feature maps of the reference image and the distorted image after self-attention enhancement; In S53, the expression of the aggregated deep feature map is: In S54, the deep feature map after preliminary processing is expressed as: Among them, BatchNorm represents batch normalization processing. is the deep feature map after aggregation, It is the deep feature map after preliminary processing; The expression of the i-th layer depth feature map after the discard layer is: Among them, Dropout represents the discard layer operation, is the i-th layer deep feature map that is initially processed; is the i-th layer deep feature map after aggregation, is the depth feature map of the i-th layer after the discard layer processing; The final concatenated depth feature map expression after input connection is: in, is the final concatenated depth feature map after connecting the inputs, They are the 4th and 5th layer depth feature maps after the discard layer processing; In S55, the expression of the fused deep feature map is: Among them, F fusion is the fused deep feature map. The fusion process includes the final spliced deep feature map Perform 3X3 convolution, Relu activation function, two consecutive 3X3 convolutions and Relu function output; In S56, the final deep feature quality score expression is: d 45 =Conv(F fusion )⊙Sigmoid(Conv(F fusion )) Among them, d 45 is the final deep feature quality score, F fusion is the fused deep feature map, which is input into the dual-branch prediction module, one branch is used for pixel prediction and the other branch is used for weighted score prediction; ⊙ represents the dot product.
10. The method for evaluating the quality of deep and shallow layers of full-reference images based on visual perception according to claim 1, characterized in that: In S6, the final deep and shallow full reference image quality score expression after the deep learning model is integrated is: D=α1d1+α2d2+α3d3+α 45 d 45 Among them, α1, α2, α3, α 45 They are four learnable parameters in the deep learning model, indicating the criticality of different layers to the final quality evaluation score. d1, d2, and d3 are the shallow feature quality scores of the first three layers of feature maps, respectively. 45 is the final deep feature quality score.
Citation Information
Patent Citations
Non-uniform distortion panoramic image blind quality evaluation method and system
CN117237279A
Full-reference frame insertion video quality evaluation method and system fusing preamble features and postamble features
CN117478974A