Content-Guided and Self-Attention Based Blind Reference Image Quality Assessment Method

By introducing content guidance and self-attention mechanisms in image quality evaluation, the content understanding network and self-attention network are constructed, and the problem of distortion perception weight adjustment in different areas in real distortion blind reference image quality evaluation is solved, and a more accurate image quality evaluation is achieved.

CN115222996BActive Publication Date: 2025-07-22TIANJIN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211011200.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-07-22
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

The prior art fails to fully utilize image content features to adjust the distortion perception weights in different regions in real distortion blind reference image quality evaluation, resulting in limited perception effects.

Method used

The content guidance and self-attention mechanism is adopted to construct a content understanding network and a content guidance self-attention network, extract image content features through a linear mapping module, and adjust distortion feature weights using a self-attention encoder to construct a dual branch quality prediction network for fusion and mapping.

Benefits of technology

It improves the network's ability to perceive true distortion and improves the accuracy and effectiveness of blind reference image quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222996B_ABST
    Figure CN115222996B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of multimedia processing. To propose a blind reference image quality assessment method for real distortion based on content guidance and self-attention mechanism to improve the network's perception ability of real distortion, the technical solution adopted by the present invention is a blind reference image quality assessment method for content guidance and self-attention real distortion. A linear mapping module is introduced on the basis of the EfficientNet-B0 network to construct a content understanding network; a content-guided self-attention network is constructed to adjust the distortion weights of different regions in the image; a dual-branch quality prediction network is constructed, and the distortion features obtained by the content-guided self-attention network are fused and mapped with the content features to obtain a quality score. The present invention is mainly applied to multimedia processing scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimedia processing, and specifically designs a blind reference image quality evaluation method based on content guidance and self-attention mechanism. Background Art

[0002] Digital images play an important role in people's daily lives. However, during the process of transmitting image data to users, distortion will be introduced into the images, resulting in a decline in the perceived image quality. With the explosive growth of digital images in recent years, image quality assessment, as a technology that can automatically predict and perceive image quality, has attracted extensive attention from researchers. Image quality can generally be divided into full-reference image quality evaluation, reduced-reference image quality evaluation, and blind-reference image quality evaluation. Due to the lack of reference images in many practical situations, the field of blind-reference image quality evaluation has received great attention.

[0003] Zhang et al. [Zhang W, Ma K, Yan J, et al. Blind image quality assessment using a deep bilinear convolutional neural network [J]. IEEE Transactions on Circuits and Systems for Video Technology, 2018, 30(1): 36-47.] proposed a blind reference image quality assessment method based on a bilinear convolutional neural network. This method uses the pre-trained VGG-16 to extract synthetic and real distortions in the image and performs bilinear fusion on the two distortion features. The fused features are finally used to regress the quality score. Liu et al. [Liu X, Van De Weijer J, Bagdanov A D. Rankiqa: Learning from rankings for no-reference image quality assessment [C] / / Proceedings of the IEEE international conference on computer vision. 2017: 1040-1049.] proposed a ranking-based blind reference image quality assessment method. This method includes dataset augmentation and quality prediction tasks, and it uses a siamese neural network to well solve the problem of the limited size of the image quality assessment dataset. Pan et al. [Pan Z, Yuan F, Lei J, et al. VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment [J]. IEEE Transactions on Image Processing, 2022, 31: 1613-1627.] proposed a blind reference image quality assessment method based on the free energy theory. This method includes a visual compensation network to simulate the repair process of distorted images in the brain and uses multi-level features in the visual compensation network to regress the quality score. Shi Ping et al. [Shi Ping, Pan Da, Ying Zefeng, Hou Ming, Zhong Dixiu, Han Mingliang. An objective assessment method for no-reference image quality based on a multi-scale generative adversarial network [P]. Beijing: CN108090902B, 2021-12-31.] proposed to generate a distortion map similar to the quality of the distorted image through a multi-scale adversarial generative network, and regress the image score by passing the generated distortion map through a convolutional neural network.Although the above methods have achieved certain success in synthetic distortion, compared with the problem of synthetic distortion, there are two major problems with real distortion: 1) The degree of distortion varies in different regions; 2) The human eye's perception of distortion is affected by content features. These two problems limit the perceptual evaluation effect of real distortion. Therefore, the above methods fail to fully utilize the image content features to adjust the distortion perception weights of different regions, thus limiting practical applications. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention aims to propose a blind reference image quality evaluation method for real distortion based on content guidance and self-attention mechanism to improve the network's perception ability of real distortion. To this end, the technical solution adopted by the present invention is a blind reference image quality evaluation method for content-guided and self-attention real distortion, which introduces a linear mapping module on the basis of the EfficientNet-B0 network to construct a content understanding network; constructs a content-guided self-attention network to adjust the distortion weights of different regions in the image; constructs a dual-branch quality prediction network, and fuses and maps the distortion features and content features obtained by the content-guided self-attention network to obtain a quality score.

[0005] The specific steps are as follows:

[0006] (1) Introduce a linear mapping module on the basis of the EfficientNet-B0 network to construct a content understanding network;

[0007] (2) Construct a content-guided self-attention network to adjust the distortion weights of different regions in the image, and further extract the distortion features and content features after weight adjustment;

[0008] (3) Construct a dual-branch quality prediction network, and fuse and map the distortion features and content features respectively to obtain a quality score;

[0009] (4) Input the distorted image and the content features of different scales extracted by the content understanding network into the content-guided self-attention network, and connect the quality regression network constructed in (3) with the content-guided self-attention network to construct a blind reference image quality evaluation network for real distortion;

[0010] (5) Select and process the distortion quality evaluation data set for verifying the blind reference image quality evaluation network for real distortion in (4);

[0011] (6) Train the blind reference image quality evaluation network to obtain the optimal blind reference image quality evaluation network parameters.

[0012] The content understanding network consists of a backbone architecture based on EfficientNet-B0 and a linear mapping module. In the content understanding network, an image is input into EfficientNet-B0 to extract content features at different scales of the image, and the extracted content features at different scales are input into the linear projection module to obtain vectors with consistent scales. The linear projection module is composed of a series connection of a 3×3 convolution, a 1×1 convolution, an average pooling, and a fully connected layer.

[0013] The content-guided self-attention network includes a content-guided position module and a self-attention encoder. The content-guided position module is composed of an average pooling, two fully connected layers, a rectified linear unit (ReLU) activation layer, and a sigmoid activation layer. The input of the self-attention encoder is obtained by merging the content features output by the content understanding network and the distorted features output by the content-guided position module.

[0014] The quality prediction network includes a distorted feature regression network and a content feature regression network. The distorted feature regression network is composed of an average pooling and a series connection of three fully connected layers; the content feature regression network is composed of a series connection of three fully connected layers.

[0015] The mean absolute error is used as its loss function to calculate the loss between the predicted score and the mean opinion score. Subsequently, the network parameters are updated through the backpropagation mechanism until the Spearman rank order correlation coefficient (SROCC) of the test set reaches the highest value, and the training is terminated. To optimize the network parameters, the adaptive moment estimation (Adam) optimization algorithm is used as the optimizer.

[0016] The content-guided self-attention network is composed of a block & transform operation and a series connection of 4 content-guided self-attention encoders. The block & transform divides a distorted image with a size of H×W×3 into n distorted blocks with a size of m×m, and then through the transform, the distorted blocks with a size of m×m are transformed into vectors. The content-guided self-attention network includes a content-guided position module and a self-attention encoder; the vector output in the content understanding network is defined as the content token f con ∈R d , and the vector obtained through the block & transform operation is defined as the distorted token f dis ∈R d , and a total of n distorted tokens are concatenated into a distorted token sequence F dis ∈R n×d , in the content-guided position module, the distorted token sequence F disDot product with content token f con Then, use two fully connected modules to calculate the weight scores of distortion tokens at different positions relative to the content token. The first fully connected module includes a global average pooling, a fully connected layer, and a ReLU activation function; the second fully connected module includes a fully connected layer and a Sigmoid activation function. The weight scores are then multiplied by the input distortion token sequence to obtain weighted distortion tokens. The overall process of the content-guided position module is defined as:

[0017] F dot = f con ⊙ F dis ,

[0018] f weight = Avgpool(F dis ),

[0019] F wdis = Sigmoid(FC(ReLU(FC(f weight )))) ⊙ F dis ,

[0020] where FC represents the fully connected layer, Avgpool represents the average pooling, and F wdis represents the weighted distortion token sequence.

[0021] To enhance feature analysis, a self-attention encoder is set up to further extract content features and distortion features. First, the content token is concatenated with the distortion token sequence to obtain the input of the self-attention encoder. The process is as follows:

[0022] F input = [f con , F dis ∈ R (n+1)×d .

[0023] The self-attention encoder consists of a multi-head self-attention and a feed-forward multi-layer perceptron. The feed-forward multi-layer perceptron consists of two fully connected layers. Both the multi-head self-attention and the feed-forward multi-layer perceptron include residual connections and layer normalization. The operation process of the self-attention encoder is expressed as:

[0024] F resMHsA = F input + MHSA(F input ),

[0025] F out = F res_MHSA + MLP(F res_MHSA ),

[0026]

[0027] Among them, MHSA() represents the multi-head self-attention mechanism; MLP() represents the feed-forward multi-layer perceptron; represents the content encoding token after feature extraction by the self-attention encoder; represents the distortion encoding token sequence after feature extraction by the self-attention encoder.

[0028] In the quality prediction network, the content encoding token and the distortion encoding token sequence extracted by the content-guided self-attention network are used as the inputs of the dual-branch quality prediction network. The quality prediction network includes a content prediction branch and a distortion prediction branch. The content prediction branch includes 3 fully connected layers, and the content score s is obtained by regressing the content encoding token through the content prediction branch con , the distortion prediction branch includes 1 average pooling layer and 3 fully connected layers, and the distortion score s is obtained by regressing the distortion encoding token sequence through the distortion prediction branch dis , and finally the content score s con and the distortion score s dis are averaged to obtain the final image quality score s,

[0029] s = 0.5 * (s con + s dis ).

[0030] Through the parallel operation, the 4-layer features extracted from the MB2_Conv3 layer, MB2_Conv4 layer, MB2_Conv6 layer, and Conv9 layer in the content understanding network are remapped into 4 content vectors through the linear mapping module, and are respectively input into the 4 content-guided self-attention encoders in the content-guided self-attention network. The different content encoding tokens obtained after feature extraction by the content-guided self-attention encoder are concatenated. At the same time, the distorted image is input into the content-guided self-attention network for feature extraction and the extracted distortion encoding token sequence is concatenated. Finally, the concatenated content encoding token and the concatenated distortion encoding token sequence are respectively input into the quality prediction network constructed in step three, so as to output the final prediction value.

[0031] The features and beneficial effects of the present invention are:

[0032] The present invention uses a content understanding network pre-trained on the ImageNet dataset to extract image content features, and a content-guided self-attention network to learn the distortion feature weights of different regions in the distorted image. In addition, a dual-branch quality prediction network is designed, which predicts the content features and distortion features respectively, and sums the final prediction results to obtain the actual prediction score. The above scheme improves the network's perception ability of real distortion, making the network achieve good results. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is the structure diagram of the content-guided self-attention network described in the present invention.

[0034] Figure 2 It is the content-guided position module described in the present invention.

[0035] Figure 3 It is the structure diagram of the blind reference image quality evaluation network for real distortion described in the present invention. Detailed implementation manners

[0036] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a method for blind reference image quality evaluation of real distortion based on content guidance and self-attention mechanism. This method uses a content understanding network pre-trained on a large-scale image classification dataset ImageNet to extract image content features, and a content-guided self-attention network to learn the distortion feature weights of different regions in the distortion map. In addition, a dual-branch quality prediction network is designed, which predicts the content features and distortion features separately and sums the final prediction results to obtain the actual prediction score. The above scheme improves the network's perception ability of real distortion, enabling the network to achieve good results.

[0037] To achieve the above purpose, the technical solution adopted by the present invention is: a method for blind reference image quality evaluation of real distortion based on content guidance and self-attention mechanism, the method comprising the following steps:

[0038] (1) Introduce a linear mapping module on the basis of the image classification pre-trained network EfficientNet-B0 network to construct a content understanding network;

[0039] (2) Construct a content-guided self-attention network to adjust the distortion weights of different regions in the image, and further extract the distortion features and content features after weight adjustment;

[0040] (3) Construct a dual-branch quality prediction network, and fuse and map the distortion features and content features respectively to obtain quality scores;

[0041] (4) Input the distortion image and the content features of different scales extracted by the content understanding network into the content-guided self-attention network, and connect the quality regression network constructed in (3) with the content-guided self-attention network to construct a blind reference image quality evaluation network for real distortion;

[0042] (5) Select and process the distortion quality evaluation dataset for verifying the blind reference image quality evaluation network for real distortion in (4);

[0043] (6) Train the blind reference image quality evaluation network to obtain the optimal blind reference image quality evaluation network parameters.

[0044] Furthermore, the content understanding network includes a backbone architecture based on EfficientNet-B0 and a linear mapping module, which includes a convolutional layer with a convolution kernel of 3×3, a convolutional layer with a convolution kernel of 1×1, an average pooling layer, and a fully connected layer.

[0045] Furthermore, the content-guided self-attention network includes a content-guided position module and a self-attention encoder. The content-guided position module consists of an average pooling layer, two fully connected layers, a ReLU activation layer, and a Sigmoid activation layer. The self-attention encoder is characterized by taking the content features output by the content understanding network and the distortion features calculated through the content-guided position module as inputs.

[0046] Furthermore, the quality prediction network includes a distortion feature regression network and a content feature regression network. The distortion feature regression network is composed of an average pooling layer and three fully connected layers connected in series; the content feature regression network is composed of three fully connected layers connected in series.

[0047] Furthermore, the method uses the mean absolute error as its loss function to calculate the loss between the predicted score and the mean opinion score, and then updates the network parameters through the backpropagation mechanism until the SROCC of the test set reaches the highest value and the training terminates. To optimize the network parameters, the method uses Adam as the optimizer, the number of training epochs is 25, the learning rate is initialized with 0.0001, and the network learning rate is adjusted to 0.9 times the original learning rate every 5 training epochs.

[0048] The following further describes the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification.

[0049] Step 1: Construct a content understanding network.

[0050] To extract rich content features from distorted images, this method uses EfficientNet-B0 as the backbone of the network. EfficientNet-B0 was first applied to image classification tasks and is widely used due to its powerful representation ability. In addition, since the ImageNet dataset includes more than 1 million images with diverse scenes, the EfficientNet-B0 pre-trained on ImageNet can extract rich content features from distorted images. Therefore, this method uses the pre-trained EfficientNet-B0 parameters on ImageNet to initialize the parameters of the content understanding network model. In addition, we use the 4-layer features extracted from the MB2_Conv3 layer, MB2_Conv4 layer, MB2_Conv6 layer, and Conv9 layer in EfficientNet-B0 as the output. Subsequently, this method sets up a linear projection module, which is composed of a 3×3 convolution, a 1×1 convolution, an average pooling, and a fully connected layer in series. The above-mentioned 4-layer content features of different scales are operated by this module to obtain vectors with consistent scales.

[0051] Step 2: Construct a content-guided self-attention network.

[0052] The content-guided self-attention network is as Figure 1 shown. This network is composed of a block & transformation operation and 4 content-guided self-attention encoders in series. The block & transformation divides a distorted image with a size of H×W×3 into n distorted blocks with a size of m×m. Then, through transformation, the distorted block with a size of m×m is transformed into a vector. The content-guided self-attention network includes a content-guided position module and a self-attention encoder. The content-guided position module is as Figure 2 shown. First, the vector output from the content understanding network is defined as the content token f con ∈R d , and the vector obtained through the block & transformation operation is defined as the distorted token f dis ∈R d . A total of n distorted tokens are concatenated into a distorted token sequence F dis ∈R n×d . In the content-guided position module, the distorted token sequence F dis is dot-multiplied with the content token f con , and then 2 fully connected modules are used to calculate the weight scores of the distorted tokens at different positions relative to the content token. The first fully connected module includes a global average pooling, a fully connected layer, and a ReLU activation function; the second fully connected module includes a fully connected layer and a Sigmoid activation function. Finally, the weight scores are multiplied by the input distorted token sequence again to obtain weighted distorted tokens. The overall process of the content-guided position module is defined as:

[0053] F dot = f con ⊙F dis ,

[0054] f weight = Avgpool(F dis ),

[0055] F wdis = Sigmoid(FC(ReLU(FC(f weight ))))⊙F dis ,

[0056] where FC represents the fully connected layer, Avgpool represents average pooling, and F wdis represents the weighted distortion token sequence.

[0057] To enhance feature analysis, this patent sets up one self-attention encoder to further extract content features and distortion features. First, the content tokens are concatenated with the distortion token sequence to obtain the input of the self-attention encoder. The process is as follows:

[0058] F input = [f con , F dis ∈ R (n+1)×d .

[0059] The self-attention encoder consists of a multi-head self-attention and a feed-forward multi-layer perceptron. The feed-forward multi-layer perceptron consists of two fully connected layers. Both the multi-head self-attention and the feed-forward multi-layer perceptron include residual connections and layer normalization. The operation process of the self-attention encoder can be expressed as:

[0060] F resMHSA = F input + MHSA(F input ),

[0061] F out = F res_MHSA + MLP(F res_MHSA ),

[0062]

[0063] where MHSA() represents the multi-head self-attention mechanism; MLP() represents the feed-forward multi-layer perceptron; represents the content encoding tokens after feature extraction by the self-attention encoder; represents the distortion encoding token sequence after feature extraction by the self-attention encoder.

[0064] Step 3, construct the quality prediction network.

[0065] The content encoding tokens extracted in Step 2 through the content-guided self-attention network and the distortion encoding token sequence are used as the inputs of the dual-branch quality prediction network, which includes a content prediction branch and a distortion prediction branch. Among them, the content prediction branch includes 3 fully-connected layers, and the content score s is obtained by regressing the content encoding tokens through the content prediction branch. con The distortion prediction branch includes 1 average pooling layer and 3 fully-connected layers, and the distortion score s is obtained by regressing the distortion encoding token sequence through the distortion prediction branch. dis Finally, the content score s con and the distortion score s dis are averaged to obtain the final image quality score s.

[0066] s = 0.5 * (s con + s dis ).

[0067] Step 4 constructs a blind reference image quality evaluation network for real distortion.

[0068] Through a parallel operation, the 4-layer features extracted from the MB2_Conv3 layer, MB2_Conv4 layer, MB2_Conv6 layer, and Conv9 layer in the content understanding network are remapped through a linear mapping module to obtain 4 content vectors, which are input into the 4 content-guided self-attention encoders in the content-guided self-attention network, and the 4 content encoding tokens obtained after feature extraction by the content-guided self-attention encoders are concatenated. At the same time, the distorted image is input into the content-guided self-attention network for feature extraction and the extracted distortion encoding token sequence is concatenated. Finally, the concatenated content encoding tokens and the concatenated distortion encoding token sequence are respectively input into the quality prediction network constructed in Step 3, so as to output the final predicted value. Its structure is as Figure 3 shown.

[0069] Step 5 collects and processes the image quality evaluation dataset.

[0070] In order to verify that the present invention can effectively handle real distortion, the present invention uses 5 real distortion datasets, namely LIVEC, BID, CID2013, KonIQ-10K, and SPAQ, to verify the blind reference image quality evaluation network for real distortion. Subsequently, on the premise of ensuring that the aspect ratio of the images in the real distortion dataset remains basically unchanged, the over-sized distorted images are scaled, and 25 image patches with a resolution of 224×224 are randomly cut out from the scaled images.

[0071] Step 6 trains the blind reference image quality evaluation network for real distortion using the quality evaluation results.

[0072] In this method, the real distortion dataset obtained by processing in step five is randomly divided into a training set and a test set at a ratio of 4:1. To weaken the error impact caused by dataset division, the present invention randomly divides the dataset ten times, and uses the average value of the Pearson Linear Correlation Coefficient (PLCC) and SROCC of 10 operations to evaluate the performance of the network. When training the network, Adam is selected as the optimizer, the number of training rounds is 25, the initial learning rate is 0.0001, and every 5 training rounds, the learning rate is adjusted to 0.9 times the original learning rate. The blind reference image quality assessment network for real distortion uses the mean absolute error as its loss function to calculate the loss between the predicted score and the mean opinion score, and then updates the network parameters through backpropagation until the SROCC of the test set reaches the highest and then stops training.

[0073] For the blind reference image quality assessment network obtained in this embodiment, the LIVEC, BID, CID2013, KonIQ-10K, and SPAQ datasets are used for testing. The hardware platform for testing is an NVIDIA 3090 GPU with 24GB video memory, and the operating environment is Ubuntu20.04 LTS. The PLCC and SROCC results are shown in Table 1:

[0074] Table 1 Comparison of SROCC and PLCC of different schemes on 5 real distortion datasets

[0075]

[0076] In the above table, the comparison scheme 1 is the scheme proposed by Bosse et al. [Bosse S, Maniry D, Müller K R, et al. Deep neural networks for no-reference and full-reference image quality assessment[J]. IEEE Transactions on image processing, 2017, 27(1):206-219.], the comparison scheme 2 is the scheme proposed by Yang et al. [Yan Q, Gong D, Zhang Y. Two-stream convolutional networks for blind image quality assessment[J]. IEEE Transactions on Image Processing, 2018, 28(5):2200-2211.], the comparison scheme 3 is the scheme proposed by Wu et al. [Wu J, Ma J, Liang F, et al. End-to-end blind image quality prediction with cascaded deep neural network[J]. IEEE Transactions on image processing, 2020, 29:7414-7426.], the comparison scheme 4 is the scheme proposed by Li et al. [Zhu H, Li L, Wu J, et al. MetaIQA: Deep meta-learning for no-reference image quality assessment[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020:14143-14152.], the comparison scheme 5 is the scheme proposed by Zhang et al. [Zhang W, Ma K, Yan J, et al. Blind image quality assessment using a deep bilinear convolutional neural network[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2018, 30(1):36-47.], the comparison scheme 6 is the scheme proposed by You et al. [You J, Korhonen J.Transformer for image quality assessment[C] / / 2021 IEEE International Conference on Image Processing(ICIP). IEEE, 2021: 1389-1393.]. It can be seen from observing Table 1 that, in terms of SROCC, the present invention achieves the optimal results on 4 datasets, namely BID, CID2013, KonIQ-10K, and SPAQ. And it achieves the sub-optimal results on the LIVEC dataset and is compared with Comparative Scheme 5. In terms of PLCC, the present invention achieves the optimal results on 5 datasets, namely LIVEC, BID, CID2013, KonIQ-10K, and SPAQ. This proves that the present invention can effectively perceive real distortions.

[0077] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A content-guided and self-attention true distortion blind reference image quality assessment method, characterized in that, Introduce a linear mapping module based on the EfficientNet-B0 network to construct a content understanding network; Construct a content-guided self-attention network to adjust the distortion weights of different regions in the image; Construct a dual-branch quality prediction network, and fuse and map the distortion features and content features obtained by the content-guided self-attention network to obtain a quality score.

2. The content-guided and self-attention true distortion blind reference image quality evaluation method according to claim 1, characterized in that The specific steps are as follows: (1) Introduce a linear mapping module based on the EfficientNet-B0 network to construct a content understanding network; (2) Construct a content-guided self-attention network to adjust the distortion weights of different regions in the image, and further extract the distortion features and content features after weight adjustment; (3) Construct a dual-branch quality prediction network, and fuse and map the distortion features and content features respectively to obtain a quality score; (4) Input the distorted image and the content features of different scales extracted by the content understanding network into the content-guided self-attention network, and connect the quality regression network constructed in (3) with the content-guided self-attention network to construct a blind reference image quality evaluation network for real distortion; (5) Select and process the distortion quality evaluation dataset for verifying the blind reference image quality evaluation network for real distortion in (4); (6) Train the blind reference image quality evaluation network to obtain the optimal blind reference image quality evaluation network parameters.

3. The content-guided and self-attention-based true distortion blind reference image quality assessment method according to claim 2, wherein the content The understanding network consists of a backbone architecture based on EfficientNet-B0 and a linear mapping module. In the content understanding network, the image is input into EfficientNet-B0 to extract the content features of different scales of the image, and the extracted content features of different scales are input into the linear projection module to obtain vectors with consistent scales. The linear projection module consists of a series connection of a 3×3 convolution, a 1×1 convolution, an average pooling, and a fully connected layer.

4. The content-guided and self-attention true distortion blind reference image quality evaluation method according to claim 2, characterized in that, The content-guided self-attention network includes a content-guided position module and a self-attention encoder. The content-guided position module consists of a series connection of an average pooling, two fully connected layers, a rectified linear unit (ReLU) activation layer, and a sigmoid activation layer. The input of the self-attention encoder is obtained by merging the content features output by the content understanding network and the distortion features output by the content-guided position module.

5. The content-guided and self-attention true distortion blind reference image quality evaluation method according to claim 2, wherein The quality prediction network includes a distortion feature regression network and a content feature regression network. The distortion feature regression network consists of a series connection of an average pooling and three fully connected layers; the content feature regression network consists of a series connection of three fully connected layers.

6. The content-guided and self-attention true distortion blind reference image quality evaluation method according to claim 2, characterized in that, Use the mean absolute error as its loss function to calculate the loss between the predicted score and the mean opinion score, and then update the network parameters through the backpropagation mechanism until the Spearman rank-order correlation coefficient (SROCC) of the test set reaches the highest. To optimize the network parameters, use the adaptive moment estimation (Adam) optimization algorithm as the optimizer.

7. The content-guided and self-attention-based true distortion blind reference image quality evaluation method according to claim 2, characterized in that the content-guided self-attention network is composed of a block & transformation operation and 4 content-guided self-attention encoders connected in series. The block & transformation divides a distorted image with a size of H×W×3 into n distorted blocks with a size of m×m, and then transforms the distorted block with a size of m×m into a vector through transformation. The content-guided self-attention network includes a content-guided position module and a self-attention encoder; the vector output in the content understanding network is defined as the content token f con ∈R d , and the vector obtained through the block & transformation operation is defined as the distortion token f dis ∈R d . A total of n distortion tokens are concatenated into a distortion token sequence F dis ∈R n×d . In the content-guided position module, the distortion token sequence F dis and the content token f con are multiplied pointwise, and then 2 fully connected modules are used to calculate the weight scores of the distortion tokens at different positions relative to the content token. The first fully connected module includes 1 global average pooling, 1 fully connected layer, and 1 ReLU activation function; the second fully connected module includes 1 fully connected layer and a Sigmoid activation function. The weight scores are then multiplied by the input distortion token sequence to obtain weighted distortion tokens. The overall process of the content-guided position module is defined as: F dot = f con ⊙F dis , f weight = Avgpool(F dis ), F wdis = Sigmoid(FC(ReLU(FC(f weight )))) ⊙ F dis , Among them, FC represents the fully connected layer, Avgpool represents average pooling, and F wdis represents the weighted distortion token sequence; To strengthen feature analysis, one self-attention encoder is set up to further extract content features and distortion features. First, the content tokens are concatenated with the distortion token sequence to obtain the input of the self-attention encoder. The process is as follows: F input = [f con , F dis ∈ R (n+1)×d ; The self-attention encoder consists of a multi-head self-attention and a feed-forward multi-layer perceptron. The feed-forward multi-layer perceptron consists of two fully-connected layers. Both the multi-head self-attention and the feed-forward multi-layer perceptron include residual connections and layer normalization. The operation process of the self-attention encoder is expressed as: F out = F res_MHSA + MLP(F res_MHSA ), Among them, MHSA() represents the multi-head self-attention mechanism; MLP() represents the feed-forward multi-layer perceptron; represents the content encoding token after feature extraction by the self-attention encoder; represents the distortion encoding token sequence after feature extraction by the self-attention encoder.

8. The content-guided and self-attention true distortion blind reference image quality assessment method according to claim 2, characterized in that, In the quality prediction network, the content encoding tokens extracted by the content-guided self-attention network and the distortion encoding token sequence are used as the inputs of the dual-branch quality prediction network. The quality prediction network includes a content prediction branch and a distortion prediction branch. The content prediction branch includes three fully connected layers, and the content encoding tokens are regressed through the content prediction branch to obtain the content score s con , the distortion prediction branch includes one layer of average pooling and three fully connected layers, and the distortion encoding token sequence is regressed through the distortion prediction branch to obtain the distortion score s dis , finally, the content score s con and the distortion score s dis are averaged to obtain the final image quality score s. s = 0.5*(s con + s dis ) Through parallel operations, the four-layer features extracted from the MB2_Conv3 layer, MB2_Conv4 layer, MB2_Conv6 layer, and Conv9 layer in the content understanding network are remapped into four content vectors through a linear mapping module, and are respectively input into the four content-guided self-attention encoders in the content-guided self-attention network. The different content encoding tokens obtained after feature extraction by the content-guided self-attention encoder are concatenated. At the same time, the distorted image is input into the content-guided self-attention network for feature extraction and the extracted distorted encoding token sequence is concatenated. Finally, the concatenated content encoding tokens and the concatenated distorted encoding token sequence are respectively input into the quality prediction network constructed in step three, so as to output the final predicted value.

Citation Information

Patent Citations

  • Quality objective evaluation method with no reference images based on multi-scale generative adversarial network

    CN108090902A

  • Non-reference image quality evaluation method based on deep feature transfer learning

    CN113421237A

  • No-reference image quality evaluation method based on spatial attention mechanism

    CN114066812A