No-Reference Image Quality Assessment Method and System Based on Multi-Dimensional Feature Fusion
By introducing a self-attention mechanism and a multi-dimensional feature fusion module in the reference-free image quality evaluation method, the problems of feature design limitations and impacts of convolutional neural network preprocessing in the prior art are solved, and more accurate and efficient image quality evaluation is achieved.
Patent Information
- Application Number
- CN202211513003.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-28
AI Technical Summary
The existing reference-free image quality evaluation method relies on the characteristics of manual design, and has great limitations. The cropping and proportional changes of convolutional neural networks during the preprocessing stage affect image quality, and the global image connection cannot be fully considered.
The self-attention mechanism is introduced, and by designing a global feature extraction subnet, a multi-scale feature fusion module, a multi-dimensional feature fusion module and a local attention module, a reference-free image quality evaluation method based on multi-dimensional feature fusion can be constructed. The characteristics of the global and local areas can be considered without cropping or changing the image proportions and improve evaluation performance.
This method can effectively integrate multi-dimensional image features, improve the performance of reference-free image quality evaluation algorithm, provide more accurate image quality scores, and do not affect the quality of the original image.
Smart Images

Figure CN115937121B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and computer vision, and particularly relates to a no-reference image quality evaluation method and system based on multi-dimensional feature fusion. Background Art
[0002] The progress of technology has made the sharing and use of multimedia content a part of our daily lives. Digital images and videos have become ubiquitous. On social media websites such as Facebook, Instagram, and Tumblr, hundreds of thousands of photos and videos are uploaded and shared every year. Streaming services such as Netflix, Amazon Prime Video, and YouTube account for 60% of all downstream Internet traffic. In such an era of information explosion, since millions of cameras generate a large amount of images and videos every moment, the goal of image quality evaluation is to measure the quality of images and determine whether the images meet certain specific application requirements. Moreover, the results of image quality assessment can be used as auxiliary reference information for some image restoration and enhancement technologies. Therefore, image quality evaluation methods are highly needed, and they can also provide a feasible way for designing and optimizing advanced image / video processing algorithms.
[0003] Traditional no-reference image quality evaluation methods rely on manually designed features, and most of them attempt to detect specific types of distortions, such as blur, blocking artifacts, various forms of noise, etc. For example, for the evaluation of image blur, there are methods based on edge analysis and methods based on the transform domain. For the evaluation of image noise, there are methods based on filtering, methods based on wavelet transform, and some other methods in the transform domain. For the evaluation of image blocking artifacts, there are methods based on block boundaries and the transform domain. There are also some general types of no-reference image quality evaluation methods. These algorithms do not detect specific types of distortions. They usually transform the no-reference image quality evaluation problem into a classification or regression problem, where classification and regression are trained using specific features. However, manually designed features have their limitations because different types of image content have different image features, which have a great impact on the quality evaluation scores.
[0004] Currently, the research work on no-reference image quality evaluation has entered the era of deep learning. Compared with manually designed features, the features extracted by convolutional neural networks are more suitable for image quality evaluation and are more powerful. However, there are still problems with using convolutional neural networks to evaluate the quality of images. First, operations such as cropping or changing the original ratio of pictures in the preprocessing stage of training convolutional neural networks will affect the quality of the pictures, resulting in incorrect evaluation results. Second, although convolutional neural networks can provide more powerful image features, they are limited by the receptive field of convolution and cannot consider the global connection of images. Summary of the Invention
[0005] To make up for the defects and deficiencies existing in the prior art, the present invention adds a self-attention mechanism to the algorithmic solution. Its characteristic of long-range dependence can make up for the deficiencies of convolution. Thus, a no-reference image quality evaluation method based on multi-dimensional feature fusion is proposed. For the input image, no operations that affect the image quality are performed, and its details and proportions are retained. And while considering the global and local regions, different degrees of attention can also be paid to the local regions, improving the performance of the no-reference image quality evaluation method.
[0006] The present invention can pay different degrees of attention to the local region while considering the global and local regions, and does not need to crop the original image or change its original proportion, improving the performance of the no-reference image quality evaluation algorithm.
[0007] The solution includes the following steps:
[0008] Step S1: Perform data preprocessing on the data in the distorted image dataset. First, pair the data, then perform data augmentation on it, and divide the dataset into a training set and a test set; Step S2: Design a global feature extraction sub-network; Step S3: Design a multi-scale feature fusion module; Step S4: Design a multi-dimensional feature fusion module; Step S5: Design a local attention module; Step S6: Design a no-reference image quality score prediction network based on multi-dimensional feature fusion, and use the designed network to train the no-reference image quality score prediction network model based on multi-dimensional feature fusion; Step S7: Input the image into the trained no-reference image quality score prediction network model based on multi-dimensional feature fusion, and output the corresponding image quality score. This algorithm can effectively fuse the features of multi-dimensional images, and does not need to crop the original image or change its original proportion, and predicts the image quality score, improving the performance of the no-reference image quality evaluation algorithm.
[0009] The present invention specifically adopts the following technical solutions:
[0010] A no-reference image quality evaluation method based on multi-dimensional feature fusion, characterized by including the following steps:
[0011] Step S1: Perform data preprocessing on the data in the distorted image dataset: First, pair the data, then perform data augmentation on it, and divide the dataset into a training set and a test set;
[0012] Step S2: Train a no-reference image quality score prediction network model based on multi-dimensional feature fusion; the training process is based on a no-reference image quality score prediction network based on multi-dimensional feature fusion, which at least includes: a global feature extraction sub-network, a multi-scale feature fusion module, a multi-dimensional feature fusion module, and a local attention module;
[0013] Step S3: Input the image to be tested into the trained no-reference image quality score prediction network model based on multi-dimensional feature fusion, and output the corresponding image quality score.
[0014] Furthermore, step S1 specifically includes the following steps:
[0015] Step S11: Pair the images in the distorted image dataset with their corresponding labels;
[0016] Step S12: Divide the images in the distorted image dataset into a training set and a test set according to a certain ratio;
[0017] Step S13: Resize the images in the training set and the test set to a fixed size H×W;
[0018] Step S14: Perform random flipping operations on the images in the training set for training set data augmentation;
[0019] Step S15: Normalize the images in the training set and the test set.
[0020] Furthermore, the global feature extraction sub-network is specifically:
[0021] Let the input of the global feature extraction sub-network be image I in , whose dimension is 3×H×W; first, use a 32×32 convolution to downsample the input to F v_d , whose dimension is c×h×w, where Then add the learnable position embedding information P v_d and the dimension type embedding information T ve to F ve to obtain F v_p , where P ve , T ve and F v_p have the dimension of c×h×w, T ve and F v_p are randomly initialized, and the calculation formula of F v_p is:
[0022] F v_d =Conv 32X32 I in
[0023] F v_p= F v_d + P ve + T ve
[0024] Among them, Conv 32×32 * represents a convolutional layer for dimensionality reduction with a convolutional kernel size of 32×32;
[0025] Suppose the number of autoencoders is N. Reshape the dimension of F v_p from c×h×w to c×l through the Reshape operation, where l = h×w. Then, pass through N autoencoders in sequence to obtain the output F v_e of the global feature extraction sub-network, with a dimension of c×l. In the i-th autoencoder, let the input be z i-1 . First, perform layer normalization on it, denoted as LN 1 . Then, input it into the multi-head self-attention. The output of the multi-head self-attention is added to z i-1 to obtain the intermediate output feature of the autoencoder Next, perform layer normalization on , denoted as LN 2 . Then, input it into two fully connected layers, denoted as MLP 1 . The output of the two fully connected layers is added to to obtain the output feature z i of the autoencoder, where i ∈ 1, 2, …, N. The calculation formula of the i-th autoencoder is:
[0026]
[0027]
[0028] Among them, MHSA(*) represents multi-head self-attention; finally, take the output z N of the N-th autoencoder as the output feature F v_e of the global feature extraction sub-network.
[0029] Furthermore, the multi-scale feature fusion module is specifically:
[0030] Construct the operation S, which consists of three convolutions. Let the input of the operation S be x, and the dimension of x is c x ×h x ×w x . First, x passes through a 1×1 convolution to reduce the number of channels c x to Then, use a 3×3 convolution to reduce h x and w x to and After that, use a 1×1 convolution to increase the number of channels to 2c x, obtain The dimension of The calculation formula is:
[0031]
[0032] where Conv 1×1 * and Conv 3×3 * respectively represent convolutional layers with a convolutional kernel size of 1×1 and 3×3;
[0033] Let the input of the multi-scale feature fusion module be F c_i , where i ∈ {1, 2, 3, 4}, and the dimension of F c_i is C i ×H i ×W i , where C i = 2C i-1 , First, pass F c_1 through operation S and add it to F c_2 to obtain F c1_d1 , then pass F c1_d1 through operation S and add it to F c_3 to obtain F c1_d2 , and then pass F c1_d2 through operation S to obtain F c1_d3 ; then pass F c_2 through operation S and add it to F c_3 to obtain F c2_d1 , then pass F c2_d1 through operation S to obtain F c2_d2 ; then pass F c_3 through operation S to obtain F c3_d1 , and finally add F c_4 , F c3_d1 , F c2_d2 and F c1_d3 together to obtain the output F s of the scale feature fusion module. The dimension of F s is C 4 ×H 4 ×W 4 , and the calculation formula of F s is:
[0034] F c1_d1 = S(F c_1 ) + F c_2
[0035] F c1_d2 = S(F c1_d1 ) + F c_3
[0036] Fc1_d3 = S(F c1_d2 )
[0037] F c2_d1 = S(F c_2 ) + F c_3
[0038] F c2_d2 = S(F c2_d1 )
[0039] F c3_d1 = S(F c_3 )
[0040] F s = F c_4 + F c3_d1 + F c2_d2 + F c1_d3
[0041] Among them, S(*) represents the operation S.
[0042] Furthermore, the multi-dimensional feature fusion module is specifically:
[0043] Let the input of the multi-dimensional feature fusion module be F v_e and F c , the dimension of F u_e is c×l, and the dimension of F c is C×h×w; first, use a 1×1 convolution to reduce the number of channels C of F c to c, and then add the position embedding information P c and the dimension type embedding information T ce to F ce respectively to obtain F c_p where the dimensions of P ce , T ce and F c_p are all c×h×w; adopt the Reshape operation, denoted as Reshape 1 to change the dimension of F c_p , from c×h×w to c×l, where l = h×w, and then input it into an autoencoder to get F c_e , with the dimension of c×l; the calculation formula of F c_e is:
[0044] F c_p = Conv 1×1 (F c ) + P ce + T ce
[0045] F c_e = SEncoder(Reshape 1 (Fc_p ))
[0046] Among them, SEncoder(*) represents an autoencoder, and Conv 1×1 (*) represents a convolutional layer for dimensionality reduction with a kernel size of 1×1;
[0047] Input F c_e and F v_e into the cross encoder for multi-dimensional feature fusion to obtain F fusion , whose dimension is c×l. In the cross encoder, first perform layer normalization on the input F v_e and F c_e , denoted as LN 3 and LN 4 , then input the multi-head cross attention. The output of the multi-head cross attention is added to F v_e to obtain the intermediate output feature of the cross encoder Then perform layer normalization on , denoted as LN 5 , then input it into two fully connected layers, denoted as MLP 2 . The output of the two fully connected layers is added to to obtain the output feature F fusion . The calculation formula of F fusion is:
[0048]
[0049]
[0050] Among them, MHCA(*,*) represents multi-head cross attention; F fusion is the output feature of the multi-dimensional feature fusion module.
[0051] Furthermore, the local attention module is specifically:
[0052] Suppose the input of the local attention mechanism module is F in , with a dimension of c×l. Input F in into the channel pooling layer to obtain the output F channel , whose dimension is 1×l. The calculation formula of F channel is:
[0053] F channel = FC(Concat(CMaxpool(F in ), CAvgpool(F in )))
[0054] Among them, CMaxpool(*) represents a channel max pooling layer with a stride of 1, CAvgpool(*) represents a channel average pooling layer with a stride of 1, Concat(*) represents feature concatenation in the channel dimension, and FC(*) represents a fully connected layer;
[0055] Let F channel Through the Reshape operation, denoted as Reshape 2 , change the dimension from 1×l to l, and then input F channel into two fully connected layers, denoted as MLP 3 , and adopt an attention mechanism to obtain the importance degree of different local regions of the image learned by the model, so as to determine the different impacts of different regions in the local region on the overall image quality evaluation; then map the value to (0,1) through the sigmoid function to obtain the feature weight w patch , for w patch Through the Reshape operation, denoted as Reshape 3 , change its dimension from l to 1×l, and then use this feature weight as the guiding weight for the local region, that is, multiply the initially input image feature F in by the weight w patch and then add F in , and the final output of the local attention module is F patch , with a dimension of c×l, and the calculation formula of F patch is:
[0056] w patch = Sigmoid(MLP 3 (Reshape 2 F channel )))
[0057] F patch = F in + (F in × Reshape 3 (w patch ))
[0058] Further, in step S2, training to obtain a no-reference image quality score prediction network model based on multi-dimensional feature fusion specifically includes the following steps:
[0059] Step S21: Select an image classification network, and use it as a local feature extraction sub-network after removing its last layer;
[0060] Step S22: Input the images in a certain batch of the training set after step S1 into the local feature extraction sub-network and the global feature extraction sub-network respectively, and obtain the outputs F of the local feature extraction sub-network and the global feature extraction sub-network cand F v_e and input F c into the multi-scale feature fusion module to obtain the output F s ;
[0061] Step S23: Input F s and F v_e from Step S22 into the multi-dimensional feature fusion module to obtain the output F fusion of the multi-dimensional feature fusion module. Then input F fusion into the local attention module to obtain the output F patch ;
[0062] Step S24: For the output F patch of Step S23, first perform a Reshape operation, denoted as Reshape 4 to change the dimension from c×l to P, where P = c×l. Then input F patch into the last two fully connected layers, denoted as MLP 4 to obtain the final image quality evaluation score F out , whose dimension is 1, representing the quality score of the image. Its calculation formula is:
[0063] F out = MLP 4 Reshape 4 F patch
[0064] The loss function of the no-reference image quality evaluation network based on multi-dimensional feature fusion is as follows:
[0065]
[0066] where m is the number of samples, y i represents the true quality score of the image, represents the quality score obtained by the image passing through the no-reference image quality evaluation network based on multi-dimensional feature fusion;
[0067] Step S26: Repeat Steps S22 to S24 in batches until the loss value calculated in Step S24 converges and stabilizes. Then save the network parameters to complete the training process of the no-reference image quality evaluation network based on multi-dimensional feature fusion.
[0068] Furthermore, in Step S7, input the images in the test set into the trained no-reference image quality evaluation network model based on multi-dimensional feature fusion to output the corresponding image quality scores.
[0069] Also, a no-reference image quality evaluation system based on multi-dimensional feature fusion, characterized in that: it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the no-reference image quality evaluation method based on multi-dimensional feature fusion as described above.
[0070] Also, a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, it implements the no-reference image quality evaluation method based on multi-dimensional feature fusion as described above.
[0071] Compared with the prior art, the present invention and its preferred solutions can effectively fuse the features of multi-dimensional images, without cropping the original image or changing its original ratio, and perform image quality score prediction, improving the performance of the no-reference image quality evaluation algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0073] Figure 1 It is a flowchart of the overall design process and implementation process of the embodiment solution of the present invention.
[0074] Figure 2 It is a network model structure diagram in the embodiment of the present invention.
[0075] Figure 3 It is a global feature extraction sub-network structure diagram in the embodiment of the present invention.
[0076] Figure 4 It is a multi-scale feature fusion module structure diagram in the embodiment of the present invention.
[0077] Figure 5 It is a local attention module structure diagram in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0079] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0080] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0081] As Figures 1 - 5 shown, the embodiments of the present invention provide an overall design process and an implementation process of a no-reference image quality evaluation method based on multi-dimensional feature fusion, including the following steps:
[0082] Step 1: Perform data preprocessing on the data in the distorted image dataset. First, pair the data, then perform data augmentation on it, and divide the dataset into a training set and a test set;
[0083] Step 2: Design a global feature extraction sub-network;
[0084] Step 3: Design a multi-scale feature fusion module;
[0085] Step 4: Design a multi-dimensional feature fusion module;
[0086] Step 5: Design a local attention module;
[0087] Step 6: Design a no-reference image quality score prediction network based on multi-dimensional feature fusion, and use the designed network to train a no-reference image quality score prediction network model based on multi-dimensional feature fusion;
[0088] Step 7: Input the image into the trained no-reference image quality score prediction network model based on multi-dimensional feature fusion, and output the corresponding image quality score.
[0089] The following is the specific implementation process of the present invention.
[0090] In this embodiment, Step 1 specifically includes the following steps:
[0091] Step 11: Pair the images in the distorted image dataset with their corresponding labels.
[0092] Step 12: Divide the images in the distorted image dataset into a training set and a test set according to a certain ratio.
[0093] Step 13: Scale the images in the training set and the test set to a fixed size H×W.
[0094] Step 14: Perform random flipping operations on the images in the training set for training set data augmentation.
[0095] Step 15: Normalize the images in the training set and the test set.
[0096] In this embodiment, step 2 specifically includes the following steps:
[0097] Step 21: Let the input of the global feature extraction sub-network be image I in , whose dimension is 3×H×W. First, use a 32×32 convolution to downsample the input to F v_d , whose dimension is c×h×w, where Then add the learnable position embedding information P v_d and the dimension type embedding information T ve to obtain F ve , P v_p , T ve , and T ve and F v_p all have the dimension of c×h×w, T ve and F v_p are randomly initialized, and the calculation formula of F v_p is:
[0098] F v_d =Conv 32X32 I in
[0099] F v_p =F v_d +P ve +T ve
[0100] where Conv 32×32 * represents a convolutional layer for dimensionality reduction with a kernel size of 32×32.
[0101] Step 22: Let the number of autoencoders be N. Change the dimension of F v_p in step 21 through the Reshape operation from c×h×w to c×l, where l = h×w. Then, pass through N autoencoders in sequence to obtain the output F v_e of the global feature extraction sub-network, whose dimension is c×l. In the i-th autoencoder, let the input be z i-1 . First, perform layer normalization on it (denoted as LN 1 ), then input the multi-head self-attention. The output of the multi-head self-attention is added to z i-1 to obtain the intermediate output feature of the autoencoder Then perform layer normalization on (denoted as LN 2 ), then input it into two fully connected layers (denoted as MLP 1 ). The output of the two fully connected layers is added to Add to obtain the output feature z of the autoencoder i , i ∈ 1, 2, …, N, the calculation formula of the i-th autoencoder is:
[0102]
[0103]
[0104] Among them, MHSA(*) represents multi-head self-attention. Finally, take the output z of the N-th autoencoder N as the output feature F of the global feature extraction sub-network v_e .
[0105] In this embodiment, step 3 specifically includes the following steps:
[0106] Step 31: Construct operation S, which consists of three convolutions. Let the input of operation S be x, and the dimension of x be c x ×h x ×w x , x will first pass through a 1×1 convolution to reduce the number of channels c x to Then use a 3×3 convolution to reduce h x and w x to and After that, use a 1×1 convolution to increase the number of channels to 2c x , obtaining with the dimension of The calculation formula is:
[0107]
[0108] Among them, Conv 1×1 (*) and Conv 3×3 (*) respectively represent convolution layers with a convolution kernel size of 1×1 and 3×3.
[0109] Step 32: Let the input of the multi-scale feature fusion module be F c_i , where i ∈ {1, 2, 3, 4}, and the dimension of F c_i is C i ×H i ×W i , where C i = 2C i-1 , First, pass F c_1 through operation S and add it to F c_2 to obtain F c1_d1 , then pass F c1_d1 through operation S and add it to F c_3Add them together to get F c1_d2 , and then F c1_d2 goes through operation S to get F c1_d3 . Then F c_2 goes through operation S and is added to F c_3 to get F c2_d1 . Then F c2_d1 goes through operation S to get F c2_d2 . Then F c_3 goes through operation S to get F c3_d1 . Finally, F c_4 , F c3_d1 , F c2_d2 and F c1_d3 are added together to get the output F of the scale feature fusion module s , and the dimension of F s is C 4 ×H 4 ×W 4 . The calculation formula of F s is:
[0110] F c1_d1 = S(F c_1 ) + F c_2
[0111] F c1_d2 = S(F c1_d1 ) + F c_3
[0112] F c1_d3 = S(F c1_d2 )
[0113] F c2_d1 = S(F c_2 ) + F c_3
[0114] F c2_d2 = S(F c2_d1 )
[0115] F c3_d1 = S(F c_3 )
[0116] F s = F c_4 + F c3_d1 + F c2_d2 + F c1_d3
[0117] Among them, S(*) represents operation S.
[0118] In this embodiment, step 4 specifically includes the following steps:
[0119] Step 41: Let the input of the multi-dimensional feature fusion module be Fv_e and F c , F v_e has a dimension of c×l, and F c has a dimension of C×h×w. First, use a 1×1 convolution to reduce the number of channels C of F c to c, and then add positional embedding information P c and dimensional type embedding information T ce to F ce respectively to obtain F c_p where P ce , T ce and F c_p all have a dimension of c×h×w. Use the Reshape operation (denoted as Reshape 1 ) to change the dimension of F c_p from c×h×w to c×l, where l = h×w, and then input it into an autoencoder to obtain F c_e , whose dimension is c×l. The calculation formula of F c_e is:
[0120] F c_p = Conv 1×1 (F c ) + P ce + T ce
[0121] F c_e = SEncoder(Reshape 1 (F c_p ))
[0122] where SEncoder(*) represents the autoencoder, and Conv 1×1 (*) represents the convolutional layer for dimensionality reduction with a kernel size of 1×1.
[0123] Step 42: Input F c_e and F v_e from Step 41 into a cross encoder for multi-dimensional feature fusion to obtain F fusion , whose dimension is c×l. In the cross encoder, first perform layer normalization on the input F v_e and F c_e (denoted as LN 3 and LN 4 ), then input multi-head cross-attention. The output of the multi-head cross-attention is added to F v_e to obtain the intermediate output feature of the cross encoder Then perform layer normalization on (denoted as LN 5 ), and then input it into two fully connected layers (denoted as MLP 2 ). The output of the two fully connected layers is added to Add them to obtain the output feature F of the cross encoder fusion . F fusion The calculation formula of is as follows:
[0124]
[0125]
[0126] Among them, MHCA(*,*) represents multi-head cross attention. F fusion is also the output feature of the multi-dimensional feature fusion module.
[0127] In this embodiment, step 5 specifically includes the following steps:
[0128] Step 51: Let the input of the local attention mechanism module be F in , whose dimension is c×l. Input F in into the channel pooling layer to obtain the output F channel , whose dimension is 1×l. F channel The calculation formula of is as follows:
[0129] F channel = FC(Concat(CMaxpool(F in ), CAvgpool(F in )))
[0130] Among them, CMaxpool(*) represents the channel maximum pooling layer with a stride of 1, CAvgpool(*) represents the channel average pooling layer with a stride of 1, Concat(*) represents feature concatenation in the channel dimension, and FC(*) represents the fully connected layer.
[0131] Step 52: Apply the Reshape operation (denoted as Reshape channel ) to F 2 obtained in step 51 to change its dimension from 1×l to l. Then input F channel into two layers of fully connected layers (denoted as MLP 3 ). Use the attention mechanism to obtain the importance degree of different local regions of the image learned by the model, so as to determine which local regions have a greater impact on the quality evaluation of the overall image. Then, map the values to (0,1) through the sigmoid function to obtain the feature weight w patch . Apply the Reshape operation (denoted as Reshape patch ) to w 3 to change its dimension from l to 1×l. Then use this feature weight as the guiding weight for the local region, that is, multiply the initially input image feature F in by the weight w patchPlus F in to obtain the final output of the local attention module as F patch , with the dimension of c×l, and the calculation formula of F patch is as follows:
[0132] w patch = SigmoidMLP 3 Reshape 2 F channel
[0133] F patch = F in + F in × Reshape 3 w patch
[0134] In this embodiment, step 6 specifically includes the following steps:
[0135] Step 61: Select one network from image classification networks such as ResNet50 and ResNet101, and use it as the local feature extraction sub-network after removing the last layer of the network.
[0136] Step 62: Input the images of a certain batch in the training set after step 1 into the models in step 61 and step 2 at the same time to obtain the outputs F c and F v_e of the local feature extraction sub-network and the global feature extraction sub-network, and input F c into the multi-scale feature fusion module designed in step 3 to obtain the output F s .
[0137] Step 63: Input F s and F v_e in step 62 into the multi-dimensional feature fusion module designed in step 4 to obtain the output F fusion of the multi-dimensional feature fusion module. Then, input F fusion into the local attention module designed in step 5 to obtain the output F patch of the local attention module.
[0138] Step 64: For the output F patch in step 63, first use the Reshape operation (denoted as Reshape 4 ) to change its dimension from c×l to P, where P = c×l. Then, input F patch into the last two fully connected layers (denoted as MLP 4 ) to obtain the final image quality evaluation score F out , whose dimension is 1, representing the quality score of the image, and its calculation formula is:
[0139] F out = MLP 4 Reshape 4 F patch
[0140] Step 65: The loss function of the no-reference image quality evaluation network based on multi-dimensional feature fusion is as follows:
[0141]
[0142] where m is the number of samples, y i represents the true quality score of the image, represents the quality score obtained by the no-reference image quality evaluation network based on multi-dimensional feature fusion for the image.
[0143] Step 66: Repeat the above steps 62 to 65 in batches until the loss value calculated in step 65 converges and stabilizes, save the network parameters, and complete the training process of the no-reference image quality evaluation network based on multi-dimensional feature fusion.
[0144] In this embodiment, step 7 specifically includes the following steps:
[0145] Step 71: Input the images in the test set into the trained no-reference image quality evaluation network model based on multi-dimensional feature fusion, and output the corresponding image quality scores.
[0146] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0147] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0148] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks specified in the block.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks specified in the block.
[0150] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0151] This patent is not limited to the above best mode. Anyone inspired by this patent can obtain various other forms of no-reference image quality assessment methods and systems based on multi-dimensional feature fusion. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A no-reference image quality assessment method based on multi-dimensional feature fusion, characterized in that, it includes the following steps: Step S1: Preprocess the data in the distorted image dataset: First, pair the data, then perform data augmentation on it, and divide the dataset into a training set and a test set; Step S2: Train a no-reference image quality score prediction network model based on multi-dimensional feature fusion; The training process is based on a no-reference image quality score prediction network with multi-dimensional feature fusion, which at least includes: a global feature extraction sub-network, a multi-scale feature fusion module, a multi-dimensional feature fusion module, and a local attention module; Step S3: Input the image to be tested into the trained no-reference image quality score prediction network model based on multi-dimensional feature fusion, and output the corresponding image quality score; Step S1 specifically includes the following steps: Step S11: Pair the images in the distorted image dataset with their corresponding labels; Step S12: Divide the images in the distorted image dataset into a training set and a test set; Step S13: Scale the images in the training set and the test set to a fixed size H×W; Step S14: Perform a random flipping operation on the images in the training set for training set data augmentation; Step S15: Normalize the images in the training set and the test set; The global feature extraction sub-network is specifically: Let the input of the global feature extraction sub-network be the image I in , whose dimension is 3×H×W; First, use a 32×32 convolution to downsample the input to F v_d , whose dimension is c×h×w, where Then add the learnable position embedding information P v_d and the dimension type embedding information T ve to F ve to obtain F v_p , where P ve , T ve and F v_p all have the dimension of c×h×w, T ve and F v_p are randomly initialized, and the calculation formula of F v_p is: F v_d = Conv 32X32 (I in ) F v_p = F v_d + P ve + T ve Among them, Conv 32×32 (*) represents a convolutional layer for dimensionality reduction with a convolutional kernel size of 32×32; Let the number of autoencoders be N, and F v_p The dimension is changed by the Reshape operation from c×h×w to c×l, where l = h×w. Then, it passes through N autoencoders in sequence to obtain the output F of the global feature extraction sub-network v_e , with the dimension of c×l. In the i-th autoencoder, let the input be z i-1 . First, perform layer normalization on it, denoted as LN 1 . Then, input it into the multi-head self-attention. The output of the multi-head self-attention is added to z i-1 to obtain the intermediate output feature of the autoencoder Then, perform layer normalization on , denoted as LN 2 . Then, input it into two fully connected layers, denoted as MLP 1 . The output of the two fully connected layers is added to to obtain the output feature z of the autoencoder i , i ∈ [1, 2,..., N]. The calculation formula of the i-th autoencoder is: Among them, MHSA(*) represents multi-head self-attention; finally, the output z of the Nth autoencoder is taken N as the output feature F of the global feature extraction sub-network v_e ; The multi-scale feature fusion module is specifically: Build operation S, which consists of three convolutions. Let the input of operation S be x, and the dimension of x be c x ×h x ×w x , x first passes through a 1×1 convolution to reduce the number of channels c x to Then use a 3×3 convolution to reduce h x and w x to and After that, use a 1×1 convolution to increase the number of channels to 2c x , obtaining with the dimension of The calculation formula is: Among them, Conv 1×1 (*) and Conv 3×3 (*) represent convolutional layers with convolutional kernel sizes of 1×1 and 3×3 respectively; Let the input of the multi-scale feature fusion module be F c_i , where i ∈ {1, 2, 3, 4}, and the dimension of F c_i is C i ×H i ×W i , where C i = 2C i-1 . First, pass F c_1 through operation S and add it to F c_2 to obtain F c1_d1 . Then, pass F c1_d1 through operation S and add it to F c_3 to obtain F c1_d2 . After that, pass F c1_d2 through operation S to obtain F c1_d3 ; Next, pass F c_2 through operation S and add it to F c_3 to obtain F c2_d1 . Then, pass F c2_d1 through operation S to obtain F c2_d2 ; Then, pass f c_3 through operation S to obtain F c3_d1 . Finally, add F c_4 , F c3_d1 , F c2_d2 and F c1_d3 together to obtain the output F s of the scale feature fusion module. The dimension of F s is C 4 ×H 4 ×W 4 . The calculation formula of F s is: F c1_d1 = S(F c_1 ) + F c_2 F c1_d2 = S(F c1_d1 ) + F c_3 F c1_d3 = S(F c1_d2 ) F c2_d1 = S(F c_2 ) + F c_3 F c2_d2 = S(F c2_d1 ) F c3_d1 = S(F c_3 ) F s = F c_4 + F c3_d1 + F c2_d2 + F c1_d3 where S(*) represents operation S; The multi-dimensional feature fusion module is specifically: Let the input of the multi-dimensional feature fusion module be F v_e and F c , the dimension of F v_e is c×l, and the dimension of F c is C×h×w; first, use a 1×1 convolution to reduce the number of channels C of F c to c, and then add the position embedding information P c and the dimension type embedding information T ce to F ce respectively to obtain F c_p where the dimensions of P ce , T ce and F c_p are all c×h×w; adopt a Reshape operation, denoted as Reshape 1 to change the dimension of F c_p from c×h×w to c×l, where l = h×w, and then input it into an autoencoder to obtain F c_e with the dimension of c×l; The calculation formula of F c_e is: F c_p = Conv 1×1 (F c ) + P ce + T ce F c_e = SEncoder(Reshape 1 (F c_p )) Among them, SEncoder(*) represents the autoencoder, and Conv 1×1 (*) represents a convolutional layer for dimensionality reduction with a convolution kernel size of 1×1; Input F c_e and F v_e into a cross encoder for multi-dimensional feature fusion to obtain F fusion , whose dimension is c×l; in the cross encoder, first perform layer normalization on the input F v_e and F c_e , denoted as LN 3 and LN 4 , then input into multi-head cross-attention, and the output of the multi-head cross-attention is added to F v_e to obtain the intermediate output feature of the cross encoder Then perform layer normalization on , denoted as LN 5 , and then input into two fully connected layers, denoted as MLP 2 , and the output of the two fully connected layers is added to to obtain the output feature F fusion , The calculation formula of F fusion is: Among them, MHCA(*,*) represents multi-head cross-attention; F fusion is the output feature of the multi-dimensional feature fusion module; The local attention module is specifically: Let the input of the local attention mechanism module be F in , with a dimension of c×l. Input F in into the channel pooling layer to obtain the output F channel , whose dimension is 1×l. The calculation formula of F channel is as follows: F channel = FC(Concat(CMaxpool(F in ), CAvgpool(F in ))) where CMaxpool(*) represents a channel maximum pooling layer with a stride of 1, CAvgpool(*) represents a channel average pooling layer with a stride of 1, Concat(*) represents feature concatenation in the channel dimension, and FC(*) represents a fully connected layer; Let F channel Through the Reshape operation, denoted as Reshape 2 , change the dimension from 1×l to l, and then input F channel into two fully-connected layers, denoted as MLP 3 , adopt the attention mechanism to obtain the importance degree of different local regions of the image learned by the model, so as to determine the different influences of different regions in the local region on the overall image quality evaluation; then map the numerical value to (0,1) through the sigmoid function to obtain the feature weight w patch , for w patch Through the Reshape operation, denoted as Reshape 3 , change its dimension from l to 1×l, and then use this feature weight as the guiding weight for the local region, that is, multiply the initially input image feature F in by the weight w patch and then add F in , and the final output of the local attention module is F patch , with the dimension of c×l, and the calculation formula of F patch is: w patch = SImgoid(MLP 3 (Reshape 2 (F channel ))) F patch = F in + (F in × Reshape 3 (w patch ))。 2. The no-reference image quality assessment method based on multi-dimensional feature fusion according to claim 1, characterized in that: In step S2, training to obtain a no-reference image quality score prediction network model based on multi-dimensional feature fusion specifically includes the following steps: Step S21: Select an image classification network, and remove the last layer of this network as a local feature extraction sub-network; Step S22: Input images of a certain batch in the training set that has undergone Step S1 into the local feature extraction sub-network and the global feature extraction sub-network respectively, to obtain the outputs F c and F v_e , and input F c into the multi-scale feature fusion module to obtain the output F s ; Step S23: Input F of step S22 s and F v_e into the multi-dimensional feature fusion module to obtain the output F fusion of the multi-dimensional feature fusion module. Then input F fusion into the local attention module to obtain the output F patch ; Step S24: For the output F of step S23 patch , first perform a Reshape operation, denoted as Reshape 4 to change the dimension from c×l to P, where P = c×l, and then input F patch into the last two fully connected layers, denoted as MLP 4 to obtain the final image quality evaluation score F out , whose dimension is 1, representing the quality score of the image, and its calculation formula is: F out = MLP 4 (Reshape 4 (F patch )) The loss function of the no-reference image quality assessment network based on multi-dimensional feature fusion is as follows: where m is the number of samples, and y i represents the true quality score of the image, represents the quality score obtained by the no-reference image quality assessment network based on multi-dimensional feature fusion; Step S26: Repeat steps S22 to S24 in batches until the loss value calculated in step S24 converges and stabilizes, save the network parameters, and complete the training process of the no-reference image quality assessment network based on multi-dimensional feature fusion.
3. The no-reference image quality assessment method based on multi-dimensional feature fusion according to claim 1, characterized in that, In step S7, input the images in the test set into the trained no-reference image quality assessment network model based on multi-dimensional feature fusion, and output the corresponding image quality scores.
4. A no-reference image quality assessment system based on multi-dimensional feature fusion, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. It is characterized in that when the processor executes the computer program, it implements the no-reference image quality evaluation method based on multi-dimensional feature fusion according to any one of claims 1-3.
5. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, it implements the no-reference image quality evaluation method based on multi-dimensional feature fusion according to any one of claims 1-3.
Citation Information
Patent Citations
Conic cornea recognition method and system based on multi-dimensional feature adaptive fusion
CN111340776A
Self-attention feature embedded night airport runway foreign matter detection method and system
CN114973116A