No-Reference Screen Content Image Quality Assessment Method Based on Multi-Region Feature Fusion
By adopting a multi-region feature fusion method in the screen content image quality evaluation model, and using attention mechanism to fuse text and image features, the quality score deviation problem caused by chunking training is solved, and higher evaluation accuracy and visual perception consistency are achieved.
Patent Information
- Application Number
- CN202310398032.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-14
AI Technical Summary
Existing screen content image quality evaluation models are prone to quality score deviations during chunking training, and it is difficult to effectively integrate the characteristics of image areas and text areas, affecting the accuracy of evaluation.
Using a multi-region feature fusion method, by designing an adaptive feature extraction module and a local image information interaction module, the attention mechanism is used to fuse text and image features of different scales, and each image block has a different attention weight.
Improve the accuracy of the image quality evaluation model of the reference-free screen content, reduce quality score deviation, and improve subjective and objective visual perception consistency.
Smart Images

Figure CN116403063B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and computer vision, and particularly to a no-reference screen content image quality assessment method based on multi-region feature fusion. Background Art
[0002] In recent years, with the rapid development of mobile devices, multimedia applications, and information dissemination technologies, remote computing and communication have been widely used, and real-time changing screen content images generated by various electronic terminal devices increasingly frequently appear in people's daily lives. Compared with traditional natural images, screen content images are a new type of image, which have more text, icons, graphics, and special structural layout information and statistical features. The number of sharp edges in screen content images far exceeds that of natural images, while the color information is less than that of natural images. During the process of screen content image encoding, compression, and transmission, due to reasons such as technical or hardware limitations, various degrees of distortion will inevitably be introduced, resulting in a decline in image quality, and further affecting the user experience and system interaction performance. Considering people's demand for clear and high-quality images, an image quality assessment method that can effectively evaluate the perceptual quality of screen content images is needed. This can not only serve as auxiliary reference information for some image restoration and enhancement technologies, but also provide a feasible way for designing and optimizing advanced image / video processing algorithms.
[0003] Traditional image quality evaluation methods can be divided into two categories: subjective evaluation and objective evaluation according to the different evaluation subjects. Subjective evaluation refers to that humans, as the ultimate recipients of image information, score the image quality. However, subjective evaluation methods are easily affected by the subjective consciousness of the subjects, and in the current era of explosive growth of data volume, it is unrealistic to conduct subjective evaluations on tens of thousands of image data. Therefore, it is difficult to meet the needs of real-time applications. Objective image quality evaluation methods, on the other hand, extract and analyze the relevant features of distorted images, and then construct corresponding mathematical models to calculate the quality assessment scores of distorted images. Compared with subjective evaluation methods, this process does not require a large number of subjects to score the images, but uses a computer to replace the human visual system to achieve automatic and efficient quality assessment, so it can better meet the application needs under the background of big data. According to the different degrees of dependence on reference image information in the process of image perceptual quality assessment, objective quality evaluation methods can be divided into three categories: full-reference, semi-reference, and no-reference methods. The dependence of these three methods on reference image information decreases in turn. However, in practical applications, it is often difficult to obtain a reference image without distortion. Therefore, no-reference image quality evaluation methods are more practical and have a broader development prospect.
[0004] With the continuous development of deep learning technology, many screen content image quality assessment models based on convolutional neural networks have emerged in recent years. Considering that deep learning is a data-driven method and the existing screen content image datasets often only contain a small number of distorted screen content images, currently, data augmentation is mainly achieved by dividing screen content images into blocks. However, a single image block cannot fully represent the quality of the entire distorted image. Moreover, since there are a large number of text and image regions in screen content images, even under the same distortion effect, image blocks with different contents will have significant quality differences. Therefore, it is necessary to fully consider the impact differences of image regions and text regions on the overall visual quality of the image, design a weight strategy based on the degree of difference, and effectively fuse the image features of image regions and text regions to further improve the accuracy of the no-reference screen content image quality assessment model. Summary of the Invention
[0005] The purpose of the present invention is to provide a no-reference screen content image quality assessment method based on multi-region feature fusion. Based on the idea of multi-region local feature fusion, this method characterizes the overall quality of an image by fusing the features of multiple local image blocks, thereby reducing the quality score deviation caused by training with a single image block. At the same time, when designing the convolutional layer of the convolutional neural network, different-sized convolutional kernels are used to more effectively extract different features of text regions and image regions in screen content images. The attention mechanism is used not only to effectively fuse text features and image features of different scales but also to assign different degrees of attention weights to each image block. Overall, this method can achieve higher subjective and objective visual perception consistency compared with other methods.
[0006] To achieve the above purpose, the technical solution of the present invention is: a no-reference screen content image quality assessment method based on multi-region feature fusion, including the following steps:
[0007] Step S1: Perform data preprocessing on the data in the distorted screen content image dataset. First, crop image blocks from each distorted screen content image, then divide the dataset into a training set and a test set, and finally perform data augmentation on the data in the training set;
[0008] Step S2: Design an adaptive feature extraction module, which can adaptively extract different-scale features of text regions and image regions in distorted screen content image blocks and fuse the text region features and image region features based on the attention mechanism;
[0009] Step S3: Design a local image information interaction module, which enhances the information interaction between any two image blocks in the distorted screen content image by introducing a self-attention mechanism and assigns different attention weights to each image block;
[0010] Step S4: Design a no-reference image quality assessment network based on multi-region feature fusion, and train a no-reference screen content image quality assessment model based on multi-region feature fusion;
[0011] Step S5: Input the distorted screen content image to be measured into the trained no-reference screen content image quality assessment model based on multi-region feature fusion, and output the corresponding quality assessment score.
[0012] In an embodiment of the present invention, the specific implementation of the step S1 is as follows:
[0013] Step S11: First, crop image patches from each distorted screen content image I in the distorted screen content image dataset; specifically, divide each distorted screen content image I into four regions: upper left, upper right, lower left, and lower right, and then randomly crop an image patch of size H×W from each region, denoted as I 1 , I 2 , I 3 and I 4 , where H and W respectively represent the height and width of the image patch;
[0014] Step S12: Divide the distorted screen content images in the distorted screen content image dataset into a training set and a test set according to a predetermined ratio;
[0015] Step S13: Perform unified horizontal random flipping and normalization processing on the four cropped image patches of each distorted screen content image I train in the training set to complete the data augmentation operation, and obtain the distorted screen content image patches for training and Perform the same normalization processing on the four cropped image patches of each distorted screen content image I test in the test set to obtain the distorted screen content image patches for testing and
[0016] In an embodiment of the present invention, the specific implementation of the step S2 is as follows:
[0017] Step S21: Design a text feature extraction sub-module, which consists of a convolutional layer with a convolutional kernel size of 3×3, two convolutional layers with a convolutional kernel size of 1×1, two LeakyReLU activation functions, and three batch normalization layers; use the convolutional layer with a convolutional kernel size of 3×3 to extract features from the text region in the distorted screen content image patch, and denote the feature map input to the text feature extraction sub-module as F t , whose size is H t ×Wt ×C t where H t , W t and C t represent the height, width, and number of channels of the input feature map F t respectively; specifically, first, the feature map F t is sequentially input into a convolutional layer with a kernel size of 3×3, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer for preliminary feature extraction to obtain an intermediate feature map F' t1 with dimensions H t ×W t ×C' t where H t , W t and C' t represent the height, width, and number of channels of the intermediate feature map F' t1 respectively; then, the input feature map F t is sequentially input into a convolutional layer with a kernel size of 1×1 and a batch normalization layer for residual feature extraction to obtain an intermediate feature map F' t2 with dimensions H t ×W t ×C' t , which has the same dimension as the intermediate feature map F' t1 ; finally, through residual connection, the intermediate feature map F' t1 is added to the intermediate feature map F' t2 , and then passed through the LeakyReLU activation function to obtain the output feature F' t of the text feature extraction sub-module, with dimensions H t ×W t ×C' t ; the specific calculation formula is as follows:
[0018] F′ t1 = BN(Conv2(LeakyReLU(BN(Conv1(F t )))))
[0019] F′ t2 = BN(Conv3(F t )))
[0020]
[0021] where Conv1(*) represents a convolutional layer with a kernel size of 3×3, and Conv2(*) and Conv3(*) represent two convolutional layers with a kernel size of 1×1, denotes matrix addition operation, LeakyReLU(·) denotes the LeakyReLU activation function, and BN(·) denotes batch normalization operation;
[0022] Step S22: Design an image feature extraction sub-module, which consists of a convolutional layer with a convolutional kernel size of 5×5, two convolutional layers with a convolutional kernel size of 1×1, two LeakyReLU activation functions, and three batch normalization layers; Use the convolutional layer with a convolutional kernel size of 5×5 to extract features from the image region in the distorted screen content image block; Denote the feature map input to the image feature extraction sub-module as F p , whose size is H p ×W p ×C p , where H p , W p and C p respectively represent the height, width, and number of channels of the input feature map F p ; Specifically, first input the feature map F p into a convolutional layer with a convolutional kernel size of 5×5, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a convolutional kernel size of 1×1, and a batch normalization layer in sequence for preliminary feature extraction to obtain an intermediate feature map F' p1 , whose dimension is H p ×W p ×C' p , where H p , W p and C' p respectively represent the height, width, and number of channels of the intermediate feature map F' p1 ; Then input the input feature map F p into a convolutional layer with a convolutional kernel size of 1×1 and a batch normalization layer in sequence for residual feature extraction to obtain an intermediate feature map F' p2 , whose dimension is H p ×W p ×C' p , which has the same dimension size as the intermediate feature map F' p1 ; Finally, through residual connection, add the intermediate feature map F' p1 and the intermediate feature map F' p2 , and then pass through the LeakyReLU activation function to obtain the output feature F' p of the image feature extraction sub-module, whose dimension is H p ×W p ×C' p ; The specific calculation formula is as follows:
[0023] F′ p1= BN(Conv2(LeakyReLU(BN(Conv1(F p )))))
[0024] F' p2 = BN(Conv3(F p )))
[0025]
[0026] Wherein, Conv1'(*) represents a convolutional layer with a convolutional kernel size of 5×5, Conv2(*) and Conv3(*) represent two convolutional layers with a convolutional kernel size of 1×1, represents matrix addition operation, LeakyReLU(·) represents the LeakyReLU activation function, and BN(·) represents batch normalization operation;
[0027] Step S23: Design an attention feature fusion sub-module, which consists of four convolutional layers with a convolutional kernel size of 1×1, a global average pooling layer, two ReLU activation functions, a Sigmoid activation function, and four batch normalization layers; the attention feature fusion sub-module can fuse text features and image features of different scales through learning. Denote the two features input to the attention feature fusion sub-module as F' t and F' p , both of which have a size of H a × W a × C a , where H a , W a and C a represent the height, width, and number of channels of the input feature maps F' t and F' p respectively; specifically, first add the two input features pixel by pixel to obtain an intermediate feature map F b , whose size is H a × W a × C a ; then input the intermediate feature map F b into the local attention extraction branch and the global attention extraction branch respectively for different attention feature extractions. The local attention extraction branch is sequentially composed of a convolutional layer with a convolutional kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a convolutional kernel size of 1×1, and a batch normalization layer connected in series. The global attention extraction branch is sequentially composed of a global average pooling layer, a convolutional layer with a convolutional kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a convolutional kernel size of 1×1, and a batch normalization layer connected in series; denote the intermediate feature map F bThe feature output after passing through the local attention extraction branch is F local 、The feature output after passing through the global attention extraction branch is F global ,and their sizes are both H a ×W a ×C a ; Then, the feature F local is added pixel by pixel to the feature F global , and then the corresponding learnable weight λ is obtained through the Sigmoid function; Finally, the learnable weight λ is weighted and fused with the input features F' t and F' p to obtain the final output F' b of the attention feature fusion sub-module, and its size is H a ×W a ×C a ; The specific calculation formula is as follows:
[0028]
[0029] F local = BN(Conv 2_a (ReLU(BN(Conv 1_a (F b )))))
[0030] F global = BN(Conv 4_a (ReLU(BN(Conv 3_a (GAP(F b )))))
[0031]
[0032]
[0033] Among them, Conv 1_a (*), Conv 2_a (*), Conv 3_a (*) and Conv 4_a (*) represent four convolutional layers with a convolutional kernel size of 1×1, represents matrix addition operation, BN(·) represents batch normalization operation, GAP(*) represents global average pooling operation, ReLU(·) represents ReLU activation function, Sigmoid(·) represents Sigmoid activation function, and λ is the learnable weight output by the network;
[0034] Step S24: Design a channel attention sub-module to enhance feature representation and obtain the key feature channel information of the input features. This module consists of two convolutional layers with a kernel size of 1×1, a ReLU activation function, and a Sigmoid activation function. Denote the feature map input to the channel attention sub-module as F c , whose size is H c ×W c ×C c , where H c , W c and C c respectively represent the height, width, and number of channels of the input feature map F c ; specifically, first use global average pooling operation to aggregate the input feature F c , then perform a dimensionality reduction operation through a convolutional layer with a size of 1×1, then perform a dimensionality increase operation through a convolutional layer with a size of 1×1, then obtain the corresponding channel attention weights through the Sigmoid function, and finally multiply the channel attention weights with the input feature F c element-wise to obtain the final output F' c of the channel attention sub-module, whose size is H c ×W c ×C c , which has the same dimension as the input feature map F c ; the specific calculation formula is as follows:
[0035] F′ c =Sigmoid(Conv 2_b (ReLU(Conv 1_b (GQP(F c )))))⊙F c
[0036] where, GAP(*) represents the global average pooling operation, Conv 1_b (*) and Conv 2_b (*) represent two convolutional layers with a kernel size of 1×1, “⊙” represents matrix multiplication operation, Sigmoid(·) represents the Sigmoid activation function, and ReLU(·) represents the ReLU activation function;
[0037] Step S25: Design an adaptive feature extraction module, which consists of four text feature extraction sub-modules described in Step S21, four image feature extraction sub-modules described in Step S22, four attention feature fusion sub-modules described in Step S23, one channel attention sub-module described in Step S24, and eight spatial average pooling layers with a stride of 2; the adaptive feature extraction module adaptively performs multi-scale feature extraction on the text privilege and image features of the input distorted screen content image block through the text feature extraction branch and the image feature extraction branch therein, and fuses different types of image features through an attention mechanism; specifically, the text feature extraction branch is serially composed of a combination of one text feature extraction sub-module and one spatial pooling layer repeated four times in sequence, and the image feature extraction branch is serially composed of a combination of one image feature extraction sub-module and one spatial pooling layer repeated four times in sequence.
[0038] Denote the size of the input distorted screen content image block as H×W×3. First, input it into the text feature extraction branch and the image feature extraction branch respectively to perform multi-scale text feature and image feature extraction. Denote the multi-level text features output after the input distorted screen content image block passes through the text feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Denote the multi-level image features output after the input distorted screen content image block passes through the image feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Then, input the multi-level text features and the corresponding multi-level image features into four attention feature fusion sub-modules respectively to perform fusion of text features and image features, and obtain the fused multi-level backbone features where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Then, for the multi-level backbone features perform global average pooling operations respectively, and then perform feature concatenation along the channel direction to obtain the multi-scale text and image feature representation F' of the input image patch tp , whose size is 1×1×15C', and the specific calculation formula is as follows:
[0039]
[0040] where Concat(·) represents the feature concatenation operation, and GAP(*) represents the global average pooling operation; Finally, input the fused multi-scale text and image feature F' tp into the channel attention sub-module to capture the key information between different channels, and obtain the final output feature F of the adaptive feature extraction module tp , and then flatten the feature F tp into a one-dimensional vector, whose dimension size is 1×D, where D represents the dimension of each image patch, and D = 960.
[0041] In an embodiment of the present invention, the specific implementation of step S3 is as follows:
[0042] Design a local image information interaction module, which consists of four fully connected layers and a Softmax function; The local image information interaction module uses the self-attention mechanism to enhance the information interaction between the features of different image patches, so that each image patch is given different degrees of attention to better aggregate the local features of each image patch; Specifically, denote the input feature of the local image information interaction module as F l , whose size is N×D, where N represents the number of image patches, N = 4; D represents the dimension of each image patch, D = 960; First, input the input feature F l into three fully connected layers to generate three new intermediate features F Q , F K and F V , whose dimension sizes are all N×D', where D' represents the dimension of the second dimension of the intermediate features F Q , F K and F V , D' = 480; Then, perform matrix multiplication between the transposes of the intermediate features F Q and F K , and use the Softmax function to generate the attention map A, whose dimension size is N×N; Then perform matrix multiplication on the intermediate feature F V and the attention map A to obtain the two-dimensional feature matrix S, whose dimension size is N×D'; Then input the two-dimensional feature matrix S into a fully connected layer to obtain the feature matrix F s , whose dimension size is N×D; Finally, the feature matrix Fs Multiply by the proportional parameter α and add it to the input feature F through a residual connection l to obtain the final output F' of the local image information interaction module l ; The specific calculation formula is as follows:
[0043] F Q = Linear1(F l )
[0044] F K = Linear2(F l )
[0045] F V = Linear3(F l )
[0046]
[0047]
[0048] F s = Linear4(S)
[0049]
[0050] where Linear1(*), Linear2(*), Linear3(*) and Linear4(*) represent four fully connected layers, Softmax(·) represents the Softmax function, Transpose(·) represents the transpose operation of a two-dimensional matrix, represents matrix multiplication operation, represents matrix addition operation, α represents a learnable proportional parameter for fusion, and F' l represents the output feature of the local image information interaction module, with a size of N×D, having the same dimension as the input feature F l
[0051] In an embodiment of the present invention, the specific implementation of step S4 is as follows:
[0052] Step S41, design a no-reference image quality assessment network based on multi-region feature fusion, which consists of four adaptive feature extraction modules described in step S25, one local image information interaction module described in step S31, and a fully connected layer; take the four image patches corresponding to each distorted screen content image in the training set obtained through step S13 and As the input of the network, their dimensional sizes are all H×W×3; specifically, first, four image patches are respectively input into four adaptive feature extraction modules to extract the multi-scale text and image features of each image patch. Denote the i-th input image patch The output feature after passing through the i-th adaptive feature extraction module is F i , and their dimensional sizes are all 1×D; then, the four one-dimensional output features F i are concatenated into a two-dimensional feature vector to obtain the initial fusion feature F, whose dimensional size is N×D; then, the initial fusion feature F is input into the local image information interaction module to strengthen the information interaction between each image patch, and the final output feature F out of the network is obtained, and its size is N×D, which is the same as the dimensionality of the initial fusion feature F;
[0053] Step S42: Perform a dimensional transformation operation on the network output feature F out obtained in Step S41, flatten it into a one-dimensional feature vector, and its dimensional size changes from N×D to 1×C, where C = N×D; then, the flattened one-dimensional feature vector is input into the fully connected layer to obtain the quality evaluation score F score of the distorted screen content image; the specific calculation formula is as follows:
[0054] F score = Linear(Reshape(F out ))
[0055] where, Linear(*) represents a fully connected layer, and Reshape(·) represents a dimensional transformation operation;
[0056] Step S43: Design the loss function of the no-reference image quality assessment network based on multi-region feature fusion, which is specifically as follows:
[0057]
[0058] where, n is the number of samples in the training set, y i represents the true quality score of the i-th distorted screen content image, represents the predicted quality score of the i-th distorted screen content image output by the network;
[0059] Step S44: Repeat Steps S41 to S43 in batches until the loss value calculated in Step S43 converges and stabilizes, save the network parameters, and complete the training process of the no-reference image quality assessment network based on multi-region feature fusion.
[0060] In an embodiment of the present invention, the specific implementation of Step S5 is as follows:
[0061] For each of the distorted screen content images in the test set obtained through step S13, the corresponding four image patches and are input into the trained no-reference screen content image quality assessment model based on multi-region feature fusion, and the corresponding quality assessment scores are output.
[0062] Compared with the prior art, the present invention has the following beneficial effects: The object of the present invention is to solve the problem of quality score deviation caused by block training of the image quality assessment model based on convolutional neural network for distorted screen content images, as well as the problem of feature extraction and fusion of different types of regions in screen content images. To solve these problems, the present invention proposes a no-reference screen content image quality assessment method based on multi-region feature fusion. This method characterizes the overall quality of the image by fusing local features of multiple regions of the distorted image, and uses convolutional kernels of different sizes for adaptive feature extraction for different types of regions. In addition, this method also uses an attention mechanism to adaptively fuse different types of image features and enhance the information interaction between local image patches, thereby effectively improving the accuracy of the no-reference screen content image quality assessment model. Description of the Drawings
[0063] Figure 1 is the flowchart of the method of the embodiment of the present invention.
[0064] Figure 2 is the network model structure diagram of the embodiment of the present invention.
[0065] Figure 3 is the structure diagram of the text feature extraction sub-module of the embodiment of the present invention.
[0066] Figure 4 is the structure diagram of the image feature extraction sub-module of the embodiment of the present invention.
[0067] Figure 5 is the structure diagram of the attention feature fusion sub-module of the embodiment of the present invention.
[0068] Figure 6 is the structure diagram of the adaptive feature extraction module of the embodiment of the present invention.
[0069] Figure 7 is the structure diagram of the local image information interaction module of the embodiment of the present invention. Detailed Embodiments
[0070] The technical solutions of the present invention will be specifically described below with reference to the drawings.
[0071] The present invention provides a no-reference screen content image quality assessment method based on multi-region feature fusion. As Figure 1 shown, it includes the following steps:
[0072] Step S1: Perform data preprocessing on the data in the distorted screen content image dataset. First, crop image patches from each distorted screen content image, then divide the dataset into a training set and a test set, and finally perform data augmentation on the data in the training set.
[0073] Step S2: Design an adaptive feature extraction module, which can adaptively extract different-scale features of the text region and the image region in the distorted screen content image patches, and fuse the text region features and the image region features based on the attention mechanism.
[0074] Step S3: Design a local image information interaction module, which enhances the information interaction between any two image patches in the distorted screen content image by introducing the self-attention mechanism, and assigns different attention weights to each image patch.
[0075] Step S4: Design a no-reference image quality assessment network based on multi-region feature fusion, and train to obtain a no-reference screen content image quality assessment model.
[0076] Step S5: Input the distorted screen content image to be measured into the trained no-reference screen content image quality assessment model based on multi-region feature fusion, and output the corresponding quality assessment score.
[0077] Figure 2 The network model structure diagram constructed by the method of the present invention.
[0078] Further, step S1 includes the following steps:
[0079] Step S11: First, crop image patches from each distorted image I in the distorted screen content image dataset. Specifically, divide each distorted image I into four regions: upper left, upper right, lower left, and lower right, and then randomly crop an image patch of size H×W from each region, denoted as I 1 、I 2 、I 3 and I 4 , where H and W represent the height and width of the image patch respectively;
[0080] Step S12: Divide the images in the distorted screen content image dataset into a training set and a test set according to a certain ratio;
[0081] Step S13: Perform unified horizontal random flipping and normalization processing on the four cropped image patches of each distorted screen content image I train in the training set, so as to complete the data augmentation operation and obtain the distorted screen content image patches for training and For each distorted screen content image I in the test set test The four cropped image patches are subjected to the same normalization process to obtain the distorted screen content image patches for testing and
[0082] Furthermore, step S2 includes the following steps:
[0083] Step S21, design a text feature extraction sub-module, as Figure 3 shown. This module consists of a convolutional layer with a kernel size of 3×3, two convolutional layers with a kernel size of 1×1, two LeakyReLU activation functions, and three batch normalization layers. Since a smaller kernel can better capture local detail information in the image, such as characters, lines, etc., a convolutional layer with a kernel size of 3×3 is used to extract features from the text region in the image. Denote the input feature map of this module as F t , whose size is H t ×W t ×C t , where H t , W t and C t respectively represent the height, width, and number of channels of the input feature map F t . Specifically, first, the feature map F t is sequentially input into a convolutional layer with a kernel size of 3×3, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer for preliminary feature extraction to obtain an intermediate feature map F' t1 , whose dimension is H t ×W t ×C' t , where H t , W t and C' t respectively represent the height, width, and number of channels of the intermediate feature map F' t1 ; then the input feature map F t is sequentially input into a convolutional layer with a kernel size of 1×1 and a batch normalization layer for residual feature extraction to obtain an intermediate feature map F' t2 , whose dimension is H t ×W t ×C' t , which has the same dimension size as the intermediate feature map F' t1 ; finally, through residual connection, the intermediate feature map F' t1 is added to the feature map F' t2 , and then through the LeakyReLU activation function to obtain the output feature F' of the text feature extraction sub-module t, with dimension H t ×W t ×C' t . The specific calculation formula is as follows:
[0084] F′ t1 = BN(Conv2(LeakyReLU(BN(Conv1(F t )))))
[0085] F′ t2 = BN(Conv3(F t )))
[0086]
[0087] Among them, Conv1(*) represents a convolutional layer with a kernel size of 3×3, Conv2(*) and Conv3(*) represent two convolutional layers with a kernel size of 1×1, represents matrix addition operation, LeakyReLU(·) represents the LeakyReLU activation function, and BN(·) represents batch normalization operation;
[0088] Step S22, design an image feature extraction sub-module, as Figure 4 shown. This module consists of a convolutional layer with a kernel size of 5×5, two convolutional layers with a kernel size of 1×1, two LeakyReLU activation functions, and three batch normalization layers. Since a larger kernel is more suitable for capturing visual features of an image under a larger receptive field, such as the overall texture and color of the image, etc., a convolutional layer with a kernel size of 5×5 is used to extract features from the image region. Denote the input feature map of this module as F p , with size H p ×W p ×C p , where H p , W p and C p respectively represent the height, width, and number of channels of the input feature map F p . Specifically, first input the feature map F p sequentially into a convolutional layer with a kernel size of 5×5, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer for preliminary feature extraction to obtain an intermediate feature map F' p1 , with dimension H p ×W p ×C' p , where H p , W p and C' p respectively represent the intermediate feature map F'p1 The height, width, and number of channels; then the input feature map F p is successively input into a convolutional layer with a kernel size of 1×1 and a batch normalization layer for residual feature extraction to obtain an intermediate feature map F' p2 , whose dimension is H p ×W p ×C' p , which is the same as the dimension size of the intermediate feature map F' p1 ; finally, through residual connection, the intermediate feature map F' p1 is added to the feature map F' p2 , and then passed through the LeakyReLU activation function to obtain the output feature F' of the image feature extraction sub-module p , whose dimension is H p ×W p ×C' p . The specific calculation formula is as follows:
[0089] F′ p1 = BN(Conv2(LeakyReLU(BN(Conv1′(F p )))))
[0090] F′ p2 = BN((Conv3(F p )))
[0091]
[0092] where BN(Conv1'(*)) represents a convolutional layer with a kernel size of 5×5, BN(Conv2(*)) and Conv3(*) represent two convolutional layers with a kernel size of 1×1, represents matrix addition operation, LeakyReLU(·) represents the LeakyReLU activation function, and BN(·) represents batch normalization operation;
[0093] Step S23, design an attention feature fusion sub-module, as Figure 5 shown. This module consists of four convolutional layers with a kernel size of 1×1, a global average pooling layer, two ReLU activation functions, a Sigmoid activation function, and four batch normalization layers. The attention feature fusion sub-module can effectively fuse text features and image features at different scales through learning, making the model pay more attention to the key information of the feature map, thereby improving the generalization performance of the model. Denote the two input feature maps of this module as F' t and F' p , both of which have a size of H a ×W a ×C a , where Ha , W a , and C a represent the height, width, and number of channels of the input feature maps F' t and F' p respectively. Specifically, first, the two input feature maps are added pixel by pixel to obtain the intermediate feature map F b , whose size is H a ×W a ×C a ; then the feature map F b is respectively input into the local attention extraction branch and the global attention extraction branch for different attention feature extractions. The local attention extraction branch is sequentially composed of a convolutional layer with a kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer connected in series. The global attention extraction branch is sequentially composed of a global average pooling layer, a convolutional layer with a kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer connected in series. Denote the feature output after the intermediate feature map F b passes through the local attention extraction branch as F local , and the feature output after passing through the global attention extraction branch as F global , whose sizes are both H a ×W a ×C a ; then the feature map F local is added pixel by pixel to the feature map F global , and then the corresponding learnable weight λ is obtained through the Sigmoid function; finally, the weight λ is weighted and fused with the input feature maps F' t and F' p to obtain the final output F' b of the attention feature fusion sub-module, whose size is H a ×W a ×C a . The specific calculation formula is as follows:
[0094]
[0095] F local = BN(Conv 2_a (ReLU(BN(Conv 1_a (F b )))))
[0096] F global = BN(Conv 4_a (ReLU(BN(Conv 3_a (GAP(F b))))))
[0097]
[0098]
[0099] Among them, Conv 1_a (*), Conv 2_a (*), Conv 3_a (*) and Conv 4_a (*) represent four convolutional layers with a convolutional kernel size of 1×1. represents matrix addition operation, BN(·) represents batch normalization operation, GAP(*) represents global average pooling operation, ReLU(·) represents ReLU activation function, Sigmoid(·) represents Sigmoid activation function, and λ is the learnable weight output by the network;
[0100] Step S24: Design a channel attention sub-module to enhance feature representation and obtain key feature channel information of the input features. This module consists of two convolutional layers with a convolutional kernel size of 1×1, a ReLU activation function, and a Sigmoid activation function. Denote the input feature map of this module as F c , with a size of H c × W c × C c , where H c , W c and C c respectively represent the height, width, and number of channels of the input feature map F c . Specifically, first use the global average pooling operation to aggregate the input feature F c , then perform a dimensionality reduction operation through a convolutional layer with a size of 1×1, then perform a dimensionality increase operation through a convolutional layer with a size of 1×1, then obtain the corresponding channel attention weight through the Sigmoid function, and finally multiply the channel attention weight with the input feature F c element-wise to obtain the final output F' c of the channel attention sub-module, with a size of H c × W c × C c , which has the same dimension as the input feature map F c . The specific calculation formula is as follows:
[0101] F′ c = Sigmoid(Conv 2_b (ReLU(Conv 1_b (GAP(F c ))))) ⊙ F c
[0102] Among them, GAP(*) represents the global average pooling operation, Conv 1_b (*), and Conv 2_b (*) represent two convolutional layers with a convolutional kernel size of 1×1. "⊙" represents the matrix multiplication operation, Sigmoid(·) represents the Sigmoid activation function, and ReLU(·) represents the ReLU activation function;
[0103] Step S25: Design an adaptive feature extraction module, as Figure 6 shown. This module consists of four text feature extraction sub-modules described in step S21, four image feature extraction sub-modules described in step S22, four attention feature fusion sub-modules described in step S23, one channel attention sub-module described in step S24, and eight spatial average pooling layers with a stride of 2. The adaptive feature extraction module can perform adaptive multi-scale feature extraction on the text privileges and image features of the input screen content image block through the text feature extraction branch and the image feature extraction branch therein, and effectively fuse different types of image features through the attention mechanism to improve the modeling ability of the model. Specifically, the text feature extraction branch is serially composed of the combination of one text feature extraction sub-module and one spatial pooling layer repeated four times in sequence, and the image feature extraction branch is serially composed of the combination of one image feature extraction sub-module and one spatial pooling layer repeated four times in sequence.
[0104] Denote the size of the input screen content image block as H×W×3. First, input it into the text feature extraction branch and the image feature extraction branch respectively for multi-scale text feature and image feature extraction. Denote the multi-level text features output after the input image block passes through the text feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Denote the multi-level image features output after the input image block passes through the image feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Then, the text features and the corresponding image features They are respectively input into four attention feature fusion sub-modules to fuse text features and image features, obtaining the fused multi-level backbone features where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; then, global average pooling operations are respectively performed on the multi-level backbone features and then feature concatenation is performed along the channel direction, thereby obtaining the multi-scale text-image feature representation F' of the input image patch tp , whose size is 1×1×15C', and the specific calculation formula is as follows:
[0105]
[0106] where Concat(·) represents the feature concatenation operation and GAP(*) represents the global average pooling operation. Finally, the fused multi-scale text-image feature F' tp is input into the channel attention sub-module to capture the key information between different channels, thereby obtaining the final output feature F of the adaptive feature extraction module tp , and then the feature F tp is flattened into a one-dimensional vector with a dimension size of 1×D (where D = 960).
[0107] Furthermore, step S3 includes the following steps:
[0108] Design a local image information interaction module, as shown in Figure 7 . This module consists of four fully connected layers and a Softmax function. The local image information interaction module uses the self-attention mechanism to enhance the information interaction between the features of different image patches, so that each image patch is given different degrees of attention to better aggregate the local features of each image patch. Specifically, denote the input feature of this module as F l , whose size is N×D (where N represents the number of image patches, N = 4; D represents the dimension of each image patch, D = 960). First, the input feature F l is input into three fully connected layers, thereby generating three new intermediate features F Q , F K and F V , whose dimension sizes are all V×D' (where D' represents the second dimension of the intermediate features F Q , F K and F V , D' = 480); then, among the intermediate features FQ and F K Perform a matrix multiplication operation between and the transpose of F, and apply the Softmax function to generate the attention map A, whose dimension size is N×N; then perform a matrix multiplication operation on the intermediate feature F V and the attention map A to obtain the two-dimensional feature matrix S, whose dimension size is N×D'; then input the feature S into a fully connected layer to obtain the feature matrix F s , whose dimension size is N×d; finally multiply the feature F s by the scaling parameter α, and add it to the input feature F l through a residual connection to obtain the final output F' of the local image information interaction module l . The specific calculation formula is as follows:
[0109] F Q = Linear1(F l )
[0110] F K = Linear2(F l )
[0111] F V = Linear3(F l )
[0112]
[0113]
[0114] F s = Linear4(S)
[0115]
[0116] where, Linear1(*), Linear2(*), Linear3(*) and Linear4(*) represent four fully connected layers, Softmax(·) represents the Softmax function, Transpose(·) represents the transpose operation of a two-dimensional matrix, represents the matrix multiplication operation, represents the matrix addition operation, α represents the learnable scaling parameter for fusion, F' l represents the output feature of the local image information interaction module, whose size is N×D, and has the same dimension as the input feature F l .
[0117] Furthermore, step S4 includes the following steps:
[0118] Step S41: Design a no-reference image quality assessment network based on multi-region feature fusion. This network consists of four adaptive feature extraction modules described in Step S25, one local image information interaction module described in Step S31, and a fully connected layer. Take the four image patches corresponding to each distorted screen content image in the training set obtained through Step S13 and as the inputs of the network, and their dimension sizes are all H×W×3. Specifically, first input the four distorted image patches into the four adaptive feature extraction modules respectively to extract the multi-scale text and image features of each image patch. Denote the output feature of the i-th input image patch after passing through the i-th adaptive feature extraction module as F i (i = 1, 2, 3, 4), and their dimension sizes are all 1×D; then splice the four one-dimensional output features F i into a two-dimensional feature vector to obtain the initial fusion feature F, and its dimension size is N×D (where N represents the number of image patches, N = 4); then input the initial fusion feature F into the local image information interaction module to strengthen the information interaction between each image patch, so as to obtain the final output feature F out of the network, and its size is N×D, which is the same as the dimension of the initial fusion feature F.
[0119] Step S42: Perform a dimension transformation operation on the network output feature F out obtained through Step S41, flatten it into a one-dimensional feature vector, and its dimension size changes from N×D to 1×C (where C = N×D). Then input the flattened one-dimensional feature vector into the fully connected layer to obtain the quality evaluation score F score of the distorted screen content image. The specific calculation formula is as follows:
[0120] F score = Linear(Reshape(F out ))
[0121] where, Linear(*) represents a fully connected layer, and Reshape(·) represents a dimension transformation operation.
[0122] Step S43: Design the loss function of the no-reference image quality assessment network based on multi-region feature fusion, as shown below:
[0123]
[0124] where, n is the number of samples in the training set, y i represents the true quality score of the i-th distorted screen content image, represents the predicted quality score of the i-th distorted screen content image output by the network.
[0125] Step S44: Repeat the above steps S41 to S43 in batches until the loss value calculated in step S43 converges and stabilizes, save the network parameters, and complete the training process of the no-reference image quality assessment network based on multi-region feature fusion.
[0126] Further, step S5 is implemented as follows:
[0127] Input the four image patches corresponding to each distorted screen content image in the test set obtained through step S13 and into the trained no-reference screen content image quality assessment model based on multi-region feature fusion, and output the corresponding quality assessment scores.
[0128] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.
Claims
1. A no-reference screen content image quality assessment method based on multi-region feature fusion, characterized in that, it includes the following steps: Step S1: Perform data preprocessing on the data in the distorted screen content image dataset. First, crop image patches from each distorted screen content image, then divide the dataset into a training set and a test set, and finally perform data augmentation on the data in the training set; Step S2: Design an adaptive feature extraction module, which can adaptively extract different scale features of the text region and the image region in the distorted screen content image patch, and fuse the text region features and the image region features based on the attention mechanism; Step S3: Design a local image information interaction module, which enhances the information interaction between any two image patches in the distorted screen content image by introducing a self-attention mechanism, and assigns different attention weights to each image patch; Step S4: Design a no-reference image quality assessment network based on multi-region feature fusion, and train to obtain a no-reference screen content image quality assessment model based on multi-region feature fusion; Step S5: Input the distorted screen content image to be measured into the trained no-reference screen content image quality assessment model based on multi-region feature fusion, and output the corresponding quality assessment score; The specific implementation of step S2 is as follows: Step S21: Design a text feature extraction sub-module, which consists of a convolutional layer with a 3×3 convolutional kernel, two convolutional layers with 1×1 convolutional kernels, two LeakyReLU activation functions, and three batch normalization layers; use the convolutional layer with a 3×3 convolutional kernel to extract features from the text region in the distorted screen content image block, and denote the feature map input to the text feature extraction sub-module as F t , whose size is H t ×W t ×C t , where H t , W t and C t respectively represent the height, width, and number of channels of the input feature map F t ; specifically, first input the feature map F t into a convolutional layer with a 3×3 convolutional kernel, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a 1×1 convolutional kernel, and a batch normalization layer in sequence for preliminary feature extraction to obtain an intermediate feature map F' t1 , whose dimension is H t ×W t ×C' t , where H t , W t and C' t respectively represent the height, width, and number of channels of the intermediate feature map F' t1 ; then input the input feature map F t into a convolutional layer with a 1×1 convolutional kernel and a batch normalization layer in sequence for residual feature extraction to obtain an intermediate feature map F' t2 , whose dimension is H t ×W t ×C' t , which has the same dimension size as the intermediate feature map F' t1 ; finally, add the intermediate feature map F' t1 and the intermediate feature map F' t2 through residual connection, and then pass through the LeakyReLU activation function to obtain the output feature F' t of the text feature extraction sub-module, whose dimension is H t ×W t ×C' t ; the specific calculation formula is as follows: F′ t1 = BN(Conv2(LeakyReLU(BN(Conv1(F t ))))) F′ t2 = BN((Conv3(F t ))) Among them, Conv1(*) represents a convolutional layer with a convolutional kernel size of 3×3, Conv2(*) and Conv3(*) represent two convolutional layers with a convolutional kernel size of 1×1, represents matrix addition operation, LeakyReLU(·) represents the LeakyReLU activation function, and BN(·) represents batch normalization operation; Step S22: Design an image feature extraction sub-module, which consists of a convolutional layer with a convolutional kernel size of 5×5, two convolutional layers with a convolutional kernel size of 1×1, two LeakyReLU activation functions, and three batch normalization layers; use the convolutional layer with a convolutional kernel size of 5×5 to extract features from the image region in the distorted screen content image block; denote the feature map input to the image feature extraction sub-module as F p , whose size is H p ×W p ×C p , where H p , W p , and C p respectively represent the height, width, and number of channels of the input feature map F p ; specifically, first input the feature map F p into a convolutional layer with a convolutional kernel size of 5×5, a batch normalization layer, a LeakyReLU activation function, a convolutional layer with a convolutional kernel size of 1×1, and a batch normalization layer in sequence for preliminary feature extraction to obtain an intermediate feature map F' p1 , whose dimension is H p ×W p ×C' p , where H p , W p , and C' p respectively represent the height, width, and number of channels of the intermediate feature map F' p1 ; then input the input feature map F p into a convolutional layer with a convolutional kernel size of 1×1 and a batch normalization layer in sequence for residual feature extraction to obtain an intermediate feature map F' p2 , whose dimension is H p ×W p ×C' p , which has the same dimension size as the intermediate feature map F' p1 ; finally, add the intermediate feature map F' p1 and the intermediate feature map F' p2 through residual connection, and then obtain the output feature F' of the image feature extraction sub-module after passing through the LeakyReLU activation function p , whose dimension is H p ×W p ×C' p ; the specific calculation formula is as follows: F′ p1 = BN(Conv2(LeakyReLU(BN(Conv1′(F p ))))) F′ p2 = BN((Conv3(F p ))) Among them, Conv1'(*) represents a convolutional layer with a convolutional kernel size of 5×5, and Conv2(*) and Conv3(*) represent two convolutional layers with a convolutional kernel size of 1×1. represents matrix addition operation, LeakyReLU(·) represents the LeakyReLU activation function, and BN(·) represents batch normalization operation; Step S23: Design an attention feature fusion sub-module, which consists of four convolutional layers with a kernel size of 1×1, a global average pooling layer, two ReLU activation functions, a Sigmoid activation function, and four batch normalization layers; the attention feature fusion sub-module can learn to fuse text features and image features at different scales. Denote the two features input to the attention feature fusion sub-module as F' t and F' p , both of which have a size of H a ×W a ×C a , where H a , W a , and C a represent the height, width, and number of channels of the input feature maps F' t and F' p respectively; specifically, first add the two input features pixel by pixel to obtain an intermediate feature map F b , whose size is H a ×W a ×C a ; then input the intermediate feature map F b into the local attention extraction branch and the global attention extraction branch respectively for different attention feature extractions. The local attention extraction branch is serially connected by a convolutional layer with a kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer in sequence. The global attention extraction branch is serially connected by a global average pooling layer, a convolutional layer with a kernel size of 1×1, a batch normalization layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a batch normalization layer in sequence; denote the feature output after the intermediate feature map F b passes through the local attention extraction branch as F local , and the feature output after passing through the global attention extraction branch as F global , both of which have a size of H a ×W a ×C a ; then add the feature F local and the feature F global pixel by pixel, and then obtain the corresponding learnable weight λ through the Sigmoid function; finally, perform weighted fusion on the learnable weight λ with the input features F' t and F' p to obtain the final output F' b of the attention feature fusion sub-module, whose size is H a ×W a ×C a ; the specific calculation formula is as follows: F local = BN(Conv 2_a (ReLU(BN(Conv 1_a (F b ))))) F global = BN(Conv 4_a (ReLU(BN(Conv 3_a (GAP(F b )))))) Among them, Conv 1_a (*), Conv 2_a (*), Conv 3_a (*) and Conv 4_a (*) represent four convolutional layers with a kernel size of 1×1. represents matrix addition operation, BN(·) represents batch normalization operation, GAP(*) represents global average pooling operation, ReLU(·) represents ReLU activation function, Sigmoid(·) represents Sigmoid activation function, and λ is the learnable weight output by the network; Step S24: Design a channel attention sub-module to enhance feature representation and obtain the key feature channel information of the input features. This module consists of two convolutional layers with a kernel size of 1×1, a ReLU activation function, and a Sigmoid activation function. Denote the feature map input to the channel attention sub-module as F c , whose size is H c ×W c ×C c , where H c , W c , and C c represent the height, width, and number of channels of the input feature map F c respectively; specifically, first use global average pooling operation to aggregate the input feature F c , then perform a dimensionality reduction operation through a convolutional layer with a size of 1×1, then perform a dimensionality increase operation through a convolutional layer with a size of 1×1, then obtain the corresponding channel attention weights through the Sigmoid function, and finally multiply the channel attention weights element-wise with the input feature F c to obtain the final output F' c of the channel attention sub-module, whose size is H c ×W c ×C c , which has the same dimension as the input feature map F c ; the specific calculation formula is as follows: F′ c = Sigmoid(Conv 2_b (ReLU(Conv 1_b (GAP(F c ))))) ⊙ F c Among them, GAP(*) represents the global average pooling operation, Conv 1_b (*) and Conv 2_b (*) represent two convolutional layers with a convolutional kernel size of 1×1, "⊙" represents the matrix multiplication operation, Sigmoid(·) represents the Sigmoid activation function, and ReLU(·) represents the ReLU activation function; Step S25: Design an adaptive feature extraction module, which consists of four text feature extraction sub-modules described in step S21, four image feature extraction sub-modules described in step S22, four attention feature fusion sub-modules described in step S23, one channel attention sub-module described in step S24, and eight spatial average pooling layers with a stride of 2; The adaptive feature extraction module adaptively performs multi-scale feature extraction on the text privilege and image features of the input distorted screen content image patch through the text feature extraction branch and the image feature extraction branch therein, and fuses different types of image features through the attention mechanism; Specifically, the text feature extraction branch is sequentially composed of the serial repetition of the combination of a text feature extraction sub-module and a spatial pooling layer four times, and the image feature extraction branch is sequentially composed of the serial repetition of the combination of an image feature extraction sub-module and a spatial pooling layer four times; Denote the size of the input distorted screen content image block as H×W×3. First, input it into the text feature extraction branch and the image feature extraction branch respectively to extract multi-scale text features and image features. Denote the multi-level text features output after the input distorted screen content image block passes through the text feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Denote the multi-level image features output after the input distorted screen content image block passes through the image feature extraction branch as where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Then input the multi-level text features and the corresponding multi-level image features into four attention feature fusion sub-modules respectively to fuse the text features and image features, and obtain the fused multi-level backbone features where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64; Then perform global average pooling operations on the multi-level backbone features respectively, and then perform feature concatenation along the channel direction to obtain the multi-scale text and image feature representation F' of the input image block tp , whose size is 1×1×15C', and the specific calculation formula is as follows: Among them, Concat(·) represents the operation of feature concatenation, and GAP(*) represents the global average pooling operation; finally, the fused multi-scale text and image feature F' tp is input into the channel attention sub-module to capture the key information between different channels, and the final output feature F of the adaptive feature extraction module is obtained tp . Then, the feature F tp is flattened into a one-dimensional vector with a dimension size of 1×D, where D represents the dimension of each image patch, and D = 960.
2. The no-reference screen content image quality assessment method based on multi-region feature fusion according to claim 1, characterized in that, the specific implementation of step S1 is as follows: Step S11: First, crop image patches from each distorted screen content image I in the distorted screen content image dataset; specifically, divide each distorted screen content image I into four equal regions: upper left, upper right, lower left, and lower right, and then randomly crop an image patch of size H×W from each region, denoted as I 1 、I 2 、I 3 and I 4 , where H and W represent the height and width of the image patch, respectively; Step S12: Divide the distorted screen content images in the distorted screen content image dataset into a training set and a test set according to a predetermined ratio; Step S13: For each distorted screen content image I in the training set train Perform unified horizontal random flipping and normalization on the four cropped image patches to complete the data augmentation operation, obtaining the distorted screen content image patches for training and For each distorted screen content image I in the test set test Perform the same normalization on the four cropped image patches to obtain the distorted screen content image patches for testing and 3. The no-reference screen content image quality assessment method based on multi-region feature fusion according to claim 1, characterized in that, the specific implementation of step S3 is as follows: Design a local image information interaction module, which consists of four fully connected layers and a Softmax function; the local image information interaction module uses a self-attention mechanism to enhance the information interaction between different image patch features, so that each image patch is given different attention levels to better aggregate the local features of each image patch; specifically, denote the input feature of the local image information interaction module as F l , whose size is N×D, where N represents the number of image patches, N = 4; D represents the dimension of each image patch, D = 960; first, input the input feature F l into three fully connected layers to generate three new intermediate features F Q , F K and F V , whose dimension sizes are all N×D', where D' represents the dimension of the second dimension of the intermediate features F Q , F K and F V , D' = 480; then, perform a matrix multiplication operation between the transposes of the intermediate features F Q and F K , and use the Softmax function to generate an attention map A, whose dimension size is N×N; then perform a matrix multiplication operation on the intermediate feature F V and the attention map A to obtain a two-dimensional feature matrix S, whose dimension size is N×D'; then input the two-dimensional feature matrix S into a fully connected layer to obtain a feature matrix F s , whose dimension size is N×D; finally, multiply the feature matrix F s by the scaling parameter α and add it to the input feature Fl through a residual connection to obtain the final output F' of the local image information interaction module l ; the specific calculation formula is as follows: F Q = Linear1(F l ) F K = Linear2(F l ) F V = Linear3(F l ) F s = Linear4(S) Among them, Linear1(*), Linear2(*), Linear3(*) and Linear4(*) represent four fully connected layers, Softmax(·) represents the Softmax function, Transpose(·) represents the transpose operation of a two-dimensional matrix, represents matrix multiplication operation, represents matrix addition operation, α represents a learnable proportional parameter for fusion, F' l represents the output feature of the local image information interaction module, with a size of N×D, which is the same as the dimension of the input feature F l is the same.
4. The no-reference screen content image quality assessment method based on multi-region feature fusion according to claim 3, characterized in that, the specific implementation of step S4 is as follows: Step S41: Design a no-reference image quality assessment network based on multi-region feature fusion. This network consists of four adaptive feature extraction modules described in step S25, one local image information interaction module described in step S31, and a fully connected layer. Take the four image patches corresponding to each distorted screen content image in the training set obtained through step S13 and as the inputs of the network, and their dimension sizes are all H×W×3. Specifically, first input the four image patches into the four adaptive feature extraction modules respectively to extract the multi-scale text and image features of each image patch. Denote the output feature of the i-th input image patch after passing through the i-th adaptive feature extraction module as F i , and their dimension sizes are all 1×D. Then splice the four one-dimensional output features F i into a two-dimensional feature vector to obtain the initial fusion feature F, whose dimension size is N×D. Then input the initial fusion feature F into the local image information interaction module to strengthen the information interaction between each image patch, and obtain the final output feature F out of the network, whose size is N×D and has the same dimension as the initial fusion feature F Step S42: For the network output feature F obtained in step S41 out perform a dimensionality transformation operation to flatten it into a one-dimensional feature vector, whose dimensionality changes from N×D to 1×C, where C = N×D; then input the flattened one-dimensional feature vector into a fully connected layer to obtain the quality evaluation score F of the distorted screen content image score ; The specific calculation formula is as follows: F score = Linear(Reshape(F out )) wherein, Linear(*) represents a fully connected layer, and Reshape(·) represents a dimension transformation operation; Step S43: Design the loss function of the no-reference image quality assessment network based on multi-region feature fusion, which is specifically as follows: where n is the number of samples in the training set, and y i represents the true quality score of the i-th distorted screen content image, and represents the predicted quality score of the i-th distorted screen content image output by the network; Step S44: Repeat Steps S41 to S43 in batches until the loss value calculated in Step S43 converges and stabilizes. Save the network parameters to complete the training process of the no-reference image quality assessment network based on multi-region feature fusion.
5. The no-reference screen content image quality assessment method based on multi-region feature fusion according to claim 2, wherein, the specific implementation of Step S5 is as follows: The four image patches corresponding to each distorted screen content image in the test set obtained through step S13 and are input into the trained no-reference screen content image quality assessment model based on multi-region feature fusion to output the corresponding quality assessment scores.