Edge Feature Guided No-Reference Screen Content Image Quality Assessment Method

Through the edge feature guidance method, the Gaussian Laplace operator is used to generate edge structure diagrams, combined with multi-scale edge feature networks and feature fusion modules, the problem of shallow edge information degradation in screen content image quality evaluation is solved, and a more accurate quality evaluation is achieved.

CN115797304BActive Publication Date: 2025-07-04FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211575974.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-07-04
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing screen content image quality evaluation model based on convolutional neural networks is difficult to effectively extract the low-level detailed information and deep semantic information of screen content images during image feature extraction.

Method used

By designing a reference-free screen content image quality evaluation method based on edge feature guidance, a Gaussian Laplace operator is used to generate an edge structure diagram, combining a multi-scale edge feature guidance network, a position attention module and a progressive feature fusion module, and gradually aggregate features from top to bottom to form a multi-scale feature representation, enhancing the extraction and fusion of edge information.

Benefits of technology

The performance of the image quality evaluation model of the reference-free screen content is improved, and the quality evaluation score of the distorted screen content image can be accurately and effectively predicted, making up for the defects of the convolutional neural network in feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797304B_ABST
    Figure CN115797304B_ABST
Patent Text Reader

Abstract

The present invention relates to a no-reference screen content image quality assessment method guided by edge features, which comprises the following steps: S1: First, use the Laplacian of Gaussian operator to generate an edge structure map dataset corresponding to the distorted screen content image dataset; S2: Design a multi-scale edge feature guidance network; S3: Design a position attention module to form a global information representation of features at different scales; S4: Design a progressive feature fusion module, which gradually aggregates features at each scale in a top-down manner to form a multi-scale feature representation of the distorted image; S5: Design an image quality assessment network guided by edge features; S6: Output the quality assessment score of the distorted image. Applying the technical solution of the present invention can not only effectively extract the low-level detail information and high-level semantic information of the screen content image through networks of different depths, but also supplement the shallow edge information of the image through the edge structure map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and computer vision, and particularly to a no-reference screen content image quality assessment method guided by edge features. Background Art

[0002] In recent years, with the rapid development of multimedia technology and cloud computing, as well as the wide use of various terminal devices in daily life, a large number of screen content images generated thereby have received extensive attention. Different from traditional natural images captured from real scenes, screen content images are a type of data that contains both text and image information generated by a computer, and include various multimedia forms such as text, tables, graphics, and animations. Usually, screen content images will inevitably be interfered by various distortion factors during the processes of acquisition, compression, transmission, encoding, and display, which causes the image quality to degrade to varying degrees, ultimately affecting the user experience and the interactive performance of the system. Therefore, it is of great significance to design an effective quality evaluation method for screen content images in various processing application scenarios.

[0003] According to the distinction from the perspective of whether there is human participation in the quality evaluation process, traditional screen content image quality evaluation methods can be divided into two categories: subjective quality evaluation and objective quality evaluation. Subjective quality evaluation is that a person, as an observer, directly scores the quality of an image. This method relies on human subjective visual perception, so the results obtained are the most accurate and real. However, subjective quality evaluation also has disadvantages such as high cost and long time consumption, and because it does not have the ability to automatically score, it is often difficult to be applied to actual production. The objective quality evaluation method is that a computer automatically quantifies the feature information of an image and a video according to an algorithm designed by humans, so as to calculate the corresponding quality evaluation score, and there is no interference from human factors in this process. According to the difference in the amount of reference image information required in the objective quality evaluation process, the objective quality evaluation method can be further divided into three categories: full-reference, semi-reference, and no-reference methods, and the dependence levels of these three methods on the amount of reference image information decrease in turn. Since it is often difficult to obtain a distortion-free reference image in the real world, the no-reference method among them has stronger practicability and development prospects.

[0004] With the continuous development of deep learning technology, many screen content image quality assessment models based on convolutional neural networks have emerged. Research shows that the prediction performance of image quality evaluation models based on deep neural networks is far higher than that of traditional quality assessment methods. At the same time, using convolutional neural networks for feature extraction can break the limitations of no-reference image quality assessment for specific distortion types. Usually, quality assessment models using convolutional neural networks mostly obtain better training results by increasing the depth of the network. However, as the depth of the network continues to increase, the problem of degradation of image detail information is often inevitable. Due to the particularity of the source of screen content images, they usually have strong edge features. Therefore, it is of great significance to extract the shallow and significant edge structure information of distorted images and effectively integrate it into the convolutional neural network to provide additional information gain for model training. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a no-reference screen content image quality assessment method based on edge feature guidance. This method can use the edge structure features corresponding to distorted images as clues to guide the network to better focus on edge information, and comprehensively extract and learn image features from shallow to deep and at multiple levels. Overall, this method can effectively extract the low-level detail information and high-level semantic information of screen content images through networks of different depths, and can also use the edge structure diagram to supplement the shallow edge information of the image, making up for the defects of convolutional neural networks in feature extraction of screen content images and improving the performance of no-reference screen content image quality assessment methods.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A no-reference screen content image quality assessment method based on edge feature guidance, including the following steps:

[0007] Step S1: First, use the Laplacian of Gaussian operator to generate an edge structure diagram dataset corresponding to the distorted screen content image dataset, and then perform the same data preprocessing on the two datasets and divide them into a training set and a test set;

[0008] Step S2: Design a multi-scale edge feature guidance network. First, extract multi-level features of the distorted screen content image, and then adaptively integrate the shallow feature information from the edge structure diagram with the multi-level backbone features through multiple edge guidance feature modules;

[0009] Step S3: Design a position attention module to form a global information representation of features at different scales;

[0010] Step S4: Design a progressive feature fusion module. This module gradually aggregates features at each scale in a top-down manner to form a multi-scale feature representation of the distorted image;

[0011] Step S5: Design an edge feature-guided image quality assessment network and train it to obtain a no-reference screen content image quality assessment model;

[0012] Step S6: Input the distorted image to be measured and the corresponding edge structure diagram into the trained edge feature-guided image quality assessment model, and output the quality assessment score of the distorted image.

[0013] In a preferred embodiment, step S1 specifically includes the following steps:

[0014] Step S11: Convert each distorted image I in the distorted screen content image dataset into a grayscale image G, then construct a Laplacian of Gaussian convolution kernel with a convolution kernel size of 13×13 and a standard deviation of 1, and use the convolution kernel to perform convolution operation on the grayscale image G to generate an intermediate result image G';

[0015] Step S12: Perform thresholding on the intermediate result image G' obtained in step S11, determine the pixel points with pixel values greater than the specified threshold as edge points, set the values of the edge points to 255, and set the values of the pixel points with pixel values less than or equal to the specified threshold to 0, so as to generate an edge structure diagram I corresponding to each distorted image I edge ;

[0016] Step S13: Repeat step S11 and step S12 to obtain an edge structure diagram dataset corresponding to the distorted screen content image dataset; then perform unified random cropping, horizontal random flipping and normalization processing on the images in the two datasets, and divide the two datasets into a training set and a test set in a unified manner, that is, the corresponding images in the two datasets belong to the training set or the test set at the same time.

[0017] In a preferred embodiment, step S2 specifically includes the following steps:

[0018] Step S21: Use Res2Net-50 as the backbone network to perform multi-scale feature extraction on the distorted screen content image I with an input size of H×W×3; since the low-level image features extracted by the shallow network contain rich edge detail information, and the high-level image features extracted by the deep network contain more semantic location information, the feature maps output by the four stages in the backbone network Res2Net-50 for the distorted image I are taken as the multi-level backbone features of the distorted image; specifically, the feature maps output by the distorted image I passing through the first stage, the second stage, the third stage and the fourth stage are denoted as F1, F2, F3 and F4 respectively, where the size of the feature map F1 is The size of the feature map F2 is The size of the feature map F3 is The size of the feature map F4 is

[0019] Step S22: Obtain the edge structure diagram I corresponding to the input distorted image I edge , with a size of H×W×1. Input it into a residual module to extract shallow feature information. The residual module consists of a convolutional layer with a 3×3 convolutional kernel, a ReLU activation function, and a Sigmoid activation function; this module does not change the size of the input features, thus retaining more detailed edge information. The specific calculation formula is as follows:

[0020]

[0021] where Conv(*) represents a convolutional layer with a 3×3 convolutional kernel, represents matrix addition operation, ReLU(·) represents the ReLU activation function, Sigmoid(·) represents the Sigmoid activation function, and F edge represents the feature map obtained after passing through the residual module, with a size of H×W×1;

[0022] Step S23: Design a feature fusion sub-module, which consists of a convolutional layer with a 3×3 convolutional kernel and a bilinear interpolation downsampling layer; denote the two input feature maps of this module as F a and F b , where the size of the feature map F a is H a ×W a ×C a , and the size of the feature map F b is H b ×W b ×1; specifically, first input the feature map F b into the bilinear interpolation downsampling layer to obtain an intermediate feature map F b ′, whose dimensional size is H a ×W a ×1, having the same height and width as the input feature map F a ; then multiply the obtained intermediate feature map F b ′ element-wise with the feature map F a and add it to the input feature map F a through a residual connection to obtain an intermediate feature map F b , with a size of H a ×W a ×C a ; finally, input the feature map F b ″ into a convolutional layer with a 3×3 convolutional kernel to obtain the output feature F c of the feature fusion sub-module, with a size of H a ×Wa ×C a , which has the same dimension as the input feature map F a ; the specific calculation formula is as follows:

[0023] F b ' = D(F b )

[0024]

[0025] F c = Conv(F b )

[0026] where Conv(*) represents a convolutional layer with a convolutional kernel size of 3×3, "⊙" represents element-wise multiplication operation, represents matrix addition operation, D(*) represents bilinear interpolation downsampling layer, F b and F b " represent the intermediate feature maps of the feature fusion sub-module, and F c represents the output feature of the feature fusion sub-module;

[0027] Step S24: Design a local channel attention sub-module to enhance feature representation and obtain key feature channel information of the input feature; this module consists of a one-dimensional convolutional layer with a convolutional kernel size of 3×3, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, and a Sigmoid activation function; denote the input feature map of this module as F d , with a size of H×W×C; specifically, first use global average pooling operation to aggregate the input feature F d , then obtain the corresponding channel attention weights through one-dimensional convolution and Sigmoid function, then multiply the channel attention weights with the input feature F d element-wise, and finally reduce the number of channels through a convolutional layer with a convolutional kernel size of 1×1 to obtain the final output F d ' of the local channel attention sub-module, with a size of The specific calculation formula is as follows:

[0028] F d ' = Conv2(sigmoid(Conv1(GAP(F d )))⊙F d )

[0029] where GAP(*) represents global average pooling operation, Conv1(*) represents a one-dimensional convolutional layer with a convolutional kernel size of 3×3, Conv2(*) represents a two-dimensional convolutional layer with a convolutional kernel size of 1×1, "⊙" represents element-wise multiplication operation, Sigmoid(·) represents Sigmoid activation function, and F d' represents the output feature of the local channel attention sub-module;

[0030] Step S25: Design an edge-guided feature module, which is composed of the feature fusion sub-module described in step S23 and the local channel attention sub-module described in step S24 connected in series;

[0031] Step S26: Use the edge-guided feature module designed in step S25 to adaptively integrate the shallow feature information from the edge structure diagram and the multi-level backbone features of the distorted image, thereby enhancing the edge information representation of the distorted screen content image and guiding the network to better focus on the edge information; specifically, the multi-level backbone features F of the distorted image obtained through step S21 i , i = 1, 2, 3, 4, and the edge structure features F obtained through step S22 edge are respectively input into four edge-guided feature modules to obtain multi-level backbone features that fuse shallow edge feature information where the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of C' = 64;

[0032] In a preferred embodiment, in step S3, a position attention module is designed to form a global information representation of features at different scales; it includes the following steps:

[0033] Step S31: Design a position attention module, which consists of three convolutional layers with a kernel size of 1×1 and a Softmax function; denote the input feature map of this module as F p , with a size of H×W×C; specifically, first input the feature map F p into three convolutional layers with a kernel size of 1×1 to generate three new intermediate feature maps F Q , F K and F V ; then perform a dimensional transformation operation on the intermediate feature maps F Q , F K and F V to change the feature dimensions, and their dimension sizes are all changed from H×W×C to C×N, where N = H×W; then, perform a matrix multiplication operation between the transposes of the feature maps F K and F Q and use the Softmax function to generate a spatial attention map S with a size of N×N; then in the intermediate feature map F VPerform a matrix multiplication operation between the spatial attention map S and the two-dimensional feature matrix S is obtained i , which has a size of C×N. Then, the feature dimension is changed through a dimension transformation operation, and its dimension size changes from C×N to H×W×C. Finally, the feature S i is multiplied by the scaling parameter α and added to the input feature map F p through a residual connection to obtain the final output F′ of the position attention module. The specific calculation formula is as follows:

[0034] F Q = Reshape(Conv1(F p ))

[0035] F K = Reshape(Conv2(F p ))

[0036] F V = Reshape(Conv3(F p ))

[0037]

[0038]

[0039]

[0040] Among them, Conv1(*), Conv2(*), and Conv3(*) represent three convolutional layers with a kernel size of 1×1. Reshape(·) represents a dimension transformation operation, Softmax(·) represents the Softmax function, Transpose(·) represents the transpose operation of a two-dimensional matrix, represents a matrix multiplication operation, represents a matrix addition operation, α represents the fusion scaling parameter, F′ represents the output feature of the position attention module, which has a size of H×W×C and the same dimension as the input feature map F p ;

[0041] Step S32: Use the position attention module designed in step S31 to model the rich context relationship of the input features, thereby enhancing its representation ability. Specifically, first, the multi-level backbone features F i a (i = 1, 2, 3, 4) obtained in step S26 are respectively input into four position attention modules to obtain multi-level backbone features F i ′(i = 1, 2, 3, 4) with global context information. Among them, the feature F i' has the same dimensionality as the input multi-level backbone feature The dimensionality is the same.

[0042] 4. According to the method for evaluating the quality of a screen content image without reference guided by edge features as described above, it is characterized in that step S4 specifically includes the following steps:

[0043] Step S41: Input the first-level backbone feature F1' obtained through step S32 into a convolutional layer with a convolutional kernel size of 3×3, and adjust the width and height of the feature map F1' to half of the original. The specific calculation formula is as follows:

[0044]

[0045] where Conv1_1(*) represents a convolutional layer with a convolutional kernel size of 3×3, represents the output after the first-level backbone feature F1' passes through the convolutional layer, and its size is

[0046] Step S42: Input the second-level backbone feature F2' obtained through step S32 into a convolutional layer with a convolutional kernel size of 3×3 to obtain a feature whose dimensionality is the same as that of F2'; then input the feature obtained through step S41 and the feature into the context aggregation sub-module to fuse features from different scales; the context aggregation sub-module consists of a convolutional layer with a convolutional kernel size of 1×1, a batch normalization layer, and a ReLU activation function; finally, input the fused feature into a convolutional layer with a convolutional kernel size of 3×3 to adjust the width and height of the feature map to half of the original. The specific calculation formula is as follows:

[0047]

[0048]

[0049]

[0050] where Conv2_1(*) and Conv2_3(*) represent two convolutional layers with a convolutional kernel size of 3×3, Conv2_2(*) represents a convolutional layer with a convolutional kernel size of 1×1, represents matrix addition operation, BN(*) represents batch normalization operation, ReLU(·) represents ReLU activation function, represents the output after the second-level backbone feature F2' passes through two convolutional layers and the context aggregation sub-module, and its size is

[0051] Step S43: Input the third-level backbone feature F3′ obtained in Step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimensional size is the same as that of F3′; then input the feature obtained in Step S42 and the feature into the above context aggregation sub-module to obtain a fused feature Finally, input the feature into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map to half of the original. The specific calculation formula is as follows:

[0052]

[0053]

[0054]

[0055] where Conv3_1(*) and Conv3_3(*) represent two convolutional layers with a kernel size of 3×3, Conv3_2(*) represents a convolutional layer with a kernel size of 1×1, represents matrix addition operation, BN(*) represents batch normalization operation, ReLU(·) represents ReLU activation function, represents the output after the third-level backbone feature F3′ passes through two convolutional layers and the context aggregation sub-module, and its size is

[0056] Step S44: Input the fourth-level backbone feature F4 obtained in Step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimensional size is the same as that of F4′; then input the feature obtained in Step S43 and the feature into the above context aggregation sub-module to obtain a fused feature Finally, input the feature into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map to half of the original. The specific calculation formula is as follows:

[0057]

[0058]

[0059]

[0060] Among them, Conv4_1(*) and Conv4_3(*) represent two convolutional layers with a convolutional kernel size of 3×3, and Conv4_2(*) represents a convolutional layer with a convolutional kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. represents the output after the fourth-level backbone feature F4′ passes through two convolutional layers and the context aggregation sub-module, and its size is

[0061] Step S45: For the feature vector obtained through step S41 the feature vector obtained through step S42 the feature vector obtained through step S43 and the feature vector obtained through step S44 perform feature fusion to form a multi-scale feature representation of the input distorted image; specifically, for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×2C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×4C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×8C′, and then splice the feature vectors and obtained after performing the global average pooling operation along the channels to obtain a multi-scale feature representation of the input distorted image, and its size is 1×1×15C′. The calculation formula is as follows:

[0062]

[0063] Among them, Concat(·) represents the splicing operation of features, GAP(*) represents the global average pooling operation, and F represents the multi-scale feature representation of the input distorted image.

[0064] In a preferred embodiment, step S5 specifically includes the following steps:

[0065] Step S51: Input the multi-scale encoded representation F of the input distorted image obtained through step S45 into the fully connected layer MLP to obtain the quality evaluation score F of the distorted image score , and its calculation formula is:

[0066] F score= MLP(F)

[0067] Step S52: Design the loss function of the image quality assessment network guided by edge features as follows:

[0068]

[0069] where m is the number of samples in each training batch, and y i represents the true quality score of the i-th image sample, and represents the predicted quality score of the i-th image sample obtained through the network;

[0070] Step S53: Repeat the above steps S51 to S52 in batches until the loss value calculated in step S52 converges and stabilizes. Save the network parameters to complete the training process of the image quality assessment network guided by edge features and obtain the image quality assessment model guided by edge features.

[0071] In a preferred embodiment, step S6 specifically includes the following steps:

[0072] Step Sδ1: Input the distorted screen content image and the corresponding edge structure diagram in the test set into the trained image quality assessment model guided by edge features to output the corresponding quality assessment score.

[0073] Compared with the prior art, the present invention has the following beneficial effects: The objective of the present invention is to solve the problem of the degradation of shallow edge information caused by the deepening of the network layer in the image feature extraction process of the screen content image quality assessment model based on the convolutional neural network. By using the edge structure diagram features corresponding to the distorted image as a clue to guide the network to better focus on the edge information, the image features are comprehensively extracted and learned from shallow to deep and multi-level, which can improve the performance of the no-reference screen content image quality assessment model. The present invention proposes a no-reference screen content image quality assessment method guided by edge features. This method can effectively extract the low-level detail information and high-level semantic information of the screen content image through networks of different depths, and can also supplement the shallow edge information of the image through the edge structure diagram, so as to accurately and effectively predict the quality assessment score of the distorted screen content image. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 is the flowchart of the method of the preferred embodiment of the present invention.

[0075] Figure 2 is the network model structure diagram of the preferred embodiment of the present invention.

[0076] Figure 3 is the edge-guided feature module structure diagram of the preferred embodiment of the present invention.

[0077] Figure 4 This is the structural diagram of the position attention module in the preferred embodiment of the present invention.

[0078] Figure 5 This is the structural diagram of the context aggregation sub-module in the progressive feature fusion module of the preferred embodiment of the present invention. Detailed implementation manners

[0079] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0080] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0081] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0082] The present invention provides a no-reference screen content image quality assessment method based on edge feature guidance, as Figures 1-5 shown, including the following steps:

[0083] Step S1: First, use the Laplacian of Gaussian operator to generate an edge structure diagram dataset corresponding to the distorted screen content image dataset, and then perform the same data preprocessing on the two datasets and divide them into a training set and a test set;

[0084] Step S2: Design a multi-scale edge feature guidance network. First, extract multi-level features of the distorted screen content image, and then adaptively integrate the shallow feature information from the edge structure diagram with the multi-level backbone features through multiple edge guidance feature modules;

[0085] Step S3: Design a position attention module to form a global information representation of features at different scales;

[0086] Step S4: Design a progressive feature fusion module, which gradually aggregates features at each scale in a top-down manner to form a multi-scale feature representation of the distorted image;

[0087] Step S5: Design an image quality assessment network based on edge feature guidance and train it to obtain a no-reference screen content image quality assessment model;

[0088] Step S6: Input the distorted image to be measured and the corresponding edge structure diagram into the trained image quality assessment model guided by edge features, and output the quality assessment score of the distorted image.

[0089] Further, step S1 includes the following steps:

[0090] Step S11: Convert each distorted image I in the distorted screen content image dataset into a grayscale image G, then construct a Laplacian of Gaussian convolution kernel with a convolution kernel size of 13×13 and a standard deviation of 1, and use the convolution kernel to perform a convolution operation on the grayscale image G to generate an intermediate result image G';

[0091] Step S12: Perform thresholding on the intermediate result image G' obtained in step S11, determine the pixel points with pixel values greater than the specified threshold as edge points, set the values of the edge points to 255, and set the values of the pixel points with pixel values less than or equal to the specified threshold to 0, so as to generate the edge structure diagram I corresponding to each distorted image I edge ;

[0092] Step S13: Repeat step S11 and step S12 to obtain an edge structure diagram dataset corresponding to the distorted screen content image dataset. Then perform unified random cropping, horizontal random flipping, and normalization processing on the images in the two datasets, and divide the two datasets into a training set and a test set in a unified manner, that is, the corresponding images in the two datasets belong to the training set or the test set at the same time.

[0093] Further, step S2 includes the following steps:

[0094] Step S21: Use Res2Net-50 as the backbone network to perform multi-scale feature extraction on the distorted screen content image I with an input size of H×W×3. Since the low-level image features extracted by the shallow network contain rich edge detail information, and the high-level image features extracted by the deep network contain more semantic location information, the feature maps output by the four stages in the backbone network Res2Net-50 for the distorted image I are taken as the multi-level backbone features of the distorted image. Specifically, the feature maps output by the distorted image I passing through the first stage, the second stage, the third stage, and the fourth stage are denoted as F1, F2, F3, and F4 respectively, where the size of the feature map F1 is The size of the feature map F2 is The size of the feature map F3 is The size of the feature map F4 is

[0095] Step S22: Take the edge structure diagram I corresponding to the input distorted image I edge, with a size of H×W×1, is input into a residual module to extract shallow feature information. The residual module consists of a convolutional layer with a 3×3 convolutional kernel, a ReLU activation function, and a Sigmoid activation function. This module does not change the size of the input features, thus retaining more detailed edge information. The specific calculation formula is as follows:

[0096]

[0097] Among them, Conv(*) represents a convolutional layer with a 3×3 convolutional kernel, represents matrix addition operation, ReLU(·) represents the ReLU activation function, Sigmoid(·) represents the Sigmoid activation function, and F edge represents the feature map obtained after passing through the residual module, with a size of H×W×1;

[0098] Step S23: Design a feature fusion sub-module, which consists of a convolutional layer with a 3×3 convolutional kernel and a bilinear interpolation downsampling layer. Denote the two input feature maps of this module as F a and F b , where the size of the feature map F a is H a ×W a ×C a , and the size of the feature map F b is H b ×W b ×1. Specifically, first input the feature map F b into the bilinear interpolation downsampling layer to obtain an intermediate feature map F′ b , whose dimensional size is H a ×W a ×1, having the same height and width as the input feature map F a ; then multiply the obtained intermediate feature map F′ b element-wise with the feature map F a and add it to the input feature map F a through a residual connection to obtain an intermediate feature map F″ b , with a size of H a ×W a ×C a ; finally, input the feature map F″ b into a convolutional layer with a 3×3 convolutional kernel to obtain the output feature F of the feature fusion sub-module c , with a size of H a ×W a ×C a , having the same dimension as the input feature map F a . The specific calculation formula is as follows:

[0099] F′ b = D(F b )

[0100]

[0101] F c = Conv(F″ b )

[0102] where Conv(*) represents a convolutional layer with a convolutional kernel size of 3×3, and "⊙" represents an element-wise multiplication operation. represents matrix addition operation, D(*) represents a bilinear interpolation downsampling layer, F′ b and F″ b represent intermediate feature maps of the feature fusion sub-module, and F c represents the output feature of the feature fusion sub-module;

[0103] Step S24: Design a local channel attention sub-module to enhance feature representation and obtain key feature channel information of the input feature. This module consists of a one-dimensional convolutional layer with a convolutional kernel size of 3×3, a two-dimensional convolutional layer with a convolutional kernel size of 1×1, and a Sigmoid activation function. Denote the input feature map of this module as F d , with a size of H×W×C. Specifically, first use global average pooling operation to aggregate the input feature F d , then obtain the corresponding channel attention weights through one-dimensional convolution and the Sigmoid function, then multiply the channel attention weights with the input feature F d element-wise, and finally reduce the number of channels through a convolutional layer with a convolutional kernel size of 1×1 to obtain the final output F′ d of the local channel attention sub-module, with a size of The specific calculation formula is as follows:

[0104] F′ d = Conv2(Sigmoid(Conv1(GAP(F d )))⊙F d )

[0105] where GAP(*) represents global average pooling operation, Conv1(*) represents a one-dimensional convolutional layer with a convolutional kernel size of 3×3, Conv2(*) represents a two-dimensional convolutional layer with a convolutional kernel size of 1×1, "⊙" represents element-wise multiplication operation, Sigmoid(·) represents the Sigmoid activation function, and F′ d represents the output feature of the local channel attention sub-module;

[0106] Step S25: Design an edge guidance feature module, which is composed of the feature fusion sub-module described in step S23 and the local channel attention sub-module described in step S24 connected in series;

[0107] Step S26: Use the edge guidance feature module designed in step S25 to adaptively integrate the shallow feature information from the edge structure diagram and the multi-level backbone features of the distorted image, thereby enhancing the edge information representation of the distorted screen content image and guiding the network to better focus on the edge information. Specifically, input the multi-level backbone features F i (i = 1, 2, 3, 4) of the distorted image obtained through step S21 and the edge structure features F edge obtained through step S22 into four edge guidance feature modules respectively to obtain the multi-level backbone features that fuse the shallow edge feature information where the feature map has a size of the feature map has a size of the feature map has a size of the feature map has a size of

[0108] Furthermore, step S3 includes the following steps:

[0109] Step S31: Design a position attention module, which is composed of three convolutional layers with a kernel size of 1×1 and a Softmax function. Denote the input feature map of this module as F p , with a size of H×W×C. Specifically, first input the feature map F p into the three convolutional layers with a kernel size of 1×1 to generate three new intermediate feature maps F Q , F K and F V ; then perform a dimension transformation operation on the intermediate feature maps F Q , F K and F V to change the feature dimension, and their dimension sizes change from H×W×C to C×N (where N = H×W). After that, perform a matrix multiplication operation between the transposes of the feature maps F K and F Q , and use the Softmax function to generate a spatial attention map S, with a size of N×N. Then perform a matrix multiplication operation between the intermediate feature map F V and the spatial attention map S to obtain a two-dimensional feature matrix S i , with a size of C×N. Then, change the feature dimension through a dimension transformation operation, and its dimension size changes from C×N to H×W×C. Finally, the feature S iMultiply by the proportionality parameter α and add it to the input feature map F through a residual connection p to obtain the final output F′ of the position attention module. The specific calculation formula is as follows:

[0110] F Q = Reshape(Conv1(F p ))

[0111] F K = Reshape(Conv2(F p ))

[0112] F V = Reshape(Conv3(F p ))

[0113]

[0114]

[0115]

[0116] where Conv1(*), Conv2(*), and Conv3(*) represent three convolutional layers with a kernel size of 1×1, Reshape(·) represents a dimensional transformation operation, Softmax(·) represents the Softmax function, Transpose(·) represents the transpose operation of a two-dimensional matrix, represents matrix multiplication operation, represents matrix addition operation, α represents the fusion proportionality parameter, F′ represents the output feature of the position attention module, with a size of H×W×C, having the same dimension as the input feature map F p ;

[0117] Step S32: Use the position attention module designed in Step S31 to model the rich context relationship of the input features, thereby enhancing its representation ability. Specifically, first input the multi-level backbone features that fuse the shallow edge feature information obtained in Step S26 into four position attention modules respectively to obtain the multi-level backbone features F′ i (i = 1, 2, 3, 4), where the feature F′ i output by the position attention module has the same dimension size as the input multi-level backbone features .

[0118] Furthermore, Step S4 includes the following steps:

[0119] Step S41: Input the first-level backbone feature F′1 obtained in step S32 into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map F′1 to half of the original. The specific calculation formula is as follows:

[0120]

[0121] where Conv1_1(*) represents a convolutional layer with a kernel size of 3×3, represents the output after the first-level backbone feature F′1 passes through the convolutional layer, and its size is

[0122] Step S42: Input the second-level backbone feature F′2 obtained in step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimension size is the same as that of F′2; then input the feature obtained in step S41 and the feature into the context aggregation sub-module to fuse features from different scales. The context aggregation sub-module consists of a convolutional layer with a kernel size of 1×1, a batch normalization layer, and a ReLU activation function. Finally, input the fused feature into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map to half of the original. The specific calculation formula is as follows:

[0123]

[0124]

[0125]

[0126] where Conv2_1(*) and Conv2_3(*) represent two convolutional layers with a kernel size of 3×3, Conv2_2(*) represents a convolutional layer with a kernel size of 1×1, represents matrix addition operation, BN(*) represents batch normalization operation, ReLU(·) represents ReLU activation function, represents the output after the second-level backbone feature F′2 passes through two convolutional layers and the context aggregation sub-module, and its size is

[0127] Step S43: Input the third-level backbone feature F′3 obtained in step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimension size is the same as that of F′3; then input the feature obtained in step S42 and the feature Input it into the above context aggregation sub-module to obtain the fused features Finally, the features are input into a convolutional layer with a kernel size of 3×3 to adjust the feature map to half of its original width and height. The specific calculation formula is as follows:

[0128]

[0129]

[0130]

[0131] Among them, Conv3_1(*) and Conv3_3(*) represent two convolutional layers with a kernel size of 3×3, and Conv3_2(*) represents a convolutional layer with a kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. represents the output of the third-level backbone feature F′3 after passing through two convolutional layers and the context aggregation sub-module, and its size is

[0132] Step S44: Input the fourth-level backbone feature F′4 obtained in step S32 into a convolutional layer with a kernel size of 3×3 to obtain features whose dimensional size is the same as F′4; then input the features obtained in step S43 and the features into the above context aggregation sub-module to obtain the fused features Finally, the features are input into a convolutional layer with a kernel size of 3×3 to adjust the feature map to half of its original width and height. The specific calculation formula is as follows:

[0133]

[0134]

[0135]

[0136] Among them, Conv4_1(*) and Conv4_3(*) represent two convolutional layers with a kernel size of 3×3, and Conv4_2(*) represents a convolutional layer with a kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. It represents the output of the fourth-level backbone feature F′4 after passing through two convolutional layers and the context aggregation sub-module, and its size is

[0137] Step S45: For the feature vectors obtained through Step S41 the feature vectors obtained through Step S42 the feature vectors obtained through Step S43 and the feature vectors obtained through Step S44 perform feature fusion to form a multi-scale feature representation of the input distorted image. Specifically, for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×2C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×4C′; for the feature vector perform global average pooling operation, and its dimension size changes from to 1×1×8C′, and then concatenate the feature vectors and along the channels to obtain a multi-scale feature representation of the input distorted image, and its size is 1×1×15C′. The calculation formula is as follows:

[0138]

[0139] where Concat(·) represents the feature concatenation operation, GAP(*) represents the global average pooling operation, and F represents the multi-scale feature representation of the input distorted image.

[0140] Furthermore, Step S5 includes the following steps:

[0141] Step S51: Input the multi-scale encoded representation F of the input distorted image obtained through Step S45 into a fully connected layer (denoted as MLP) to obtain the quality evaluation score F score of the distorted image, and its calculation formula is:

[0142] F score = MLP(F)

[0143] Step S52: Design the loss function of the image quality evaluation network based on edge feature guidance as follows:

[0144]

[0145] where m is the number of samples in each training batch, and y i represents the true quality score of the i-th image sample, and

[0146] represents the predicted quality score obtained by the network for the i-th image sample;

[0147] Further, step S6 includes the following steps:

[0148] Input the distorted screen content images and the corresponding edge structure diagrams in the test set into the trained edge feature-guided image quality assessment model, and output the corresponding quality assessment scores.

[0149] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functions and effects produced do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.

Claims

1. A no-reference screen content image quality assessment method guided by edge features, characterized in that It includes the following steps: Step S1: First, use the Laplacian of Gaussian operator to generate an edge structure map dataset corresponding to the distorted screen content image dataset. Then, perform the same data preprocessing on the two datasets and divide them into a training set and a test set; Step S2: Design a multi-scale edge feature-guided network. First, extract multi-level features of the distorted screen content image, and then adaptively integrate the shallow feature information from the edge structure map with the multi-level backbone features through multiple edge-guided feature modules; Step S3: Design a position attention module to form a global information representation of features at different scales; Step S4: Design a progressive feature fusion module, which gradually aggregates features at each scale in a top-down manner to form a multi-scale feature representation of the distorted image; Step S5: Design an edge feature-guided image quality assessment network and train it to obtain a no-reference screen content image quality assessment model; Step S6: Input the distorted image to be measured and the corresponding edge structure map into the trained edge feature-guided image quality assessment model, and output the quality assessment score of the distorted image.

2. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 1, characterized in that The specific steps of step S1 include the following steps: Step S11: Convert each distorted image I in the distorted screen content image dataset into a grayscale image G. Then, construct a Laplacian of Gaussian convolution kernel with a convolution kernel size of 13×13 and a standard deviation of 1, and use the convolution kernel to perform a convolution operation on the grayscale image G to generate an intermediate result image G'; Step S12: Perform thresholding on the intermediate result graph G' obtained in step S11. Determine the pixel points with pixel values greater than the specified threshold as edge points, set the value of the edge points to 255, and set the value of the pixel points with pixel values less than or equal to the specified threshold to 0, so as to generate the edge structure diagram I corresponding to each distorted image I edge ; Step S13: Repeat step S11 and step S12 to obtain an edge structure map dataset corresponding to the distorted screen content image dataset; then perform unified random cropping, horizontal random flipping, and normalization processing on the images in the two datasets, and divide the two datasets into a training set and a test set in a unified manner, that is, the corresponding images in the two datasets belong to the training set or the test set at the same time.

3. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 1, wherein The specific steps of step S2 include the following steps: Step S21: Using Res2Net-50 as the backbone network, perform multi-scale feature extraction on the distorted screen content image I with an input size of H×W×3. Since the low-level image features extracted by the shallow network contain rich edge detail information, while the high-level image features extracted by the deep network contain more semantic location information, the feature maps output by the four stages in the backbone network Res2Net-50 for the distorted image I are taken as the multi-level backbone features of the distorted image. Specifically, denote the feature maps output by the distorted image I after the first stage, the second stage, the third stage, and the fourth stage as F1, F2, F3, and F4 respectively, where the size of the feature map F1 is The size of the feature map F2 is The size of the feature map F3 is The size of the feature map F4 is C = 256; Step S22: Obtain the edge structure diagram I corresponding to the input distorted image I edge , with a size of H×W×1, and input it into a residual module to extract shallow feature information. The residual module consists of a convolutional layer with a convolutional kernel size of 3×3, a ReLU activation function, and a Sigmoid activation function; this module does not change the size of the input features, thereby retaining more detailed edge information. The specific calculation formula is as follows: Among them, Conv(*) represents a convolutional layer with a convolutional kernel size of 3×3, represents matrix addition operation, ReLU(·) represents ReLU activation function, Sigmoid(·) represents Sigmoid activation function, F edge represents the feature map obtained after passing through the residual module, and its size is H×W×1; Step S23: Design a feature fusion sub-module, which consists of a convolutional layer with a 3×3 convolutional kernel and a bilinear interpolation downsampling layer. Denote the two input feature maps of this module as F a and F b , where the size of feature map F a is H a ×W a ×C a , and the size of feature map F b is H b ×W b ×1. Specifically, first input feature map F b into the bilinear interpolation downsampling layer to obtain an intermediate feature map F' b , whose dimension size is H a ×W a ×1, having the same height and width as the input feature map F a . Then multiply the obtained intermediate feature map F' b element-wise with feature map F a , and add it to the input feature map F a through a residual connection to obtain an intermediate feature map F'' b , whose size is H a ×W a ×C a . Finally, input feature map F'' b into a convolutional layer with a 3×3 convolutional kernel to obtain the output feature F c of the feature fusion sub-module, whose size is H a ×W a ×C a , having the same dimension as the input feature map F a . The specific calculation formula is as follows: F′ b = D(F b ) F c = Conv(F″ b ) Among them, Conv(*) represents a convolutional layer with a convolutional kernel size of 3×3, "⊙" represents the element-wise multiplication operation, represents the matrix addition operation, D(*) represents the bilinear interpolation downsampling layer, F′ b and F″ b represent the intermediate feature maps of the feature fusion sub-module, F c represents the output feature of the feature fusion sub-module; Step S24: Design a local channel attention sub-module to enhance feature representation and obtain the key feature channel information of the input features. This module consists of a one-dimensional convolutional layer with a 3×3 convolutional kernel, a two-dimensional convolutional layer with a 1×1 convolutional kernel, and a Sigmoid activation function. Denote the feature map input to this module as F d , whose size is H×W×C. Specifically, first use global average pooling operation to aggregate the input feature F d , then obtain the corresponding channel attention weights through one-dimensional convolution and the Sigmoid function, then multiply the channel attention weights with the input feature F d element-wise, and finally reduce the number of channels through a convolutional layer with a 1×1 convolutional kernel to obtain the final output F′ of the local channel attention sub-module d , whose size is The specific calculation formula is as follows: F′ d = Conv1_3(Sigmoid(Conv1_2(GAP(F d ))) ⊙ F d ) Among them, GAP(*) represents the global average pooling operation, Conv1_2(*) represents a one-dimensional convolutional layer with a convolutional kernel size of 3×3, Conv1_3(*) represents a two-dimensional convolutional layer with a convolutional kernel size of 1×1, "⊙" represents the element-wise multiplication operation, Sigmoid(·) represents the Sigmoid activation function, and F′ d represents the output feature of the local channel attention sub-module; Step S25: Design an edge-guided feature module, which is composed of the feature fusion sub-module described in step S23 and the local channel attention sub-module described in step S24 connected in series; Step S26: Use the edge guidance feature module designed in Step S25 to adaptively integrate the shallow feature information from the edge structure diagram with the multi-level backbone features of the distorted image, thereby enhancing the edge information representation of the distorted screen content image and guiding the network to better focus on edge information. Specifically, the multi-level backbone features F i , where i = 1, 2, 3, 4, obtained in Step S21, and the edge structure features F edge obtained in Step S22 are respectively input into four edge guidance feature modules to obtain the multi-level backbone features that fuse the shallow edge feature information where the size of the feature map is the size of the feature map is the size of the feature map is the size of the feature map is C' = 64.

4. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 3, wherein The specific steps of step S3 include the following steps: Step S31: Design a location attention module, which consists of three convolutional layers with a convolutional kernel size of 1×1 and a Softmax function; denote the feature map input to this module as F p , with a size of H×W×C; specifically, first input the feature map F p into three convolutional layers with a convolutional kernel size of 1×1 to generate three new intermediate feature maps F Q , F K and F V ; then perform a dimension transformation operation on the intermediate feature maps F Q , F K and F V to change the feature dimensions, and their dimension sizes are all changed from H×W×C to C×N, where N = H×W; then, perform a matrix multiplication operation between the transpose of the feature maps F K and F Q , and use the Softmax function to generate a spatial attention map S, with a size of N×N; then perform a matrix multiplication operation between the intermediate feature map F V and the spatial attention map S to obtain a two-dimensional feature matrix S i , with a size of C×N, then change the feature dimensions through a dimension transformation operation, and its dimension size is changed from C×N to H×W×C. Finally, multiply the feature S i by the scaling parameter α and add it to the input feature map F p through a residual connection to obtain the final output F′ of the location attention module; the specific calculation formula is as follows: F Q = Reshape(Conv1_4(F p )) F K = Reshape(Conv1_5(F p )) F V = Reshape(Conv1_6(F p )) Among them, Conv1_4(*), Conv1_5(*), and Conv1_6(*) represent three convolutional layers with a convolutional kernel size of 1×1, Resh ape(·) represents a dimensional transformation operation, Softmax(·) represents the Softmax function, and Transpose(·) represents the transpose operation of a two-dimensional matrix. represents matrix multiplication operation. represents matrix addition operation. α represents the fusion ratio parameter, and F′ represents the output feature of the position attention module, with a size of H×W×C, which has the same dimension as the input feature map F. p is the same. Step S32: Use the position attention module designed in Step S31 to model the rich context relationship of the input features, thereby enhancing its representation ability; specifically, first, the multi-level backbone features that fuse the shallow edge feature information obtained in Step S26 are respectively input into four position attention modules to obtain multi-level backbone features F′ i with global context information, where i = 1, 2, 3, 4. Among them, the feature F′ i output by the position attention module has the same dimensional size as the input multi-level backbone features.

5. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 4, wherein The specific steps of step S4 include the following steps: Step S41: Input the first-level backbone feature F'1 obtained after step S32 into a convolutional layer with a convolution kernel size of 3×3 to adjust the width and height of the feature map F'1 to half of the original. The specific calculation formula is as follows: Among them, Conv1_1(*) represents a convolutional layer with a convolutional kernel size of 3×3, represents the output after the first-level backbone feature F′1 passes through the convolutional layer, and its size is Step S42: Input the second-level backbone feature F′2 obtained in step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimensionality is the same as that of F′2; then input the feature obtained in step S41 and the feature into the context aggregation sub-module to fuse features from different scales; the context aggregation sub-module consists of a convolutional layer with a kernel size of 1×1, a batch normalization layer, and a ReLU activation function; finally, input the fused feature into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map to half of the original, and the specific calculation formula is as follows: Among them, Conv2_1(*) and Conv2_3(*) represent two convolutional layers with a convolutional kernel size of 3×3, and Conv2_2(*) represents a convolutional layer with a convolutional kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. represents the output of the second-level backbone feature F′2 after passing through two convolutional layers and the context aggregation sub-module, and its size is Step S43: Input the third-level backbone feature F′3 obtained in Step S32 into a convolutional layer with a kernel size of 3×3 to obtain a feature whose dimensionality is the same as that of F′3; then input the feature obtained in Step S42 and the feature into the above context aggregation sub-module to obtain a fused feature Finally, input the feature into a convolutional layer with a kernel size of 3×3 to adjust the width and height of the feature map to half of the original, and the specific calculation formula is as follows: Among them, Conv3_1(*) and Conv3_3(*) represent two convolutional layers with a convolutional kernel size of 3×3, and Conv3_2(*) represents a convolutional layer with a convolutional kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. represents the output of the third-stage backbone feature F′3 after passing through two convolutional layers and the context aggregation sub-module, and its size is Step S44: Input the fourth-level backbone feature F′4 obtained in step S32 into a convolutional layer with a convolutional kernel size of 3×3 to obtain a feature whose dimensional size is the same as that of F′4; then input the feature obtained in step S43 and the feature into the above context aggregation sub-module to obtain a fused feature Finally, input the feature into a convolutional layer with a convolutional kernel size of 3×3 to adjust the width and height of the feature map Among them, Conv4_1(*) and Conv4_3(*) represent two convolutional layers with a convolutional kernel size of 3×3, and Conv4_2(*) represents a convolutional layer with a convolutional kernel size of 1×1. represents matrix addition operation, BN(*) represents batch normalization operation, and ReLU(·) represents ReLU activation function. represents the output of the fourth-level backbone feature F′4 after passing through two convolutional layers and the context aggregation sub-module, and its size is Step S45: Perform feature fusion on the feature vectors obtained through step S41 the feature vectors obtained through step S42 the feature vectors obtained through step S43 and the feature vectors obtained through step S44 to form a multi-scale feature representation of the input distorted image; specifically, perform global average pooling on the feature vector such that its dimension size changes from to 1×1×C′; perform global average pooling on the feature vector such that its dimension size changes from to 1×1×2C′; perform global average pooling on the feature vector such that its dimension size changes from to 1×1×4C′; perform global average pooling on the feature vector such that its dimension size changes from to 1×1×8C', and then concatenate the feature vectors and along the channels to obtain a multi-scale feature representation of the input distorted image, with a size of 1×1×15C', and the calculation formula is as follows: Where Concat(·) represents the operation of feature concatenation, GAP(*) represents global average pooling operation, and F represents the multi-scale feature representation of the input distorted image.

6. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 5, wherein The specific steps of step S5 include the following steps: Step S51: Input the multi-scale encoded representation F of the input distorted image obtained through step S45 into the fully connected layer MLP to obtain the quality evaluation score F of the distorted image score , and its calculation formula is: F score = MLP(F) Step S52: Design the loss function of the edge feature-guided image quality assessment network as follows: where m is the number of samples in each training batch, and y i represents the true quality score of the i-th image sample, and represents the predicted quality score obtained by passing the i-th image sample through the network; Step S53: Repeat the above steps S51 to S52 in batches until the loss value calculated in step S52 converges and stabilizes, save the network parameters, complete the training process of the edge feature-guided image quality assessment network, and obtain an edge feature-guided image quality assessment model.

7. The no-reference screen content image quality assessment method based on edge feature guidance according to claim 1, wherein The specific steps of step S6 include the following steps: Step S61: Input the distorted screen content images and the corresponding edge structure diagrams in the test set into the trained image quality assessment model guided by edge features, and output the corresponding quality assessment scores.