Screen content video quality evaluation method and device based on deep and shallow spatio-temporal features
By constructing and training a dual-branch screen content video quality evaluation model containing spatial and temporal feature extraction branches, the problem of difficulty in evaluating the video quality of screen content in the prior art is solved, and higher evaluation accuracy and visual quality optimization are achieved.
Patent Information
- Application Number
- CN202510495985.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art is difficult to effectively evaluate and optimize the visual quality of screen content videos, especially when facing various noise interferences and different types of content.
The video quality evaluation model of dual-branch screen content based on deep and shallow spatial and temporal characteristics is adopted, and the video quality of screen content is effectively evaluated through spatial feature extraction branch and temporal feature extraction branch, airspace time domain fusion module and quality regression module.
It improves the accuracy and adaptability of video quality evaluation, optimizes the visual quality of video, and reduces the computational complexity.
Smart Images

Figure CN120031869B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video evaluation, and particularly to a method and system for evaluating the quality of screen content videos based on deep and shallow spatio-temporal features. Background Art
[0002] With the rapid development of artificial intelligence and digital multimedia technologies and the popularization of various portable devices, screen content videos are increasingly widely used in the field of digital multimedia applications (such as online education, video conferencing, webcasting, game videos, etc.), providing users with a more flexible and free multimedia experience. Different from natural videos, screen content videos are a mixture of images and computer-generated text / image regions.
[0003] MSCN (Mean Subtracted Contrast Normalized) is a technique widely used in the fields of image processing and quality assessment, mainly for feature extraction and preprocessing stages. It reduces the impact of factors such as illumination changes or noise interference on image quality assessment by eliminating the mean within the local region of the image and normalizing it. Specifically, MSCN analyzes the neighborhood information of each pixel point, calculates the local mean of this point, subtracts this local mean from the original pixel value, and then performs contrast normalization on the result. This technique is particularly suitable for enhancing the structural information in the image while suppressing unstructured noise components, making the subsequent feature extraction process more effective. In the quality assessment of screen content videos, MSCN is used to process the color moment feature vectors of cartoon images to eliminate the mutual interference between neighborhood information, thereby improving the accuracy and reliability of color feature extraction. This method is crucial for improving the accuracy of overall video quality assessment.
[0004] 3D LOG (Three-Dimensional Laplacian of Gaussian) is a technique for capturing the spatio-temporal features of screen content in videos. It is extended from the two-dimensional Laplacian of Gaussian (LoG) operator, which is widely used in the field of image processing to detect edge and texture information in images. In the context of video processing, 3D LOG extends the time dimension for consecutive frames, enabling it to not only extract spatial features within a single frame but also capture the changing features between frames. This makes 3D LOG particularly suitable for analyzing the fast-changing text, icons, and other computer-generated content unique to screen content videos, helping to identify quality degradation caused by compression or transmission.
[0005] 3D NSS (3D Natural Scene Statistics) is a method specifically designed to capture the spatio-temporal characteristics of natural scenes in videos. Different from 3D LOG which focuses on screen content, 3D NSS focuses on the dynamic elements in natural scenes, such as landscapes, people, etc., which usually have visual characteristics and motion patterns different from computer-generated content. By analyzing the spatial structure and its temporal evolution of natural scenes in a video sequence, 3D NSS can effectively simulate the perception of the human visual system on the quality of natural videos. This method is particularly useful for evaluating the quality of videos containing a large number of natural scenes, because it can quantify the changes in visual experience caused by distortion, thus providing an important basis for video quality evaluation.
[0006] At different processing stages of screen content videos (such as acquisition, transmission, display, etc.), they will inevitably be interfered by various noises, resulting in varying degrees of decline in perceptual quality and seriously affecting the user experience. Therefore, there is an urgent need to develop an effective video quality evaluation method to optimize the visual quality of screen content videos and improve the performance of related tasks. Summary of the Invention
[0007] To solve the above problems, the present invention proposes a method and system for evaluating the quality of screen content videos based on deep and shallow spatio-temporal features. By constructing and training a dual-branch screen content video quality evaluation model including spatial and temporal feature extraction branches, an effective evaluation of the quality of screen content videos is achieved. At the same time, different strategies are adopted to extract features for different types of content, improving the evaluation accuracy and optimizing the visual quality of the videos.
[0008] The specific solutions are as follows:
[0009] On the one hand, the method for evaluating the quality of screen content videos based on deep and shallow spatio-temporal features includes:
[0010] Obtain a screen content video, and extract a number of video blocks and a number of key frames from the screen content video;
[0011] Construct a screen content video quality evaluation model and train it to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatio-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on the Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on the SlowFast network, a three-dimensional Laplacian of Gaussian module, and a three-dimensional natural scene statistics module;
[0012] Input the several video blocks and several key frames into the trained screen content video quality evaluation model; specifically, input each key frame into the deep feature extraction module, edge feature extraction module, color extraction module, and natural scene statistics module respectively to obtain the deep semantic features, edge features of text images, color features of cartoon images, and natural scene statistics features in each key frame; splice the deep semantic features, edge features, color features, and natural scene statistics features to obtain spatial features; input the video blocks into the temporal feature extraction module, three-dimensional Laplacian of Gaussian module, and natural scene statistics module respectively to obtain temporal features, natural spatio-temporal features, and time features; splice the spatial features and time features and input them into the spatial-temporal fusion module to obtain the fused spatial features and time features; input the fused spatial features and time features into the quality regression module to obtain the final video quality score.
[0013] Further, input each key frame into the deep feature extraction module to obtain the deep semantic features in each key frame, and the calculation formula is as follows:
[0014] ;
[0015] ;
[0016] ;
[0017] ;
[0018] Among them, = , which represents the feature map extracted based on the key frame x at the k-th stage, , represents the feature map represents the height of the feature map represents the feature map represents the width of the feature map represents the feature map represents the number of channels of the feature map; represents the global mean of the feature map ; represents the standard deviation of the feature map ; k is the identifier of the stage number; represents the spatial feature representation at the k-th stage; represents the deep semantic features; represents the feature splicing function; represents global average pooling; represents global standard deviation pooling.
[0019] Further, each key frame is input into an edge feature extraction module to obtain the edge features of the Chinese text images in each key frame. Specifically: the edge feature extraction module uses a two-dimensional Laplacian of Gaussian operator to extract the edge features of the text images; among them, the kernel function calculation formula of the two-dimensional Laplacian of Gaussian operator is as follows:
[0020] ;
[0021] ;
[0022] ;
[0023] Among them, represents the abscissa of the input text image pixel; represents the ordinate of the input text image pixel; represents the two-dimensional Gaussian function; represents the input image at the coordinate pixel value at; represents the standard deviation of the Gaussian function; represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction ; represents performing the Laplacian operator operation; represents the kernel function of the two-dimensional Laplacian of Gaussian operator; represents the result after the two-dimensional Laplacian of Gaussian operator processes the image; * represents the convolution operation.
[0024] Further, each key frame is input into a color extraction module to obtain the color features of the cartoon images in each key frame, specifically including:
[0025] The first color moment, the second color moment, and the third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formulas are as follows:
[0026] ;
[0027] ;
[0028] ;
[0029] Among them, I represents the channel map; M() represents the mean operation; represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; Represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map;
[0030] Obtain the color moment feature vector of the cartoon image in each key frame based on the first color moment, the second color moment, and the third color moment ; where Represents the first color moment on the hue channel; Represents the second color moment on the hue channel; Represents the third color moment on the hue channel; Represents the first color moment on the saturation channel; Represents the second color moment on the saturation channel; Represents the third color moment on the saturation channel; Represents the first color moment on the luminance channel; Represents the second color moment on the luminance channel; Represents the third color moment on the luminance channel;
[0031] Use the mean subtracted contrast normalization coefficient MSCN to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows:
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] where Represents the weight value of the two-dimensional circular symmetric Gaussian weighting filter at the position (k, l). k represents the offset of the two-dimensional circular symmetric Gaussian weighting filter in the vertical direction, and l represents the offset of the two-dimensional circular symmetric Gaussian weighting filter in the horizontal direction; Represents the pixel value of the normalized key frame at the position (x, y); Represents the pixel value of the key frame at the position ; Represents the local average; Represents the local standard deviation; Represents the maximum offset in the horizontal direction; Represents the maximum offset in the vertical direction; K = L = 3;; Represents the Pixel value at the position (x, y) after applying average filtering; Represents the offset in the horizontal direction when the average filter moves; Indicates the offset of the average filter in the vertical direction when moving; (h, w) represents the weight value of the average filter at the position (h, w); Indicates the height of the average filter, Indicates the width of the average filter, H = W = 3;
[0037] Calculate and entropy to obtain the color entropy of each channel map , and the calculation formula is as follows:
[0038] ;
[0039] Among them, Indicates the color entropy of each channel map; Indicates the frequency of the intensity value u; i represents different HSV channels; Indicates or ;
[0040] Obtain the color entropy feature vector based on the color entropy of each channel map ;
[0041] Indicates the color entropy obtained by processing the key frame in the hue channel through MSCN; Indicates the color entropy of the key frame in the hue channel after MSCN and average filtering; Indicates the color entropy obtained by processing the key frame in the saturation channel through MSCN; Indicates the color entropy of the key frame in the saturation channel after MSCN and average filtering; Indicates the color entropy obtained by processing the key frame in the brightness channel through MSCN; Indicates the color entropy of the key frame in the brightness channel after MSCN and average filtering;
[0042] Concatenate F CE and F CM to obtain the color feature of the cartoon image in each key frame .
[0043] Furthermore, input each key frame into the natural scene statistics module to obtain the natural scene statistics features in each key frame, specifically including:
[0044] Preprocess each key frame by adopting the method of local mean removal and normalization, and the calculation formula is as follows:
[0045] ;
[0046] ;
[0047] ;
[0048] Among them, represents the pixel value of each key frame at position ; represents the local mean; represents the offset in the vertical direction; represents the offset in the horizontal direction; represents the maximum offset in the horizontal direction; represents the maximum offset in the vertical direction; K = L = 3; represents the local standard deviation; represents the pixel value of the processed key frame at position ; represents the weight value of the two-dimensional circular symmetric Gaussian weighted filter at position ;
[0049] Divide the preprocessed key frame into 96*96 image blocks; Set 0.7 times the maximum sharpness value in the image block as the sharpness threshold, and select the image blocks with sharpness higher than this threshold as the image blocks to be processed;
[0050] Calculate the statistical distribution of the image blocks to be processed, and fit the statistical distribution of the image blocks to be processed through the Generalized Gaussian Distribution GGD; The calculation formula of the Generalized Gaussian Distribution GGD is as follows:
[0051] ;
[0052] Among them, z represents the preprocessed image coefficient; represents the shape parameter, which is used to describe the sharpness of the distribution; represents the scale parameter, which is used to describe the width of the distribution; represents the gamma function; represents the probability density function of the generalized Gaussian distribution of z;
[0053] Extract the correlation features between adjacent coefficients in the image blocks to be processed through the Asymmetric Generalized Gaussian Distribution AGGD of the product of adjacent coefficients. The specific formula is as follows:
[0054] ;
[0055] Among them, γ represents the shape parameter of the asymmetric generalized Gaussian distribution; represents the scale parameter on the left side of the asymmetric generalized Gaussian distribution; βr represents the scale parameter on the right side of the asymmetric generalized Gaussian distribution; m represents the product of adjacent coefficients; The probability density function representing the asymmetric generalized Gaussian distribution of m;
[0056] Calculate the mean of the distribution, and the formula is as follows:
[0057] ;
[0058] Represents the mean of the distribution;
[0059] Extract the natural scene statistical features based on the fitted statistical distribution, the correlation characteristics between adjacent coefficients, and the mean of the asymmetric generalized Gaussian distribution.
[0060] Furthermore, after splicing the spatial features and temporal features, input them into the spatio-temporal fusion module to obtain the fused spatial and temporal features, specifically including:
[0061] The spatial features extracted from the i-th key frame , and extract the temporal features from the i-th video segment , and the spatial features and the temporal features After performing learnable feature fusion, input them into the multi-layer perceptron MLP to obtain the fused spatial and temporal features, and the calculation formula is as follows:
[0062] ;
[0063] Among them, Represents the learnable feature fusion function. The learnable feature fusion function is specifically: first perform feature mapping through a fully connected layer containing 1024 neurons, and then apply the RELU activation function for non-linear transformation; Represents the fused feature of the i-th video segment; Represents feature splicing.
[0064] Furthermore, input the fused spatial and temporal features into the quality regression module to obtain the final video quality score, specifically: represent the fused spatial and temporal features as the quality score of the video clip through two fully connected layers of the quality regression module, and the calculation formula is as follows:
[0065] ;
[0066] Among them, Represents the quality score of the i-th video clip; FC() represents the fully connected layer; Represents the fused feature of the i-th video segment;
[0067] Average the quality scores of the video clips of k video segments to obtain the final video quality score:
[0068] ;
[0069] Among them, Q represents the final video quality score, and k represents the number of video segments.
[0070] On the other hand, a screen content video quality evaluation device based on deep and shallow spatio-temporal features, characterized by comprising:
[0071] A screen content video extraction module, configured to obtain a screen content video and extract a plurality of video blocks and a plurality of key frames in the screen content video;
[0072] A model construction and training module, configured to construct and train a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatio-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on the Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on the SlowFast network, a three-dimensional Laplacian of Gaussian module, and a three-dimensional natural scene statistics module;
[0073] A video quality score evaluation module, configured to input the plurality of video blocks and the plurality of key frames into the trained screen content video quality evaluation model; specifically, input each key frame into the deep feature extraction module, the edge feature extraction module, the color extraction module, and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image, and the natural scene statistics features in each key frame; splice the deep semantic features, the edge features, the color features, and the natural scene statistics features to obtain spatial features; input the video blocks into the temporal feature extraction module, the three-dimensional Laplacian of Gaussian module, and the natural scene statistics module respectively to obtain temporal features, natural spatio-temporal features, and time features; splice the spatial features and the time features and input them into the spatio-temporal fusion module to obtain fused spatial features and time features; input the fused spatial features and time features into the quality regression module to obtain the final video quality score.
[0074] The present invention adopts the above technical solutions and has the following beneficial effects:
[0075] (1) By extracting deep and shallow spatio-temporal features, the present invention effectively reduces the computational complexity, and at the same time adopts different strategies to extract spatial and temporal features for different types of content (text, cartoon, natural image), improving the accuracy of quality assessment;
[0076] (2) The present invention uses a dual-branch screen content video quality evaluation model including a spatial feature extraction branch, a temporal feature extraction branch, a spatio-temporal fusion module, and a quality regression module to effectively evaluate the quality of screen content videos and optimize the visual quality of videos.
[0077] (3) The present invention extracts spatial and temporal features by combining the Conv Next network and the SlowFast network respectively, and uses a multi-layer perceptron and a fully connected layer for feature fusion and quality scoring, further enhancing the adaptability to different types of screen content videos and the accuracy of the evaluation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 is a flowchart of a screen content video quality evaluation method based on deep and shallow spatio-temporal features according to an embodiment of the present invention;
[0079] Figure 2 is a schematic flowchart of screen content video quality evaluation according to an embodiment of the present invention;
[0080] Figure 3 is a diagram of a screen content video quality evaluation device based on deep and shallow spatio-temporal features according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto. As Figure 1 shown, the screen content video quality evaluation method based on deep and shallow spatio-temporal features of the present invention includes:
[0082] S1, obtaining a screen content video, and extracting a plurality of video blocks and a plurality of key frames from the screen content video.
[0083] Specifically, in this embodiment, given a screen content video V, the number of frames and the frame rate of the video are l and r respectively. The video V is divided into consecutive video blocks, denoted as represents the jth frame in the ith block. The first frame in each video block is selected as a key frame to extract spatial features. Considering that the time information is not sensitive to the resolution, for all frames in the video block
[0084] S2. Construct a screen content video quality evaluation model and train it to obtain a trained screen content video quality evaluation model. The screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatio-temporal fusion module, and a quality regression module. The spatial feature extraction branch includes a deep feature extraction module based on the Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module. The temporal feature extraction branch includes a temporal feature extraction module based on the SlowFast network, a three-dimensional Laplacian of Gaussian module, and a three-dimensional natural scene statistics module.
[0085] Specifically, as Figure 2 shown, the screen content video quality evaluation model proposed in this embodiment has two branches, namely a spatial feature extraction branch and a temporal feature extraction branch, which are respectively used to calculate spatial features and temporal features. First, construct a Conv Next neural network for the key frames, and design corresponding modules for text images, cartoon images, and natural images in the key frames respectively, and finally obtain spatial features. Construct a SlowFast (3D CNN) network for the video segments, and at the same time use 3D-LOG and 3D-NSS to process the video segments, and finally obtain temporal features. Combine the spatial features and temporal features, and obtain the quality score through quality regression, which can effectively realize the quality evaluation of screen content videos.
[0086] S3. Input the several video blocks and several key frames into the trained screen content video quality evaluation model. Specifically, input each key frame into the deep feature extraction module, the edge feature extraction module, the color extraction module, and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of text images, the color features of cartoon images, and the natural scene statistics features in each key frame. Concatenate the deep semantic features, edge features, color features, and natural scene statistics features to obtain spatial features. Input the video blocks into the temporal feature extraction module, the three-dimensional Laplacian of Gaussian module, and the natural scene statistics module respectively to obtain temporal features, natural spatio-temporal features, and time features. Concatenate the spatial features and temporal features and input them into the spatio-temporal fusion module to obtain the fused spatial features and temporal features. Input the fused spatial features and temporal features into the quality regression module to obtain the final video quality score.
[0087] Specifically, input each key frame into the deep feature extraction module to obtain the deep semantic features in each key frame, and the calculation formula is as follows:
[0088] ;
[0089] ;
[0090] ;
[0091] ;
[0092] Among them, = , which represents the feature map extracted based on the key frame x at the k-th stage. , represents the feature map 's height. represents the feature map 's width. represents the feature map 's number of channels; represents the global mean of the feature map ; represents the standard deviation of the feature map ; k is the identifier of the stage number; represents the spatial feature representation of the k-th stage; represents the deep semantic feature; represents the feature concatenation function; represents global average pooling; represents global standard deviation pooling.
[0093] Specifically, each key frame is input into the edge feature extraction module to obtain the edge features of the text image in each key frame. Specifically: the edge feature extraction module uses the two-dimensional Laplacian of Gaussian operator to extract the edge features of the text image; among them, the kernel function calculation formula of the two-dimensional Laplacian of Gaussian operator is as follows:
[0094] ;
[0095] ;
[0096] ;
[0097] Among them, represents the abscissa of the input text image pixel; represents the ordinate of the input text image pixel; represents the two-dimensional Gaussian function; represents the pixel value of the input image at the coordinate ; represents the standard deviation of the Gaussian function; represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction; represents performing the Laplacian operator operation; represents the kernel function of the two-dimensional Laplacian of Gaussian operator. Indicates the result after processing the image with a two-dimensional Laplacian of Gaussian operator; * represents the convolution operation.
[0098] Specifically, each key frame is input into the color extraction module to obtain the color features of the cartoon image in each key frame, specifically including:
[0099] The first color moment, the second color moment, and the third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formulas are as follows:
[0100] ;
[0101] ;
[0102] ;
[0103] where, I represents the channel map; M() represents the mean operation; represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map;
[0104] Based on the first color moment, the second color moment, and the third color moment, a color moment feature vector of the cartoon image in each key frame is obtained ; where, represents the first color moment on the hue channel; represents the second color moment on the hue channel; represents the third color moment on the hue channel; represents the first color moment on the saturation channel; represents the second color moment on the saturation channel; represents the third color moment on the saturation channel; represents the first color moment on the brightness channel; represents the second color moment on the brightness channel; represents the third color moment on the brightness channel;
[0105] The mean subtraction contrast normalization coefficient MSCN is used to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows:
[0106] ;
[0107] ;
[0108] ;
[0109] ;
[0110] Among them, represents the weight value of the two-dimensional circular symmetric Gaussian weighted filter at the position (k, l), k represents the offset of the two-dimensional circular symmetric Gaussian weighted filter in the vertical direction, and l represents the offset of the two-dimensional circular symmetric Gaussian weighted filter in the horizontal direction; represents the pixel value of the normalized key frame at the position (x, y); represents the key frame at the position ; represents the local average value; represents the local standard deviation; represents the maximum offset in the horizontal direction; represents the maximum offset in the vertical direction; K = L = 3; represents the result of applying average filtering at the position (x, y); represents the offset of the average filter in the horizontal direction during movement; represents the offset of the average filter in the vertical direction during movement; (h, w) represents the weight value of the average filter at the position (h, w); represents the height of the average filter, represents the width of the average filter, H = W = 3;
[0111] Calculate and of the entropy to obtain the color entropy of each channel map , and the calculation formula is as follows:
[0112] ;
[0113] Among them, represents the color entropy of each channel map; represents the frequency of the intensity value u; i represents different HSV channels; represents or ;
[0114] Obtain the color entropy feature vector based on the color entropy of each channel map ;
[0115] represents the color entropy obtained by processing the key frame in the hue channel through MSCN; Denotes the color entropy of the key frame after MSCN and average filtering in the hue channel; Denotes the color entropy of the key frame obtained by MSCN processing in the saturation channel; Denotes the color entropy of the key frame after MSCN and average filtering in the saturation channel; Denotes the color entropy of the key frame obtained by MSCN processing in the luminance channel Denotes the color entropy of the key frame after MSCN and average filtering in the luminance channel;
[0116] Concatenate F CE and to obtain the color features of the cartoon image in each key frame .
[0117] Specifically, in this embodiment, for the natural scene images in the screen content video, the natural scene statistics module is used to obtain natural scene features. Specifically, for the key frames, the natural scene statistics module extracts two features from the Generalized Gaussian Distribution (GGD) under the conditions of the original resolution of the image and the resolution reduced by two times ( , ); Extract (γ, , ) three features from the Asymmetric Generalized Gaussian Distribution (AGGD) in four directions: horizontal, vertical, positive diagonal, and anti-diagonal, for a total of 12 features; Calculate the mean of the four-direction asymmetric generalized Gaussian distribution, for a total of 4 features; Therefore, a total of 36 features are extracted. These 36 features are the natural scene statistics features.
[0118] Specifically, extracting the natural scene statistics features in each key frame through the natural scene statistics module for natural images specifically includes:
[0119] Preprocess each key frame by adopting the method of local mean removal and normalization, and the calculation formula is as follows:
[0120] ;
[0121] ;
[0122] ;
[0123] Among them, Denotes the pixel value of each key frame at position ; Denotes the local mean; Denotes the offset in the vertical direction; Denotes the offset in the horizontal direction; Represents the maximum offset in the horizontal direction; Represents the maximum offset in the vertical direction; K = L = 3; Represents the local standard deviation Represents the pixel value of the processed key frame at the position ; Represents the weight value of the two-dimensional circular symmetric Gaussian weighted filter at the position (k, l).
[0124] Divide the preprocessed key frame into 96*96 image blocks; Set 0.7 times the maximum sharpness value in the image block as the sharpness threshold, and select the image blocks with sharpness higher than this threshold as the image blocks to be processed;
[0125] Calculate the statistical distribution of the image blocks to be processed, and fit the statistical distribution of the image blocks to be processed through the Generalized Gaussian Distribution (GGD); The calculation formula of the Generalized Gaussian Distribution (GGD) is as follows:
[0126] ;
[0127] Among them, z represents the preprocessed image coefficient; Represents the shape parameter, which is used to describe the sharpness of the distribution; Represents the scale parameter, which is used to describe the width of the distribution; Represents the gamma function; Represents the probability density function of the Generalized Gaussian Distribution of z;
[0128] Extract the correlation features between adjacent coefficients in the image blocks to be processed through the Asymmetric Generalized Gaussian Distribution (AGGD) of the product of adjacent coefficients. The specific formula is as follows:
[0129] ;
[0130] Among them, γ represents the shape parameter of the Asymmetric Generalized Gaussian Distribution; Represents the scale parameter on the left side of the Asymmetric Generalized Gaussian Distribution; Represents the scale parameter on the right side of the Asymmetric Generalized Gaussian Distribution; m represents the product of adjacent coefficients; Represents the probability density function of the Asymmetric Generalized Gaussian Distribution of m;
[0131] Calculate the mean of the distribution. The formula is as follows:
[0132] ;
[0133] Represents the mean of the distribution;
[0134] Extract the natural scene statistical features based on the fitted statistical distribution, the correlation features between adjacent coefficients, and the mean of the Asymmetric Generalized Gaussian Distribution.
[0135] Specifically, in this embodiment, for the text image in the screen content video, the LOG operator is used to extract the edge features of the text image. The full name of the LOG operator is the Laplacian of Gaussian operator. The Laplacian edge detection operator does not smooth the image, so it is very sensitive to noise. Therefore, the image can be first subjected to Gaussian smoothing processing and then convolved with the Laplacian operator.
[0136] Specifically, in this embodiment, in order to retain sufficient time information and at the same time reduce the computational complexity, the video is uniformly segmented into lower-resolution segments to extract time features. Specifically, two branches are designed to respectively obtain deep learning-based features and handcrafted features. For the deep learning-based features, a pre-trained SlowFast network is used to extract the motion features of each video segment:
[0137] ;
[0138] Among them, represents the th video segment, Φ(·) represents the operation of extracting motion features, represents the motion features extracted from the th video segment;
[0139] Based on 3D-LOG for screen spatio-temporal feature extraction, 2D-LOG is extended to 3D-LOG and used to capture screen spatio-temporal features. The calculation formula of 3D-LOG is derived from 2D-LOG:
[0140] ;
[0141] Among them, represents the response value of the three-dimensional Laplacian of Gaussian filter at the position (x, y, t), where x, y, and t respectively represent the spatial abscissa, spatial ordinate, and time coordinate of the pixel in the screen content video. exp() represents the exponential function, represents the standard deviation of the Gaussian kernel. Here takes 1.485. By convolving the screen content video segment and the three-dimensional Laplacian of Gaussian filter, the screen spatio-temporal features are extracted, that is:
[0142] ;
[0143] Among them, represents the pixel value of the screen content video segment at the position (x, y, t), represents the screen spatio-temporal feature map, Denotes the response value of the 3D Gaussian Laplacian filter at the position (x, y, t). x, y, and t represent the spatial abscissa, spatial ordinate, and time coordinate of the pixel in the screen content video respectively. * denotes the convolution operation.
[0144] Based on 3D-NSS for natural spatio-temporal feature extraction, the MSCN coefficients are extended to 3D space to describe the natural spatio-temporal features of screen content video segments, similar to the aforementioned screen spatio-temporal feature extraction. The brightness of the screen content video segment is normalized to calculate the 3D-MSCN coefficients, and it is used to capture the natural spatio-temporal quality degradation, that is:
[0145] ;
[0146] ;
[0147] ;
[0148] ;
[0149] Where Denotes the pixel value at the position (x, y, t) of the screen content video segment, and c = 1 is a normal constant to ensure the stability of the formula. Denotes the mean of the local region centered at (x, y, t), Denotes the standard deviation of the local region centered at (x, y, t), c represents a constant, which is taken as 1 here, Denotes the value at the position (x, y, t) after the video segment undergoes mean subtraction and contrast normalization, Denotes the difference between the pixel value at (x, y, t) and the local mean. x, y, and t represent the spatial abscissa, spatial ordinate, and time coordinate of the pixel in the screen content video respectively. Denotes the weight value of the 3D circularly symmetric Gaussian filter at the position (p, u, v). p represents the offset of the 3D circularly symmetric Gaussian filter on the x-axis, u represents the offset of the 3D circularly symmetric Gaussian filter on the y-axis, v represents the offset of the 3D circularly symmetric Gaussian filter on the time axis, P represents the maximum offset on the x-axis, U represents the maximum offset on the y-axis, and V represents the maximum offset on the time axis, 。 Denotes the pixel value at the position offset by (p, u, v) relative to the center position (x, y, t) in the local window.
[0150] Specifically, after splicing the spatial feature and the temporal feature, they are input into the spatio-temporal fusion module to obtain the fused spatial feature and temporal feature, which specifically includes:
[0151] Spatial features extracted from the i-th key frame , and temporal features are extracted from the i-th video segment . The spatial features and temporal features are subjected to learnable feature fusion and then input into a multi-layer perceptron MLP to obtain fused spatial and temporal features. The calculation formula is as follows:
[0152]
[0153] where represents the learnable feature fusion function, and the learnable feature fusion function is specifically: first, feature mapping is performed through a fully connected layer containing 1024 neurons, and then the RELU activation function is applied for non-linear transformation; represents the feature after fusion of the i-th video segment; represents feature concatenation.
[0154] Specifically, the fused spatial and temporal features are input into the quality regression module to obtain the final video quality score, specifically: the fused spatial and temporal features are represented as the quality score of the video clip through two fully connected layers of the quality regression module, and the calculation formula is as follows:
[0155] ;
[0156] where represents the quality score of the i-th video clip; FC() represents the fully connected layer; represents the feature after fusion of the i-th video segment;
[0157] The quality scores of the video clips of k video segments are averaged to obtain the final video quality score:
[0158] ;
[0159] where Q represents the final video quality score, and k represents the number of video segments.
[0160] In this embodiment, the proposed screen content video quality evaluation model based on deep and shallow spatio-temporal features is built in a Pytorch environment and experiments are conducted using an NVIDIA RTX A6000 GPU; the minimum resolution size of the key frames is adjusted to 520 while maintaining the original aspect ratio. During training, the key frames are randomly cropped to 448*448. For video blocks, during training, the video block resolution is adjusted to 224*224, the batch size is 64, the number of training epochs is 100, the initial learning rate is set to 0.00001, and training is performed using the Adam optimizer. The experimental dataset is SCVD, which contains 16 reference screen content videos with different content scenarios and 800 distorted screen content videos, 80% of which are used for training and the remaining 20% are used for testing. The Spearman rank correlation coefficient (SROCC), Pearson linear correlation coefficient (PLCC), and root mean square error (RMSE) are selected to evaluate the performance of the above model.
[0161] As Figure 3 shown, this embodiment also discloses a screen content video quality evaluation device based on deep and shallow spatio-temporal features, including:
[0162] A screen content video extraction module 31, configured to obtain a screen content video and extract a plurality of video blocks and a plurality of key frames from the screen content video;
[0163] A model construction and training module 32, configured to construct and train a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatio-temporal domain fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on the Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on the SlowFast network, a three-dimensional Laplacian of Gaussian module, and a three-dimensional natural scene statistics module;
[0164] The video quality score evaluation module 33 is used to input the several video blocks and several key frames into the trained screen content video quality evaluation model. Specifically, each key frame is respectively input into the deep feature extraction module, the edge feature extraction module, the color extraction module, and the natural scene statistics module to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image, and the natural scene statistics features in each key frame. The deep semantic features, the edge features, the color features, and the natural scene statistics features are spliced to obtain the spatial features. The video blocks are respectively input into the temporal feature extraction module, the three-dimensional Laplacian of Gaussian module, and the natural scene statistics module to obtain the temporal features, the natural spatio-temporal features, and the time features. The spatial features and the time features are spliced and then input into the spatio-temporal fusion module to obtain the fused spatial features and time features. The fused spatial features and time features are input into the quality regression module to obtain the final video quality score.
[0165] The specific implementation of the screen content video quality evaluation device based on the deep and shallow spatio-temporal features is the same as that of the screen content video quality evaluation method based on the deep and shallow spatio-temporal features, and this embodiment will not be repeated.
[0166] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them are within the protection scope of the present invention.
Claims
1. A method for evaluating screen content video quality based on deep and shallow spatiotemporal features, characterized in that: include: Acquire a screen content video, and extract a plurality of video blocks and a plurality of key frames from the screen content video; Constructing and training a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module; Input the plurality of video blocks and the plurality of key frames into a trained screen content video quality assessment model; input each key frame into a deep feature extraction module, an edge feature extraction module, a color extraction module and a natural scene statistics module respectively, and obtain deep semantic features, edge features of text images, color features of cartoon images and natural scene statistics features in each key frame; The deep semantic features, edge features, color features and natural scene statistical features are spliced to obtain spatial features; the video blocks are input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistical module respectively to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced and input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score; After splicing the spatial features and temporal features, they are input into the spatial and temporal fusion module to obtain the fused spatial and temporal features, which include: The spatial feature SI extracted from the i-th key frame i , extract the temporal feature TI from the i-th video segment i , the spatial feature SI i and time characteristics TI i After the learnable features are fused, they are input into the multi-layer perceptron MLP to obtain the fused spatial features and temporal features. The calculation formula is as follows: Among them, Ψ(·) represents the learnable feature fusion function, which is as follows: firstly, feature mapping is performed through a fully connected layer containing 1024 neurons, and then the RELU activation function is applied for nonlinear transformation; FF i represents the fused features of the i-th video clip; Represents feature splicing; The fused spatial features and temporal features are input into the quality regression module to obtain the final video quality score. Specifically, the fused spatial features and temporal features are regressed into the quality score of the video clip through the two fully connected layers of the quality regression module. The calculation formula is as follows: Q i =FC(FF i ); Among them, Q i represents the quality score of the i-th video clip; FC() represents the fully connected layer; FF i represents the fused features of the i-th video clip; Average the quality scores of the video clips of k video segments to get the final video quality score: Among them, Q represents the final video quality score, and k represents the number of video clips.
2. The screen content video quality evaluation method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the deep feature extraction module to obtain the deep semantic features in each key frame. The calculation formula is as follows: in, represents the feature map extracted based on key frame x at the kth stage, H k Representation feature map Height, W k Representation feature map Width, C k Representation feature map The number of channels; Representation feature map The global mean of Representation feature map The standard deviation of ; k is the identifier of the stage number; represents the spatial feature representation of the kth stage; F S represents deep semantic features; cat() represents feature concatenation function; GP avg () indicates global average pooling; GP std () represents global standard error pooling.
3. The screen content video quality assessment method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the edge feature extraction module to obtain the edge features of the Chinese text image in each key frame. Specifically, the edge feature extraction module uses a two-dimensional Gaussian Laplacian operator to extract edge features from the text image; wherein the kernel function calculation formula of the two-dimensional Gaussian Laplacian operator is as follows: Where x represents the horizontal coordinate of the input text image pixel; y represents the vertical coordinate of the input text image pixel; G σ (x, y) represents a two-dimensional Gaussian function; I(x, y) represents the pixel value of the input image at the coordinate (x, y); σ represents the standard deviation of the Gaussian function; exp() represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction; Indicates Laplacian operation; LOG indicates the kernel function of the two-dimensional Laplacian of Gaussian operator; LOG(x,y) indicates the result of the two-dimensional Laplacian of Gaussian operator processing the image; * indicates the convolution operation.
4. The screen content video quality assessment method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the color extraction module to obtain the color features of the cartoon image in each key frame, including: The first color moment, second color moment, and third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formula is as follows: m(I)=M(I); Where, I represents the channel map; M() represents the mean operation; m(I) represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; d(I) represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; s(I) represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map; The color moment feature vector F of the cartoon image in each key frame is obtained based on the first color moment, the second color moment and the third color moment. CM =[m H ,d H ,s H ,m S ,d S ,s S ,m V ,d V ,s V ]; where m H Represents the first color moment on the hue channel; d H Represents the second color moment on the hue channel; s H Represents the third color moment on the hue channel; m S Represents the first color moment on the saturation channel; d S Represents the second color moment on the saturation channel; s S Represents the third color moment on the saturation channel; m V Represents the first color moment on the brightness channel; d V Represents the second color moment on the brightness channel; s V Represents the third color moment on the brightness channel; The contrast normalization coefficient MSCN is used to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows: Among them, w k,l I represents the weight value of the two-dimensional circularly symmetric Gaussian weighted filter at the position (k, l), k represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the vertical direction, and l represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the horizontal direction; I C (x, y) represents the normalized pixel value of the key frame at position (x, y); I(x, y) represents the pixel value of the key frame at position (x, y); μ(x, y) represents the local mean; σ(x, y) represents the local standard deviation; L represents the maximum offset in the horizontal direction; K represents the maximum offset in the vertical direction; K = L = 3; I CA (x,y) represents the C (x, y) is the pixel value at position (x, y) after applying the average filter; w represents the horizontal offset of the average filter when it moves; h represents the vertical offset of the average filter when it moves; represents the weight value of the average filter at position (h, w); H represents the height of the average filter, W represents the width of the average filter, H = W = 3; Computation I C (x,y) and I CA The entropy of (x,y) is used to obtain the color entropy of each channel map. The calculation formula is as follows: in, Represents the color entropy of each channel map; p u represents the frequency of the intensity value u; i represents different HSV channels; I j Indicates I C (x,y) or I CA (x,y); Get the color entropy feature vector based on the color entropy of each channel image It represents the color entropy of the key frame after MSCN processing in the hue channel; Represents the color entropy of the key frame after MSCN and average filtering in the hue channel; It represents the color entropy of the key frame after MSCN processing in the saturation channel; Represents the color entropy of the key frame after MSCN and average filtering in the saturation channel; It represents the color entropy of the key frame after MSCN processing in the brightness channel; Represents the color entropy of the key frame after MSCN and average filtering in the brightness channel; F CE and F CM Splice to obtain the color feature F of the cartoon image in each key frame C =[F CM ,F CE ].
5. The method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to claim 1, characterized in that: Each key frame is input into the natural scene statistics module to obtain the natural scene statistical features in each key frame, including: The local mean removal and normalization method is used to preprocess each key frame. The calculation formula is as follows: Where I(x,y) represents the pixel value of each key frame at position (x,y); μ(x,y) represents the local mean; k represents the offset in the vertical direction; l represents the offset in the horizontal direction; L represents the maximum offset in the horizontal direction; K represents the maximum offset in the vertical direction; K = L = 3; σ(x,y) represents the local standard deviation; represents the pixel value of the key frame at position (x, y) after processing; w k,l Represents the weight value of the two-dimensional circularly symmetric Gaussian weighted filter at position k, l; The preprocessed key frame is divided into 96*96 image blocks; 0.7 times of the maximum sharpness value in the image block is set as the sharpness threshold, and the image blocks with sharpness higher than this threshold are selected as the image blocks to be processed; The statistical distribution of the image block to be processed is calculated, and the statistical distribution of the image block to be processed is fitted by the generalized Gaussian distribution GGD; the calculation formula of the generalized Gaussian distribution GGD is as follows: Wherein, z represents the image coefficient after preprocessing; α represents the shape parameter, which is used to describe the sharpness of the distribution; β represents the scale parameter, which is used to describe the width of the distribution; Γ() represents the gamma function; f(z; α, β) represents the probability density function of the generalized Gaussian distribution of z; The correlation characteristics between adjacent coefficients in the image block to be processed are extracted through the asymmetric generalized Gaussian distribution AGGD of the product of adjacent coefficients. The specific formula is as follows: Where γ represents the shape parameter of the asymmetric generalized Gaussian distribution; β l represents the scale parameter on the left side of the asymmetric generalized Gaussian distribution; βr represents the scale parameter on the right side of the asymmetric generalized Gaussian distribution; m represents the product of adjacent coefficients; f(x; γ, β l ,β r ) represents the probability density function of the asymmetric generalized Gaussian distribution of m; Calculate the mean of the distribution using the following formula: η represents the mean of the distribution; The statistical features of natural scenes are extracted based on the fitted statistical distribution, the correlation characteristics between adjacent coefficients and the mean of the asymmetric generalized Gaussian distribution.
6. A device for evaluating screen content video quality based on deep and shallow spatiotemporal features according to the method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to any one of claims 1 to 5, characterized in that: include: A screen content video extraction module, used to obtain a screen content video, and to extract a number of video blocks and a number of key frames from the screen content video; A model construction and training module, used to construct and train a screen content video quality evaluation model, and to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module; A video quality score evaluation module is used to input the plurality of video blocks and the plurality of key frames into a trained screen content video quality evaluation model; each key frame is respectively input into a deep feature extraction module, an edge feature extraction module, a color extraction module and a natural scene statistics module to obtain deep semantic features, edge features of text images, color features of cartoon images and natural scene statistics features in each key frame; The deep semantic features, edge features, color features and natural scene statistical features are spliced to obtain spatial features; the video blocks are respectively input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistical module to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced and input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.
Citation Information
Patent Citations
Image feature enhancement method and device, and storage medium
CN113947545A
Video quality evaluation method and system
CN119135879A