Screen content video quality evaluation method and device based on deep and shallow layer spatial-temporal characteristics

By constructing and training a dual-branch screen content video quality evaluation model containing spatial and temporal feature extraction branches, the problem of difficulty in evaluating the video quality of screen content in the prior art is solved, and higher evaluation accuracy and visual quality optimization are achieved.

CN120031869AActive Publication Date: 2025-05-23HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510495985.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate and optimize the visual quality of screen content videos, especially when facing various noise interferences and different types of content.

Method used

The video quality evaluation model of dual-branch screen content based on deep and shallow spatial and temporal characteristics is adopted, and the video quality of screen content is effectively evaluated through spatial feature extraction branch and temporal feature extraction branch, airspace time domain fusion module and quality regression module.

Benefits of technology

It improves the accuracy and adaptability of video quality evaluation, optimizes the visual quality of video, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031869A_ABST
    Figure CN120031869A_ABST
Patent Text Reader

Abstract

The invention discloses a screen content video quality evaluation method and device based on deep and shallow layer spatial-temporal characteristics, and relates to the field of video evaluation, and the method comprises the steps: obtaining a screen content video, and extracting a video block and a key frame from the screen content video, constructing a double-branch screen content video quality evaluation model comprising a spatial feature extraction branch, a time feature extraction branch, a space domain and time domain fusion module and a quality regression module; a key frame is input into a spatial feature extraction branch to obtain spatial features, a video block is input into a time feature extraction branch to obtain time features, the spatial features and the time features are spliced and then integrated through a spatial domain and time domain fusion module, and finally a video quality score is output through a quality regression module. According to the method, the double-branch screen content video quality evaluation model containing the space and time feature extraction branches is constructed and trained, so that the screen content video quality is effectively evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video evaluation, and in particular to a method and system for evaluating the quality of screen content video based on deep and shallow spatiotemporal features. Background Art

[0002] With the rapid development of artificial intelligence and digital multimedia technology and the popularity of various portable devices, screen content videos are increasingly used in digital multimedia application fields (such as online education, video conferencing, webcasting, game videos, etc.), providing users with a more flexible and free multimedia experience. Unlike natural videos, screen content videos are a mixture of images and computer-generated text / image areas.

[0003] MSCN (Mean Subtracted Contrast Normalized) is a widely used technology in the field of image processing and quality assessment, mainly used in the feature extraction and preprocessing stages. It eliminates the mean in the local area of ​​the image and normalizes it, thereby reducing the impact of factors such as illumination changes or noise interference on image quality assessment. Specifically, MSCN analyzes the neighborhood information of each pixel, calculates the local mean of the point, subtracts this local mean from the original pixel value, and then performs contrast normalization on the result. This technology is particularly suitable for enhancing structural information in images while suppressing unstructured noise components, making the subsequent feature extraction process more effective. In screen content video quality assessment, MSCN is used to process the color moment feature vector of cartoon images to eliminate mutual interference between neighborhood information, thereby improving the accuracy and reliability of color feature extraction. This method is crucial to improving the accuracy of overall video quality assessment.

[0004] 3D LOG (3D Laplace of Gaussian) is a technique for capturing the spatiotemporal features of screen content in videos. It is based on an extension of the 2D Laplace of Gaussian (LoG) operator, which is widely used in image processing to detect edge and texture information in images. In the context of video processing, 3D LOG expands the temporal dimension of consecutive frames, allowing it to not only extract spatial features within a single frame, but also capture the changing features between frames. This makes 3D LOG particularly suitable for analyzing fast-changing text, icons, and other computer-generated content that is unique to screen content videos, and helps identify quality degradation due to compression or transmission.

[0005] 3D NSS (3D Natural Scene Statistics) is a method specifically designed to capture the spatiotemporal characteristics of natural scenes in videos. Unlike 3D LOG, which focuses on screen content, 3D NSS looks at dynamic elements in natural scenes, such as landscapes, people, etc., which usually have visual characteristics and motion patterns different from computer-generated content. By analyzing the spatial structure of natural scenes in video sequences and their temporal evolution, 3D NSS can effectively simulate the human visual system's perception of natural video quality. This method is particularly useful for evaluating the quality of videos containing a large number of natural scenes, because it can quantify the changes in visual experience caused by distortion, thus providing an important basis for video quality evaluation.

[0006] In different processing stages of screen content videos (such as acquisition, transmission, display, etc.), they will inevitably be disturbed by various noises, resulting in varying degrees of degradation of perceived quality, which seriously affects the user experience. Therefore, it is urgent to develop an effective video quality assessment method to optimize the visual quality of screen content videos and improve the performance of related tasks. Summary of the invention

[0007] In order to solve the above problems, the present invention proposes a screen content video quality evaluation method and system based on deep and shallow spatiotemporal features. By constructing and training a dual-branch screen content video quality evaluation model including spatial and temporal feature extraction branches, effective evaluation of screen content video quality is achieved. At the same time, different strategies are adopted to extract features for different types of content, which improves the evaluation accuracy and optimizes the visual quality of the video.

[0008] The specific plan is as follows:

[0009] On the one hand, the screen content video quality evaluation method based on deep and shallow spatiotemporal features includes:

[0010] Acquire a screen content video, and extract a plurality of video blocks and a plurality of key frames from the screen content video;

[0011] Constructing and training a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module;

[0012] The several video blocks and several key frames are input into the trained screen content video quality evaluation model; specifically, each key frame is input into the deep feature extraction module, the edge feature extraction module, the color extraction module and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image and the natural scene statistics features in each key frame; the deep semantic features, edge features, color features and natural scene statistics features are spliced ​​to obtain spatial features; the video blocks are input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistics module respectively to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced ​​and then input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.

[0013] Furthermore, each key frame is input into the deep feature extraction module to obtain the deep semantic features in each key frame. The calculation formula is as follows:

[0014] ;

[0015] ;

[0016] ;

[0017] ;

[0018] in, = , represents the feature map extracted based on key frame x at the kth stage, , Representation feature map Height, Representation feature map The width of Representation feature map The number of channels; Representation feature map The global mean of Representation feature map The standard deviation of ; k is the identifier of the stage number; represents the spatial feature representation of the kth stage; Represents deep semantic features; represents the feature concatenation function; represents global average pooling; Represents global standard deviation pooling.

[0019] Furthermore, each key frame is input into the edge feature extraction module to obtain the edge features of the text image in each key frame. Specifically, the edge feature extraction module uses a two-dimensional Gaussian Laplace operator to extract edge features from the text image; wherein the kernel function calculation formula of the two-dimensional Gaussian Laplace operator is as follows:

[0020] ;

[0021] ;

[0022] ;

[0023] in, Represents the horizontal coordinate of the pixel of the input text image; Represents the vertical coordinate of the pixel of the input text image; represents a two-dimensional Gaussian function; Indicates that the input image is at coordinates The pixel value at ; represents the standard deviation of the Gaussian function; represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction; Indicates Laplace operator operation; represents the kernel function of the two-dimensional Laplacian of Gaussian; Represents the result of processing the image by the two-dimensional Laplacian of Gaussian operator; * represents the convolution operation.

[0024] Furthermore, each key frame is input into the color extraction module to obtain the color features of the cartoon image in each key frame, specifically including:

[0025] The first color moment, second color moment, and third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formula is as follows:

[0026] ;

[0027] ;

[0028] ;

[0029] Among them, I represents the channel map; M() represents the mean operation; Represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; Represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; Represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map;

[0030] The color moment feature vector of the cartoon image in each key frame is obtained based on the first color moment, the second color moment and the third color moment. ;in, Represents the first color moment on the hue channel; Represents the second color moment on the hue channel; Represents the third color moment on the hue channel; Represents the first color moment on the saturation channel; Represents the second color moment on the saturation channel; Represents the third color moment on the saturation channel; Represents the first color moment on the brightness channel; Represents the second color moment on the brightness channel; Represents the third color moment on the brightness channel;

[0031] The contrast normalization coefficient MSCN is used to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows:

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] in, represents the weight value of the two-dimensional circularly symmetric Gaussian weighted filter at the position (k, l), k represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the vertical direction, and l represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the horizontal direction; Represents the normalized pixel value of the key frame at position (x, y); Indicates that the keyframe is at position The pixel value at ; represents the local average; represents the local standard deviation; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; Indicates that The pixel value at position (x,y) after applying the average filter; Indicates the horizontal offset when the average filter moves; Indicates the vertical offset of the average filter when it moves; (h,w) represents the weight value of the average filter at position (h,w); represents the height of the averaging filter, represents the width of the average filter, H=W=3;

[0037] calculate and Entropy to obtain the color entropy of each channel map , the calculation formula is as follows:

[0038] ;

[0039] in, Represents the color entropy of each channel map; represents the frequency of the intensity value u; i represents different HSV channels; express or ;

[0040] Get the color entropy feature vector based on the color entropy of each channel image ;

[0041] It represents the color entropy of the key frame after MSCN processing in the hue channel; Represents the color entropy of the key frame after MSCN and average filtering in the hue channel; It represents the color entropy of the key frame after MSCN processing in the saturation channel; Represents the color entropy of the key frame after MSCN and average filtering in the saturation channel; It represents the color entropy of the key frame after MSCN processing in the brightness channel; Represents the color entropy of the key frame after MSCN and average filtering in the brightness channel;

[0042] F CE and F CM Perform stitching to obtain the color features of the cartoon image in each key frame .

[0043] Furthermore, each key frame is input into the natural scene statistics module to obtain the natural scene statistics features in each key frame, including:

[0044] The local mean removal and normalization method is used to preprocess each key frame. The calculation formula is as follows:

[0045] ;

[0046] ;

[0047] ;

[0048] in, Indicates that each key frame is at position The pixel value of represents the local mean; Indicates the offset in the vertical direction; Indicates the offset in the horizontal direction; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; represents the local standard deviation; Indicates that the key frame is at position after processing The pixel value of represents a two-dimensional circularly symmetric Gaussian weighted filter at position The weight value at .

[0049] The preprocessed key frame is divided into 96*96 image blocks; 0.7 times of the maximum sharpness value in the image block is set as the sharpness threshold, and the image blocks with sharpness higher than this threshold are selected as the image blocks to be processed;

[0050] The statistical distribution of the image block to be processed is calculated, and the statistical distribution of the image block to be processed is fitted by the generalized Gaussian distribution GGD; the calculation formula of the generalized Gaussian distribution GGD is as follows:

[0051] ;

[0052] Where z represents the image coefficient after preprocessing; represents the shape parameter, which is used to describe the sharpness of the distribution; represents the scale parameter, which is used to describe the width of the distribution; represents the gamma function; represents the probability density function of the generalized Gaussian distribution of z;

[0053] The correlation characteristics between adjacent coefficients in the image block to be processed are extracted through the asymmetric generalized Gaussian distribution AGGD of the product of adjacent coefficients. The specific formula is as follows:

[0054] ;

[0055] Where γ represents the shape parameter of the asymmetric generalized Gaussian distribution; represents the scale parameter on the left side of the asymmetric generalized Gaussian distribution; βr represents the scale parameter on the right side of the asymmetric generalized Gaussian distribution; m represents the product of adjacent coefficients; represents the probability density function of the asymmetric generalized Gaussian distribution of m;

[0056] Calculate the mean of the distribution using the following formula:

[0057] ;

[0058] represents the mean of the distribution;

[0059] The statistical features of natural scenes are extracted based on the fitted statistical distribution, the correlation characteristics between adjacent coefficients and the mean of the asymmetric generalized Gaussian distribution.

[0060] Furthermore, the spatial features and temporal features are spliced ​​and input into the spatial and temporal fusion module to obtain fused spatial features and temporal features, which specifically include:

[0061] Spatial features extracted from the i-th keyframe , extract temporal features from the i-th video segment , the spatial features and time characteristics After the learnable features are fused, they are input into the multi-layer perceptron MLP to obtain the fused spatial features and temporal features. The calculation formula is as follows:

[0062] ;

[0063] in, It represents a learnable feature fusion function. Specifically, the learnable feature fusion function is as follows: feature mapping is first performed through a fully connected layer containing 1024 neurons, and then a RELU activation function is applied for nonlinear transformation; Represents the features after fusion of the i-th video clip; Represents feature concatenation.

[0064] Furthermore, the fused spatial features and temporal features are input into the quality regression module to obtain the final video quality score. Specifically, the fused spatial features and temporal features are regressed into the quality score of the video clip through the two fully connected layers of the quality regression module. The calculation formula is as follows:

[0065] ;

[0066] in, represents the quality score of the i-th video clip; FC() represents the fully connected layer; Represents the features after fusion of the i-th video clip;

[0067] Average the quality scores of the video clips of k video segments to get the final video quality score:

[0068] ;

[0069] Among them, Q represents the final video quality score, and k represents the number of video clips.

[0070] On the other hand, a device for evaluating screen content video quality based on deep and shallow spatiotemporal features is characterized by comprising:

[0071] A screen content video extraction module, used to obtain a screen content video, and to extract a number of video blocks and a number of key frames from the screen content video;

[0072] A model construction and training module, used to construct and train a screen content video quality evaluation model, and to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module;

[0073] The video quality score evaluation module is used to input the several video blocks and several key frames into the trained screen content video quality evaluation model; specifically, each key frame is input into the deep feature extraction module, the edge feature extraction module, the color extraction module and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image and the natural scene statistics features in each key frame; the deep semantic features, edge features, color features and natural scene statistics features are spliced ​​to obtain spatial features; the video blocks are input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistics module respectively to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced ​​and then input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.

[0074] The present invention adopts the above technical solution and has the following beneficial effects:

[0075] (1) The present invention effectively reduces the computational complexity by extracting deep and shallow spatiotemporal features. It also adopts different strategies to extract spatial and temporal features for different types of content (text, cartoons, and natural images), thereby improving the accuracy of quality assessment.

[0076] (2) The present invention utilizes a dual-branch screen content video quality evaluation model including a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module to effectively evaluate the quality of the screen content video and optimize the visual quality of the video;

[0077] (3) The present invention combines the Conv Next network and the SlowFast network to extract spatial and temporal features respectively, and uses a multi-layer perceptron and a fully connected layer for feature fusion and quality scoring, thereby further enhancing the adaptability to videos with different types of screen content and the accuracy of the evaluation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a flow chart of a method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to an embodiment of the present invention;

[0079] Figure 2 A schematic diagram of a process for evaluating screen content video quality according to an embodiment of the present invention;

[0080] Figure 3 A diagram of a device for evaluating screen content video quality based on deep and shallow spatiotemporal features according to an embodiment of the present invention. DETAILED DESCRIPTION

[0081] The present invention is further described in detail below in conjunction with the embodiments and drawings, but the embodiments of the present invention are not limited thereto. Figure 1 As shown, the screen content video quality evaluation method based on deep and shallow spatiotemporal features of the present invention includes:

[0082] S1, obtaining a screen content video, and extracting a plurality of video blocks and a plurality of key frames from the screen content video.

[0083] Specifically, in this embodiment, given a screen content video V, the number of frames and the frame rate of the video are l and r respectively. The video V is divided into two time intervals: Divide consecutive video blocks, denoted as ,in , in each video block There are Frame, denoted as ,in, represents the jth frame in the i-th block, select each video block The first frame is used as the key frame to extract spatial features. Considering that temporal information is not sensitive to resolution, the video blocks are For all frames in the video, the video resolution is adjusted to a low resolution to extract temporal features.

[0084] S2, constructing and training a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module and a three-dimensional natural scene statistics module.

[0085] Specifically, Figure 2 As shown, the screen content video quality assessment model proposed in this embodiment has two branches, namely a spatial feature extraction branch and a temporal feature extraction branch, which are used to calculate spatial features and temporal features respectively. First, a Conv Next neural network is constructed for the key frames, and corresponding modules are designed for the text images, cartoon images, and natural images in the key frames, and finally the spatial features are obtained. A SlowFast (3D CNN) network is constructed for the video clips, and the video clips are processed using 3D-LOG and 3D-NSS to finally obtain the temporal features. Combining spatial features and temporal features, and obtaining a quality score through quality regression can effectively realize the quality assessment of screen content videos.

[0086] S3, input the several video blocks and several key frames into the trained screen content video quality evaluation model; specifically, input each key frame into the deep feature extraction module, the edge feature extraction module, the color extraction module and the natural scene statistics module respectively, and obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image and the natural scene statistics features in each key frame; splice the deep semantic features, edge features, color features and natural scene statistics features to obtain spatial features; input the video blocks into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistics module respectively, and obtain temporal features, natural spatiotemporal features and time features; splice the spatial features and time features and input them into the spatial-temporal fusion module to obtain fused spatial features and time features; input the fused spatial features and time features into the quality regression module to obtain the final video quality score.

[0087] Specifically, each key frame is input into the deep feature extraction module to obtain the deep semantic features in each key frame. The calculation formula is as follows:

[0088] ;

[0089] ;

[0090] ;

[0091] ;

[0092] in, = , represents the feature map extracted based on key frame x at the kth stage, , Representation feature map Height, Representation feature map The width of Representation feature map The number of channels; Representation feature map The global mean of Representation feature map The standard deviation of ; k is the identifier of the stage number; represents the spatial feature representation of the kth stage; Represents deep semantic features; represents the feature concatenation function; represents global average pooling; Represents global standard deviation pooling.

[0093] Specifically, each key frame is input into the edge feature extraction module to obtain the edge features of the text image in each key frame. Specifically, the edge feature extraction module uses a two-dimensional Gaussian Laplace operator to extract edge features from the text image; wherein the kernel function calculation formula of the two-dimensional Gaussian Laplace operator is as follows:

[0094] ;

[0095] ;

[0096] ;

[0097] in, Represents the horizontal coordinate of the input text image pixel; Represents the vertical coordinate of the pixel of the input text image; represents a two-dimensional Gaussian function; Indicates that the input image is at coordinates The pixel value at ; represents the standard deviation of the Gaussian function; represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction; Indicates Laplace operator operation; represents the kernel function of the two-dimensional Laplacian of Gaussian; Represents the result of processing the image by the two-dimensional Laplacian of Gaussian operator; * represents the convolution operation.

[0098] Specifically, each key frame is input into the color extraction module to obtain the color features of the cartoon image in each key frame, including:

[0099] The first color moment, second color moment, and third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formula is as follows:

[0100] ;

[0101] ;

[0102] ;

[0103] Among them, I represents the channel map; M() represents the mean operation; Represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; Represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; Represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map;

[0104] The color moment feature vector of the cartoon image in each key frame is obtained based on the first color moment, the second color moment and the third color moment. ;in, Represents the first color moment on the hue channel; Represents the second color moment on the hue channel; Represents the third color moment on the hue channel; Represents the first color moment on the saturation channel; Represents the second color moment on the saturation channel; Represents the third color moment on the saturation channel; Represents the first color moment on the brightness channel; Represents the second color moment on the brightness channel; Represents the third color moment on the brightness channel;

[0105] The contrast normalization coefficient MSCN is used to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows:

[0106] ;

[0107] ;

[0108] ;

[0109] ;

[0110] in, represents the weight value of the two-dimensional circularly symmetric Gaussian weighted filter at the position (k, l), k represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the vertical direction, and l represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the horizontal direction; Represents the normalized pixel value of the key frame at position (x, y); Indicates that the keyframe is at position The pixel value at ; represents the local average; represents the local standard deviation; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; Indicates that The pixel value at position (x,y) after applying the average filter; Indicates the horizontal offset when the average filter moves; Indicates the vertical offset of the average filter when it moves; (h,w) represents the weight value of the average filter at position (h,w); represents the height of the averaging filter, represents the width of the average filter, H=W=3;

[0111] calculate and Entropy to obtain the color entropy of each channel map , the calculation formula is as follows:

[0112] ;

[0113] in, Represents the color entropy of each channel map; represents the frequency of the intensity value u; i represents different HSV channels; express or ;

[0114] Get the color entropy feature vector based on the color entropy of each channel image ;

[0115] It represents the color entropy of the key frame after MSCN processing in the hue channel; Represents the color entropy of the key frame after MSCN and average filtering in the hue channel; It represents the color entropy of the key frame after MSCN processing in the saturation channel; Represents the color entropy of the key frame after MSCN and average filtering in the saturation channel; Represents the color entropy of the key frame in the brightness channel after MSCN processing Represents the color entropy of the key frame after MSCN and average filtering in the brightness channel;

[0116] F CE and Perform stitching to obtain the color features of the cartoon image in each key frame .

[0117] Specifically, in this embodiment, for the natural scene image in the screen content video, the natural scene statistics module is used to obtain the natural scene features. Specifically, for the key frame, the natural scene statistics module extracts two features ( , ); extract from the non-generalized Gaussian distribution (AGGD) in four directions: horizontal, vertical, diagonal, and anti-diagonal (γ, , ) for a total of 12 features; the mean of the asymmetric generalized Gaussian distribution in four directions is calculated, for a total of 4 features; therefore, a total of 36 features are extracted. These 36 features are the statistical features of natural scenes.

[0118] Specifically, extracting natural scene statistical features in each key frame through the natural scene statistical module for natural images specifically includes:

[0119] The local mean removal and normalization method is used to preprocess each key frame. The calculation formula is as follows:

[0120] ;

[0121] ;

[0122] ;

[0123] in, Indicates that each key frame is at position The pixel value of represents the local mean; Indicates the offset in the vertical direction; Indicates the offset in the horizontal direction; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; Represents the local standard deviation Indicates that the key frame is at position after processing The pixel value of Represents the weight value of a two-dimensional circularly symmetric Gaussian weighted filter at position (k, l).

[0124] The preprocessed key frame is divided into 96*96 image blocks; 0.7 times of the maximum sharpness value in the image block is set as the sharpness threshold, and the image blocks with sharpness higher than this threshold are selected as the image blocks to be processed;

[0125] The statistical distribution of the image block to be processed is calculated, and the statistical distribution of the image block to be processed is fitted by the generalized Gaussian distribution GGD; the calculation formula of the generalized Gaussian distribution GGD is as follows:

[0126] ;

[0127] Where z represents the image coefficient after preprocessing; represents the shape parameter, which is used to describe the sharpness of the distribution; represents the scale parameter, which is used to describe the width of the distribution; represents the gamma function; represents the probability density function of the generalized Gaussian distribution of z;

[0128] The correlation characteristics between adjacent coefficients in the image block to be processed are extracted through the asymmetric generalized Gaussian distribution AGGD of the product of adjacent coefficients. The specific formula is as follows:

[0129] ;

[0130] Where γ represents the shape parameter of the asymmetric generalized Gaussian distribution; represents the scale parameter on the left side of the asymmetric generalized Gaussian distribution; represents the scale parameter on the right side of the asymmetric generalized Gaussian distribution; m represents the product of adjacent coefficients; represents the probability density function of the asymmetric generalized Gaussian distribution of m;

[0131] Calculate the mean of the distribution using the following formula:

[0132] ;

[0133] represents the mean of the distribution;

[0134] The statistical features of natural scenes are extracted based on the fitted statistical distribution, the correlation characteristics between adjacent coefficients and the mean of the asymmetric generalized Gaussian distribution.

[0135] Specifically, in this embodiment, for the text image in the screen content video, the LOG operator is used to extract edge features of the text image. The full name of the LOG operator is Gaussian Laplacian operator. The Laplacian edge detection operator does not smooth the image, so it is very sensitive to noise. Therefore, the image can be first Gaussian smoothed and then convolved with the Laplacian operator.

[0136] Specifically, in this embodiment, in order to retain sufficient temporal information and reduce computational complexity, the video is uniformly segmented into segments with lower resolution to extract temporal features. Specifically, two branches are designed to obtain deep learning-based features and handcrafted features respectively. For deep learning-based features, the pre-trained SlowFast network is used to extract motion features for each video segment:

[0137] ;

[0138] in, Indicates video clips, Φ(·) represents the motion feature extraction operation, Indicates that from Motion features extracted from video clips;

[0139] Based on 3D-LOG, screen spatiotemporal features are extracted, 2D-LOG is extended to 3D-LOG, and it is used to capture screen spatiotemporal features. The calculation formula of 3D-LOG is derived from 2D-LOG:

[0140] ;

[0141] in, Represents the response value of the three-dimensional Gaussian Laplacian filter at position (x, y, t), where x, y, and t represent the spatial horizontal coordinate, spatial vertical coordinate, and time coordinate of the pixel in the screen content video. exp() represents the exponential function. represents the standard deviation of the Gaussian kernel. Take 1.485. By convolving the screen content video clip with the three-dimensional Gaussian Laplacian filter, the screen spatiotemporal features are extracted, namely:

[0142] ;

[0143] in, Represents the pixel value of the screen content video fragment at position (x, y, t), Represents the screen spatiotemporal feature map, Represents the response value of the three-dimensional Gaussian Laplacian filter at position (x, y, t). x, y, t represent the spatial horizontal coordinate, spatial vertical coordinate, and time coordinate of the pixel in the screen content video. * represents the convolution operation.

[0144] Based on 3D-NSS, natural spatiotemporal feature extraction is performed, and the MSCN coefficients are extended to 3D space to describe the natural spatiotemporal features of screen content video clips, similar to the aforementioned screen spatiotemporal feature extraction. The brightness of the screen content video clips is normalized to calculate the 3D-MSCN coefficients, and it is used to capture the natural spatiotemporal quality degradation, that is:

[0145] ;

[0146] ;

[0147] ;

[0148] ;

[0149] in, Represents the pixel value of the screen content video fragment at position (x, y, t), and c=1 is a positive constant to ensure the stability of the formula. represents the mean of the local area centered at (x, y, t), represents the standard deviation of the local area centered at (x, y, t), c represents a constant, which is 1 here, represents the value of the video clip at position (x, y, t) after mean subtraction and contrast normalization. Represents the difference between the pixel value at (x, y, t) and the local mean, where x, y, and t represent the spatial horizontal coordinate, spatial vertical coordinate, and time coordinate of the pixel in the screen content video, respectively. Represents the weight value of the three-dimensional circularly symmetric Gaussian filter at the position (p,u,v), p represents the offset of the three-dimensional circularly symmetric Gaussian filter on the x-axis, u represents the offset of the three-dimensional circularly symmetric Gaussian filter on the y-axis, v represents the offset of the three-dimensional circularly symmetric Gaussian filter on the time axis, P represents the maximum offset on the x-axis, U represents the maximum offset on the y-axis, and V represents the maximum offset on the time axis. . Represents the pixel value at the offset (p,u,v) relative to the center position (x,y,t) in the local window.

[0150] Specifically, the spatial features and the temporal features are spliced ​​and input into the spatial-temporal fusion module to obtain fused spatial features and temporal features, which specifically includes:

[0151] Spatial features extracted from the i-th keyframe , extract temporal features from the i-th video segment , the spatial features and time characteristics After the learnable features are fused, they are input into the multi-layer perceptron MLP to obtain the fused spatial features and temporal features; the calculation formula is as follows:

[0152]

[0153] in, It represents a learnable feature fusion function. Specifically, the learnable feature fusion function is as follows: feature mapping is first performed through a fully connected layer containing 1024 neurons, and then a RELU activation function is applied for nonlinear transformation; represents the fused features of the i-th video clip; Represents feature concatenation.

[0154] Specifically, the fused spatial features and temporal features are input into the quality regression module to obtain the final video quality score. Specifically, the fused spatial features and temporal features are regressed into the quality score of the video clip through the two fully connected layers of the quality regression module. The calculation formula is as follows:

[0155] ;

[0156] in, represents the quality score of the i-th video clip; FC() represents the fully connected layer; represents the fused features of the i-th video clip;

[0157] Average the quality scores of the video clips of k video segments to get the final video quality score:

[0158] ;

[0159] Among them, Q represents the final video quality score, and k represents the number of video clips.

[0160] In this embodiment, the proposed screen content video quality evaluation model based on deep and shallow spatiotemporal features uses Pytorch to build an environment, and uses NVIDIA RTX A6000 GPU for experiments; the minimum resolution size of the key frame is adjusted to 520 while maintaining the original aspect ratio. During training, the key frame is randomly cropped to 448*448. For video blocks, the video block resolution is adjusted to 224*224 during training, the batch size is 64, the number of training rounds is 100, the initial learning rate is set to 0.00001, and the Adam optimizer is used for training. The experimental dataset is SCVD, which contains 16 reference screen content videos of different content scenes and 800 distorted screen content videos, 80% of which are used for training and the remaining 20% ​​are used for testing. The Spearman rank correlation coefficient (SROCC), Pearson linear correlation coefficient (PLCC) and root mean square error (RMSE) are selected to evaluate the performance of the above model.

[0161] like Figure 3 As shown, this embodiment also discloses a screen content video quality evaluation device based on deep and shallow spatiotemporal features, including:

[0162] A screen content video extraction module 31, used to obtain a screen content video, and to extract a number of video blocks and a number of key frames from the screen content video;

[0163] A model building and training module 32 is used to build and train a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module;

[0164] The video quality score evaluation module 33 is used to input the several video blocks and several key frames into the trained screen content video quality evaluation model; specifically, each key frame is input into the deep feature extraction module, the edge feature extraction module, the color extraction module and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image and the natural scene statistics features in each key frame; the deep semantic features, edge features, color features and natural scene statistics features are spliced ​​to obtain spatial features; the video blocks are input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistics module respectively to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced ​​and input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.

[0165] The specific implementation of the screen content video quality assessment device based on deep and shallow spatiotemporal features is the same as the screen content video quality assessment method based on deep and shallow spatiotemporal features, and will not be repeated in this embodiment.

[0166] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, it should be understood by those skilled in the art that various changes may be made to the present invention in form and details without departing from the spirit and scope of the present invention as defined by the appended claims, all of which are within the scope of protection of the present invention.

Claims

1. A method for evaluating screen content video quality based on deep and shallow spatiotemporal features, characterized in that: include: Acquire a screen content video, and extract a plurality of video blocks and a plurality of key frames from the screen content video; Constructing and training a screen content video quality evaluation model to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module; Input the several video blocks and several key frames into the trained screen content video quality evaluation model; specifically, input each key frame into the deep feature extraction module, the edge feature extraction module, the color extraction module and the natural scene statistics module respectively to obtain the deep semantic features, the edge features of the text image, the color features of the cartoon image and the natural scene statistics features in each key frame; The deep semantic features, edge features, color features and natural scene statistical features are spliced ​​to obtain spatial features; the video blocks are respectively input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistical module to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced ​​and input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.

2. The screen content video quality evaluation method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the deep feature extraction module to obtain the deep semantic features in each key frame. The calculation formula is as follows: ; ; ; ; in, = , represents the feature map extracted based on key frame x at the kth stage, , Representation feature map Height, Representation feature map The width of Representation feature map The number of channels; Representation feature map The global mean of Representation feature map The standard deviation of ; k is the identifier of the stage number; represents the spatial feature representation of the kth stage; Represents deep semantic features; represents the feature concatenation function; represents global average pooling; Represents global standard deviation pooling.

3. The screen content video quality assessment method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the edge feature extraction module to obtain the edge features of the Chinese text image in each key frame. Specifically, the edge feature extraction module uses a two-dimensional Gaussian Laplacian operator to extract edge features from the text image; wherein the kernel function calculation formula of the two-dimensional Gaussian Laplacian operator is as follows: ; ; ; in, Represents the horizontal coordinate of the pixel of the input text image; Represents the vertical coordinate of the pixel of the input text image; represents a two-dimensional Gaussian function; Indicates that the input image is at coordinates The pixel value at ; represents the standard deviation of the Gaussian function; represents the exponential function; represents the second-order partial derivative of the Gaussian function in the x direction; represents the second-order partial derivative of the Gaussian function in the y direction; Indicates Laplace operator operation; represents the kernel function of the two-dimensional Laplacian of Gaussian; Represents the result of processing the image by the two-dimensional Laplacian of Gaussian operator; * represents the convolution operation.

4. The screen content video quality assessment method based on deep and shallow spatiotemporal features according to claim 1 is characterized in that: Each key frame is input into the color extraction module to obtain the color features of the cartoon image in each key frame, including: The first color moment, second color moment, and third color moment of the cartoon image in each key frame are obtained through the color extraction module to describe the color changes caused by different distortions. The calculation formula is as follows: ; ; ; Among them, I represents the channel map; M() represents the mean operation; Represents the first color moment, which is used to describe the overall color average of the cartoon image on the channel map; Represents the second color moment, which is used to describe the local contrast change of the cartoon image on the channel map; Represents the third color moment, which is used to describe the asymmetry of the color distribution of the cartoon image on the channel map; The color moment feature vector of the cartoon image in each key frame is obtained based on the first color moment, the second color moment and the third color moment. ;in, Represents the first color moment on the hue channel; Represents the second color moment on the hue channel; Represents the third color moment on the hue channel; Represents the first color moment on the saturation channel; Represents the second color moment on the saturation channel; Represents the third color moment on the saturation channel; Represents the first color moment on the brightness channel; Represents the second color moment on the brightness channel; Represents the third color moment on the brightness channel; The contrast normalization coefficient MSCN is used to eliminate the mutual interference of neighborhood information in the color moment feature vector of the cartoon image in each key frame. The calculation formula is as follows: ; ; ; ; in, represents the weight value of the two-dimensional circularly symmetric Gaussian weighted filter at the position (k, l), k represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the vertical direction, and l represents the offset of the two-dimensional circularly symmetric Gaussian weighted filter in the horizontal direction; Represents the normalized pixel value of the key frame at position (x, y); Indicates that the keyframe is at position The pixel value at ; represents the local average; represents the local standard deviation; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; Indicates that The pixel value at position (x,y) after applying the average filter; Indicates the horizontal offset when the average filter moves; Indicates the vertical offset of the average filter when it moves; (h,w) represents the weight value of the average filter at position (h,w); represents the height of the averaging filter, represents the width of the average filter, H=W=3; calculate and Entropy to obtain the color entropy of each channel map , the calculation formula is as follows: ; in, Represents the color entropy of each channel map; represents the frequency of the intensity value u; i represents different HSV channels; express or ; Get the color entropy feature vector based on the color entropy of each channel image ; It represents the color entropy of the key frame after MSCN processing in the hue channel; Represents the color entropy of the key frame after MSCN and average filtering in the hue channel; It represents the color entropy of the key frame after MSCN processing in the saturation channel; Represents the color entropy of the key frame after MSCN and average filtering in the saturation channel; It represents the color entropy of the key frame after MSCN processing in the brightness channel; Represents the color entropy of the key frame after MSCN and average filtering in the brightness channel; F CE and F CM Perform stitching to obtain the color features of the cartoon image in each key frame .

5. The method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to claim 1, characterized in that: Each key frame is input into the natural scene statistics module to obtain the natural scene statistical features in each key frame, including: The local mean removal and normalization method is used to preprocess each key frame. The calculation formula is as follows: ; ; ; in, Indicates that each key frame is at position The pixel value of represents the local mean; Indicates the offset in the vertical direction; Indicates the offset in the horizontal direction; Indicates the maximum offset in the horizontal direction; Indicates the maximum offset in the vertical direction; K=L=3; represents the local standard deviation; Indicates that the key frame is at position after processing The pixel value of represents a two-dimensional circularly symmetric Gaussian weighted filter at position The weight value at ; The preprocessed key frame is divided into 96*96 image blocks; 0.7 times of the maximum sharpness value in the image block is set as the sharpness threshold, and the image blocks with sharpness higher than this threshold are selected as the image blocks to be processed; The statistical distribution of the image block to be processed is calculated, and the statistical distribution of the image block to be processed is fitted by the generalized Gaussian distribution GGD; the calculation formula of the generalized Gaussian distribution GGD is as follows: ; Where z represents the image coefficient after preprocessing; represents the shape parameter, which is used to describe the sharpness of the distribution; represents the scale parameter, which is used to describe the width of the distribution; represents the gamma function; represents the probability density function of the generalized Gaussian distribution of z; The correlation characteristics between adjacent coefficients in the image block to be processed are extracted through the asymmetric generalized Gaussian distribution AGGD of the product of adjacent coefficients. The specific formula is as follows: ; Where γ represents the shape parameter of the asymmetric generalized Gaussian distribution; represents the scale parameter on the left side of the asymmetric generalized Gaussian distribution; βr represents the scale parameter on the right side of the asymmetric generalized Gaussian distribution; m represents the product of adjacent coefficients; represents the probability density function of the asymmetric generalized Gaussian distribution of m; Calculate the mean of the distribution using the following formula: ; represents the mean of the distribution; The statistical features of natural scenes are extracted based on the fitted statistical distribution, the correlation characteristics between adjacent coefficients and the mean of the asymmetric generalized Gaussian distribution.

6. The method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to claim 1, characterized in that: After splicing the spatial features and temporal features, they are input into the spatial and temporal fusion module to obtain the fused spatial and temporal features, which include: Spatial features extracted from the i-th keyframe , extract temporal features from the i-th video segment , the spatial features and time characteristics After the learnable features are fused, they are input into the multi-layer perceptron MLP to obtain the fused spatial features and temporal features. The calculation formula is as follows: ; in, It represents a learnable feature fusion function. Specifically, the learnable feature fusion function is as follows: feature mapping is first performed through a fully connected layer containing 1024 neurons, and then a RELU activation function is applied for nonlinear transformation; Represents the features after fusion of the i-th video clip; Represents feature concatenation.

7. The method for evaluating screen content video quality based on deep and shallow spatiotemporal features according to claim 1, characterized in that: The fused spatial features and temporal features are input into the quality regression module to obtain the final video quality score. Specifically, the fused spatial features and temporal features are regressed into the quality score of the video clip through the two fully connected layers of the quality regression module. The calculation formula is as follows: ; in, represents the quality score of the i-th video clip; FC() represents the fully connected layer; Represents the features after fusion of the i-th video clip; Average the quality scores of the video clips of k video segments to get the final video quality score: ; Among them, Q represents the final video quality score, and k represents the number of video clips.

8. A device for evaluating screen content video quality based on deep and shallow spatiotemporal features, characterized in that: include: A screen content video extraction module, used to obtain a screen content video, and to extract a number of video blocks and a number of key frames from the screen content video; A model construction and training module, used to construct and train a screen content video quality evaluation model, and to obtain a trained screen content video quality evaluation model; the screen content video quality evaluation model includes a spatial feature extraction branch, a temporal feature extraction branch, a spatial-temporal fusion module, and a quality regression module; the spatial feature extraction branch includes a deep feature extraction module based on a Conv Next network, an edge feature extraction module, a color extraction module, and a natural scene statistics module; the temporal feature extraction branch includes a temporal feature extraction module based on a SlowFast network, a three-dimensional Laplace-Gaussian module, and a three-dimensional natural scene statistics module; A video quality score evaluation module is used to input the plurality of video blocks and the plurality of key frames into a trained screen content video quality evaluation model; specifically, each key frame is input into a deep feature extraction module, an edge feature extraction module, a color extraction module and a natural scene statistics module, respectively, to obtain deep semantic features, edge features of text images, color features of cartoon images and natural scene statistics features in each key frame; The deep semantic features, edge features, color features and natural scene statistical features are spliced ​​to obtain spatial features; the video blocks are respectively input into the temporal feature extraction module, the three-dimensional Laplace Gaussian module and the natural scene statistical module to obtain temporal features, natural spatiotemporal features and time features; the spatial features and time features are spliced ​​and input into the spatial-temporal fusion module to obtain fused spatial features and time features; the fused spatial features and time features are input into the quality regression module to obtain the final video quality score.

Citation Information

Patent Citations

  • Image feature enhancement method and device, and storage medium

    CN113947545A

  • Video quality evaluation method and system

    CN119135879A

  • No-reference visual media assessment combining deep neural networks and models of human visual system and video content / distortion analysis

    US20210233259A1

  • Visual Quality Assessment-based Affine Transformation

    US20220182676A1

  • Method and apparatus for processing video data, training method and electronic device

    WO2024152523A1

Cited By

  • No-reference video quality evaluation method and equipment based on space-time edge feature extraction

    CN120726411A

  • Video quality evaluation method and system based on space-frequency combination and time sequence interaction

    CN121397207A