Video Quality Assessment Method Based on Semantic Changes

Through improved edge detection and multi-frequency component pooling technology, combined with the Transformer Block model, the insufficient capture of timing semantic changes in reference-free video quality evaluation is solved, and the comprehensiveness and accuracy of video quality evaluation is improved.

CN115620116BActive Publication Date: 2025-07-25FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211373056.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-07-25
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

Existing reference-free video quality evaluation methods are difficult to effectively capture the semantic changes in the video in timing, resulting in insufficient perception of phenomena such as video blur and motion distortion.

Method used

The improved edge detection operator is used to extract multi-scale features, combine multi-frequency component pooling and improved Transformer Block model, and establish the timing relationship between video frames by querying the semantic information changes of adjacent frames and predicting video quality.

Benefits of technology

It improves the ability to perceive motion distortion such as video jitter and artifacts, and enhances the comprehensiveness and accuracy of the video quality evaluation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620116B_ABST
    Figure CN115620116B_ABST
Patent Text Reader

Abstract

The present invention proposes a video quality assessment method based on semantic changes, including the following steps; Step S1: For videos of different scenes captured by a mobile device, extract edge features for each frame of the video; Step S2: Input the edges of each frame of the video and the original image into the spatial feature extraction network respectively to obtain the multi-scale spatial features of the video. At the same time, input the video into the temporal feature extraction network to obtain the multi-scale temporal features, and perform multi-frequency component pooling and standard pooling on the multi-scale features; Step S3: Combine the results after pooling to obtain the spatio-temporal features of the video, and reduce the dimension of the spatio-temporal features; Step S4: Input the dimension-reduced spatio-temporal features of the video into the quality prediction network to model the temporal relationship, and then predict the quality score of the overall video; The present invention can effectively extract the spatio-temporal features of the video and add semantic change information, so that the video distortion information obtained by the quality evaluation model is more comprehensive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a video quality assessment method based on semantic changes. Background Art

[0002] Video quality assessment is Video Quality Assessment (VQA). Its main task is to predict the perceived quality of a video clip by humans given a video clip. With the continuous development of video quality assessment in recent years, a large number of video quality assessment models with good evaluation effects and fast running speeds have emerged. Although certain results have been achieved in the research on existing full-reference and semi-reference video quality assessment, due to the dependence on the original video, these two methods are often not practical. Because in real life, the original undistorted video is often not easily obtained, and the videos we watch often need to be compressed and transmitted, and the mixed video distortion generated during this process is difficult to estimate. Therefore, no-reference video quality evaluation, and even no-reference quality evaluation based on user-generated videos, will have a wider application in the future.

[0003] Currently, the mainstream framework adopted by the academic community for the problem of no-reference video quality assessment is to extract the spatial features of the video and then use GRU (Gate Recurrent Unit) to model the temporal information. In this mainstream framework, how to make the features better represent the video distortion degree and how to establish the temporal connection of video frames are the problems to be solved by the present invention. Therefore, the present invention adopts the methods of extracting multi-scale features and edge detection to enhance the perception of video blur degree, and uses Q to query the semantic information changes of adjacent frames in the Transformer Block to improve the model's perception of motion distortions such as video jitter and blur. Summary of the Invention

[0004] The present invention proposes a video quality assessment method based on semantic changes, which can effectively extract the spatio-temporal features of the video, making the video distortion information obtained by the quality evaluation model more comprehensive.

[0005] The present invention adopts the following technical solutions.

[0006] A video quality assessment method based on semantic changes includes the following steps;

[0007] Step S1: For videos of different scenes shot by a mobile device, extract edge features for each frame of the video;

[0008] Step S2: Input the edge of each frame of the video and the original image into the spatial feature extraction network respectively to obtain the multi-scale spatial features of the video. At the same time, input the video into the temporal feature extraction network to obtain multi-scale temporal features, and perform multi-frequency component pooling and standard pooling on the multi-scale features;

[0009] Step S3: Merge the results after pooling to obtain the spatio-temporal features of the video, and reduce the dimensionality of the spatio-temporal features;

[0010] Step S4: Input the video spatio-temporal features after dimensionality reduction into a quality prediction network to model the temporal relationship, and then predict the quality score of the overall video.

[0011] The specific steps of step S1 are as follows;

[0012] Step S11: After dividing the video in step S1 into video frames, use an improved edge detection operator to extract the edge information of each video frame;

[0013] Step S12: Use an improved edge detection operator to extract the edge information to obtain the video edge R, and let R = {Canny i}, i = 1, 2,..., T, representing the set of all detection results in a video sequence, where T represents the number of frames in a video sequence, and Canny i represents the edge detection image of the i-th frame in a video sequence.

[0014] In step S12, the method for extracting edge information is specifically as follows:

[0015] First, use a relatively large 5×5 Gaussian convolution kernel for filtering to remove sharp image noise to a greater extent; then use an improved sobel operator to calculate the gradient magnitude and gradient direction of the image; then perform non-maximum suppression operation on the image edge according to the gradient magnitude and gradient direction.

[0016] The sobel operator is as follows:

[0017]

[0018] The calculation method of the gradient direction is as follows:

[0019] G left = Sobel_left * frame and G right = Sobel_right * frame Formula 2;

[0020]

[0021]

[0022] where frame represents the video frame image, Sobel_left represents the sobel operator in the left diagonal direction, Sobel_right represents the sobel operator in the right diagonal direction, and G left represents the gradient magnitude of the image in the left diagonal direction, and G rightrepresents the gradient magnitude of the image in the right diagonal direction, G represents the gradient magnitude of a video frame, and θ represents the gradient direction of the image.

[0023] Step S2 specifically includes the following steps;

[0024] Step S21: Input the edge image R of the processed video frame and the original video frame into the spatial feature extraction network respectively. According to the transfer learning concept of the model, the spatial feature extraction network is based on ConvNeXt-T pre-trained on ImageNet, removing the last pooling layer and fully connected layer. For the 4 stages of ConvNeXt-T, extract the features of each stage. The feature scales output by each stage are different, which are called multi-scale features;

[0025] Step S22: Input the video into the temporal feature extraction network. The temporal feature extraction network uses the Fast stream in the SlowFast network pre-trained on the Kinetics-400 video dataset, denoted as SlowFast F , which can generate motion features with high temporal resolution. For the expression of the video motion features within the time range obtained by SlowFast F , also extract multi-scale features;

[0026] Step S23: Use a method of multiple frequency component pooling that is better than global average pooling for the multi-scale features to compress the features of a single frame.

[0027] Step S3 specifically includes the following steps;

[0028] Step S31: Since the time stride of SlowFast F is 2, in order to match the temporal resolution of the motion pipeline, sample every two frames of the spatial and edge feature tensors, that is, i = 1, 3, 5,..., T, and then connect the spatial, edge, and temporal convolutional features along the channel dimension;

[0029] Step S32: In order to establish the correlation in the video temporal dimension, use an improved Transformer Block. First, use a fully connected layer to perform dimensionality reduction on the frame-level feature vector to a 128-dimensional tensor, denoted as Feature.

[0030] The step S4 specifically includes the following steps;

[0031] Step S41: For the 128-dimensional feature vector after dimensionality reduction, add a special token to the input sequence, denoted as Feature cls , which is used to learn the features of the entire sequence, and then add the positional encoding E pos , and the specific operation is as follows:

[0032] F0 = [Feature cls ; Feature1; Feature3; …; Feature T + E pos Formula Five;

[0033] Where Feature i represents the spatio-temporal feature of the i-th frame (i = 1, 3, 5, …, T) of the video;

[0034] Step S42: The processed vector is input into the improved Transformer Block;

[0035] Step S43: For the feature vector output by the Transformer Block, use a fully connected layer to map the state sequence to the frame-level quality score;

[0036] Step S44: For the frame-level quality score, use average pooling to temporarily aggregate the frame-level quality scores into the overall video quality score.

[0037] The operation of the improved Transformer Block is as follows:

[0038] The difference in semantic information between adjacent video frames is manifested as the change of the video. The change between adjacent frames is an important factor leading to video artifacts, jitter and other distortions.

[0039] First, calculate the feature change between adjacent frames:

[0040] Feature_difference i = Feature i+1 - Feature i , 0 ≤ i < T - 1

[0041] Feature_difference i = 0, i = T - 1 Formula Six;

[0042] Given a sequence Feature as input, the self-attention module first projects Feature and Feature_difference into the query Q, key K, value V, and feature difference D matrices as follows:

[0043] Q = Linear(Feature), K = Linear(Feature), V = Linear(Feature) Formula Seven;

[0044] D = Linear(Feature_difference) Formula Eight;

[0045] Among them, Linear represents the fully connected projection, and Feature_differnece represents the feature change between adjacent frames; then Q is used to query K and the feature difference D to obtain the query matrix M, and the elements in M represent the attention values of Q to K and D:

[0046]

[0047]

[0048]

[0049] Among them, M QK represents the attention matrix of Q to K, and M QD represents the attention matrix of Q to D.

[0050] The mobile device is a digital camera, a smart phone or a tablet computer.

[0051] The present invention has the following beneficial effects compared with the prior art:

[0052] 1. The present invention uses an improved edge detection operator to extract image edges. Compared with only focusing on the gradients in the horizontal and vertical directions, the proposed method highlights the diagonal gradients more, increases the gradient amplitude to make the edges clearer, depicts the blur degree of the video, and uses a multi-scale method to highlight the features of fine-grained and coarse-grained videos;

[0053] 2. Since global average pooling represents the lowest frequency component of 2D-DCT, in order to better compress the channel and introduce more information, the present invention generalizes global average pooling to more frequency components of 2D-DCT to compress more information;

[0054] 3. The present invention uses an improved TransformerBlock. By querying the changes in the semantic information of adjacent frames, the model can better learn motion distortions such as video jitter and artifacts. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0056] Attached Figure 1 is a schematic flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] As Figure 1 shown, the video quality assessment method based on semantic changes includes the following steps;

[0058] Step S1: For videos of different scenes captured by a mobile device, extract edge features for each frame of the video;

[0059] Step S2: Input the edges of each frame of the video and the original image into the spatial feature extraction network respectively to obtain the multi-scale spatial features of the video. At the same time, input the video into the temporal feature extraction network to obtain the multi-scale temporal features, and perform multi-frequency component pooling and standard pooling on the multi-scale features;

[0060] Step S3: Combine the results after pooling to obtain the spatio-temporal features of the video, and reduce the dimension of the spatio-temporal features;

[0061] Step S4: Input the dimension-reduced spatio-temporal features of the video into the quality prediction network to model the temporal relationship, and then predict the quality score of the overall video.

[0062] The specific steps of step S1 are as follows;

[0063] Step S11: After dividing the video in step S1 into video frames, use an improved edge detection operator to extract the edge information of each video frame;

[0064] Step S12: Use an improved edge detection operator to extract the edge information to obtain the video edge R, and let R = {Canny i}, i = 1, 2,..., T, representing the set of all detection results in a video sequence, where T represents the number of frames in a video sequence, and Canny i represents the edge detection image of the i-th frame in a video sequence.

[0065] In step S12, the method for extracting the edge information is specifically as follows:

[0066] First, use a relatively large 5×5 Gaussian convolution kernel for filtering to remove sharp image noise to a greater extent; then use the improved sobel operator to calculate the gradient magnitude and gradient direction of the image; then perform non-maximum suppression operation on the image edge according to the gradient magnitude and gradient direction.

[0067] The sobel operator is as follows:

[0068]

[0069] The calculation method of the gradient direction is as follows:

[0070] G left = Sobel_left * frame and G right = Sobel_right * frame Formula 2;

[0071]

[0072]

[0073] Where frame represents the video frame image, Sobel_left represents the sobel operator in the left diagonal direction, Sobel_right represents the sobel operator in the right diagonal direction, and G left represents the gradient magnitude of the image in the left diagonal direction, and G right represents the gradient magnitude of the image in the right diagonal direction. G represents the gradient magnitude of a video frame, and θ represents the gradient direction of the image.

[0074] Step S2 specifically includes the following steps;

[0075] Step S21: Input the edge image R of the processed video frame and the original video frame into the spatial feature extraction network respectively. According to the transfer learning concept of the model, the spatial feature extraction network is based on ConvNeXt-T pre-trained on ImageNet, removing the last pooling layer and fully connected layer. For the 4 stages of ConvNeXt-T, extract the features of each stage. The feature scales output by each stage are different, which are called multi-scale features;

[0076] Step S22: Input the video into the temporal feature extraction network. The temporal feature extraction network uses the Fast stream in the SlowFast network pre-trained on the Kinetics-400 video dataset, denoted as SlowFast F which can generate motion features with high temporal resolution. For the expression of the video motion features within the time range obtained by SlowFast F also extract multi-scale features;

[0077] Step S23: Use a method of multiple frequency component pooling that is superior to global average pooling for the multi-scale features to compress the features of a single frame.

[0078] Step S3 specifically includes the following steps;

[0079] Step S31: Since the time stride of SlowFast F is 2, in order to match the temporal resolution of the motion pipeline, sample every two frames of the spatial and edge feature tensors, that is, i = 1, 3, 5,..., T, and then connect the spatial, edge, and temporal convolutional features along the channel dimension;

[0080] Step S32: In order to establish the correlation in the video temporal dimension, use an improved TransformerBlock. First, use a fully connected layer to perform dimensionality reduction on the frame-level feature vector to reduce it to a 128-dimensional tensor, denoted as Feature.

[0081] The said step S4 specifically includes the following steps;

[0082] Step S41: For the 128-dimensional feature vector after dimensionality reduction, add a special token, denoted as Feature, to the input sequence to learn the features of the entire sequence, and then add the position encoding E, as follows: cls , for learning the features of the entire sequence, and then adding the position encoding E pos : The specific operation is as follows:

[0083] F0 = [Feature cls ; Feature1; Feature3; …; Feature T + E pos Equation Five;

[0084] where Feature i represents the spatio-temporal features of the i-th frame (i = 1, 3, 5, …, T) of the video;

[0085] Step S42: Input the processed vector into the improved Transformer Block;

[0086] Step S43: For the feature vector output by the Transformer Block, use a fully connected layer to map the state sequence to the frame-level quality score;

[0087] Step S44: For the frame-level quality score, use average pooling to temporarily aggregate the frame-level quality scores into the overall video quality score.

[0088] The operation of the improved Transformer Block is as follows:

[0089] The difference in semantic information between adjacent video frames is manifested as the change of the video. The change between adjacent frames is an important factor leading to video artifacts, jitter and other distortions.

[0090] First, calculate the feature change between adjacent frames:

[0091] Feature_difference i = Feature i+1 - Feature i , 0 ≤ i < T - 1

[0092] Feature_difference i = 0, i = T - 1 Equation Six;

[0093] Given a sequence Feature as the input, the self-attention module first projects Feature and Feature_difference into the query Q, key K, value V, and feature difference D matrices, as follows:

[0094] Q = Linear(Feature), K = Linear(Feature), V = Linear(Feature), Equation 7;

[0095] D = Linear(Feature_difference), Equation 8;

[0096] Among them, Linear represents a fully connected projection, and Feature_differnece represents the feature change between adjacent frames; then use Q to query K and the feature difference D to obtain the query matrix M, and the elements in M represent the attention values of Q to K and D:

[0097]

[0098]

[0099]

[0100] Among them, M QK represents the attention matrix of Q to K, and M QD represents the attention matrix of Q to D.

[0101] The mobile device is a digital camera, a smart phone or a tablet computer.

Claims

1. A video quality assessment method based on semantic change, characterized in that: Including the following steps; Step S1: For videos of different scenes captured by a mobile device, extract edge features for each frame of the video; Step S2: Input the edges of each frame of the video and the original image into a spatial feature extraction network respectively to obtain the multi-scale spatial features of the video. At the same time, input the video into a temporal feature extraction network to obtain multi-scale temporal features, and perform multi-frequency component pooling and standard pooling on the multi-scale features; Step S3: Combine the results after pooling to obtain the spatio-temporal features of the video, and reduce the dimensionality of the spatio-temporal features; Step S4: Input the reduced-dimensional spatio-temporal features of the video into a quality prediction network to model the temporal relationship, and then predict the quality score of the overall video; In step S1, the method for extracting edge information is specifically as follows: First, filter using a large 5×5 Gaussian convolution kernel; then use the sobel operator to calculate the gradient magnitude and gradient direction of the image; then perform non-maximum suppression operation on the image edges according to the gradient magnitude and gradient direction; The sobel operator is as follows: The calculation method of the gradient direction is as follows: G left = Sobel_left * frame and G right = Sobel_right * frame Formula 2; Among them, frame represents the video frame image, Sobel_left represents the sobel operator in the left diagonal direction, Sobel_right represents the sobel operator in the right diagonal direction, and G left represents the gradient magnitude of the image in the left diagonal direction, and G right represents the gradient magnitude of the image in the right diagonal direction. G represents the gradient magnitude of a video frame, and θ represents the gradient direction of the image; In step S3, an improved Transformer Block is adopted. First, use a fully connected layer to perform dimensionality reduction on the frame-level feature vector, denoted as Feature; The operations of the Transformer Block are as follows: First, calculate the feature changes between adjacent frames: Feature_difference i = Feature i+1 - Feature i , 0 ≤ i < T - 1 Feature_difference i = 0, i = T - 1, Formula 6; Given a sequence Feature as input, the self-attention module first projects Feature and Feature_difference into query Q, key K, value V, and feature difference D matrices as follows: Q = Linear(Feature), K = Linear(Feature), V = Linear(Feature) Equation 7; D = Linear(Feature_difference) Equation 8; Among them, Linear represents the fully connected projection, and Feature_difference i represents the feature change between adjacent frames; then Q is used to query K and the feature difference D to obtain the query matrix M, and the elements in M represent the attention values of Q to K and D: Among them, M QK represents the attention matrix of Q to K, and M QD represents the attention matrix of Q to D.

2. The video quality assessment method based on semantic change according to claim 1, characterized in that: Step S1 specifically includes the following steps; Step S11: After dividing the video in step S1 into video frames, use an improved edge detection operator to extract the edge information of each video frame; Step S12: Extract edge information using an improved edge detection operator to obtain video edge R, and let R = {Canny i}, where i = 1, 2,..., T, representing the set of all detection results in a video sequence, where T represents the number of frames in a video sequence, and Canny i represents the edge detection image of the i-th frame in a video sequence.

3. The video quality assessment method based on semantic change according to claim 1, characterized in that: Step S2 specifically Including the following steps; Step S21: Input the edge image R of the processed video frame and the original video frame into a spatial feature extraction network respectively. The spatial feature extraction network is based on ConvNeXt-T pre-trained on ImageNet, removing the last pooling layer and fully connected layer. For the 4 stages of ConvNeXt-T, extract the features of each stage. The feature scales output by each stage are different, which are called multi-scale features; Step S22: Input the video into the temporal feature extraction network, and the temporal feature extraction network uses the Fast stream in the pre-trained SlowFast network on the Kinetics-400 video dataset, denoted as SlowFast F , which generates motion features with high temporal resolution. For the expression of the video motion features within the temporal range obtained by SlowFast F , multi-scale features are also extracted; Step S23: Use a method of multiple frequency component pooling superior to global average pooling to compress the features of a single frame for the multi-scale features.

4. The video quality assessment method based on semantic change according to claim 3, characterized in that: Step S3 specifically Including the following steps; Step S31: Since the time stride of SlowFast F is 2, sample every two frames of the spatial and edge feature tensors, i.e., i = 1, 3, 5, …, T, and then concatenate the convolutional features of space, edge, and time along the channel dimension; Step S32: Use an improved Transformer Block to perform dimensionality reduction on the frame-level feature vector.

5. The video quality assessment method based on semantic change according to claim 1, characterized in that: Step S4 specifically includes the following steps; Step S41: For the 128-dimensional feature vector after dimensionality reduction, add a special token, denoted as Feature, to the input sequence to learn the features of the entire sequence, and then add the position encoding E, and the specific operation is as follows: cls , which is used to learn the features of the entire sequence, and then add the position encoding E pos , and the specific operation is as follows: F0 = [Feature cls ; Feature1; Feature3; …; Feature T + E pos Formula Five; Among them, Feature i represents the spatio-temporal feature of the i-th frame (i = 1, 3, 5, …, T) of the video; Step S42: Input the processed vector into an improved Transformer Block; Step S43: For the feature vectors output by the Transformer Block, use a fully connected layer to map the state sequence to frame-level quality scores; Step S44: For the frame-level quality scores, use average pooling to temporarily aggregate the frame-level quality scores to an overall video quality score.

6. The video quality evaluation method based on semantic change according to claim 1, characterized in that: The mobile device is a digital camera, a smartphone, or a tablet computer.

Citation Information

Patent Citations

  • Video moire removing method based on linear sparse attention Transformer

    CN114881888A

  • Full-reference video quality evaluation method based on two stages of adaptive sampling and multi-scale time sequence

    CN115239647A