Video quality evaluation method based on spatio-temporal feature fusion
Through the video quality evaluation method based on spatiotemporal feature fusion, the time and space feature extraction modules are used to combine the graph convolution network to perform feature fusion, which solves the problem of failing to effectively utilize the relationship between spatiotemporal features and inter-frame features in the prior art, and achieves a more accurate video quality evaluation.
Patent Information
- Application Number
- CN202510431395.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video quality evaluation method fails to effectively utilize the relationship between the spatial and temporal features of the video and the inter-frame features, resulting in inaccurate evaluation results and neglecting the importance of distorted information.
The video quality evaluation method based on spatiotemporal feature fusion is adopted, and the short video features and high-level semantic features are extracted respectively through the temporal feature extraction module and the spatial feature extraction module, and feature fusion is performed in combination with the graph convolution network, and mass regression is used for multi-layer perceptron to calculate the total loss of the model to adjust the model parameters.
It improves the accuracy of video quality evaluation, significantly alleviates the distribution offset problem between laboratory data and real scenes, enhances cross-domain invariance, and achieves more accurate video quality prediction.
Smart Images

Figure CN120374527A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video quality evaluation, and particularly relates to a video quality evaluation method based on spatio-temporal feature fusion. Background Art
[0002] With the rapid development of video shooting devices, wireless network technologies, portable video playback devices, and social media and video creation and sharing platforms, shooting, sharing, and watching videos have become an indispensable part of people's daily lives. This trend has not only promoted the explosive growth of video content but also caused the scale of video data to increase exponentially. A large number of studies have shown that evaluating video quality by combining the spatial information and temporal information of videos makes the evaluation results more effective. Some UGC video datasets have effectively promoted the progress and development of research on VQA (Video Quality Assessment) methods, and some VQA methods with objective performance have also been proposed in the prior art. However, these methods still have some problems. First, some previous methods use the entire video or a long video segment as input. Although this retains the necessary motion information to help extract temporal feature information, due to the high similarity between consecutive video frames, it brings a large amount of redundant information in spatial information extraction, dilutes the learning effect of the feature extraction module, and instead makes the model unable to extract spatial and temporal information more effectively. Second, for the extraction of spatial feature information, some methods only use a pre-trained model in the field of image classification as one branch for extracting spatial feature information. Obviously, due to the differences in research ideas in different fields, applying a pre-trained model in the field of image classification can only help the spatial information extraction branch of the VQA model extract the semantic information of the frame to a certain extent, but this is not enough. Many studies have shown that although semantic information is important for quality evaluation, distortion information closely related to quality is equally important. To accurately evaluate the quality of an image or video, distortion information is indispensable. Finally, after the model successfully obtains temporal feature information and spatial feature information, using a reasonable strategy to fuse the two types of information can help the model more effectively predict the quality score. However, most current methods simply concatenate the temporal feature information and spatial feature information and input them into a multi-layer perceptron for feature fusion and quality regression. Such a feature fusion method ignores the relationships that originally exist between features and is not a scientific feature fusion method.
[0003] In summary, there is an urgent need for a new video quality evaluation method to effectively utilize the mutual relationship between spatio-temporal features and inter-frame features, thereby improving the accuracy of quality evaluation. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention proposes a video quality evaluation method based on spatio-temporal feature fusion, which includes: taking the video to be evaluated and preprocessing it, and inputting the preprocessed video into a trained video quality evaluation model for processing to obtain a video quality evaluation result;
[0005] The training process of the video quality evaluation model includes:
[0006] S1: Obtaining a video quality evaluation data set and preprocessing it to obtain preprocessed video data;
[0007] S2: Using a time feature extraction module to process the preprocessed video data to obtain short video features and long video features;
[0008] S3: Using a spatial feature extraction module to process the preprocessed video data to obtain high-level semantic features and distortion features;
[0009] S4: Using a feature fusion module to fuse the short video features, long video features, high-level semantic features and distortion features to obtain fused features;
[0010] S5: Inputting the fused features into a multi-layer perceptron for processing to obtain a video quality score;
[0011] S6: Calculating the total loss of the model and continuously adjusting the model parameters according to the total loss of the model to obtain a trained video quality evaluation model.
[0012] Preferably, the process of preprocessing the video quality evaluation data set includes:
[0013] Uniformly sampling each video into 8 segments, and each segment contains 32 frames of original resolution images;
[0014] Selecting the first frame of each segment as a key frame and adjusting it to a pixel size of 224×224.
[0015] Preferably, the process of the time feature extraction module processing the preprocessed video data includes:
[0016] Taking the video segments in the preprocessed video data as long video blocks, and downsampling the long video blocks to obtain short video blocks;
[0017] Respectively extracting features from the short video blocks and the long video blocks to obtain short video features and long video features.
[0018] Preferably, the process of the spatial feature extraction module processing the preprocessed video data includes:
[0019] Select the first frame of the video frame as the key frame, and input the key frame into the global semantic path and the local distortion path respectively for processing to obtain the primary high-level semantic features and the primary distortion features; among them, the global semantic path scales the key frame to the standard size; the local distortion path divides the key frame into regular grids and samples an image from the grids;
[0020] Adopt a hierarchical pooling strategy to fuse and process the primary high-level semantic features and the primary distortion features to obtain the final high-level semantic features and distortion features.
[0021] Furthermore, the process of processing the primary high-level semantic features and the primary distortion features includes:
[0022] Input the primary high-level semantic features and the primary distortion features into the Swin-Transformer network for processing respectively; the Swin-Transformer network includes four Stage blocks; perform global standard pooling and global average pooling on the outputs of the latter three Stage blocks respectively, and splice the two pooled features to obtain the final three high-level semantic features and three distortion features.
[0023] Preferably, the process of obtaining the fusion features includes:
[0024] Take the short video features, long video features, high-level semantic features and distortion features as feature nodes;
[0025] Define three types of edge connection rules and connect the feature nodes according to the three types of edge connection rules to obtain the intra-block feature relationship graph;
[0026] Sample the graph convolutional network to process the intra-block feature relationship graph to obtain the intra-block features;
[0027] Connect the multiple intra-block features obtained from the processing of multiple video blocks in sequence to obtain the inter-block feature relationship graph;
[0028] Sample the graph convolutional network to process the inter-block feature relationship graph to obtain the fusion features.
[0029] Furthermore, the three types of edge connection rules are cross-scale adjacent relationship, same-scale spatial relationship and spatio-temporal joint relationship respectively.
[0030] Preferably, the formula for calculating the total loss of the model is:
[0031]
[0032] Among them, L plcc represents the total loss of the model, represents the normalized true value, It represents the true value after standardization, MSE represents the mean square error, ρ represents the intermediate parameter, and mean represents the mean function.
[0033] The beneficial effects of the present invention are as follows:
[0034] 1. In the spatio-temporal dual-path transformation branch of the present invention, the multi-scale design concept is used for feature extraction to comprehensively extract effective information as much as possible. The prior knowledge of the spatio-temporal continuity of the video is transformed into a quantifiable dynamic distortion representation, providing discriminative temporal feature information for subsequent quality regression. This hierarchical processing paradigm not only follows the coarse-to-fine perception characteristics of the human visual system but also conforms to the multi-scale law of video quality distortion propagation in the time domain, enabling the model to more accurately predict the video quality.
[0035] 2. The core advantage of the present invention lies in the enhancement of cross-domain invariance. By introducing the feature relationship graph of graph convolution, the feature information in different scale domains is aligned, significantly alleviating the distribution offset problem between laboratory data and real scenarios. By explicitly modeling the spatio-temporal dependence relationships across scales and modalities, the collaborative optimization of distortion representations is realized, and the relationship between feature information is better utilized for feature fusion.
[0036] The present invention effectively utilizes the mutual relationship between spatio-temporal features and inter-frame features, thereby improving the accuracy of quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a block diagram of the video quality evaluation model in the present invention;
[0038] Figure 2 It is a schematic diagram of the time feature extraction module in the present invention;
[0039] Figure 3 It is a schematic diagram of the space feature extraction module in the present invention;
[0040] Figure 4 It is a schematic diagram of the construction of the feature relationship graph of the graph convolution network in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0042] In the video quality assessment task, the multi-level feature fusion mechanism has been verified as an effective strategy to improve the discriminative ability of the model. Existing research shows that visual distortion has differential representation characteristics in feature spaces at different abstraction levels: low-level features are highly sensitive to local texture distortions such as noise interference and motion blur, while high-level features are more suitable for capturing semantic correlation distortions such as scene incoherence and abnormal object deformation. This hierarchical perception characteristic highly coincides with the multi-stage processing mechanism of the human visual system, that is, the initial processing of the retina focuses on local contrast and instantaneous changes, while the cerebral cortex is responsible for integrating spatio-temporal information to form a global quality perception. Fusing these features can help the model better handle various distortion situations. By utilizing features at different levels, the model can obtain a richer feature representation, which helps to improve the accuracy of quality assessment. To solve the problem that the prior art does not effectively utilize the mutual relationship between spatio-temporal features and inter-frame features, the present invention proposes a video quality assessment method based on spatio-temporal feature fusion, and the method includes the following content:
[0043] Obtain the video to be evaluated and preprocess it, and input the preprocessed video into a trained video quality assessment model for processing to obtain a video quality assessment result.
[0044] The training process of the video quality assessment model includes:
[0045] S1: Obtain a video quality assessment data set and preprocess it to obtain preprocessed video data.
[0046] Obtain a video quality assessment data set, for example, the UGC video quality assessment data set, and preprocess the data set: uniformly sample each video into 8 segments, and each segment contains 32 original resolution images; to balance computational efficiency and information integrity, select the first frame of each segment as the key frame and adjust it to an input size of 224×224 pixels. This resolution setting takes into account the receptive field requirements of the feature extraction network and the GPU memory limit, and is also compatible with the input specifications of mainstream pre-trained models.
[0047] S2: Use a time feature extraction module to process the preprocessed video data to obtain short video features and long video features.
[0048] The present invention constructs a multi-scale network based on graph convolutional spatio-temporal feature fusion, namely the video quality assessment model; as Figure 1 shown, the model architecture mainly consists of three modules: a time feature extraction module, a spatial feature, and a feature fusion and quality regression module.
[0049] First, the preprocessed video data is input into the temporal feature extraction module to extract temporal features. During the temporal feature extraction process, the module introduces a multi-scale temporal analysis paradigm. In this module, the original video segments are downsampled in the time domain to generate short video sequences. This operation essentially deconstructs the motion patterns at different time resolutions: long video blocks retain complete action cycle information to capture macroscopic motion laws, while short video blocks amplify local inter-frame relationships to sensitively capture instantaneous anomalies. Experimental verification shows that this dual-granularity feature complementary mechanism can significantly improve the model's discriminative ability for temporal scale distortion. Through the above design, the temporal feature extraction module successfully converts the prior knowledge of the spatio-temporal continuity of the video into a quantifiable dynamic distortion representation, providing discriminative temporal feature information for subsequent quality regression. This hierarchical processing paradigm not only follows the coarse-to-fine perception characteristics of the human visual system but also conforms to the multi-scale laws of video quality distortion propagation in the time domain. As Figure 2 shown, the specific process is as follows:
[0050] Take the video segments in the preprocessed video data as long video blocks, and the long video blocks are denoted as V L ; Extract video features through motion features. Preferably, the SlowFast model is used to extract features from short video blocks and long video blocks, and the temporal features of the long video block, i.e., the long video features F TL :
[0051] F TL = SlowFast(V L )
[0052] While extracting the motion information of the video, introduce a downsampling strategy in the time dimension to help the motion feature extraction module extract multi-scale information in the time dimension at different time frequencies of the video. First, evenly halve the long video block for downsampling to obtain the short video block V S :
[0053] V S = Sub(V L )
[0054] Then use the SlowFast model to extract features to obtain the temporal features of the short video block, i.e., the short video features F TS :
[0055] F TS = SlowFast(V S )
[0056] S3: Use the spatial feature extraction module to process the preprocessed video data to obtain high-level semantic features and distortion features.
[0057] As Figure 3As shown, in this module, first, the spatio-temporal information and static semantics are separated through the key-frame sampling strategy. The first frame of the video block is selected as the key frame because it carries the semantic information of the scene main body and is least affected by temporal distortion. Subsequently, a two-way spatial transformation is performed on the key frame, including a global semantic path and a local distortion path. Among them, the global semantic path scales the key frame to a standard size to retain the complete scene layout information. The local distortion path divides the key frame into regular grids and samples an image from the grids, thereby magnifying the local area details and removing the influence of semantic information.
[0058] After the key frame undergoes a two-way spatial transformation, the primary high-level semantic feature f R and the primary distortion feature f G are obtained; which is expressed as:
[0059] f R = Resize(Frame K )
[0060] f G = GMS(Frame K )
[0061] Among them, Resize represents scaling, Frame K is the key frame K, and GMS is the Grid Mini-patch Sampling (GMS) strategy.
[0062] The hierarchical pooling strategy is adopted to fuse multi-scale information, that is, the hierarchical pooling strategy is used to process the primary high-level semantic feature and the primary distortion feature to obtain the final high-level semantic feature and distortion feature; specifically:
[0063] The primary high-level semantic feature and the primary distortion feature are respectively input into the Swin-Transformer network for processing; the Swin-Transformer network includes four Stage blocks; the outputs of the last three Stage blocks are respectively subjected to global standard pooling and global average pooling processing, and the two pooled features are concatenated to obtain the final three high-level semantic features and three distortion features. It is expressed as:
[0064]
[0065] Among them, the global semantic feature encodes high-level semantics such as scene categories and object compositions, and the local distortion feature focuses on low-level abnormal patterns such as block effects and noise distributions. Experiments show that this explicit feature decoupling design enables the model to focus on semantic-related and distortion-related gradient signals respectively during the training process, effectively improving the accuracy of video quality prediction in the case of feature space confusion. Among them and The expression is:
[0066]
[0067]
[0068] Among them, SwinBlock are the respective Stages of the baseline model Swin-Transformer for feature processing.
[0069] S4: Use the feature fusion module to fuse the short video features, long video features, high-level semantic features, and distortion features to obtain the fused features.
[0070] The feature fusion and quality regression module includes a feature fusion module and a quality regression module. The feature fusion module is based on the feature fusion paradigm of graph structure learning, and realizes the collaborative optimization of distortion representation by explicitly modeling the spatio-temporal dependence relationships across scales and modalities. As Figure 4 shown, the present invention constructs a double-layer graph convolutional network (GCN) architecture to process the intra-video-block feature relationships and the inter-video-block temporal relationships respectively. Specifically:
[0071] Intra-video-block feature relationship graph:
[0072] Take the short video features, long video features, high-level semantic features, and distortion features as feature nodes;
[0073] Define three types of edge connection rules and connect the feature nodes according to the three types of edge connection rules to obtain the intra-video-block feature relationship graph. Among them, the three types of edge connection rules are cross-scale adjacent relationships (connecting feature nodes at adjacent levels such as from Stage2 to Stage3, modeling the process of distortion from concrete to abstract), same-scale spatial relationships (connecting the global semantic path and local distortion path feature nodes within the same Stage, modeling the propagation path of distortion from local to global), and spatio-temporal joint relationships (pairwise connecting spatio-temporal feature nodes to simulate the coupled diffusion of distortion in the spatio-temporal dimension).
[0074] Sample the graph convolutional network to process the intra-video-block feature relationship graph to obtain the intra-video-block features, expressed as:
[0075]
[0076] Among them, represents the feature map output after graph convolution of the nth video block through the intra-video-block feature relationship graph Map in That is, the intra-video-block features, and n is 8.
[0077] Inter-video-block temporal relationship graph:
[0078] Taking the feature vectors of each video block as super nodes, construct edges based on temporal adjacency to connect the intra-block features obtained by processing multiple video blocks in sequence, resulting in an inter-block feature relationship graph; this process enables the model to perceive the gradual change law of distortion over a long time scale, such as the inter-frame cumulative effect of compression artifacts.
[0079] Sample a graph convolutional network to process the inter-block feature relationship graph, obtaining inter-block features, expressed as:
[0080]
[0081] where F fbt represents the feature map output after graph convolution by the inter-block temporal relationship graph Map bt i.e., the fused feature; after multi-order graph convolution, the node features encode the joint distortion information from local to global and from static to dynamic.
[0082] S5: Input the fused feature into a multi-layer perceptron for processing to obtain a video quality score.
[0083] In the quality regression module, map it to the quality score space through a lightweight multi-layer perceptron, expressed as:
[0084] S pre = W fbt ·F fbt + b fbt
[0085] where S pre is the quality prediction score output by the model, W fbt represents the weight, and b fbt represents the bias.
[0086] S6: Calculate the total loss of the model and continuously adjust the model parameters according to the total loss of the model to obtain a trained video quality evaluation model.
[0087]
[0088] where ε is a small constant to prevent division by zero, set to 10 -8 and std represents the standard deviation.
[0089] The formula for calculating the total loss of the model is:
[0090]
[0091] where L plcc represents the total loss of the model, represents the standardized true value, represents the standardized true value, MSE represents the mean squared error, ρ represents an intermediate parameter, and mean represents the mean function.
[0092] Continuously iterate and train, and adjust the model parameters through backpropagation according to the total loss of the model. Stop training when the total loss of the model converges or reaches the maximum preset number of iterations to ensure the model parameters and obtain a trained video quality evaluation model. Obtain the video to be evaluated and preprocess it, and input the preprocessed video into the trained video quality evaluation model for processing to obtain the quality prediction score, that is, the video quality evaluation result.
[0093] In summary, the present invention proposes to achieve comprehensive quality assessment by jointly utilizing multi-scale spatio-temporal features. The model architecture mainly consists of three modules: a temporal feature extraction module, a spatial feature extraction module, and a feature fusion and quality regression module. The present invention first extracts quality-related features of the input video through the temporal feature extraction module and the spatial feature extraction module, and then reorganizes the extracted features and inputs them into the feature fusion and quality regression module to fuse the spatio-temporal features and remap the features to the corresponding quality scores. It should be noted that the design idea of multi-scale is used for feature extraction on the spatio-temporal dual-path spatial transformation branch to extract effective information as comprehensively as possible. Specifically, multi-scale features in the time dimension are obtained by downsampling the video chunks, and multi-scale features in the space dimension are obtained by retaining the outputs of different stages of the spatial feature extraction module. Moreover, the decoupled extraction design of time and space features does not simply superimpose the dual-path information, but realizes it by independently modeling the distortion propagation laws of the spatio-temporal dimensions, providing a more complementary and more specific and effective solution for the subsequent fusion stage in utilizing the spatio-temporal relationship. Through the above modular design, the model realizes the all-round quality perception ability from local to global and from static to dynamic, improving the accuracy of model prediction.
[0094] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A video quality evaluation method based on spatio-temporal feature fusion, characterized in that, Including: Obtain the video to be evaluated and preprocess it, input the preprocessed video into the trained video quality evaluation model for processing, and obtain the video quality evaluation result; The training process of the video quality evaluation model includes: S1: Obtain the video quality evaluation data set and preprocess it to obtain the preprocessed video data; S2: Use the time feature extraction module to process the preprocessed video data to obtain short video features and long video features; S3: Use the spatial feature extraction module to process the preprocessed video data to obtain high-level semantic features and distortion features; S4: Use the feature fusion module to fuse the short video features, long video features, high-level semantic features and distortion features to obtain the fusion features; S5: Input the fusion features into a multi-layer perceptron for processing to obtain the video quality score; S6: Calculate the total model loss and continuously adjust the model parameters according to the total model loss to obtain the trained video quality evaluation model.
2. The video quality evaluation method based on spatio-temporal feature fusion according to claim 1, wherein The process of preprocessing the video quality evaluation data set includes: Uniformly sample each video into 8 segments, and each segment contains 32 frames of original resolution images; Select the first frame of each segment as the key frame and adjust it to a size of 224×224 pixels.
3. A video quality evaluation method based on spatio-temporal feature fusion according to claim 1, characterized in that The process of the time feature extraction module processing the preprocessed video data includes: Take the video segments in the preprocessed video data as long video blocks, and downsample the long video blocks to obtain short video blocks; Extract features from the short video blocks and long video blocks respectively to obtain short video features and long video features.
4. A video quality evaluation method based on spatio-temporal feature fusion according to claim 1, characterized in that, The process of the spatial feature extraction module processing the preprocessed video data includes: Select the first frame of the video frame as the key frame, and input the key frame into the global semantic path and the local distortion path respectively for processing to obtain the primary high-level semantic features and the primary distortion features; among them, the global semantic path scales the key frame to the standard size; the local distortion path divides the key frame into regular grids and samples an image from the grids; Adopt a hierarchical pooling strategy to fuse and process the primary high-level semantic features and the primary distortion features to obtain the final high-level semantic features and distortion features.
5. A video quality evaluation method based on spatio-temporal feature fusion according to claim 4, characterized in that, The process of processing the primary high-level semantic features and the primary distortion features includes: Input the primary high-level semantic features and the primary distortion features into the Swin-Transformer network for processing respectively; the Swin-Transformer network includes four Stage blocks; perform global standard pooling and global average pooling processing on the outputs of the latter three Stage blocks respectively, and splice the two pooled features to obtain the final three high-level semantic features and three distortion features.
6. The video quality evaluation method based on spatio-temporal feature fusion according to claim 1, characterized in that, The process of obtaining the fusion features includes: Take the short video features, long video features, high-level semantic features and distortion features as feature nodes; Define three types of edge connection rules and connect the feature nodes according to the three types of edge connection rules to obtain the in-block feature relationship graph; Sample the graph convolutional network to process the in-block feature relationship graph to obtain the in-block features; Connect the multiple in-block features obtained by processing multiple video blocks in sequence to obtain the inter-block feature relationship graph; The sampled graph convolutional network processes the inter-block feature relationship graph to obtain the fused features.
7. A video quality evaluation method based on spatio-temporal feature fusion according to claim 6, characterized in that, The three types of edge connection rules are cross-scale adjacent relationship, same-scale spatial relationship, and spatio-temporal joint relationship.
8. The video quality evaluation method based on spatio-temporal feature fusion according to claim 1, wherein, The formula for calculating the total loss of the model is: Among them, L plcc represents the total loss of the model, represents the normalized true value, represents the normalized true value, MSE represents the mean square error, ρ represents the intermediate parameter, and mean represents the mean value function.
Citation Information
Cited By
Training method of video quality evaluation model, video quality evaluation method and model
CN121330577A
Screen content video quality evaluation method and device based on frequency-space complementation and semantics
CN121392718A
Screen content video quality evaluation method and device based on frequency space complementarity and semantics
CN121392718B