Video quality evaluation method and system based on space-frequency combination and time sequence interaction

By employing a spatial-frequency joint and temporal interaction approach, a referenceless video quality assessment system was constructed, which solved the problem of missing frequency domain information and temporal correlation in immersive video quality assessment, and achieved high-precision and robust video quality assessment.

CN121397207AActive Publication Date: 2026-01-23HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202511972389.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-01-23
Estimated Expiration
2045-12-25

AI Technical Summary

Technical Problem

Existing no-reference video quality assessment methods mainly focus on the structural and textural information of images, ignoring the energy distribution in the frequency domain and the dynamic correlation between frames, resulting in insufficient accuracy in the quality assessment of immersive videos under various degradation factors.

Method used

By employing a spatial-frequency joint and temporal interaction approach, a no-reference video quality assessment system is constructed through spatial feature extraction, frequency feature extraction, bidirectional Cross-Attention fusion, and Transformer temporal modeling, thereby achieving dynamic fusion of spatial-frequency features and capture of temporally correlated features.

Benefits of technology

It improves the accuracy and robustness of immersive video quality assessment, enabling high-precision quality assessment that conforms to human visual perception under no-reference conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397207A_ABST
    Figure CN121397207A_ABST
Patent Text Reader

Abstract

The invention discloses a video quality evaluation method and system based on space-frequency combination and time sequence interaction, and relates to the technical field of computer vision, and the method comprises the steps: carrying out the standardization and data enhancement of an immersive video key frame; then, spatial domain and frequency domain double-path features are constructed through learnable Gabor convolution, two-dimensional Fourier transform and a logarithmic magnitude spectrum; space-frequency dynamic fusion is realized through bidirectional Cross-Attention, and a joint representation is generated; then, inputting Transform architecture modeling inter-frame time sequence dynamic and global dependence, and obtaining time sequence correlation characteristics; and finally, outputting frame-by-frame quality scores through learning weighted MLP regression, and aggregating the frame-by-frame quality scores into an overall quality score. According to the method, high-fidelity no-reference quality evaluation of immersive video distortion is realized through space-frequency joint modeling and time sequence interactive learning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a video quality evaluation method and system based on space-frequency joint and time sequence interaction. BACKGROUND

[0002] With the wide application of virtual reality and immersive video technology, people have higher requirements for subjective visual experience in the process of watching multi-scene multi-view video. Immersive video has the characteristics of high resolution, wide view angle and complex space-time structure compared with traditional video, but it is easily affected by various degradation factors in the process of acquisition, encoding, transmission and rendering, such as compression artifacts, noise pollution, blur, color drift, motion inconsistency and frame rate fluctuation, etc. These will cause the user's subjective perception quality to decrease significantly. The existing video quality evaluation methods are mainly divided into: (1) full reference (FR) method: relying on the original video as reference, but it is difficult to obtain in actual scene; (2) reduced reference (RR) method: only using part of the feature information, still needing part of the reference signal; (3) no reference (NR) method: completely relying on the distorted video itself, which is the key and difficulty of immersive video quality perception research.

[0003] Most of the existing no-reference models are based on spatial domain convolution network, only focusing on the structure and texture information of the image, ignoring the energy distribution in the frequency domain and the inter-frame dynamic correlation. The degradation of immersive video often exists in spatial texture degradation, spectral energy disturbance and time sequence continuity destruction at the same time. Therefore, how to construct a unified modeling framework integrating spatial, frequency and time sequence information has become a key problem to improve the accuracy of immersive video quality evaluation. SUMMARY

[0004] In order to solve the above problems, the present application provides a video quality evaluation method and system based on space-frequency joint and time sequence interaction, which realizes high-precision, high-robustness and human-eye-perception immersive video quality evaluation under no-reference condition through space-frequency dual-path feature extraction, bidirectional Cross-Attention fusion, Transformer time sequence modeling and learnable weighted regression.

[0005] On the one hand, the video quality evaluation method based on space-frequency joint and time sequence interaction comprises:

[0006] S1, acquiring an immersive video sequence, performing key frame extraction, size normalization and color standardization processing on the immersive video sequence to obtain standardized key frames, and performing data enhancement processing on the standardized key frames to obtain enhanced key frames;

[0007] S2, input the enhanced key frame into the spatial feature extraction branch, extract multi-directional texture and edge structure information in the enhanced key frame through a learnable Gabor convolution kernel, obtain the spatial feature of the key frame, and use a multi-stage residual network to perform multi-scale modeling on the spatial feature of the key frame to obtain a deep and shallow layer fused spatial feature representation;

[0008] S3, input the enhanced key frame into the frequency domain feature extraction branch, perform two-dimensional fast Fourier transform on the enhanced key frame to obtain a log amplitude spectrum, capture the energy distribution and periodic characteristics of the key frame through the log amplitude spectrum, then input the log amplitude spectrum into a convolution network to extract the frequency spectrum energy and noise structure features of the key frame, and obtain the frequency domain feature of the key frame;

[0009] S4, input the deep and shallow layer fused spatial feature representation and the frequency domain feature into the spatial-frequency interaction module, realize dynamic fusion of the spatial-frequency features through a bidirectional Cross-Attention mechanism, and obtain a joint spatial-frequency feature representation;

[0010] S5, input the joint spatial-frequency feature into a multi-head self-attention network based on a Transformer architecture in a time sequence, model the dynamic changes and global temporal dependencies of the joint spatial-frequency feature, and obtain inter-frame temporal correlation features;

[0011] S6, input the inter-frame temporal correlation features into a quality regression module, predict the quality score of each frame of video through a multi-layer perception (MLP), then perform weighted average on the quality scores of all frames based on a learnable weight, and obtain a video overall quality score.

[0012] Further, in S2, the spatial feature extraction branch includes a learnable Gabor convolution kernel.

[0013] The learnable Gabor convolution kernel is defined as follows:

[0014] ;

[0015] wherein, and represent spatial pixel coordinates; exp() represents an exponential function; cos() represents a cosine function, , represent coordinate components after rotation of Gabor coordinates; represents a wavelength; represents a direction parameter; represents a phase shift; represents a Gaussian envelope width; represents a spatial aspect ratio.

[0016] Further, in S3, a two-dimensional fast Fourier transform is performed on the enhanced key frame to obtain a log amplitude spectrum, and the calculation formula is as follows:

[0017] ;

[0018] ;

[0019] wherein, and represent the row and column indexes of the pixel in the spatial domain of the input frame; represents the input frame; represents the height of the frame; represents the width of the frame; represents the frequency domain coordinate; represents the frequency spectrum in the form of a complex number; represents the log amplitude spectrum, and log(.) represents the logarithmic function.

[0020] Further, in S4, dynamic fusion of spatial and frequency features is realized through a bidirectional Cross-Attention mechanism to obtain a joint spatial and frequency feature representation, and the calculation process is as follows:

[0021] ;

[0022] ;

[0023] wherein, Softmax(.) represents a standardization exponential function; represents a layer normalization operation; and represent bias terms for attention calculation between source and target positions; and represent scaling factors between source and target positions; represents a noise term from the spatial domain to the frequency domain; represents a noise term from the frequency domain to the spatial domain; represents a dropout operation, is a dropout rate; and represent query mapping matrices of spatial and frequency domain features; and represent key mapping matrices of spatial and frequency domain features; and represent value mapping matrices of spatial and frequency domain features; and respectively represent bidirectional attention weight matrices, and subscript represents the spatial domain, and subscript represents the frequency domain, represents an embedding dimension; This represents the joint space-frequency characteristic representation; These are learnable fusion weights.

[0024] Furthermore, in S5, the joint spatial-frequency features are input sequentially into a multi-head self-attention network based on the Transformer architecture. By modeling the dynamic changes and global temporal dependencies of the joint spatial-frequency features, inter-frame temporal correlation features are obtained, calculated as follows:

[0025] ;

[0026] in, Indicates the first Layer Self-attention calculation for each head; MHA(.) represents multi-head self-attention operation; Concat(.) represents concatenation operation; Presentation layer normalization operation; This indicates a discard operation. For discard rate; MHA ( ) represents the inter-frame temporal correlation feature; This represents the feature matrix input into the multi-head self-attention network; Indicates the first Layer Learnable bias terms for each head; Indicates the first Layer Scaling factor for each head; Indicates the first Layer Noise items for each head; This represents a globally learnable bias term; , and Indicates the first The query, key, and value matrix of each attention head; Indicates the number of heads of attention; Represents a linear output mapping matrix; This represents the dimension of the key vector.

[0027] Furthermore, in S6, after predicting the quality score of each video frame using a multilayer perceptron (MLP), the quality scores are weighted and averaged based on learnable weights to obtain the overall video quality score. The calculation formula is as follows:

[0028] ;

[0029] ;

[0030] in, This represents the fusion feature of frame t; denotes a frame-level score; denotes a trainable parameter; denotes a nonlinear activation function; denotes a learnable weight; is a video overall quality score; is a weight calculation vector; is a total number of video frames.

[0031] In another aspect, a video quality evaluation system based on space-frequency joint and temporal interaction includes:

[0032] A key frame enhancement module is configured to obtain an immersive video sequence, perform key frame extraction, size normalization and color standardization processing on the immersive video sequence, obtain standardized key frames, perform brightness adjustment, random cropping and horizontal flipping on the standardized key frames, and obtain enhanced key frames.

[0033] A spatial feature representation acquisition module is configured to input the enhanced key frames into a spatial feature extraction branch, extract multi-directional texture and edge structure information in the enhanced key frames through a learnable Gabor convolution kernel, obtain spatial features of the key frames, and use a multi-stage residual network to perform multi-scale modeling on the spatial features of the key frames to obtain spatial feature representations of deep and shallow layers.

[0034] A frequency domain feature acquisition module is configured to input the enhanced key frames into a frequency domain feature extraction branch, perform two-dimensional fast Fourier transform on the enhanced key frames to obtain a log amplitude spectrum, capture energy distribution and periodicity features of the key frames through the log amplitude spectrum, and then input the log amplitude spectrum into a convolution network to extract frequency spectrum energy and noise structure features of the key frames to obtain frequency domain features of the key frames.

[0035] A space-frequency feature fusion module is configured to input the spatial feature representations of deep and shallow layers and the frequency domain features into a space-frequency interaction module, realize dynamic fusion of the space-frequency features through a bidirectional Cross-Attention mechanism, and obtain joint space-frequency feature representations.

[0036] An inter-frame temporal correlation feature acquisition module is configured to input the joint space-frequency features into a multi-head self-attention network based on a Transformer architecture in a time sequence, model dynamic changes and global temporal dependencies of the joint space-frequency features, and obtain inter-frame temporal correlation features.

[0037] A quality score module is configured to input the inter-frame temporal correlation features into a quality regression module, predict quality scores of each frame of video through a multi-layer perceptron (MLP), and then perform weighted averaging on the quality scores based on learnable weights to obtain a video overall quality score.​​​

[0038] The application adopts the above technical scheme and has beneficial effects:

[0039] (1) The application extracts spatial texture edge features through a learnable Gabor convolution and a multi-stage residual network, extracts frequency domain energy and noise features by combining a two-dimensional FFT logarithmic amplitude spectrum and a frequency spectrum convolution, and realizes dynamic interaction and fusion of spatial and frequency domains by using a bidirectional Cross-Attention, thereby improving the comprehensiveness and fine-grained representation ability of quality evaluation;

[0040] (2) The application models the joint spatial and frequency domain features in time sequence through a multi-head self-attention mechanism based on a Transformer architecture, captures long-range inter-frame dependencies and dynamic change rules, and realizes effective discrimination of motion blur, time sequence distortion and cross-frame artifacts;

[0041] (3) The application directly maps the spatial and frequency domain joint features and the time sequence correlation features to quality scores, and introduces a learnable weighted average strategy, realizes adaptive quality prediction independent of reference videos, takes into account high precision, strong robustness and real-time applicability, and significantly improves the consistency with human subjective perception. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 The application is based on a video quality evaluation method flowchart of spatial and frequency domain joint and time sequence interaction;

[0043] Figure 2 The application is a general principle diagram of a video quality evaluation method;

[0044] Figure 3 The application is a principle diagram of a spatial and frequency domain interaction attention module;

[0045] Figure 4 The application is a video quality evaluation system diagram based on spatial and frequency domain joint and time sequence interaction. DETAILED DESCRIPTION

[0046] The application will be further described in detail below in combination with embodiments and drawings, but the embodiments of the application are not limited thereto.

[0047] As shown in the drawings, Figure 1 The application is based on a video quality evaluation method of spatial and frequency domain joint and time sequence interaction, which comprises:

[0048] S1, an immersive video sequence is obtained, key frames of the immersive video sequence are extracted, size normalization and color standardization processing are performed, standardized key frames are obtained, brightness adjustment, random cropping and horizontal flipping are performed on the standardized key frames, and enhanced key frames are obtained.

[0049] Specifically, in the embodiment, the original video sequence is obtained from the immersive video source:

[0050] ;

[0051] wherein, represents the t-th frame image, is the total number of video frames, a key frame extraction algorithm is used to select representative frames from each video segment, covering the main content and scene changes; each frame image is subjected to size normalization, color standardization and data enhancement processing, including brightness adjustment, random cropping, horizontal flipping, etc., to unify the input format and enhance the model robustness.

[0052] S2, input the enhanced key frame into the spatial feature extraction branch, extract the multi-directional texture and edge structure information in the enhanced key frame through the learnable Gabor convolution kernel, obtain the spatial feature of the key frame, and use the multi-stage residual network to model the spatial feature of the key frame in multiple scales, and obtain the spatial feature representation of the deep and shallow layers.

[0053] Specifically, the spatial feature extraction branch includes a learnable Gabor convolution kernel;

[0054] The learnable Gabor convolution kernel is defined as follows:

[0055] ;

[0056] wherein, and represent the spatial pixel coordinates; exp() represents the exponential function; cos() represents the cosine function, , represents the coordinate component after rotation of the Gabor coordinates; represents the wavelength; represents the direction parameter; represents the phase shift; represents the Gaussian envelope width; represents the spatial aspect ratio.

[0057] S3, input the enhanced key frame into the frequency domain feature extraction branch, perform two-dimensional fast Fourier transform on the enhanced key frame to obtain the log amplitude spectrum, capture the energy distribution and periodic characteristics of the key frame through the log amplitude spectrum, and then input the log amplitude spectrum into the convolution network to extract the frequency spectrum energy and noise structure features of the key frame, to obtain the frequency domain feature of the key frame.

[0058] Specifically, the two-dimensional fast Fourier transform is performed on the enhanced key frame to obtain the log amplitude spectrum, and the calculation formula is as follows:

[0059] ;

[0060] ;

[0061] wherein, and denote the row and column indices of the pixel in the spatial domain of the input frame; denote the input frame; denote the height of the frame; denote the width of the frame; denote the frequency domain coordinates; denote the spectrum in complex form; denote the log-amplitude spectrum, log(.) denotes the logarithm function.

[0062] Specifically, in the embodiment, high-dimensional semantic features are extracted through a multi-layer residual network to obtain a spatial feature set:

[0063] ;

[0064] wherein, denote the spatial feature set, denote the spatial feature output by the spatial feature extraction function at the kth stage, and each frame of image is input into a frequency domain feature extraction module, which performs two-dimensional fast Fourier transform (FFT) on each frame of image.

[0065] S4, the spatial feature representation and the frequency domain feature fused by deep and shallow layers are input into a space-frequency interaction module, and dynamic fusion of the space-frequency features is realized through a bidirectional Cross-Attention mechanism to obtain a joint space-frequency feature representation.

[0066] Specifically, dynamic fusion of the space-frequency features is realized through a bidirectional Cross-Attention mechanism to obtain a joint space-frequency feature representation, and the calculation process is as follows:

[0067] ;

[0068] ;

[0069] wherein, Softmax(.) denotes a standardization exponential function; denote the layer normalization operation; and denote the bias term of attention calculation between the source position and the target position; and denote the scaling factor between the source and target positions; denote the noise term from the spatial domain to the frequency domain; denote the noise term from the frequency domain to the spatial domain; denote the dropout operation, is the dropout rate. and denote the query map matrix of spatial and frequency domain features; and denote the key map matrix of spatial and frequency domain features; and denote the value map matrix of spatial and frequency domain features; and denote the bidirectional attention weight matrix, respectively, subscript denotes the spatial domain, subscript denotes the frequency domain, denotes the embedding dimension; denotes the joint spatial and frequency feature representation; is a learnable fusion weight.

[0070] Specifically, in the embodiment, the high-frequency and low-frequency mask decomposition method is adopted to extract the high-frequency details and low-frequency structure information of the image respectively:

[0071] ;

[0072] wherein, denotes the high-frequency details, denotes the low-frequency structure.

[0073] The frequency domain feature map is obtained through the convolution layer:

[0074] ;

[0075] wherein, denotes the function of convolution operation.

[0076] Further, the spatial domain feature and the frequency domain feature are input into the spatial and frequency interaction attention module, and the query, key and value matrices are constructed through linear mapping:

[0077] ;

[0078] wherein, , denote the query map matrix of spatial and frequency domain features, , denote the key map matrix of spatial and frequency domain features, , denote the value map matrix of spatial and frequency domain features, , , denote the weight matrix of query, key and value of spatial domain feature, , , A weight matrix representing queries, keys, and values of frequency domain features.

[0079] S5, input the joint space-frequency features in time sequence into the multi-head self-attention network based on the Transformer architecture, model the dynamic changes and global time sequence dependence of the joint space-frequency features, and obtain inter-frame time sequence correlation features.

[0080] Specifically, the joint space-frequency features are input in time sequence into the multi-head self-attention network based on the Transformer architecture, the dynamic changes and global time sequence dependence of the joint space-frequency features are modeled, and inter-frame time sequence correlation features are obtained, and the calculation formula is as follows:

[0081] ;

[0082] Among them, denotes the self-attention calculation of the th head of the th layer; MHA(.) denotes the multi-head self-attention operation; Concat(.) denotes the concatenation operation; denotes the layer normalization operation; denotes the dropout operation, is the dropout rate; MHA( ) denotes the inter-frame time sequence correlation feature; denotes the feature matrix input into the multi-head self-attention network; denotes the learnable bias term of the th head of the th layer; denotes the scaling factor of the th head of the th layer; denotes the noise term of the th head of the th layer; denotes the global learnable bias term; , and denote the query, key, and value matrices of the th attention head; denotes the number of attention heads; denotes the linear output mapping matrix; denotes the key vector dimension.

[0083] S6, input the inter-frame time sequence correlation features into the quality regression module, predict the quality score of each frame of video through the multi-layer perception MLP, and then weight average the quality scores based on the learnable weights to obtain the overall quality score of the video.

[0084] Specifically, after predicting the quality score of each frame of video by the multi-layer perception (MLP), the quality scores are weighted and averaged based on learnable weights to obtain the overall video quality score, and the calculation formula is as follows:

[0085] ;

[0086] ;

[0087] wherein, denotes the fusion feature of the t-th frame; denotes the frame-level score; , , and denote trainable parameters; denotes a nonlinear activation function; denotes a learnable weight; is the overall video quality score; is a weight calculation vector; is the total number of video frames.

[0088] Specifically, Figure 2 shows the overall architecture of the immersive video quality evaluation model proposed in the application, covering the complete process from key frame extraction, space-frequency dual-path feature extraction, space-frequency interaction fusion to time series modeling and quality regression. Among them, the spatial domain branch adopts a learnable Gabor convolution and a multi-stage residual network to extract texture and structure information, the frequency domain branch obtains a log amplitude spectrum through a two-dimensional FFT and combines a convolution network to mine energy distribution and noise features; after dynamic fusion through bidirectional Cross-Attention, the two are input into a time series modeling module based on a Transformer, and finally an MLP is used to regress the overall video quality score, realizing comprehensive, robust and human eye perception quality evaluation of multiple distortion phenomena. Figure 3 The space-frequency interaction module in Figure 2 is described in detail, focusing on the design and implementation of the bidirectional Cross-Attention mechanism. This module maps the spatial and frequency domain features into query, key and value vectors respectively, and through attention calculation in the spatial-to-frequency and frequency-to-spatial directions, realizes cross-domain guidance and complementary enhancement between features; multi-head parallel processing further improves the modeling capability, and the fused joint features not only retain the structural nature of spatial details, but also integrate the statistical nature of frequency energy, significantly enhancing the model's discrimination accuracy and generalization ability for complex distortions such as blur, noise and compression artifacts.

[0089] Specifically, the main purpose of the present application is to provide an immersive video quality evaluation method based on space-frequency joint and time sequence interaction modeling, which organically integrates spatial (spatial structure), frequency (energy distribution) and time sequence (dynamic dependence) information, and builds an end-to-end no-reference video quality prediction model, so as to realize automatic, high-precision and interpretable evaluation of immersive video quality without reference video. Compared with the prior art, the present application can effectively overcome the problem of precision decline of traditional methods in the absence of spectral information, insufficient time sequence perception and multi-distortion scene, and realize the consistency of video perceptual quality and human subjective score.

[0090] Specifically, in the present embodiment, the immersive video quality evaluation model based on space-frequency joint and time sequence interaction modeling is built in Pytorch environment and experiments are conducted using NVIDIA RTX A6000 GPU; the minimum resolution size of the key frame is adjusted to 520 while maintaining the original aspect ratio. During training, the key frame is randomly cropped to 448*448. For the video block, the video block resolution is adjusted to 224*224 during training, the batch size is 64, the number of training rounds is 100, the initial learning rate is set to 0.00001, and the Adam optimizer is used for training. The experimental data set is LIVE-360, CVIQ, VQA-ODV, etc. public data set, 80% of which is used for training and the remaining 20% is used for testing. Spearman rank correlation coefficient (SROCC), Pearson linear correlation coefficient (PLCC) and root mean square error (RMSE) are selected to evaluate the performance of the above model.

[0091] As shown in Figure 4 The present embodiment also discloses a video quality evaluation system based on space-frequency joint and time sequence interaction, which comprises:

[0092] The key frame enhancement module 41 is used for acquiring an immersive video sequence, performing key frame extraction, size normalization and color standardization processing on the immersive video sequence, obtaining standardized key frames, performing brightness adjustment, random cropping and horizontal flipping on the standardized key frames, and obtaining enhanced key frames;

[0093] The spatial feature representation acquisition module 42 is used for inputting the enhanced key frames into a spatial feature extraction branch, extracting multi-directional texture and edge structure information in the enhanced key frames through a learnable Gabor convolution kernel, obtaining spatial features of the key frames, and using a multi-stage residual network to perform multi-scale modeling on the spatial features of the key frames, and obtaining spatial feature representations of deep and shallow layers.

[0094] The frequency domain feature acquisition module 43 is configured to input the enhanced key frame into a frequency domain feature extraction branch, perform two-dimensional fast Fourier transform on the enhanced key frame, acquire a log amplitude spectrum, capture energy distribution and periodicity characteristics of the key frame through the log amplitude spectrum, input the log amplitude spectrum into a convolution network to extract frequency spectrum energy and noise structure characteristics of the key frame, and obtain frequency domain features of the key frame.

[0095] The spatial-frequency feature fusion module 44 is configured to input the spatial feature representation and the frequency domain features fused in the deep and shallow layers into a spatial-frequency interaction module, realize dynamic fusion of the spatial-frequency features through a bidirectional Cross-Attention mechanism, and obtain joint spatial-frequency feature representation.

[0096] The inter-frame time sequence correlation feature acquisition module 45 is configured to input the joint spatial-frequency features in time sequence into a multi-head self-attention network based on a Transformer architecture, model dynamic changes and global time sequence dependence of the joint spatial-frequency features, and acquire inter-frame time sequence correlation features.

[0097] The quality score module 46 is configured to input the inter-frame time sequence correlation features into a quality regression module, predict a quality score of each frame of video through a multi-layer perception MLP, perform weighted average on the quality scores based on learnable weights, and obtain a video overall quality score.

[0098] The specific implementation of the video quality evaluation system based on spatial-frequency joint and time sequence interaction is the same as a video quality evaluation method based on spatial-frequency joint and time sequence interaction, and the present embodiment will not be repeated.

[0099] Although the present application is specifically shown and introduced in combination with the preferred embodiments, it should be understood by those skilled in the art that various changes can be made to the present application in form and details without departing from the spirit and scope of the present application defined in the appended claims, and all such changes are within the protection scope of the present application.

Claims

1. A video quality assessment method based on spatial-frequency joint and temporal interaction, characterized in that, Includes the following steps: S1, acquire the immersive video sequence, perform keyframe extraction, size normalization and color normalization on the immersive video sequence to obtain standardized keyframes, and perform data augmentation on the standardized keyframes to obtain enhanced keyframes. S2, the enhanced keyframe is input into the spatial feature extraction branch, and multi-directional texture and edge structure information in the enhanced keyframe is extracted through learnable Gabor convolution kernels to obtain the spatial features of the keyframe. Multi-stage residual network is used to model the spatial features of the keyframe at multiple scales to obtain a spatial feature representation fused from deep and shallow layers. S3 inputs the enhanced keyframe into the frequency domain feature extraction branch, performs a two-dimensional fast Fourier transform on the enhanced keyframe to obtain the logarithmic amplitude spectrum, captures the energy distribution and periodicity features of the keyframe through the logarithmic amplitude spectrum, and then inputs the logarithmic amplitude spectrum into the convolutional network to extract the spectral energy and noise structure features of the keyframe to obtain the frequency domain features of the keyframe. S4. Input the spatial feature representation and frequency domain feature fused from the deep and shallow layers into the space-frequency interaction module. The dynamic fusion of space-frequency features is achieved through the bidirectional Cross-Attention mechanism to obtain the joint space-frequency feature representation. S5. The joint spatial frequency features are input into a multi-head self-attention network based on the Transformer architecture in chronological order. By modeling the dynamic changes and global temporal dependencies of the joint spatial frequency features, inter-frame temporal correlation features are obtained. S6. Input the inter-frame temporal correlation features into the quality regression module. After predicting the quality score of each frame of video through the multilayer perceptron (MLP), the quality scores of all frames are weighted and averaged based on the learnable weights to obtain the overall video quality score.

2. The video quality evaluation method based on spatial-frequency joint and temporal interaction according to claim 1, characterized in that, In S2, the spatial feature extraction branch includes learnable Gabor convolution kernels; The learnable Gabor convolution kernel is defined as follows: ; in, and Represents pixel coordinates in the spatial domain; exp() represents the exponential function; cos() represents the cosine function. , Represents the coordinate components after Gabor coordinate rotation; Indicates wavelength; Indicates the direction parameter; Indicates phase shift; Indicates the width of the Gaussian envelope; Indicates the aspect ratio of a space.

3. The video quality evaluation method based on space-frequency joint and temporal interaction according to claim 1, characterized in that, In S3, a two-dimensional fast Fourier transform is performed on the enhanced keyframes to obtain the logarithmic amplitude spectrum, calculated as follows: ; ; in, and This represents the row and column index of the input frame in the spatial domain; Indicates the input frame; Indicates the height of the frame; Indicates the width of the frame; Represents frequency domain coordinates; Represents the spectrum in complex form; This represents the logarithmic amplitude spectrum, and log(.) represents the logarithmic function.

4. The video quality evaluation method based on spatial-frequency joint and temporal interaction according to claim 1, characterized in that, In S4, the dynamic fusion of spatial-frequency features is achieved through a bidirectional Cross-Attention mechanism to obtain a joint spatial-frequency feature representation. The calculation process is as follows: ; ; Where Softmax(.) represents the standardized exponential function; Presentation layer normalization operation; and This represents the bias term used in the attention calculation between the source and target locations; and This represents the scaling factor between the source and target locations; This represents the noise term from the spatial domain to the frequency domain; This represents the noise term from the frequency domain to the spatial domain; This indicates a discard operation. For discard rate; and A query mapping matrix representing spatial and frequency domain features; and A key mapping matrix representing spatial and frequency domain features; and The value mapping matrix represents the spatial domain features and the frequency domain features; and These represent the bidirectional attention weight matrix, with subscripts... Indicates airspace, subscript Represents the frequency domain. Indicates the embedding dimension; This represents the joint space-frequency characteristic representation; These are learnable fusion weights.

5. The video quality evaluation method based on spatial-frequency joint and temporal interaction according to claim 4, characterized in that, In S5, the joint spatial-frequency features are input sequentially into a multi-head self-attention network based on the Transformer architecture. By modeling the dynamic changes and global temporal dependencies of the joint spatial-frequency features, inter-frame temporal correlation features are obtained. The calculation formula is as follows: ; in, Indicates the first Layer Self-attention calculation for each head; MHA(.) represents multi-head self-attention operation; Concat(.) represents concatenation operation; Presentation layer normalization operation; This indicates a discard operation. For discard rate; MHA ( ) represents the inter-frame temporal correlation feature; This represents the feature matrix input into the multi-head self-attention network; Indicates the first Layer Learnable bias terms for each head; Indicates the first Layer Scaling factor for each head; Indicates the first Layer Noise items for each head; This represents a globally learnable bias term; , and Indicates the first The query, key, and value matrix of each attention head; Indicates the number of heads of attention; Represents a linear output mapping matrix; This represents the dimension of the key vector.

6. The video quality evaluation method based on spatial-frequency joint and temporal interaction according to claim 1, characterized in that, In S6, after predicting the quality score of each frame of video using a multilayer perceptron (MLP), the quality scores are weighted and averaged based on learnable weights to obtain the overall video quality score. The calculation formula is as follows: ; ; in, This represents the fusion feature of frame t; Indicates frame-level score; , , and Indicates trainable parameters; Represents a non-linear activation function; Represents the learnable weights; The overall quality score for the video; Calculate the weight vector; This represents the total number of video frames.

7. A video quality evaluation system based on space-frequency joint and temporal interaction, characterized in that, include: The keyframe enhancement module is used to acquire immersive video sequences, extract keyframes, normalize their size, and normalize their color to obtain standardized keyframes. The standardized keyframes are then subjected to brightness adjustment, random cropping, and horizontal flipping to obtain enhanced keyframes. The spatial feature representation acquisition module is used to input the enhanced keyframe into the spatial feature extraction branch, extract multi-directional texture and edge structure information in the enhanced keyframe through learnable Gabor convolution kernels, obtain the spatial features of the keyframe, and use a multi-stage residual network to perform multi-scale modeling of the spatial features of the keyframe to obtain a deep and shallow layer fused spatial feature representation. The frequency domain feature acquisition module is used to input the enhanced keyframe into the frequency domain feature extraction branch, perform a two-dimensional fast Fourier transform on the enhanced keyframe to obtain the logarithmic amplitude spectrum, capture the energy distribution and periodicity features of the keyframe through the logarithmic amplitude spectrum, and then input the logarithmic amplitude spectrum into the convolutional network to extract the spectral energy and noise structure features of the keyframe to obtain the frequency domain features of the keyframe. The spatial-frequency feature fusion module is used to input the spatial feature representations fused from the deep and shallow layers and the frequency domain features into the spatial-frequency interaction module. The dynamic fusion of spatial-frequency features is achieved through a bidirectional Cross-Attention mechanism to obtain a joint spatial-frequency feature representation. The inter-frame temporal correlation feature acquisition module is used to input the joint spatial frequency features into a multi-head self-attention network based on the Transformer architecture in chronological order. By modeling the dynamic changes and global temporal dependencies of the joint spatial frequency features, the inter-frame temporal correlation features are obtained. The quality scoring module is used to input the inter-frame temporal correlation features into the quality regression module. After predicting the quality score of each frame of video through a multilayer perceptron (MLP), the quality scores are weighted and averaged based on learnable weights to obtain the overall video quality score.

Citation Information

Patent Citations

  • SCV coding perception code rate control method and device fusing space-frequency domain saliency characteristics

    CN118450127A

  • No-reference video quality evaluation method based on space-time perception feature fusion

    CN118968267A

  • Immersive video enhancement method and device based on frequency domain boundary collaborative optimization

    CN119850441A

  • Screen content video quality evaluation method and device based on deep and shallow layer spatial-temporal characteristics

    CN120031869A

  • Time domain compression femtosecond holographic microscopy reconstruction method based on spatial domain and frequency domain joint learning

    CN121147341A

Cited By

  • Immersive video quality evaluation method and system based on multi-modal perception

    CN122134726A