Immersive Video Quality Evaluation Method and Device Based on Viewpoint Spatiotemporal Correlation

By constructing an immersive video quality evaluation model that includes multi-viewpoint spatial interaction network and temporal non-local network, the problem of inaccurate evaluation of immersive video quality in the prior art is solved, and high-precision evaluation consistent with human visual perception is achieved.

CN120047434BActive Publication Date: 2025-07-22HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510505105.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-22
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing video quality assessment methods cannot accurately evaluate the quality of immersive videos, especially in virtual reality and augmented reality applications, which cannot effectively reflect dynamic changes in viewpoint selection and time dimensions, resulting in the impact of user experience.

Method used

An immersive video quality evaluation model including texture depth feature extraction, multi-viewpoint spatial interaction network, multi-viewpoint time non-local network, channel attention mechanism and gated loop unit is constructed. By processing multi-viewpoint texture video blocks, deep video blocks and keyframes, the quality of immersive video is achieved.

Benefits of technology

The evaluation accuracy of immersive video quality is improved, making it highly consistent with human visual perception characteristics, enhancing the perception ability and time perception ability of viewpoint spatial correlation, and providing more accurate quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047434B_ABST
    Figure CN120047434B_ABST
Patent Text Reader

Abstract

The present invention discloses an immersive video quality evaluation method and device based on view spatio-temporal correlation, which relates to the technical field of video image processing. The method includes: obtaining an immersive video containing multi-view texture videos and depth videos, and extracting texture video blocks, depth video blocks, texture key frames, and depth key frames therefrom; inputting the extracted data into a trained immersive video quality evaluation model for processing; the model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; obtaining a texture video score and a depth video score through the interactive processing of the model, and weighted aggregation of the scores to obtain a final immersive video quality score. The present invention realizes the evaluation of the quality of immersive videos by obtaining and processing multi-view texture and depth information in immersive videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image processing, and particularly to an immersive video quality evaluation method and device based on view spatio-temporal correlation. Background Art

[0002] The immersive video quality evaluation based on view spatio-temporal correlation is an important research direction for evaluating the quality of user experience in virtual reality (VR), augmented reality (AR), and other 360-degree panoramic video applications. With the development of these technologies, traditional two-dimensional video quality evaluation methods can no longer meet the requirements because immersive videos not only contain spatial information but also involve the user's view selection and dynamic changes in the time dimension.

[0003] Today, with the rapid development of digital multimedia technology, the concept of the metaverse has taken root in people's hearts and spawned a series of new applications that integrate virtual and real technologies, among which immersive videos are particularly eye-catching. Immersive video technology has shown its extensive application potential in many fields such as distance education and medical care. Compared with traditional videos, immersive videos not only contain richer texture details but also incorporate complex parallax information, bringing users an unprecedented visual experience.

[0004] However, during the acquisition, transmission, and display of immersive videos, distortion problems such as motion blur and Gaussian noise will inevitably occur. These problems seriously affect the video quality and user experience. Therefore, there is an urgent need to develop a method and device that can accurately evaluate the quality of immersive videos and conform to human visual perception, providing a solid foundation for the further development and optimization of immersive video technology. Summary of the Invention

[0005] To solve the above problems, the present invention proposes an immersive video quality evaluation method and device based on view spatio-temporal correlation. By constructing an immersive video quality evaluation model that includes texture depth feature extraction, multi-view spatial interaction network, multi-view temporal non-local network, channel attention mechanism, and gated recurrent unit, the texture video blocks, depth video blocks, and key frames of multiple views are processed to achieve the evaluation of immersive video quality.

[0006] The specific solutions are as follows:

[0007] On the one hand, an immersive video quality evaluation method based on view spatio-temporal correlation includes:

[0008] Obtain an immersive video containing texture videos and depth videos of multiple views, extract texture video blocks and depth video blocks in the texture videos and depth videos of each view, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0009] Build an immersive video quality evaluation model and train it to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part; the spatio-temporal interaction part of texture-depth features includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0010] Input the texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model; specifically, extract features from the texture key frames and depth key frames through the texture-depth feature extraction module to obtain the first texture features and depth features of multiple viewpoints; input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and obtain multi-viewpoint spatial fusion features based on the multi-viewpoint texture-depth fusion features; process the multi-viewpoint spatial fusion features through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features; extract features from the texture video blocks through the texture video feature extraction module to obtain second texture features; input the second texture features into the channel attention module to obtain third texture features; splice the third texture features and the multi-viewpoint spatio-temporal interaction features to obtain texture fusion features; sequentially input the texture fusion features into the gated recurrent unit and the texture video quality regression module to obtain the texture video score; extract depth video block features through the depth video block feature extraction module to obtain depth video block features; input the depth video block features into the depth video quality regression module to obtain the depth video score; weighted aggregate the texture video score and the depth video score to obtain the final immersive video score.

[0011] Furthermore, extract the texture video blocks and depth video blocks in the texture video and depth video of each viewpoint, and extract the texture key frames and depth key frames based on the texture video blocks and depth video blocks, as follows:

[0012] For the texture video of each viewpoint and the depth video Perform block division to obtain K consecutive texture video blocks and K consecutive depth video blocks;

[0013] Among them, ; represents the floor function; N represents the number of frames of the texture video or depth video of each viewpoint; and respectively represent the i-th texture video frame and the i-th depth video frame; the expression of the j-th texture video block is ; the expression of the j-th depth video block is ; represents the number of video frames included in the texture video block or the depth video block;

[0014] Extract the first frame of the texture video block and the depth video block as the key frame of the corresponding block, and obtain the texture key frame and the depth key frame .

[0015] Furthermore, input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain the multi-viewpoint texture-depth fusion features. The calculation formula is as follows:

[0016] ;

[0017] ;

[0018] ;

[0019] where, , represents the -th viewpoint depth feature; represents the first texture feature of the -th viewpoint, ; H, W, and C represent the height, width, and number of channels of the feature; represents feature concatenation; represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the -th viewpoint texture-depth fusion feature; represents element-wise multiplication of features.

[0020] Furthermore, based on the multi-viewpoint texture-depth fusion features, obtain the multi-viewpoint spatial fusion features. The calculation formula is as follows:

[0021] ;

[0022] ;

[0023] ;

[0024] ;

[0025] ;

[0026] ;

[0027] Among them, represents the multi-view synthesis feature after feature splicing; represents feature splicing; represents spatial global average pooling; represents the feature after spatial global pooling; represents the fully connected layer; represents the dimensionality-reduced feature; represents the channel attention vector of different view texture features; represents the second channel attention vector; represents the multi-view spatial fusion feature; represents the convolution operation; e represents the exponential calculation; represents the texture depth fusion feature of the 1st to Nth views.

[0028] Furthermore, the multi-view spatial fusion feature is processed by the multi-view temporal non-local network to obtain the integrated multi-view spatio-temporal interaction feature. The calculation formula is as follows:

[0029] ;

[0030] ;

[0031] ;

[0032] ;

[0033] ;

[0034] ;

[0035] Among them, represents adjusting the size of the feature dimension; represents the multi-view spatial fusion feature at time k; represents the preliminary synthesis feature; represents the 3D convolution operation; represents adjusting the feature shape; , and respectively represent the query feature, key feature, and value feature; represents matrix multiplication, represents the multi-view spatio-temporal fusion feature; represents the pooling strategy along different channels; represents the multi-view spatio-temporal interaction feature at time k.

[0036] Furthermore, the third texture feature and the multi-view spatio-temporal interaction feature are spliced to obtain a texture fusion feature, and the calculation formula is as follows:

[0037] ;

[0038] ;

[0039] ;

[0040] Among them, represents n-channel attention modules; represents the second texture feature; represents the third texture feature; represents the multi-view average pooling strategy; represents the multi-view pooling feature; represents the texture fusion feature at the k-th moment.

[0041] Furthermore, the texture fusion feature is sequentially input into a gated recurrent unit and a texture video quality regression module to obtain a texture video score, and the specific formula is as follows:

[0042] ;

[0043] ;

[0044] ;

[0045] ;

[0046] ;

[0047] Among them, represents the gated recurrent unit; represents the minimum selection operation; represents the Gaussian weight corresponding to the frame; represents the hyperparameter for adjusting the direct and indirect influence effects; represents the indirect influence element; represents the direct influence element; represents the texture video frame quality score obtained by weighted summation after quality regression; represents the global average pooling of all texture video frame quality scores to obtain the texture video score; represents the time; represents the time length of the direct influence; represents the time of the frame with the worst quality in the time length of the direct influence; Indicates the length of the texture video frame.

[0048] Furthermore, the weighted aggregated texture video score and depth video score are used to obtain the final immersive video score, and the calculation formula is as follows:

[0049] ;

[0050] ;

[0051] ;

[0052] Among them, represents the multi-layer perceptron in the quality regression module; represents the depth video block feature of the i-th viewpoint at the same moment; represents the depth feature after multi-viewpoint pooling at time k; represents the hyperparameter for adjusting the influence effect of texture and depth video scores; represents the final immersive video quality score; represents the depth video score.

[0053] On the other hand, an immersive video quality evaluation device based on view spatio-temporal correlation includes:

[0054] An immersive video extraction module for obtaining an immersive video containing texture videos and depth videos of multiple viewpoints, extracting texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extracting texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0055] A model construction and training module for constructing and training an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0056] The immersive video score evaluation module is used to input texture video blocks, depth video blocks, texture key frames, and depth key frames into a trained immersive video quality evaluation model. Specifically, the texture depth feature extraction module extracts features from the texture key frames and depth key frames to obtain the first texture features and depth features of multiple viewpoints. The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture depth fusion features. The multi-viewpoint spatial fusion features are processed by the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. The texture video feature extraction module extracts features from the texture video blocks to obtain second texture features. The second texture features are input into the channel attention module to obtain third texture features. The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features. The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The depth video block feature extraction module extracts depth video block features to obtain depth video block features. The depth video block features are input into the depth video quality regression module to obtain the depth video score. The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

[0057] The present invention adopts the above technical solutions and has the following beneficial effects:

[0058] (1) In the present invention, the multi-viewpoint spatial interaction network obtains disparity features through the disparity estimation module to enhance the texture key frame features, and the multi-viewpoint selective feature integration module interacts and integrates the multi-viewpoint features at the same moment to enhance the features with higher viewpoint spatial correlation.

[0059] (2) The present invention realizes a three-dimensional attention mechanism through the multi-viewpoint temporal non-local attention network to focus on the temporal correlation of multi-viewpoint features, so as to improve the temporal perception ability of different viewpoint features.

[0060] (3) The present invention proposes to perform quality perception on immersive videos through three parts: the quality of texture video blocks, the quality of depth video blocks, and the texture-depth key frame interaction quality. Through the designed weighted aggregation strategy, the overall quality of immersive videos is accurately perceived and evaluated, and the evaluation results are highly consistent with the human visual perception characteristics. Description of the Drawings

[0061] Figure 1 It is a flowchart of the immersive video quality evaluation method based on viewpoint spatio-temporal correlation according to an embodiment of the present invention;

[0062] Figure 2 It is a schematic structural diagram of the immersive video quality evaluation model according to an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the structure of the multi-viewpoint spatial interaction network according to an embodiment of the present invention;

[0064] Figure 4 Schematic diagram of the structure of the multi-viewpoint temporal non-local network according to an embodiment of the present invention;

[0065] Figure 5 Diagram of the immersive video quality evaluation device based on viewpoint spatio-temporal correlation according to an embodiment of the present invention. Detailed implementation manners

[0066] The present invention will be further described in detail below with reference to embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto. As Figure 1 shown, the immersive video quality evaluation method based on viewpoint spatio-temporal correlation of the present invention includes:

[0067] S1. Obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks.

[0068] Specifically, extracting texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extracting texture key frames and depth key frames based on the texture video blocks and depth video blocks are as follows:

[0069] For the texture video of each viewpoint and the depth video perform block division to obtain K consecutive texture video blocks and K consecutive depth video blocks;

[0070] wherein, ; represents rounding down; N represents the number of frames of the texture video or depth video of each viewpoint; and respectively represent the i-th texture video frame and the i-th depth video frame; the expression of the j-th texture video block is ; the expression of the j-th depth video block is ; represents the number of video frames included in the texture video block or depth video block;

[0071] Extract the first frame of the texture video block and depth video block as the key frame of the corresponding block to obtain the texture key frame and the depth key frame .

[0072] S2. Construct an immersive video quality evaluation model and train it to obtain a trained immersive video quality evaluation model. The immersive video quality evaluation model includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part. The spatio-temporal interaction part of texture-depth features includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network. The texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module. The depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module.

[0073] Specifically, as Figure 2 shown, the immersive video quality evaluation model based on viewpoint spatio-temporal correlation proposed in the embodiment of the present application includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part. The spatio-temporal interaction part of texture-depth features includes a feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network. The texture video quality evaluation part includes a texture video block feature extraction module, a channel attention module, a gated recurrent unit, and a texture video frame quality regression module. The depth video quality evaluation part includes a depth video block feature extraction module and a quality regression module.

[0074] S3. Input texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model. Specifically, use the texture-depth feature extraction module to extract features from the texture key frames and depth key frames to obtain the first texture features and depth features of multiple viewpoints. Input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and obtain multi-viewpoint spatial fusion features based on the multi-viewpoint texture-depth fusion features. Process the multi-viewpoint spatial fusion features through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. Use the texture video feature extraction module to extract features from the texture video blocks to obtain second texture features. Input the second texture features into the channel attention module to obtain third texture features. Concatenate the third texture features and the multi-viewpoint spatio-temporal interaction features to obtain texture fusion features. Input the texture fusion features into the gated recurrent unit and the texture video quality regression module in sequence to obtain the texture video score. Extract depth video block features through the depth video block feature extraction module to obtain depth video block features. Input the depth video block features into the depth video quality regression module to obtain the depth video score. Weightedly aggregate the texture video score and the depth video score to obtain the final immersive video score.

[0075] Specifically, input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features. The calculation formula is as follows:

[0076] ;

[0077] ;

[0078] ;

[0079] Among them, , representing the th viewpoint depth feature; represents the first texture feature of the th viewpoint, ; H, W, and C represent the height, width, and number of channels of the feature; represents feature concatenation; represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the th viewpoint's texture-depth fusion feature; represents element-wise multiplication of features.

[0080] Specifically, the multi-viewpoint spatial fusion feature is obtained based on the multi-viewpoint texture-depth fusion feature, and the calculation formula is as follows:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] ;

[0087] Among them, represents the multi-viewpoint synthesis feature after feature concatenation; represents feature concatenation; represents spatial global average pooling; represents the feature after spatial global pooling; represents the fully connected layer; represents the dimensionality-reduced feature; represents the channel attention vector of different viewpoint texture features; represents the second channel attention vector; represents the multi-viewpoint spatial fusion feature; represents the convolution operation; e represents the exponential calculation; Represents the texture depth fusion features of the 1st to Nth viewpoints.

[0088] Specifically, the multi-view spatial fusion features are processed by the multi-view temporal non-local network to obtain the integrated multi-view spatiotemporal interaction features. The calculation formula is as follows:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] in, Indicates adjusting the size of the feature dimension; Represents the multi-view spatial fusion features at time k; represents preliminary synthetic features; Represents a three-dimensional convolution operation; Indicates adjustment of feature shape; , and Represent query features, key features, and value features respectively; represents matrix multiplication, Represents the spatiotemporal fusion features of multiple viewpoints; represents the pooling strategies along different channels; Represents the multi-view spatiotemporal interaction features at time k.

[0096] Specifically, the third texture feature and the multi-view spatiotemporal interaction feature are spliced to obtain the texture fusion feature, and the calculation formula is as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] in, represents n channel attention modules; represents the second texture feature; represents the third texture feature; Represents the multi-view average pooling strategy; Represents the multi-viewpoint pooling feature; Represents the texture fusion feature at time k.

[0101] Specifically, the texture fusion feature is sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The specific formula is as follows:

[0102] ;

[0103] ;

[0104] ;

[0105] ;

[0106] ;

[0107] Among them, Represents the gated recurrent unit; Represents the minimum selection operation; Represents the Gaussian weight corresponding to the Represents the hyperparameter for adjusting the direct and indirect influence effects; Represents the indirect influence element; Represents the direct influence element; Represents the texture video frame quality score after weighted summation for quality regression; Represents the texture video score obtained by performing global average pooling on all texture video frame quality scores; Represents the time; Represents the time length of the direct influence; Represents the time length of the indirect influence; Represents the time of the frame with the worst quality in the time length of the direct influence; Represents the length of the texture video frame.

[0108] Specifically, as Figure 3 shown, the multi-viewpoint spatial interaction network includes a disparity estimation module and a viewpoint selective feature integration module. The texture key frame feature and the depth key frame feature are input into the multi-viewpoint spatial interaction network. The i-th depth key frame feature is concatenated with the depth key frame features of its adjacent viewpoints to obtain the input feature of the disparity module, which is input into the disparity module to obtain the disparity feature of the i-th viewpoint. The disparity feature of each viewpoint is multiplied by the corresponding viewpoint's texture key frame feature to obtain the enhanced texture feature (texture-depth fusion feature).

[0109] Specifically, as Figure 4As shown, it represents inputting the multi-viewpoint space fusion features at different times into the multi-viewpoint temporal non-local network, obtaining the preliminary synthesized features through resizing and feature concatenation, obtaining the corresponding query, key, and value features through different 3D convolution operations, then adjusting the feature shape to adapt to matrix multiplication, outputting the multi-viewpoint spatio-temporal fusion features, and adjusting them into the multi-viewpoint spatio-temporal interaction features at different times through the pooling strategy.

[0110] Specifically, the weighted aggregated texture video score and depth video score are obtained to get the final immersive video score, and the calculation formula is as follows:

[0111] ;

[0112] ;

[0113] ;

[0114] Among them, represents the multi-layer perceptron in the quality regression module; represents the depth video block feature of the i-th viewpoint at the same time; represents the depth feature after multi-viewpoint pooling at time k; represents the hyperparameter that adjusts the influence effect of the texture and depth video scores; represents the final immersive video quality score; represents the depth video score.

[0115] Specifically, in this embodiment, the texture video quality evaluation part extracts features from several texture video blocks of multiple viewpoints to obtain the texture features of multiple viewpoints , inputs the texture features into the channel attention module to obtain more representative texture features, and through resizing and average pooling, obtains the multi-viewpoint pooling features, and concatenates them with the multi-viewpoint spatio-temporal interaction features to obtain the texture fusion features; obtains the initial quality score at each time through the gated recurrent unit , inputs it into the texture video quality regression module. Specifically, for the feature at time k, by applying the minimum pooling operation to the previous frames, an element that can represent the direct effect of the current frame is obtained , that is, the quality of the th frame with the worst quality. The period of the direct effect is set to the past frame sequence. If , the frame sequence is , otherwise it is ; perform weighted processing on several frames after the th frame with the worst quality to obtain the indirect influence element , which is directly related to the element The weighted sum is used to obtain the quality score of the texture video frame after quality regression. Global average pooling is performed on all texture video frame quality scores to obtain the texture video score. .

[0116] Specifically, the immersive video quality evaluation model based on view spatio-temporal correlation proposed in this embodiment is built using the Python 3.8.18 programming language, CUDA 12.2, and torch 2.3.1, and experiments are conducted using the NVIDIA RTX A6000 GPU. The experimental dataset is IMVD, 80% of which is used for training, and the test set and validation set are 10% respectively. This model is trained on input videos with a size of 448×448, with a batch size of 8 and 20 epochs; the initial learning rate is set to 0.00001, and the Adam optimizer is used for training.

[0117] As Figure 5 shown, this embodiment also discloses an immersive video quality evaluation device based on view spatio-temporal correlation, including:

[0118] An immersive video extraction module 51, configured to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0119] A model construction and training module 52, configured to construct and train an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0120] The immersive video score evaluation module 53 is used to input the texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model. Specifically, the texture depth feature extraction module extracts features from the texture key frames and depth key frames to obtain the first texture features and depth features of multiple viewpoints. The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture depth fusion features. The multi-viewpoint spatial fusion features are processed by the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. The texture video feature extraction module extracts features from the texture video blocks to obtain second texture features. The second texture features are input into the channel attention module to obtain third texture features. The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features. The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The depth video block feature extraction module extracts depth video block features to obtain depth video block features. The depth video block features are input into the depth video quality regression module to obtain the depth video score. The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

[0121] The specific implementation of the immersive video quality evaluation device based on viewpoint spatio-temporal correlation is the same as that of the immersive video quality evaluation method based on viewpoint spatio-temporal correlation, and will not be repeated in this embodiment.

[0122] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all such changes are within the protection scope of the present invention.

Claims

1. An immersive video quality evaluation method based on the spatio-temporal correlation of viewpoints, characterized in that, Including: Obtain an immersive video including texture videos and depth videos with multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks; Construct and train an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module; Input the texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model; extract features of the texture key frames and depth key frames through the texture-depth feature extraction module to obtain first texture features and depth features of multiple viewpoints; Input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and obtain multi-viewpoint spatial fusion features based on the multi-viewpoint texture-depth fusion features; process the multi-viewpoint spatial fusion features through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features; extract features of the texture video blocks through the texture video feature extraction module to obtain second texture features; input the second texture features into the channel attention module to obtain third texture features; Concatenate the third texture features and the multi-viewpoint spatio-temporal interaction features to obtain texture fusion features; sequentially input the texture fusion features into the gated recurrent unit and the texture video quality regression module to obtain a texture video score; extract depth video block features through the depth video block feature extraction module to obtain depth video block features; input the depth video block features into the depth video quality regression module to obtain a depth video score; Weightedly aggregate the texture video score and the depth video score to obtain a final immersive video score; Input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and the calculation formula is as follows: Among them, represents the depth feature of the i-th viewpoint; F t i represents the first texture feature of the i-th viewpoint, H, W, and C represent the height, width, and number of channels of the feature; Concat(·) represents feature concatenation; DEM() represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the texture-depth fusion feature of the i-th viewpoint; represents element-wise multiplication of features.

2. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 1, wherein, Extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks, specifically as follows: Texture videos for each viewpoint and depth videos are segmented to obtain K consecutive texture video blocks and K consecutive depth video blocks; Among them, represents the floor function; N represents the number of frames of the texture video or depth video for each viewpoint; x i and y i respectively represent the i-th texture video frame and the i-th depth video frame; the expression for the j-th texture video block is The expression for the j-th depth video block is τ represents the number of video frames included in the texture video block or depth video block; Extract the first frame of the texture video block and the depth video block as the key frame of the corresponding block, and obtain the texture key frame f tj = x τj and the depth key frame f dj = y τj .

3. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 1, characterized in that, Obtain multi-viewpoint spatial fusion features based on the multi-viewpoint texture-depth fusion features, and the calculation formula is as follows: F p = GAP(F s ) F p′ = FC(F p ); W i = CNN i (F p′ ); Among them, F s represents the multi-viewpoint synthesis feature after feature concatenation; Concat(·) represents feature concatenation; GAP(·) represents spatial global average pooling; F p represents the feature after spatial global pooling; FC(·) represents the fully connected layer; F p′ represents the dimensionality-reduced feature; W i represents the channel attention vector of texture features of different viewpoints; W i ' represents the second channel attention vector; F SI represents the multi-viewpoint spatial fusion feature; CNN i represents the convolution operation; e represents the exponential calculation; represents the texture depth fusion feature of the 1st to Nth viewpoints.

4. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 3, wherein, Process the multi-viewpoint spatial fusion features through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features, and the calculation formula is as follows: V Q = Per(3DCNN q (V)); V K = Per(3DCNN k (V)); V V = Per(3DCNN v (V)); V′ = V + 3DCNN(Per(V Q ×V K ×V V )); Among them, Reshape(·) represents resizing the feature dimension; represents the multi-viewpoint space fusion feature at time k; V represents the preliminary composite feature; 3DCNN(·) represents the three-dimensional convolution operation; Per(·) represents adjusting the feature shape; V Q 、V K and V V respectively represent the query feature, key feature, and value feature; × represents matrix multiplication, and V′ represents the multi-viewpoint spatio-temporal fusion feature; Ρ vk (·) represents the pooling strategy along different channels; represents the multi-viewpoint spatio-temporal interaction feature at time k.

5. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 4, characterized in that Concatenate the third texture features and the multi-viewpoint spatio-temporal interaction features to obtain texture fusion features, and the calculation formula is as follows: Among them, CA(·) n represents n channel attention modules; represents the second texture feature; represents the third texture feature; represents the multi-viewpoint average pooling strategy; F GP represents the multi-viewpoint pooling feature; represents the texture fusion feature at time k.

6. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 5, wherein The texture fusion features are sequentially input into a gated recurrent unit and a texture video quality regression module to obtain a texture video score. The specific formula is as follows: Among them, GRU(·) represents the gated recurrent unit; min{·} represents the minimum selection operation; w z represents the Gaussian weight corresponding to the z-th frame; γ represents the hyperparameter for adjusting the direct and indirect influence effects; represents the indirect influence element; represents the direct influence element; Q k represents the texture video frame quality score after weighted summation to obtain quality regression; Q tex represents the texture video score obtained by performing global average pooling on all texture video frame quality scores; k represents the time moment; l represents the time length of the direct influence; T represents the time length of the indirect influence; r represents the time moment of the frame with the worst quality in the time length of the direct influence; K represents the length of the texture video frame.

7. The immersive video quality evaluation method based on view point spatio-temporal correlation according to claim 6, characterized in that The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score. The calculation formula is as follows: Q = λQ tex +(1 - λ)Q dep ; Among them, MLP(·) represents the multi-layer perceptron in the quality regression module; represents the depth video block feature of the i-th viewpoint at the same moment; represents the depth feature after multi-viewpoint pooling at the k-th moment; λ represents the hyperparameter for adjusting the influence effect of texture and depth video scores; Q represents the final immersive video quality score; Q dep represents the depth video score.

8. An immersive video quality evaluation device based on the spatio-temporal correlation of viewpoints, characterized in that, Including: An immersive video extraction module, configured to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and the depth video blocks; A model construction and training module, configured to construct and train an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module; An immersive video score evaluation module, configured to input the texture video blocks, the depth video blocks, the texture key frames, and the depth key frames into the trained immersive video quality evaluation model; specifically, feature extraction is performed on the texture key frames and the depth key frames through the texture-depth feature extraction module to obtain first texture features and depth features of multiple viewpoints; The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture-depth fusion features; the multi-viewpoint spatial fusion features are processed through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features; feature extraction is performed on the texture video blocks through the texture video feature extraction module to obtain second texture features; the second texture features are input into the channel attention module to obtain third texture features; The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features; the texture fusion features are sequentially input into a gated recurrent unit and a texture video quality regression module to obtain a texture video score; depth video block features are extracted through the depth video block feature extraction module to obtain depth video block features; the depth video block features are input into the depth video quality regression module to obtain a depth video score; The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score; The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features. The calculation formula is as follows: Among them, represents the depth feature of the i-th viewpoint; F t i represents the first texture feature of the i-th viewpoint, H, W, and C represent the height, width, and number of channels of the feature; Concat(·) represents feature concatenation; DEM() represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the texture-depth fusion feature of the i-th viewpoint; represents element-wise multiplication of features.

Citation Information

Patent Citations

  • Immersive video quality evaluation method and device based on multi-feature fusion

    CN118411583A

  • Immersive video quality evaluation method and device based on frame-level time aggregation strategy

    CN118609034A