Immersive video quality evaluation method and device based on viewpoint space-time correlation

By constructing an immersive video quality evaluation model that includes multi-viewpoint spatial interaction network and temporal non-local network, the problem of difficulty in evaluating immersive video quality in the prior art is solved, and accurate evaluation and improvement of video quality is achieved.

CN120047434AActive Publication Date: 2025-05-27HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510505105.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-27
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate and improve the quality of immersive videos, especially in terms of viewpoint spatiotemporal correlation, affecting user experience.

Method used

By constructing an immersive video quality evaluation model including texture depth feature extraction, multi-viewpoint spatial interaction network, multi-viewpoint time non-local network, channel attention mechanism and gated loop unit, the textured video blocks, deep video blocks and keyframes of multi-viewpoints are processed to achieve immersive video quality evaluation.

Benefits of technology

The accurate evaluation of the quality of immersive videos is achieved, which can accurately perceive and evaluate the overall quality of the video, and the evaluation results are highly consistent with the human visual perception characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047434A_ABST
    Figure CN120047434A_ABST
Patent Text Reader

Abstract

The invention discloses an immersive video quality evaluation method and device based on viewpoint space-time correlation, and relates to the technical field of video image processing, and the method comprises the steps: obtaining an immersive video containing a multi-viewpoint texture video and a multi-viewpoint depth video, and extracting a texture video block, a depth video block, a texture key frame and a depth key frame from the immersive video; inputting the extracted data into the trained immersive video quality evaluation model for processing; the model comprises a texture depth feature space-time interaction part, a texture video quality evaluation part and a depth video quality evaluation part; and obtaining a texture video score and a depth video score through model interaction processing, and carrying out weighted aggregation on the scores to obtain a final immersive video quality score. According to the method, the multi-view texture and depth information in the immersive video is acquired and processed, so that the quality of the immersive video is evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image processing, and particularly to an immersive video quality evaluation method and device based on view point spatio-temporal correlation. Background Art

[0002] The immersive video quality evaluation based on view point spatio-temporal correlation is an important research direction for evaluating the quality of user experience in virtual reality (VR), augmented reality (AR), and other 360-degree panoramic video applications. With the development of these technologies, traditional two-dimensional video quality evaluation methods can no longer meet the requirements because immersive videos not only contain spatial information but also involve the user's view point selection and dynamic changes in the time dimension.

[0003] Today, with the rapid development of digital multimedia technology, the concept of the metaverse has taken root in people's hearts and spawned a series of new applications that integrate virtual and real technologies, among which immersive videos are particularly eye-catching. Immersive video technology has shown its broad application potential in many fields such as distance education and medical treatment. Compared with traditional videos, immersive videos not only contain richer texture details but also incorporate complex parallax information, bringing users an unprecedented visual experience.

[0004] However, during the acquisition, transmission, and display of immersive videos, distortion problems such as motion blur and Gaussian noise will inevitably occur. These problems seriously affect the video quality and user experience. Therefore, there is an urgent need to develop a method and device that can accurately evaluate the quality of immersive videos and conform to human visual perception, providing a solid foundation for the further development and optimization of immersive video technology. Summary of the Invention

[0005] To solve the above problems, the present invention proposes an immersive video quality evaluation method and device based on view point spatio-temporal correlation. By constructing an immersive video quality evaluation model that includes texture depth feature extraction, multi-view point spatial interaction network, multi-view point time non-local network, channel attention mechanism, and gated recurrent unit, the texture video blocks, depth video blocks, and key frames of multiple view points are processed to achieve the evaluation of immersive video quality.

[0006] The specific solutions are as follows:

[0007] On the one hand, an immersive video quality evaluation method based on view point spatio-temporal correlation includes:

[0008] Obtain an immersive video containing texture videos and depth videos of multiple view points, extract texture video blocks and depth video blocks in the texture videos and depth videos of each view point, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0009] Construct an immersive video quality evaluation model and train it to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part; the spatio-temporal interaction part of texture-depth features includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0010] Input the texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model; specifically, extract features from the texture key frames and depth key frames through the texture-depth feature extraction module to obtain the first texture features and depth features of multiple viewpoints; input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and obtain multi-viewpoint spatial fusion features based on the multi-viewpoint texture-depth fusion features; process the multi-viewpoint spatial fusion features through the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features; extract features from the texture video blocks through the texture video feature extraction module to obtain second texture features; input the second texture features into the channel attention module to obtain third texture features; splice the third texture features and the multi-viewpoint spatio-temporal interaction features to obtain texture fusion features; sequentially input the texture fusion features into the gated recurrent unit and the texture video quality regression module to obtain the texture video score; extract depth video block features through the depth video block feature extraction module to obtain depth video block features; input the depth video block features into the depth video quality regression module to obtain the depth video score; weighted aggregate the texture video score and the depth video score to obtain the final immersive video score.

[0011] Furthermore, extract the texture video blocks and depth video blocks in the texture video and depth video of each viewpoint, and extract the texture key frames and depth key frames based on the texture video blocks and depth video blocks, specifically as follows:

[0012] For the texture video of each viewpoint and the depth video Perform chunking to obtain K consecutive texture video blocks and K consecutive depth video blocks;

[0013] Among them, ; represents the floor function; N represents the number of frames of the texture video or depth video of each viewpoint; and respectively represent the i-th texture video frame and the i-th depth video frame; the expression of the j-th texture video block is ; the expression of the j-th depth video block is ; represents the number of video frames included in the texture video block or the depth video block;

[0014] Extract the first frame of the texture video block and the depth video block as the key frame of the corresponding block, and obtain the texture key frame and the depth key frame .

[0015] Furthermore, input the first texture features and depth features of multiple viewpoints into the multi-viewpoint spatial interaction network to obtain the multi-viewpoint texture-depth fusion features. The calculation formula is as follows:

[0016] ;

[0017] ;

[0018] ;

[0019] where, , represents the depth feature of the -th viewpoint; represents the first texture feature of the -th viewpoint, ; H, W, and C represent the height, width, and number of channels of the feature; represents feature concatenation; represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the texture-depth fusion feature of the -th viewpoint; represents element-wise multiplication of features.

[0020] Furthermore, based on the multi-viewpoint texture-depth fusion features, obtain the multi-viewpoint spatial fusion features. The calculation formula is as follows:

[0021] ;

[0022] ;

[0023] ;

[0024] ;

[0025] ;

[0026] ;

[0027] Among them, represents the multi-viewpoint synthesis feature after feature splicing; represents feature splicing; represents spatial global average pooling; represents the feature after spatial global pooling; represents the fully connected layer; represents the dimensionality-reduced feature; represents the channel attention vector of different viewpoint texture features; represents the second channel attention vector; represents the multi-viewpoint spatial fusion feature; represents the convolution operation; e represents the exponential calculation; represents the texture depth fusion feature of the 1st to Nth viewpoints.

[0028] Furthermore, the multi-viewpoint spatial fusion feature is processed by the multi-viewpoint temporal non-local network to obtain the integrated multi-viewpoint spatio-temporal interaction feature. The calculation formula is as follows:

[0029] ;

[0030] ;

[0031] ;

[0032] ;

[0033] ;

[0034] ;

[0035] Among them, represents adjusting the size of the feature dimension; represents the multi-viewpoint spatial fusion feature at time k; represents the preliminary synthesis feature; represents the 3D convolution operation; represents adjusting the feature shape; , and respectively represent the query feature, key feature, and value feature; represents matrix multiplication, represents the multi-viewpoint spatio-temporal fusion feature; represents the pooling strategy along different channels; represents the multi-viewpoint spatio-temporal interaction feature at time k.

[0036] Further, the third texture feature and the multi-view spatio-temporal interaction feature are spliced to obtain a texture fusion feature, and the calculation formula is as follows:

[0037] ;

[0038] ;

[0039] ;

[0040] where, represents n-channel attention modules; represents the second texture feature; represents the third texture feature; represents the multi-view average pooling strategy; represents the multi-view pooling feature; represents the texture fusion feature at time k.

[0041] Further, the texture fusion feature is sequentially input into a gated recurrent unit and a texture video quality regression module to obtain a texture video score, and the specific formula is as follows:

[0042] ;

[0043] ;

[0044] ;

[0045] ;

[0046] ;

[0047] where, represents the gated recurrent unit; represents the minimum selection operation; represents the Gaussian weight corresponding to the frame; represents the hyperparameter for adjusting the direct and indirect influence effects; represents the indirect influence element; represents the direct influence element; represents the texture video frame quality score after weighted summation for quality regression; represents the texture video score obtained by globally averaging the texture video frame quality scores; represents the time; represents the time length of the direct influence; represents the time length of the indirect influence; represents the time of the frame with the worst quality in the time length of the direct influence; Indicates the length of the texture video frame.

[0048] Furthermore, the weighted aggregated texture video score and depth video score are used to obtain the final immersive video score, and the calculation formula is as follows:

[0049] ;

[0050] ;

[0051] ;

[0052] Among them, Indicates the multi-layer perceptron in the quality regression module; Indicates the depth video block feature of the i-th viewpoint at the same moment; Indicates the depth feature after multi-viewpoint pooling at time k; Indicates the hyperparameter that adjusts the influence effect of the texture and depth video scores; Indicates the final immersive video quality score; Indicates the depth video score.

[0053] On the other hand, an immersive video quality evaluation device based on view spatio-temporal correlation includes:

[0054] An immersive video extraction module for obtaining an immersive video containing texture videos and depth videos of multiple viewpoints, extracting texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extracting texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0055] A model construction and training module for constructing and training an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0056] The immersive video score evaluation module is used to input texture video blocks, depth video blocks, texture key frames, and depth key frames into a trained immersive video quality evaluation model. Specifically, the texture depth feature extraction module extracts features from the texture key frames and depth key frames to obtain the first texture features and depth features of multiple viewpoints. The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture depth fusion features. The multi-viewpoint spatial fusion features are processed by the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. The texture video feature extraction module extracts features from the texture video blocks to obtain second texture features. The second texture features are input into the channel attention module to obtain third texture features. The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features. The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The depth video block feature extraction module extracts depth video block features to obtain depth video block features. The depth video block features are input into the depth video quality regression module to obtain the depth video score. The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

[0057] The present invention adopts the above technical solutions and has the following beneficial effects:

[0058] (1) In the present invention, the multi-viewpoint spatial interaction network obtains disparity features through the disparity estimation module to enhance the texture key frame features, and the multi-viewpoint selective feature integration module interacts and integrates the multi-viewpoint features at the same moment to enhance the features with higher viewpoint spatial correlation.

[0059] (2) The present invention realizes a three-dimensional attention mechanism through the multi-viewpoint temporal non-local attention network to focus on the temporal correlation of multi-viewpoint features, so as to improve the temporal perception ability of different viewpoint features.

[0060] (3) The present invention proposes to perceive the quality of immersive videos through three parts: the quality of texture video blocks, the quality of depth video blocks, and the interaction quality of texture-depth key frames. Through the designed weighted aggregation strategy, the overall quality of immersive videos is accurately perceived and evaluated, and the evaluation results are highly consistent with the characteristics of human visual perception. Description of the Drawings

[0061] Figure 1 It is a flowchart of the immersive video quality evaluation method based on viewpoint spatio-temporal correlation according to an embodiment of the present invention;

[0062] Figure 2 It is a schematic structural diagram of the immersive video quality evaluation model according to an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the structure of the multi-viewpoint spatial interaction network according to an embodiment of the present invention;

[0064] Figure 4 Schematic diagram of the structure of the multi-viewpoint temporal non-local network according to an embodiment of the present invention;

[0065] Figure 5 Diagram of the immersive video quality evaluation device based on viewpoint spatio-temporal correlation according to an embodiment of the present invention. Detailed implementation manners

[0066] The present invention will be further described in detail below with reference to embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto. As Figure 1 shown, the immersive video quality evaluation method based on viewpoint spatio-temporal correlation of the present invention includes:

[0067] S1. Obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks.

[0068] Specifically, extracting texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extracting texture key frames and depth key frames based on the texture video blocks and depth video blocks are as follows:

[0069] For the texture video of each viewpoint and the depth video perform chunking to obtain K consecutive texture video blocks and K consecutive depth video blocks;

[0070] wherein, ; represents rounding down; N represents the number of frames of the texture video or depth video of each viewpoint; and respectively represent the i-th texture video frame and the i-th depth video frame; the expression of the j-th texture video block is ; the expression of the j-th depth video block is ; represents the number of video frames included in the texture video block or depth video block;

[0071] Extract the first frame of the texture video block and depth video block as the key frame of the corresponding block to obtain the texture key frame and the depth key frame .

[0072] S2. Construct an immersive video quality evaluation model and train it to obtain a trained immersive video quality evaluation model. The immersive video quality evaluation model includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part. The spatio-temporal interaction part of texture-depth features includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network. The texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module. The depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module.

[0073] Specifically, as Figure 2 shown, the immersive video quality evaluation model based on viewpoint spatio-temporal correlation proposed in the embodiment of the present application includes a spatio-temporal interaction part of texture-depth features, a texture video quality evaluation part, and a depth video quality evaluation part. The spatio-temporal interaction part of texture-depth features includes a feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network. The texture video quality evaluation part includes a texture video block feature extraction module, a channel attention module, a gated recurrent unit, and a texture video frame quality regression module. The depth video quality evaluation part includes a depth video block feature extraction module and a quality regression module.

[0074] S3. Input the texture video blocks, depth video blocks, texture key frames, and depth key frames into the trained immersive video quality evaluation model. Specifically, the texture key frames and depth key frames are feature-extracted by the texture-depth feature extraction module to obtain the first texture features and depth features of multiple viewpoints. The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture-depth fusion features. The multi-viewpoint spatial fusion features are processed by the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. The texture video blocks are feature-extracted by the texture video feature extraction module to obtain second texture features. The second texture features are input into the channel attention module to obtain third texture features. The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features. The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain texture video scores. The depth video block features are extracted by the depth video block feature extraction module to obtain depth video block features. The depth video block features are input into the depth video quality regression module to obtain depth video scores. The texture video scores and depth video scores are weighted and aggregated to obtain the final immersive video score.

[0075] Specifically, the first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture-depth fusion features. The calculation formula is as follows:

[0076] ;

[0077] ;

[0078] ;

[0079] Among them, , represents the th view depth feature; represents the first texture feature of the th view, ; H, W, and C represent the height, width, and number of channels of the feature; represents feature concatenation; represents processing the input feature through the disparity estimation module in the multi-viewpoint spatial interaction network; represents the input feature of the disparity estimation module; represents the disparity feature; represents the th view texture-depth fusion feature; represents element-wise multiplication of features.

[0080] Specifically, the multi-viewpoint spatial fusion feature is obtained based on the multi-viewpoint texture-depth fusion feature, and the calculation formula is as follows:

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] ;

[0086] ;

[0087] Among them, represents the multi-viewpoint synthesis feature after feature concatenation; represents feature concatenation; represents spatial global average pooling; represents the feature after spatial global pooling; represents the fully connected layer; represents the dimensionality-reduced feature; represents the channel attention vector of different view texture features; represents the second channel attention vector; represents the multi-viewpoint spatial fusion feature; represents the convolution operation; e represents the exponential calculation; Represents the texture depth fusion features of the 1st to Nth viewpoints.

[0088] Specifically, the multi-view spatial fusion features are processed by the multi-view temporal non-local network to obtain the integrated multi-view spatiotemporal interaction features. The calculation formula is as follows:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] in, Indicates adjusting the size of the feature dimension; Represents the multi-view spatial fusion features at time k; represents preliminary synthetic features; Represents a three-dimensional convolution operation; Indicates adjustment of feature shape; , and Represent query features, key features, and value features respectively; represents matrix multiplication, Represents the spatiotemporal fusion features of multiple viewpoints; represents the pooling strategies along different channels; Represents the multi-view spatiotemporal interaction features at time k.

[0096] Specifically, the third texture feature and the multi-view spatiotemporal interaction feature are spliced ​​to obtain the texture fusion feature, and the calculation formula is as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] in, represents n channel attention modules; represents the second texture feature; represents the third texture feature; Represents the multi-view average pooling strategy; Represents the multi-viewpoint pooling feature; Represents the texture fusion feature at time k.

[0101] Specifically, the texture fusion feature is sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The specific formula is as follows:

[0102] ;

[0103] ;

[0104] ;

[0105] ;

[0106] ;

[0107] Among them, Represents the gated recurrent unit; Represents the minimum selection operation; Represents the Gaussian weight corresponding to the Represents the hyperparameter for adjusting the direct and indirect influence effects; Represents the indirect influence element; Represents the direct influence element; Represents the texture video frame quality score after weighted summation for quality regression; Represents the texture video score obtained by performing global average pooling on all texture video frame quality scores; Represents the time; Represents the time length of the direct influence; Represents the time length of the indirect influence; Represents the time of the frame with the worst quality in the time length of the direct influence; Represents the texture video frame length.

[0108] Specifically, as Figure 3 shown, the multi-viewpoint spatial interaction network includes a disparity estimation module and a viewpoint selective feature integration module. The texture key frame feature and the depth key frame feature are input into the multi-viewpoint spatial interaction network. The i-th depth key frame feature is concatenated with the depth key frame features of its adjacent viewpoints to obtain the input feature of the disparity module, which is input into the disparity module to obtain the disparity feature of the i-th viewpoint. The disparity feature of each viewpoint is multiplied by the corresponding viewpoint's texture key frame feature to obtain the enhanced texture feature (texture-depth fusion feature).

[0109] Specifically, as Figure 4As shown, it represents inputting the multi-viewpoint space fusion features at different times into the multi-viewpoint temporal non-local network, obtaining the preliminary synthesized features through resizing and feature splicing, obtaining the corresponding query, key, and value features through different 3D convolution operations, then adjusting the feature shape to adapt to matrix multiplication, outputting the multi-viewpoint spatio-temporal fusion features, and adjusting them into the multi-viewpoint spatio-temporal interaction features at different times through the pooling strategy.

[0110] Specifically, the weighted aggregation of the texture video score and the depth video score is performed to obtain the final immersive video score, and the calculation formula is as follows:

[0111] ;

[0112] ;

[0113] ;

[0114] Among them, represents the multi-layer perceptron in the quality regression module; represents the depth video block feature of the i-th viewpoint at the same time; represents the depth feature after multi-viewpoint pooling at time k; represents the hyperparameter for adjusting the influence effect of the texture and depth video scores; represents the final immersive video quality score; represents the depth video score.

[0115] Specifically, in this embodiment, the feature extraction of several texture video blocks of multiple viewpoints is performed through the texture video quality evaluation part to obtain the texture features of multiple viewpoints , inputting the texture features into the channel attention module to obtain more representative texture features, and obtaining the multi-viewpoint pooling features through resizing and average pooling, and splicing them with the multi-viewpoint spatio-temporal interaction features to obtain the texture fusion features; obtaining the initial quality score at each time through the gated recurrent unit , inputting it into the texture video quality regression module. Specifically, for the feature at time k, by applying the minimum pooling operation to the previous frames, the element representing the direct effect of the current frame is obtained , that is, the quality of the th frame with the worst quality. The period of the direct effect is set to the past frame sequence. If , the frame sequence is , otherwise it is ; weighted processing is performed on several frames after the th frame with the worst quality to obtain the indirect influence element , which is directly related to the influencing elements The weighted sum is used to obtain the quality score of the texture video frame after quality regression Global average pooling is performed on all texture video frame quality scores to obtain the texture video score .

[0116] Specifically, the immersive video quality evaluation model based on view spatio-temporal correlation proposed in this embodiment is built using the Python 3.8.18 programming language, CUDA 12.2, and torch 2.3.1, and experiments are conducted using the NVIDIA RTX A6000 GPU. The experimental dataset is IMVD, 80% of which is used for training, and the test set and validation set are each 10%. This model is trained on input videos of size 448×448, with a batch size of 8 and 20 iterations; the initial learning rate is set to 0.00001, and the Adam optimizer is used for training.

[0117] As Figure 5 shown, this embodiment also discloses an immersive video quality evaluation device based on view spatio-temporal correlation, including:

[0118] The immersive video extraction module 51 is used to obtain an immersive video containing texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks;

[0119] The model construction and training module 52 is used to construct and train an immersive video quality evaluation model to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture-depth feature spatio-temporal interaction part, a texture video quality evaluation part, and a depth video quality evaluation part; the texture-depth feature spatio-temporal interaction part includes a texture-depth feature extraction module, a multi-viewpoint spatial interaction network, and a multi-viewpoint temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit, and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module;

[0120] The immersive video score evaluation module 53 is used to input texture video blocks, depth video blocks, texture key frames, and depth key frames into a trained immersive video quality evaluation model. Specifically, the texture depth feature extraction module extracts features from the texture key frames and depth key frames to obtain the first texture features and depth features of multiple viewpoints. The first texture features and depth features of multiple viewpoints are input into the multi-viewpoint spatial interaction network to obtain multi-viewpoint texture depth fusion features, and multi-viewpoint spatial fusion features are obtained based on the multi-viewpoint texture depth fusion features. The multi-viewpoint spatial fusion features are processed by the multi-viewpoint temporal non-local network to obtain integrated multi-viewpoint spatio-temporal interaction features. The texture video feature extraction module extracts features from the texture video blocks to obtain second texture features. The second texture features are input into the channel attention module to obtain third texture features. The third texture features and the multi-viewpoint spatio-temporal interaction features are concatenated to obtain texture fusion features. The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The depth video block feature extraction module extracts depth video block features to obtain depth video block features. The depth video block features are input into the depth video quality regression module to obtain the depth video score. The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

[0121] The specific implementation of the immersive video quality evaluation device based on viewpoint spatio-temporal correlation is the same as that of the immersive video quality evaluation method based on viewpoint spatio-temporal correlation, and will not be repeated in this embodiment.

[0122] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims, and all of them fall within the protection scope of the present invention.

Claims

1. An immersive video quality assessment method based on viewpoint spatiotemporal correlation, characterized in that: include: Acquire an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks from the texture videos and depth videos of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks; An immersive video quality evaluation model is constructed and trained to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture depth feature spatiotemporal interaction part, a texture video quality evaluation part and a depth video quality evaluation part; the texture depth feature spatiotemporal interaction part includes a texture depth feature extraction module, a multi-view spatial interaction network and a multi-view temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module; Input the texture video block, the depth video block, the texture key frame and the depth key frame into the trained immersive video quality evaluation model; specifically, extract the texture key frame and the depth key frame through the texture depth feature extraction module to obtain the first texture features and depth features of multiple viewpoints; Input the first texture features and depth features of multiple viewpoints into the multi-view spatial interaction network to obtain the multi-view texture depth fusion features, and obtain the multi-view spatial fusion features based on the multi-view texture depth fusion features; process the multi-view spatial fusion features through the multi-view temporal non-local network to obtain the integrated multi-view spatiotemporal interaction features; extract the features of the texture video blocks through the texture video feature extraction module to obtain the second texture features; input the second texture features into the channel attention module to obtain the third texture features; The third texture feature and the multi-view spatiotemporal interaction feature are spliced ​​to obtain a texture fusion feature; the texture fusion feature is sequentially input into the gated recurrent unit and the texture video quality regression module to obtain a texture video score; the deep video block feature is extracted through the deep video block feature extraction module to obtain a deep video block feature; the deep video block feature is input into the deep video quality regression module to obtain a deep video score; The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

2. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 1, characterized in that: Extract the texture video block and the depth video block in the texture video and the depth video of each viewpoint, and extract the texture key frame and the depth key frame based on the texture video block and the depth video block, as follows: Texture video for each viewpoint and deep video Divide the video into blocks to obtain K consecutive texture video blocks and K consecutive depth video blocks; in, ; Indicates the round-down symbol; N indicates the number of frames of texture video or depth video for each viewpoint; and Represent the i-th texture video frame and the i-th depth video frame respectively; the expression of the j-th texture video block is ; The expression of the j-th depth video block is ; Indicates the number of video frames contained in the texture video block or depth video block; Extract the first frame of the texture video block and the depth video block as the key frame of the corresponding block to obtain the texture key frame and depth keyframes .

3. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 1, characterized in that: The first texture features and depth features of multiple viewpoints are input into the multi-view spatial interaction network to obtain the multi-view texture depth fusion features. The calculation formula is as follows: ; ; ; in, , indicating the Viewpoint depth features; Indicates The first texture feature of the viewpoint, ; H, W and C represent the height, width and number of channels of the feature; Represents feature splicing; Indicates that the input features are processed by the disparity estimation module in the multi-view spatial interaction network; Represents the input features of the disparity estimation module; Represents disparity features; Indicates Texture depth fusion features from each viewpoint; Represents element-wise multiplication of features.

4. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 3 is characterized in that: The multi-view spatial fusion feature is obtained based on the multi-view texture depth fusion feature. The calculation formula is as follows: ; ; ; ; ; ; in, Represents the multi-view synthetic features after feature splicing; Represents feature splicing; Represents spatial global average pooling; Represents the features after spatial global pooling; represents a fully connected layer; Represents dimensionality reduction features; Channel attention vectors representing texture features from different viewpoints; represents the second channel attention vector; Represents multi-view spatial fusion features; represents the convolution operation; e represents the exponential calculation; Represents the texture depth fusion features of the 1st to Nth viewpoints.

5. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 4 is characterized in that: The multi-view spatial fusion features are processed by the multi-view temporal non-local network to obtain the integrated multi-view spatiotemporal interaction features. The calculation formula is as follows: ; ; ; ; ; ; in, Indicates adjusting the size of the feature dimension; Represents the multi-view spatial fusion features at time k; represents preliminary synthetic features; Represents a three-dimensional convolution operation; Indicates adjustment of feature shape; , and Represent query features, key features, and value features respectively; represents matrix multiplication, Represents the spatiotemporal fusion features of multiple viewpoints; represents the pooling strategies along different channels; Represents the multi-view spatiotemporal interaction features at time k.

6. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 5, characterized in that: The third texture feature and the multi-view spatiotemporal interaction feature are spliced ​​to obtain the texture fusion feature. The calculation formula is as follows: ; ; ; in, represents n channel attention modules; represents the second texture feature; represents the third texture feature; Represents the multi-view average pooling strategy; Represents multi-view pooling features; Represents the texture fusion feature at time k.

7. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 6, characterized in that: The texture fusion features are sequentially input into the gated recurrent unit and the texture video quality regression module to obtain the texture video score. The specific formula is as follows: ; ; ; ; ; in, represents a gated recurrent unit; Indicates the minimum value selection operation; Indicates Gaussian weight corresponding to the frame; represents the hyperparameters that adjust the direct and indirect effects; Indicates indirect influence elements; Indicates directly affecting elements; represents the quality score of the texture video frame after quality regression obtained by weighted summation; Indicates that the texture video score is obtained by global average pooling of all texture video frame quality scores; Indicates the moment; Indicates the length of time of direct impact; Indicates the length of time of the indirect impact; Indicates the moment of the worst quality frame in the directly affected time span; Indicates the texture video frame length.

8. The immersive video quality assessment method based on viewpoint spatiotemporal correlation according to claim 7, characterized in that: The weighted aggregate texture video score and depth video score are used to obtain the final immersive video score, which is calculated as follows: ; ; ; in, Represents the multi-layer perceptron in the quality regression module; Represents the deep video block features of the i-th viewpoint at the same time; Represents the depth feature after multi-view pooling at time k; Represents the hyperparameters that adjust the effect of texture and depth video scores; represents the final immersive video quality score; Represents the deep video score.

9. An immersive video quality assessment device based on viewpoint spatiotemporal correlation, characterized in that: include: An immersive video extraction module, used to obtain an immersive video including texture videos and depth videos of multiple viewpoints, extract texture video blocks and depth video blocks in the texture video and depth video of each viewpoint, and extract texture key frames and depth key frames based on the texture video blocks and depth video blocks; A model construction and training module is used to construct an immersive video quality evaluation model and train it to obtain a trained immersive video quality evaluation model; the immersive video quality evaluation model includes a texture depth feature spatiotemporal interaction part, a texture video quality evaluation part and a depth video quality evaluation part; the texture depth feature spatiotemporal interaction part includes a texture depth feature extraction module, a multi-view spatial interaction network and a multi-view temporal non-local network; the texture video quality evaluation part includes a texture video feature extraction module, a channel attention module, a gated recurrent unit and a texture video quality regression module, and the depth video quality evaluation part includes a depth video block feature extraction module and a depth video quality regression module; The immersive video sub-evaluation module is used to input the texture video block, the depth video block, the texture key frame and the depth key frame into the trained immersive video quality evaluation model; specifically, the texture key frame and the depth key frame are subjected to feature extraction by the texture depth feature extraction module to obtain the first texture features and depth features of multiple viewpoints; Input the first texture features and depth features of multiple viewpoints into the multi-view spatial interaction network to obtain the multi-view texture depth fusion features, and obtain the multi-view spatial fusion features based on the multi-view texture depth fusion features; process the multi-view spatial fusion features through the multi-view temporal non-local network to obtain the integrated multi-view spatiotemporal interaction features; extract the features of the texture video blocks through the texture video feature extraction module to obtain the second texture features; input the second texture features into the channel attention module to obtain the third texture features; The third texture feature and the multi-view spatiotemporal interaction feature are spliced ​​to obtain a texture fusion feature; the texture fusion feature is sequentially input into the gated recurrent unit and the texture video quality regression module to obtain a texture video score; the deep video block feature is extracted through the deep video block feature extraction module to obtain a deep video block feature; the deep video block feature is input into the deep video quality regression module to obtain a deep video score; The texture video score and the depth video score are weighted and aggregated to obtain the final immersive video score.

Citation Information

Patent Citations

  • Immersive video quality evaluation method and device based on multi-feature fusion

    CN118411583A

  • Immersive video quality evaluation method and device based on multi-feature network

    CN118506168A

  • Immersive video quality evaluation method and device based on frame-level time aggregation strategy

    CN118609034A

  • Method for generating a quality oriented significance map for assessing the quality of an image or video

    US20060233442A1

  • Saliency-weighted video quality assessment

    US20170154415A1

Cited By

  • Immersive video quality evaluation method and system based on multi-modal perception

    CN122134726A

  • Immersive video quality evaluation method and system based on multi-modal perception

    CN122134726B