A method and system for evaluating virtual reality image quality

By using Vision Transformer based on attention mechanism and multi-scale auxiliary network in VR panoramic image quality evaluation, combined with the spatial position correlation fusion module between viewports, the problem of unused correlation between viewports in the prior art is solved, and the accuracy and objectivity of quality evaluation are improved.

CN115546162BActive Publication Date: 2025-06-20ANQING NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211257173.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2025-06-20
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

The existing VR panoramic image quality evaluation method is difficult to effectively utilize the spatial position correlation between viewport images, resulting in inaccurate quality evaluation results.

Method used

Vision Transformer (ViT) based on attention mechanism is adopted as the backbone network, combining multi-scale auxiliary networks to extract and fuse multi-scale features of viewport images, and through the spatial position correlation fusion module, the model's perception of the relationship between viewports is enhanced.

Benefits of technology

It improves the accuracy and objectivity of VR panoramic image quality evaluation, more in line with the human eye's visual quality perception characteristics, and enhances the model's understanding of multi-scale visual perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546162B_ABST
    Figure CN115546162B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for evaluating the quality of virtual reality images, including: constructing a backbone network; obtaining two-dimensional ERP format images, converting them into multiple viewport images under a spherical structure, and obtaining the prediction scores of the backbone network; constructing a multi-scale auxiliary network, obtaining multi-scale features based on the multi-scale auxiliary network, and fusing them to obtain multi-scale fusion features; splicing the multi-scale fusion features and performing perceptual quality regression to obtain the prediction scores of the auxiliary network; splicing the prediction scores of the backbone network and the auxiliary network to obtain the prediction scores of the panoramic image perceptual quality; calculating the loss between the prediction scores of the panoramic image perceptual quality and the subjective quality scores of the panoramic image, and training and optimizing the overall network to obtain an optimal model, and further evaluating the quality of virtual reality images. The present application takes into account the multi-scale perception characteristics of the human eye when viewing VR, and further improves the accuracy of VR panoramic image quality prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a virtual reality image quality evaluation method and system. Background Art

[0002] Virtual reality (VR), as an emerging immersive technology, uses devices such as head-mounted devices as carriers, and can present 360° panoramic videos and image content, providing consumers with a brand-new visual information interaction method. With the increasing popularity of VR technology, it has begun to be widely used in various fields such as medical, military, and entertainment. However, due to the recording of full-view content, VR images often require very high resolutions, and coupled with their spherical representation form, it brings many difficulties to the processing (acquisition, encoding and decoding, transmission, etc.) of VR images. In order to meet the high-quality experience of users, studying the quality evaluation of VR image content is of great significance and practical value for guiding the optimization of existing algorithms and enhancing the user experience of VR.

[0003] Image quality evaluation can be divided into full-reference quality evaluation, semi-reference quality evaluation, and no-reference quality evaluation. Traditional methods for evaluating image quality based on pixels, such as PSNR and SSIM, have a low consistency with human image quality perception. Moreover, with the development of the Internet and social media, the objects of quality evaluation are often distorted images, and it is difficult to obtain the original undistorted images, especially for panoramic images. Therefore, no-reference quality evaluation based on deep learning has more extensive research value.

[0004] Most of the no-reference panoramic image quality evaluation methods based on deep learning use the ERP format after panoramic image compression as the network input. Although great improvements have been achieved compared with traditional full-reference quality evaluation methods, due to the characteristics of panoramic images, the ERP format will introduce geometric distortion, which still has a great impact on the final results. Many recent algorithms use the viewport images extracted from the ERP as the input to simulate the scene when people view VR panoramic images, thus making significant improvements in the quality evaluation results. However, this type of algorithm often only feeds the viewport images into the network in parallel, ignoring the feature correlation in the spatial position between the viewport images, because the perceived quality of each viewport is only a part of the overall perceived quality, and the relative position in space between the viewports will greatly affect people's overall perceived quality of the panoramic image. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide a brand-new VR panoramic image quality assessment method and system based on viewport input, which uses the Vision Transformer (ViT) based on the attention mechanism as the main network to replace the conventional convolutional neural network, and is more in line with the characteristics of human eye visual quality perception. This application also provides a brand-new method for fusing the spatial position correlation of viewports to improve the objective quality assessment level. On this basis, this application also simulates the multi-scale visual perception characteristics of people when watching VR, so as to further improve the model's quality perception level.

[0006] To achieve the above object, this application provides the following solutions:

[0007] A virtual reality image quality assessment method, comprising:

[0008] S1. Construct a backbone network;

[0009] S2. Obtain a two-dimensional ERP format image, and convert the two-dimensional ERP format image into multiple viewport images under a spherical structure;

[0010] S3. Based on the backbone network and the viewport images, obtain the prediction scores of the backbone network;

[0011] S4. Construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain multi-scale fusion features;

[0012] S5. Concatenate the multi-scale features and the multi-scale fusion features to obtain a third concatenated feature, and perform perceptual quality regression on the third concatenated feature to obtain the prediction scores of the auxiliary network;

[0013] S6. Concatenate the prediction scores of the backbone network and the prediction scores of the auxiliary network to obtain the panoramic image perceptual quality prediction scores;

[0014] S7. Calculate the loss between the panoramic image perceptual quality prediction scores and the panoramic image subjective quality scores, and based on the loss, train and optimize the overall network to obtain an optimal model, and based on the optimal model, perform quality assessment on virtual reality images.

[0015] Preferably, the backbone network in S1 includes a Vision Transformer based on the attention mechanism.

[0016] Preferably, the multiple viewport images in S2 include: viewport images corresponding to six directions of up, down, front, back, left, and right.

[0017] Preferably, the method for obtaining the prediction scores of the backbone network in S3 includes:

[0018] Input the viewport image into the backbone network to obtain the high-dimensional features of the viewport image;

[0019] Concatenate the high-dimensional image features to obtain the first concatenated feature, and perform feature fusion on the first concatenated feature based on the spatial position to obtain the fused viewport spatial position correlation feature;

[0020] Concatenate the high-dimensional features and the viewport spatial position correlation feature to obtain the second concatenated feature, and perform perceptual quality regression on the second concatenated feature to obtain the prediction score of the backbone network.

[0021] Preferably, the method for obtaining the fused viewport spatial position correlation feature includes:

[0022] Concatenate the high-dimensional features to obtain the first concatenated feature;

[0023] Add one-dimensional position encoding to the first concatenated feature to obtain the viewport spatial concatenated feature;

[0024] Based on the viewport spatial concatenated feature, obtain the viewport spatial fusion feature;

[0025] Adjust the dimension of the viewport spatial fusion feature, and based on the feature fusion extraction module, obtain the fused viewport spatial position correlation feature.

[0026] Preferably, the method for constructing the multi-scale auxiliary network in S4 includes:

[0027] Resize the front azimuth viewport image in the viewport image to multiple new scales as the multi-scale input;

[0028] Use Resnet50 as the basic framework, and based on the multi-scale input, set the pooling kernel size in the highest pooling layer to obtain the multi-scale auxiliary network.

[0029] Preferably, the method for calculating the loss in S7 includes: using the MAE loss function to calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score.

[0030] This application also provides a virtual reality image quality evaluation system, including: a first network construction module, a format conversion module, a first score prediction module, a second network construction module, a second score prediction module, a third score prediction module, and a model optimization module;

[0031] The first network construction module is used to construct a backbone network based on Vision Transformer with an attention mechanism;

[0032] The format conversion module is used to convert the two-dimensional ERP format image into multiple viewport images under the spherical structure;

[0033] The multiple viewport images include viewport images corresponding to six orientations: up, down, front, back, left, and right;

[0034] The first score prediction module is used to obtain the backbone network prediction score based on the backbone network and the viewport images;

[0035] The second network construction module is used to construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain multi-scale fusion features;

[0036] The second score prediction module is used to splice the multi-scale features and the multi-scale fusion features to obtain a third spliced feature, and perform perceptual quality regression on the third spliced feature to obtain the prediction score of the auxiliary network;

[0037] The third score prediction module is used to splice the prediction score of the backbone network and the prediction score of the auxiliary network to obtain the panoramic image perceptual quality prediction score;

[0038] The model optimization module is used to calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score based on the MAE loss function, and train and optimize the overall network based on the loss to obtain an optimal model, and perform quality assessment on the virtual reality image based on the optimal model.

[0039] Preferably, the method for the second network construction module to construct the multi-scale auxiliary network includes:

[0040] Resize the front-orientation viewport image in the viewport images to multiple new scales as multi-scale inputs;

[0041] Taking Resnet50 as the basic framework, based on the multi-scale inputs, set the pooling kernel size in the highest pooling layer to obtain the multi-scale auxiliary network.

[0042] Preferably, the process for the first score prediction module to obtain the backbone network prediction score includes:

[0043] Input the viewport images into the backbone network to obtain the high-dimensional features of the viewport images;

[0044] Splice the high-dimensional image features to obtain a first spliced feature, and perform feature fusion on the first spliced feature based on the spatial position to obtain the fused viewport spatial position correlation feature;

[0045] Concatenate the high-dimensional features and the viewport spatial position association features to obtain a second concatenated feature, and perform perceptual quality regression on the second concatenated feature to obtain the prediction score of the backbone network.

[0046] The beneficial effects of this application are as follows:

[0047] This application discloses a virtual reality image quality evaluation method and system. Using the viewport image as the network input, it is more in line with the visual perception effect when the human eye views VR. On this basis, the overall network architecture is built with ViT as the basic network. In the backbone network, in addition to extracting features based on attention within each viewport, the spatial position features between each viewport are innovatively further fused to effectively extract the features after viewport fusion. In addition, this invention also considers the multi-scale perception characteristics when the human eye views VR. Based on the original viewport scale, multiple scales are adjusted as auxiliary network inputs, and a multi-scale feature extraction and fusion network is established for the multi-scale inputs, thereby further improving the prediction accuracy of VR panoramic image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions of this application, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 Schematic diagram of the process of the virtual reality image quality evaluation method of this application;

[0050] Figure 2 Schematic diagram of the specific process of the virtual reality image quality evaluation method of this application:

[0051] Figure 3 Schematic diagram of the backbone network module of this application;

[0052] Figure 4 Schematic diagram of the auxiliary network module of this application;

[0053] Figure 5 Schematic diagram of the feature extraction sub-module after spatial position correlation fusion of this application;

[0054] Figure 6 Schematic diagram of the multi-scale fusion sub-module of this application;

[0055] Figure 7 Schematic diagram of the structure of the virtual reality image quality evaluation system of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0057] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0058] Embodiment 1

[0059] As Figure 1 、 Figure 2 shown, a schematic flowchart of a virtual reality image quality evaluation method of the present application includes:

[0060] S1. Construct a backbone network;

[0061] In this embodiment, the backbone network uses a hybrid network of Vision Transformer (ViT) and Resnet50 based on the attention mechanism as the main network to replace the conventional convolutional neural network.

[0062] In this embodiment, the backbone network is mainly divided into two stages: Embedded and Encoder, which will be described in detail in combination with step S3. The specific backbone network module is as Figure 3 shown.

[0063] S2. Obtain a two-dimensional ERP format image and convert the two-dimensional ERP format image into multiple viewport images in a spherical structure;

[0064] Considering that more viewpoints of people will fall near the equator and very few viewpoints will fall at the poles when viewing panoramic images, the viewport extraction scheme used in this embodiment is as follows: Taking the viewer as the center of the sphere, extract the front, back, left, right, up, and down six viewports corresponding to the current position as a group of viewport images. Considering the randomness of the starting position during the viewing process, in this example, taking the viewer as the center of the circle and the equator as the circle, with every 2° as a starting point, a group of viewport images is extracted. Thus, the number of viewport images corresponding to a panoramic image in this example is: (360 / 2)×6 = 1080.

[0065] Perform Resize operation on the viewport image: uniformly adjust the dimensions of the viewport image to 3×224×224, and then perform the Normalization operation. In this embodiment, set mean = [0.5, 0.5, 0.5] and std = [0.5, 0.5, 0.5]. After completing the normalization process, reasonably divide the viewport image into a training set and a test set. In this embodiment, the training set and the test set are divided according to 7:3.

[0066] S3. Based on the backbone network and the viewport image, obtain the backbone network prediction score;

[0067] The method for obtaining the backbone network prediction score includes:

[0068] S31. Input the viewport image into the backbone network to obtain the high-dimensional features of the viewport image;

[0069] Embedded stage: To effectively utilize the position encoding information of the tokens in ViT during the pre-training stage, before sending the features into ViT, the data dimensions need to be unified to b×196×768 (b represents Batch_size). In this example, the original data input dimension is b×3×224×224. After extracting the viewport image features through Resnet50, the feature dimension is b×1024×14×14. Flatten the third and fourth dimensions. After flattening, the feature dimension is b×1024×196. Swap the second and third dimensions. After swapping, the feature dimension is b×196×1024. Then pass the features through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are set to 1024 and 768 respectively. The output feature dimension is b×196×768. Then add a token with the same dimension, that is, a token with the dimension of b×1×768, for downstream tasks, which we call Class_token. The output feature dimension is b×197×768. Then through the sum operation, add learnable position encoding to the output features to ensure the integrity of the viewport image content information. The output encoded feature dimension is b×197×768.

[0070] Encoder Stage: Further feature fusion based on the attention mechanism is performed on the encoded features. The encoded features corresponding to each viewport image are linearly mapped to three vectors Q, K, and V of the same dimension through a fully connected layer three times. Here, the number of input and output channels of the fully connected layer is set to 768 and 768 respectively. Using the Q vector as the query vector and the K vector as the matching vector, the "similarity" between every two tokens is calculated by dot product. We call this "similarity" the attention weight: Attention_weight. Then, the attention weight is applied to the K vector by dot product to complete the extraction of the attention features between the tokens inside the viewport. The formula is expressed as:

[0071]

[0072] In the formula, T represents the transpose operation, d k represents the dimension of the K vector, and softmax represents the data normalization operation.

[0073] It should be noted that when calculating the attention weight, considering the diversity of features contained in the image, in this example, the Q, K, and V vectors corresponding to each token are split into multiple ones. Taking the Q vector as an example, after splitting, the corresponding vector groups are Q1, Q2,..., Q n . The process of performing attention feature extraction on the Q, K, and V vectors with corresponding subscripts respectively, and then merging the obtained multiple groups of attention features to obtain the viewport image features based on attention. We call this method: multi-head attention mechanism. The number of multi-head attentions set in this example is 12. The viewport image features based on attention obtained are passed through two fully connected layers. The number of input and output channels of the two fully connected layers is 768, 3072 and 3072, 768 respectively. The obtained output feature dimension is b×197×768. The above process is repeated multiple times to better fuse the features inside the viewport. The number of attention module layers set in this example is 12. Finally, the Class_token corresponding to each viewport image is extracted. The dimension of the Class_token is b×768. Through a fully connected layer, where the number of input and output channels of the fully connected layer is 768, 10, the finally obtained high-dimensional feature is b×10.

[0074] S32. Concatenate the high-dimensional image features to obtain the first concatenated feature, and perform feature fusion on the first concatenated feature based on the spatial position to obtain the fused viewport spatial position correlation feature;

[0075] Among them, the method for obtaining the fused viewport spatial position correlation feature includes:

[0076] S321. Concatenate the high-dimensional features to obtain the first concatenated feature;

[0077] Concatenate the Class_tokens corresponding to each of the above viewports to obtain a first concatenated feature with a dimension of b×6×768;

[0078] S322. Add a one-dimensional position encoding to the first concatenated feature to obtain a viewport space concatenated feature;

[0079] Add a one-dimensional position encoding through the Sum operation and output the viewport space concatenated feature with a dimension of b×6×768.

[0080] S323. Based on the viewport space concatenated feature, obtain a viewport space fusion feature;

[0081] Pass the viewport space concatenated feature through a multi-layer attention module, and extract the fusion feature between each viewport according to the attention mechanism to obtain the viewport space fusion feature. Specifically:

[0082] Pass the viewport space concatenated feature through a 3-layer attention module. Among them, the attention module implements the calculation process of Attention_weight in step S31, and outputs the viewport space fusion feature with a dimension of b×6×768.

[0083] S324. Adjust the dimension of the viewport space fusion feature, and based on the feature fusion extraction module, obtain the fused viewport space position association feature.

[0084] Adjust the dimension of the viewport space fusion feature to b×6×16×48, and then pass it through a feature extraction sub-module, specifically as Figure 5 shown.

[0085] Among them, the feature extraction sub-module includes four Blocks and a pooling layer. Each Block includes a two-dimensional convolutional layer, a BatchNorm layer, and a ReLu layer. The number of input and output channels of the convolutional layer in the first Block are 6 and 12 respectively, the convolutional kernel size is 7×7, the stride is 2×2, and the padding value is 3×3. The convolutional kernel size of the remaining Blocks is 3×3, the stride is 2×2, and the padding value is 1×1. The number of input and output channels of the convolutional layer in the second Block are 12 and 24 respectively, the number of input and output channels of the convolutional layer in the third Block are 24 and 48 respectively, and the number of input and output channels of the convolutional layer in the fourth Block are 48 and 64 respectively. The convolutional kernel size of the pooling layer is 1×3, and the stride is 1×1.

[0086] The output viewport fusion feature after passing through the feature extraction sub-module has a dimension of b×64×1×1. Adjust its feature dimension to b×64, and then pass it through a fully connected layer. Among them, the number of input and output channels of the fully connected layer is 64 and 10 respectively. Finally, obtain the viewport space position association feature with a dimension of b×10.

[0087] S33. Concatenate the high-dimensional features and the viewport spatial position correlation features to obtain the second concatenated feature, and perform perceptual quality regression on the second concatenated feature to obtain the prediction score of the backbone network.

[0088] Concatenate the high-dimensional features and the viewport spatial position correlation features, with the dimension being b×70. Pass through a fully connected layer, where the number of input and output channels of the fully connected layer is 70 and 1 respectively. Perform perceptual quality regression once to obtain the prediction score score1 of the backbone network.

[0089] S4. Construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain the multi-scale fused feature.

[0090] S41. The method for constructing the auxiliary network includes:

[0091] (1) Resize the front azimuth viewport image in the viewport image to multiple new scales as the multi-scale input.

[0092] Perform two resize operations on the front azimuth viewport image, resizing it to 448×448 and 112×112 respectively. Take the adjusted image scales together with the original image scale as the multi-scale input of the auxiliary network.

[0093] (2) Use Resnet50 as the basic framework, and based on the multi-scale input, set the pooling kernel size at the highest pooling layer to obtain the multi-scale auxiliary network.

[0094] The auxiliary network is built based on Resnet50. For different input scales, set different Average_pooling sizes. In this example, set the pooling layers with convolution kernel sizes of 14×14, 7×7, and 4×4 for 448×448, 224×224, and 112×112 respectively. We call this improved network Re_resnet50, that is, the multi-scale auxiliary network. The specific auxiliary network module is as Figure 4 shown.

[0095] The multi-scale input passes through the Re_Resnet50 network in parallel to obtain the same output dimension b×1024, and then passes through a fully connected layer, where the number of input and output channels of the fully connected layer is 1024 and 10 respectively, to obtain the multi-scale features, and the dimension of the multi-scale features is b×10.

[0096] S42. Further fuse the multi-scale features to obtain the multi-scale fused feature:

[0097] Concatenate the multi-scale features, with the dimension of b×3×10, then adjust the dimension to b×1×3×10, and then pass through the multi-scale feature fusion sub-module, such as Figure 6 As shown, the multi-scale feature fusion sub-module consists of a two-dimensional dilated convolutional layer with a convolution kernel of 3×3 and a ReLu layer, which is used to fuse the multi-scale feature dimensions. In this example, the number of multi-scales is 3, and the output multi-scale fusion feature dimension is b×1×3×10. Adjust the multi-scale fusion feature dimension to b×30, and pass through a fully connected layer. The input and output channel numbers of the fully connected layer are 30 and 10 respectively, and finally output the multi-scale fusion feature, with its dimension being b×10.

[0098] S5. Concatenate the multi-scale features and the multi-scale fusion features to obtain the third concatenated feature, and perform perceptual quality regression on the third concatenated feature to obtain the prediction score of the auxiliary network;

[0099] Concatenate the multi-scale features and the multi-scale fusion features, with the dimension of b×40, and pass through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are 40 and 1 respectively, and perform perceptual quality regression once to obtain the prediction score score2 of the auxiliary network;

[0100] S6. Concatenate the prediction score of the backbone network and the prediction score of the auxiliary network to obtain the panoramic image perceptual quality prediction score;

[0101] Concatenate the above prediction scores score1 and score2, with the dimension of b×2, and then pass through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are 2 and 1 respectively, to obtain the final panoramic image perceptual quality prediction score score.

[0102] S7. Calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score, and based on the loss, train and optimize the overall network to obtain the optimal model. Based on the optimal model, perform quality assessment on the virtual reality image.

[0103] Adopt the MAE loss function to calculate the loss between the panoramic image perceptual quality prediction score score and the subjective quality score score of the corresponding panoramic image. Train and optimize the network according to the loss function Loss=(score - score ground_truth ) ground_truth ), so that the loss gradually decreases. After training, finally obtain a VR panoramic image objective quality assessment model with better robust performance. 2 In this embodiment, within 5 training times, the optimal model will be obtained and saved in the.pkl format for quality assessment of virtual reality images.

[0104] In this embodiment, within 5 training times, the optimal model will be obtained and saved in the.pkl format for quality assessment of virtual reality images.

[0105] The virtual reality panoramic image quality prediction method proposed in this application fully considers the characteristics of human visual perception and simulates the scenario of human eyes viewing VR in a real scene. Using the viewport image as the network input, during the model construction process, ViT is used as the backbone network. On this basis, an effective viewport spatial position fusion model is established to effectively fuse the information of viewports at six different spatial positions; the auxiliary network simulates the multi-scale perception characteristics of human eyes when viewing VR, and a multi-scale quality perception network is built to effectively fuse multi-scale features. Finally, by integrating the two-branch network, the predicted quality score of the VR panoramic image is obtained. The present invention fully considers the influence of the viewport spatial position characteristics on the human eye visual quality perception, thereby improving the performance of panoramic image quality prediction.

[0106] Embodiment 2

[0107] As Figure 7 shown, this application also provides a virtual reality image quality evaluation system, including: a first network construction module, a format conversion module, a first score prediction module, a second network construction module, a second score prediction module, a third score prediction module, and a model optimization module;

[0108] The first network construction module is used to construct the backbone network. In this embodiment, the backbone network uses a hybrid network of Vision Transformer (ViT) based on the attention mechanism and Resnet50 as the main network to replace the conventional convolutional neural network.

[0109] In this embodiment, the backbone network is mainly divided into two stages: Embedded and Encoder. The specific working process will be described in detail in combination with other modules;

[0110] The format conversion module is used to convert the two-dimensional ERP format image into multiple viewport images in a spherical structure;

[0111] The specific working process includes:

[0112] Considering that more viewpoints of people will fall near the equator and very few viewpoints will fall at the poles when viewing panoramic images, the viewport extraction scheme used in this embodiment is: taking the viewer as the center of the sphere, extracting the front, back, left, right, up, and down six viewports corresponding to the current position as a group of viewport images. Considering the randomness of the starting position during the viewing process, in this example, taking the viewer as the center of the circle and the equator as the circle, with every 2° as a starting point, a group of viewport images is extracted. Thus, the number of viewport images corresponding to a panoramic image in this example is: (360 / 2)×6 = 1080.

[0113] Perform Resize operation on the viewport image: uniformly adjust the dimensions of the viewport image to 3×224×224, and then perform Normalization operation. In this embodiment, set mean = [0.5, 0.5, 0.5] and std = [0.5, 0.5, 0.5]. After completing the normalization process, reasonably divide the viewport image into a training set and a test set. In this embodiment, the training set and the test set are divided according to 7:3.

[0114] The first score prediction module is used to obtain the backbone network prediction score based on the backbone network and the viewport image;

[0115] The specific working process includes:

[0116] (1) Input the viewport image into the backbone network to obtain the high-dimensional features of the viewport image;

[0117] Embedded stage: To effectively utilize the positional encoding information of tokens in ViT during the pre-training stage, before sending the features into ViT, the data dimensions need to be unified to b×196×768 (b represents Batch_size). In this example, the original data input dimension is b×3×224×224. After extracting the viewport image features through Resnet50, the feature dimension is b×1024×14×14. Flatten the third and fourth dimensions. After flattening, the feature dimension is b×1024×196. Swap the second and third dimensions. After swapping, the feature dimension is b×196×1024. Then pass the features through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are set to 1024 and 768 respectively. The output feature dimension is b×196×768. Then add a token with the same dimension, that is, a token with a dimension of b×1×768, for downstream tasks, which we call Class_token. The output feature dimension is b×197×768. Then through the sum operation, add learnable positional encoding to the output features to ensure the integrity of the viewport image content information. The output encoded feature dimension is b×197×768.

[0118] Encoder stage: Further feature fusion based on the attention mechanism is performed on the encoded features. The encoded features corresponding to each viewport image are linearly mapped to three vectors Q, K, and V of the same dimension through a fully connected layer three times. Among them, the number of input and output channels of the fully connected layer is set to 768 and 768 respectively. Using the Q vector as the query vector and the K vector as the matching vector, the "similarity" between every two tokens is calculated by dot product. We call this "similarity" the attention weight: Attention_weight. Then, the attention weight is applied to the K vector by dot product to complete the extraction of the attention features between the tokens inside the viewport. The formula is expressed as:

[0119]

[0120] In the formula, T represents the transpose operation, d k represents the dimension of the K vector, and softmax represents the data normalization operation.

[0121] It should be noted that when calculating the attention weight, considering the diversity of features contained in the image, in this example, the Q, K, and V vectors corresponding to each token are sliced into multiple. Taking the Q vector as an example, after slicing, the corresponding vector groups are Q1, Q2,..., Q n . The process of performing attention feature extraction on the Q, K, and V vectors with corresponding subscripts respectively, and then merging the obtained multiple groups of attention features to obtain the viewport image features based on attention. We call this method: multi-head attention mechanism. The number of multi-head attentions set in this example is 12. The viewport image features based on attention obtained are passed through two fully connected layers. The number of input and output channels of the two fully connected layers is 768, 3072 and 3072, 768 respectively. The obtained output feature dimension is b×197×768. The above process is repeated multiple times to better fuse the features inside the viewport. The number of attention module layers set in this example is 12. Finally, the Class_token corresponding to each viewport image is extracted. The dimension of Class_token is b×768. Through a fully connected layer, where the number of input and output channels of the fully connected layer is 768, 10, the final high-dimensional feature is b×10.

[0122] (2) Concatenate the high-dimensional image features to obtain the first concatenated feature, and perform feature fusion on the first concatenated feature based on the spatial position to obtain the fused viewport spatial position correlation feature;

[0123] (21) Concatenate the high-dimensional features to obtain the first concatenated feature;

[0124] Concatenate the Class_tokens corresponding to each of the above viewports to obtain the first concatenated feature, with a dimension of b×6×768;

[0125] (22) Add a one-dimensional positional encoding to the first concatenated feature to obtain the concatenated feature in the viewport space;

[0126] Add a one-dimensional positional encoding through the Sum operation and output the concatenated feature in the viewport space, with a dimension of b×6×768.

[0127] (23) Based on the concatenated feature in the viewport space, obtain the fused feature in the viewport space;

[0128] Pass the concatenated feature in the viewport space through multiple attention modules, and extract the fused feature between each viewport according to the attention mechanism to obtain the fused feature in the viewport space. Specifically:

[0129] Pass the concatenated feature in the viewport space through 3 attention modules. Among them, the attention module implements the calculation process of Attention_weight in step S31, and outputs the fused feature in the viewport space with a dimension of b×6×768.

[0130] (24) Adjust the dimension of the fused feature in the viewport space, and based on the feature fusion extraction module, obtain the associated feature of the viewport space position after fusion.

[0131] Adjust the dimension of the fused feature in the viewport space to b×6×16×48, and then pass it through a feature extraction sub-module.

[0132] Among them, the feature extraction sub-module includes four Blocks and a pooling layer. Each Block includes a two-dimensional convolutional layer, a BatchNorm layer, and a ReLu layer. The input and output channel numbers of the convolutional layer in the first Block are 6 and 12 respectively, the convolutional kernel size is 7×7, the stride is 2×2, and the padding value is 3×3. The convolutional kernel size of the remaining Blocks is 3×3, the stride is 2×2, and the padding value is 1×1; the input and output channel numbers of the convolutional layer in the second Block are 12 and 24 respectively, the input and output channel numbers of the convolutional layer in the third Block are 24 and 48 respectively, and the input and output channel numbers of the convolutional layer in the fourth Block are 48 and 64 respectively. The convolutional kernel size of the pooling layer is 1×3, and the stride is 1×1.

[0133] The output viewport fusion feature after passing through the feature extraction sub-module has a dimension of b×64×1×1. Adjust its feature dimension to b×64, and then pass it through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are 64 and 10 respectively, and finally obtain the associated feature of the viewport space position, with a dimension of b×10.

[0134] (3) Concatenate the high-dimensional features and the viewport spatial position correlation features to obtain the second concatenated feature, and perform perceptual quality regression on the second concatenated feature to obtain the prediction score of the backbone network.

[0135] Concatenate the high-dimensional features and the viewport spatial position correlation features. The dimension is b×70. Pass through a fully connected layer, where the number of input and output channels of the fully connected layer is 70 and 1 respectively. Perform perceptual quality regression once to obtain the backbone network prediction score score1.

[0136] The second network construction module is used to construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain multi-scale fusion features.

[0137] The specific working process includes:

[0138] (1) Resize the front azimuth viewport image in the viewport image to multiple new scales as multi-scale inputs.

[0139] Perform two resize operations on the front azimuth viewport image, resizing it to 448×448 and 112×112 respectively. Together with the original image scale, the adjusted image scales are used as the multi-scale inputs of the auxiliary network.

[0140] (2) Use Resnet50 as the basic framework. Based on the multi-scale inputs, set the pooling kernel size in the highest pooling layer to obtain the multi-scale auxiliary network.

[0141] The auxiliary network is built based on Resnet50. For different input scales, set different Average_pooling sizes. In this example, for 448×448, 224×224, and 112×112, set pooling layers with convolution kernel sizes of 14×14, 7×7, and 4×4 respectively. We call this improved network Re_resnet50, that is, the multi-scale auxiliary network.

[0142] The multi-scale inputs pass through the Re_Resnet50 network in parallel to obtain the same output dimension b×1024, and then pass through a fully connected layer, where the number of input and output channels of the fully connected layer is 1024 and 10 respectively, to obtain multi-scale features with a dimension of b×10.

[0143] (3) Further fuse the multi-scale features to obtain multi-scale fusion features:

[0144] The multi-scale features are concatenated with a dimension of b×3×10, then the dimension is adjusted to b×1×3×10, and then passed through a multi-scale feature fusion module. The multi-scale feature fusion module consists of a two-dimensional dilated convolutional layer with a convolution kernel of 3×3 and a ReLu layer, which is used to fuse the multi-scale feature dimensions. In this example, the number of multi-scales is 3, and the output multi-scale fusion feature dimension is b×1×3×10. The multi-scale fusion feature is adjusted to a dimension of b×30, and passed through a fully connected layer. The input and output channel numbers of the fully connected layer are 30 and 10 respectively, and finally the multi-scale fusion feature is output with a dimension of b×10.

[0145] The second score prediction module is used to concatenate the multi-scale features and the multi-scale fusion features to obtain the third concatenated feature, and perform perceptual quality regression on the third concatenated feature to obtain the prediction score of the auxiliary network.

[0146] The multi-scale features and the multi-scale fusion features are concatenated with a dimension of b×40, and passed through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are 40 and 1 respectively, and a perceptual quality regression is performed once to obtain the prediction score score2 of the auxiliary network.

[0147] The third score prediction module is used to concatenate the prediction score of the backbone network and the prediction score of the auxiliary network to obtain the panoramic image perceptual quality prediction score.

[0148] The specific working process includes:

[0149] The above prediction scores score1 and score2 are concatenated with a dimension of b×2, and then passed through a fully connected layer. Among them, the input and output channel numbers of the fully connected layer are 2 and 1 respectively, to obtain the final panoramic image perceptual quality prediction score score.

[0150] The model optimization module is used to calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score based on the MAE loss function, and train and optimize the overall network based on the loss to obtain the optimal model. Based on the optimal model, the quality of the virtual reality image is evaluated.

[0151] The specific working process includes:

[0152] The MAE loss function is used to calculate the loss between the panoramic image perceptual quality prediction score score and the subjective quality score score of the corresponding panoramic image. The network is trained and optimized according to the loss function Loss=(score - score ground_truth ) ground_truth 2 , so that the loss is gradually reduced. After training, an objective quality evaluation model with better robustness for VR panoramic images is finally obtained, which is used to evaluate the quality of virtual reality images. ​

[0153] The embodiments described above are only descriptions of the preferred embodiments of the present application, and do not limit the scope of the present application. Without departing from the design spirit of the present application, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present application shall fall within the protection scope determined by the claims of the present application.

Claims

1. A virtual reality image quality evaluation method, characterized in that, Including: S1. Construct a backbone network; S2. Obtain a two-dimensional ERP format image, and convert the two-dimensional ERP format image into multiple viewport images under a spherical structure; S3. Based on the backbone network and the viewport images, obtain a backbone network prediction score; The method for obtaining the backbone network prediction score in S3 includes: Input the viewport images into the backbone network to obtain high-dimensional features of the viewport images; Perform splicing on the high-dimensional image features to obtain a first spliced feature, and perform feature fusion on the first spliced feature based on the spatial position to obtain a fused viewport spatial position correlation feature; Splice the high-dimensional features and the viewport spatial position correlation features to obtain a second spliced feature, and perform perceptual quality regression on the second spliced feature to obtain the prediction score of the backbone network; S4. Construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain multi-scale fusion features; The construction method of the multi-scale auxiliary network in S4 includes: Resize the front azimuth viewport image in the viewport images to multiple new scales as multi-scale inputs; Taking Resnet50 as the basic framework, based on the multi-scale inputs, set the pooling kernel size in the highest pooling layer to obtain the multi-scale auxiliary network; S5. Splice the multi-scale features and the multi-scale fusion features to obtain a third spliced feature, and perform perceptual quality regression on the third spliced feature to obtain the prediction score of the auxiliary network; S6. Splice the prediction score of the backbone network and the prediction score of the auxiliary network to obtain a panoramic image perceptual quality prediction score; S7. Calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score, and based on the loss, train and optimize the overall network to obtain an optimal model, and based on the optimal model, evaluate the quality of virtual reality images.

2. The virtual reality image quality evaluation method according to claim 1, characterized in that, The backbone network in S1 includes a Vision Transformer based on an attention mechanism.

3. The virtual reality image quality evaluation method according to claim 1, characterized in that, The multiple viewport images in S2 include: viewport images corresponding to six azimuths of up, down, front, back, left, and right.

4. The virtual reality image quality evaluation method according to claim 1, characterized in that, The method for obtaining the fused viewport spatial position correlation feature includes: Perform splicing on the high-dimensional features to obtain the first spliced feature; Add one-dimensional position encoding to the first spliced feature to obtain a viewport spatial splicing feature; Based on the viewport spatial splicing feature, obtain a viewport spatial fusion feature; Adjust the dimension of the viewport spatial fusion feature, and based on a feature fusion extraction module, obtain the fused viewport spatial position correlation feature.

5. The virtual reality image quality evaluation method according to claim 1, characterized in that, The calculation method of the loss in S7 includes: using an MAE loss function to calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score.

6. A virtual reality image quality evaluation system, characterized in that, Including: A first network construction module, a format conversion module, a first score prediction module, a second network construction module, a second score prediction module, a third score prediction module, and a model optimization module; The first network construction module is used to construct a backbone network based on the Vision Transformer with an attention mechanism; The format conversion module is used to convert two-dimensional ERP format images into multiple viewport images in a spherical structure; The multiple viewport images include viewport images corresponding to six directions: up, down, front, back, left, and right; The first score prediction module is used to obtain a backbone network prediction score based on the backbone network and the viewport images; The process by which the first score prediction module obtains the backbone network prediction score includes: Inputting the viewport images into the backbone network to obtain high-dimensional features of the viewport images; Concatenating the high-dimensional image features to obtain a first concatenated feature, and performing feature fusion on the first concatenated feature based on spatial positions to obtain a fused viewport spatial position correlation feature; Concatenating the high-dimensional features and the viewport spatial position correlation feature to obtain a second concatenated feature, and performing perceptual quality regression on the second concatenated feature to obtain the prediction score of the backbone network; The second network construction module is used to construct a multi-scale auxiliary network, obtain multi-scale features based on the multi-scale auxiliary network, and fuse the multi-scale features to obtain a multi-scale fusion feature; The method by which the second network construction module constructs the multi-scale auxiliary network includes: Resizing the front-direction viewport image in the viewport images to multiple new scales as multi-scale inputs; Using Resnet50 as the basic framework, setting the pooling kernel size in the highest pooling layer based on the multi-scale inputs to obtain the multi-scale auxiliary network; The second score prediction module is used to concatenate the multi-scale features and the multi-scale fusion feature to obtain a third concatenated feature, and perform perceptual quality regression on the third concatenated feature to obtain the prediction score of the auxiliary network; The third score prediction module is used to concatenate the prediction score of the backbone network and the prediction score of the auxiliary network to obtain a panoramic image perceptual quality prediction score; The model optimization module is used to calculate the loss between the panoramic image perceptual quality prediction score and the panoramic image subjective quality score based on the MAE loss function, and train and optimize the overall network based on the loss to obtain an optimal model, and perform quality assessment on virtual reality images based on the optimal model.

Citation Information

Patent Citations

  • Blind image quality evaluation method based on natural scene statistics and perceived quality propagation

    CN104282019A

  • No-reference image quality evaluation method based on spatial attention mechanism

    CN114066812A