A panoramic image quality evaluation method
By employing a panoramic image quality assessment method based on CMP projection, multi-layer feature extraction from viewport images and saliency maps is used to generate perceptual quality scores. This solves the problems of geometric distortion and dizziness in panoramic image quality assessment, improving the accuracy of image quality prediction and user experience.
Patent Information
- Application Number
- CN202211348772.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-10-31
AI Technical Summary
Existing 360° panoramic image quality assessment methods are difficult to effectively handle geometric distortion when a sphere is mapped to a plane and dizziness caused by the user's rotating viewpoint. Furthermore, traditional methods are difficult to accurately assess image quality under no-reference conditions.
A panoramic image quality assessment method based on CMP projection is adopted. By acquiring the viewport image and saliency map of the panoramic image, multi-dimensional feature vectors are extracted and combined with a weighted generation network to generate a perceptual quality score, simulating the visual perception effect of the human eye.
It improves the accuracy of panoramic image quality prediction, better matches human visual perception, and enhances the objectivity of user experience quality evaluation.
Smart Images

Figure CN115601539B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of image processing, and particularly relates to a panoramic image quality evaluation method. BACKGROUND
[0002] 360° panoramic pictures are obtained by capturing image information of a whole scene by a professional camera, splicing pictures in each direction by using software, and enabling a user to experience 360° panoramic landscape by using a professional head-mounted device. The 360° panoramic picture technology is an aspect of virtual reality technology, which simulates a two-dimensional plane picture into a real three-dimensional space. The 360° panoramic picture technology is different from text, pictures and video, and is a new information transmission medium. A user can observe a picture scene by rotating a visual angle at will, and feels real and immersive. Distortion is generated in the process of shooting, picture splicing, compression and transmission of the 360° panoramic picture, and the experience quality of a user is affected. Image quality evaluation (IQA) plays an important role in the process of user quality experience. The higher the evaluation score is, the better the picture quality is, and the stronger the user experience is.
[0003] Traditional plane image quality evaluation is divided into subjective quality evaluation and objective quality evaluation. The subject of the subjective quality evaluation is a user, the user scores the picture quality, and the result of the subjective quality evaluation directly reflects the experience quality of a person. The subject of the objective quality evaluation is a machine, the machine automatically scores the picture, the subjective quality evaluation is time-consuming and laborious, and it is important to develop the objective quality evaluation. The objective quality evaluation is divided into full-reference quality evaluation, semi-reference quality evaluation and no-reference quality evaluation. The main difference among the three is the degree of utilization of reference pictures. The no-reference quality evaluation is widely concerned because it does not need any information of reference pictures. Distortion is generated in the process of acquisition, compression and transmission of a picture, and a real picture is often difficult to obtain. It is important to develop the no-reference image quality evaluation.
[0004] The 360-degree panoramic image quality evaluation is different from the traditional plane image quality evaluation, and has the following difficulties: 1, the 360-degree panoramic image needs to be mapped from a spherical surface to a plane, and the content on the two poles of the spherical surface will be geometrically distorted after mapping. 2, the user needs to use a head-mounted device to watch the 360-degree panoramic image, and when a person observes the 360-degree panoramic image in a rotating visual angle, due to the limitation of the visual angle and the resolution of the device, dizziness and other uncomfortable feelings will be caused. At present, the 360-degree image quality evaluation mainly studies the influence of the mainstream compression format such as jpg on the picture quality. Deep learning uses a deep neural network to simulate human pattern recognition, and the end-to-end learning method often has stronger feature learning ability and expression ability than the traditional manual feature extraction method, and has achieved remarkable results in the field of computer vision such as content recognition, edge extraction and image retrieval. The projection mode of the 360-degree panoramic image has ERP, CMP, SSP and the like, and the application provides a new 360-degree panoramic image quality evaluation method based on the CMP projection mode by using deep learning. SUMMARY
[0005] The purpose of the present application is to provide a panoramic image quality evaluation method to solve the problems existing in the prior art.
[0006] To achieve the above purpose, the present application provides a panoramic image quality evaluation method, comprising:
[0007] obtaining a panoramic image, converting the panoramic image into a corresponding viewport image, and extracting a saliency map corresponding to the viewport image;
[0008] performing feature extraction on the viewport image to obtain a corresponding multi-layer dimensional feature vector and a high-dimensional feature; performing feature extraction on the saliency map to obtain a salient content area feature;
[0009] obtaining a second high-dimensional feature based on the high-dimensional feature and the salient content area feature, obtaining a second high-dimensional feature vector after average pooling of the second high-dimensional feature, and obtaining a viewport total feature based on the second high-dimensional feature;
[0010] splicing the second high-dimensional feature vector and the multi-layer dimensional feature vector to obtain a viewport total feature vector;
[0011] obtaining a weight value and a bias value based on the viewport total feature and a weight generation network;
[0012] obtaining a perceptual quality score based on the viewport total feature vector, the weight value and the bias value.
[0013] Optionally, the panoramic image is converted into six viewport images under a spherical structure, and a plurality of groups of viewport images are extracted based on a fixed angle value.
[0014] Optionally, the obtaining process of the total viewport feature vector comprises: the number of the second high-dimensional feature vectors is consistent with the number of the viewport images, the second high-dimensional feature vectors are spliced with the multi-layer dimensional feature vectors respectively to obtain third high-dimensional feature vectors, the third high-dimensional features are consistent with the number of the viewport images, and the third high-dimensional feature vectors of the six viewports are spliced to obtain the total viewport feature vector.
[0015] Optionally, the obtaining process of the multi-layer dimensional feature vector comprises: a multi-dimensional feature extraction network is constructed, multi-layer feature maps are obtained through a backbone network, the multi-layer feature maps are input into the multi-dimensional feature extraction network, and after two-dimensional convolution and average pooling dimension reduction, the multi-layer feature maps are laid into vectors, and then the multi-layer dimensional feature vector is obtained through a full connection layer; wherein the multi-dimensional feature extraction network comprises three structures, three multi-layer dimensional feature vectors are output, and the backbone network is Resnet34.
[0016] Optionally, the obtaining process of the weight value and the bias value comprises: the total viewport feature is input into four different weight generation networks to obtain four weight values and bias values respectively, the four weight generation networks all comprise a two-dimensional convolution layer, an average pooling layer and a full connection layer; the fifth weight value and the bias value are obtained by sequentially passing the total viewport feature through the average pooling layer and the full connection layer; wherein the two-dimensional convolution layer and the full connection layer of each weight generation network are different in design.
[0017] Optionally, five full connection layers are constructed based on the weight value and the bias value, and the total viewport feature vector is sequentially input into the five full connection layers to obtain the perceptual quality score, wherein the weight value and the bias value are used as the weight and the bias of the full connection layer respectively.
[0018] Optionally, the saliency map and the viewport image are also subjected to image dimension adjustment before feature extraction.
[0019] The 360° panoramic image quality evaluation method provided by the application takes six viewport images in the CMP format as network input, which is more consistent with the visual perception effect when the human eye watches. On this basis, the overall network architecture is built based on Resnet34 as the basic network. In the backbone network, in addition to using the high-dimensional information of the last layer of the network, the application also integrates the information of multiple dimensions of the network. The application also considers that the human eye only watches the salient area of the picture, and when a person processes complex picture content information, the person often observes the content of interest such as a person, and ignores other unimportant information. The salient content features of the saliency map and the high-dimensional features of the backbone network are fused to obtain weighted features. In addition, the application uses the high-dimensional feature map to generate corresponding Weight and Bias for image quality regression, thereby further improving the panoramic image quality prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application and are incorporated herein in
[0021] Figure 1 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0022] Figure 2 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0023] Figure 3 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0024] Figure 4 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0025] Figure 5 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0026] Figure 6 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image;
[0027] Figure 7 Fig. a is a 360° panoramic image projection principle, and Fig. b is a CMP projection of a 360° image; DETAILED DESCRIPTION
[0028] It should be noted that the embodiments and features in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and embodiments.
[0029] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0030] Embodiment One
[0031] As shown in the present embodiment, a panoramic image quality evaluation method based on virtual reality technology is provided, comprising: Figures 1-7
[0032] The technical problem to be solved in the embodiment is to provide a new 360-degree panoramic image quality evaluation method based on viewport input. The embodiment is based on Resnet34 as the main network because it has achieved remarkable results in the field of image recognition. The embodiment also provides a new spatial attention method based on a saliency map to simulate the fact that people only look at interesting content when viewing images, and uses the saliency map to select the most attractive areas in the image.
[0033] The 360-degree panoramic image quality evaluation method of the embodiment includes the following steps:
[0034] (1) Convert the two-dimensional ERP format image into the corresponding six CMP projection mode viewport images under the spherical structure, which are divided into V f (front viewport), V ba (back viewport), V l (left viewport), V r (right viewport), V t (top viewport), V bm (bottom viewport), and extract the saliency maps corresponding to the six viewports, which are divided into S f (front viewport saliency map), S ba (back viewport saliency map), S l (left viewport saliency map), S r (right viewport saliency map), S t (top viewport saliency map), and S bm (bottom viewport saliency map).
[0035] (2) Parallelly input the viewport images obtained in step (1) into the parallel backbone network (Resnet34), and extract the corresponding high-dimensional features of each viewport image, which are f1, f2, f3, f4, f5, and f6, respectively;
[0036] (3) Input each saliency map obtained in step (1) into the branch network 1-space saliency content extraction network to obtain the saliency content area features of the viewport images, which are S1, S2, S3, S4, S5, and S6, respectively;
[0037] (4) Input the different layer dimension feature maps of the branch network in step (2) into the branch network 2-multidimensional feature extraction network MS to obtain the multidimensional feature vectors of the viewport images, which are L1, L2, and L3, respectively;
[0038] (5) The high-dimensional features obtained in step (2) and the salient content region features obtained in step (3) are multiplied element by element to weight the salient content regions of the high-dimensional features, respectively f1xS1, f2xS2, f3xS3, f4xS4, f5xS5, f6xS6, to obtain new high-dimensional features f1, f2, f3, f4, f5, f6, and then divided into two steps: the first step, after a 7x7 average pooling layer, the feature vectors V1, V2, V3, V4, V5, V6 of the six viewports are obtained. V1 is spliced with L1, L2, L3 of step (4) to obtain a new V1, and V2, V3, V4, V5, V6 are spliced with L1, L2, L3 of step (4) to obtain new V2, V3, V4, V5, V6, respectively, and V1, V2, V3, V4, V5, V6 are spliced to obtain the total feature vector V of the six viewports; the second step, f1, f2, f3, f4, f5, f6 are spliced to obtain the total feature map f of the six viewports;
[0039] (6) The feature map f obtained in step (5) and the feature vector V are reduced in dimension, the feature map f enters the weight generation network to generate the corresponding Weight and Bias, and the feature vector V passes through five connection layers to obtain the final score score;
[0040] (7) According to the MSE loss function, the loss between the panoramic image perception quality prediction score score and the corresponding panoramic image subjective quality score scoregrond_truth is calculated, and the overall network is trained and optimized according to the loss, and finally the optimal model is obtained.
[0041] When a user observes a 360° panoramic image, the content of the sphere will be projected in the direction tangent to the user's view angle, converting the panoramic image from a two-dimensional ERP format image into a corresponding six viewport images under the structure of a sphere, which is more in line with the user's real experience. Taking the viewer as the center of the sphere, the V f , V ba , V l , V r , V t , V bm six viewports corresponding to the current position are extracted as a set of viewport images. Considering the randomness of the starting position during viewing, in the present example, the viewer is taken as the center of the circle, the equator is taken as the circle, and every 2° is taken as a starting point, and the set of viewport images is extracted. Thus, in the present example, the number of viewport images corresponding to one panoramic image is: (360 / 2) x 6 = 1080. The saliency maps corresponding to the six viewports are extracted and denoted as S f , S ba , S l , S r , S t , S bm .
[0042] Resize operation is performed on the viewport image and the saliency map: In this example, the viewport and saliency images are uniformly adjusted to have image dimensions of 3x224x224, and the number of channels is 3, the height and width are both 224. Normalization operation: In this example, the mean is set to mean = [0.485, 0.456, 0.406], and the standard deviation is set to std = [0.229, 0.224, 0.225], and then all the viewport images are divided into a training set and a test set according to a ratio of 7:3.
[0043] The set of viewport and saliency images, i.e., V f and S f , V ba and S ba , V l and S l , V r and S r , V t and S t , V bm and S bm are sent into the network in parallel. The network is specifically divided into: a backbone network, a spatial position fusion network with six viewports as input; a quality regression network that fuses the extracted features into scores; an auxiliary network, a spatial position saliency feature extraction network with a saliency map as input. As shown in Figure 5 The normalized saliency map is subjected to average pooling AvgPool2d(32, stride = 32 (step size is 32)) and two-dimensional convolution Conv2d(1, 512, kernel_size = 1 (convolution kernel size is 1x1), stride = 1 (step size is 1), padding = 0 (no padding)) to increase the dimension to obtain S1, S2, S3, S4, S5, S6 of size 512x7x7, and then Normalize(mean = 0 (mean value is 0), std = 1 (standard deviation is 1)) is performed on the data to standardize it;
[0044] In this example, the backbone network is Resnet34, and the backbone network is mainly used for feature extraction. Resnet network is first designed for image recognition, and when used for image quality evaluation, using the Resnet model trained on ImageNet as the initialization parameter can significantly improve the effect. The network structure of Resnet34 is as shown in Figure 6 The structure is divided into four parts, namely conv_2, conv_3, conv_4, and conv_5. The feature extracted by the neural network can be represented as f(x) = score, where x represents the input viewport, f represents the mapping mode, and is the mapping parameter. After the six viewports pass through the Resnet34 network, the corresponding features f1, f2, f3, f4, f5, and f6 are obtained.
[0045] ResNet networks are designed for image recognition, but the model uses global features for classification while ignoring local information. Human vision is highly sensitive to local distortions. In this invention, we fuse multi-dimensional features to better integrate local and global information. For convenience, we extract information from four dimensions: conv2_6, conv3_8, and conv4_12, and input them into the branch network 2-multi-layer feature extraction network. For example... Figure 4 As shown, the 64×56×56 conv2_6 is reduced to a 16×8×8 feature map by 2D convolution Conv2d(64,16,kernel_size=1, stride=1, padding=0) and average pooling AvgPool2d(7,stride=7)). This feature map is then flattened into a vector and finally passed through a fully connected layer Linear(16×64,16) to generate L1. The 128×28×28 conv3_8 is further reduced by 2D convolution Conv2d(128,32,kernel_size=1, stride=1, padding=0). The dimensionality of the conv4_12 with size 128×14×14 is reduced by 2D convolution Conv2d(128,64,kernel_size=1,stride=1,padding=0) and average pooling AvgPool2d(7,stride=7) to obtain a 32×4×4 feature map, which is flattened into a vector and finally passed through a fully connected layer Linear(32×16,16) to generate L2;
[0046] The f1 in (4) is average-pooled using AvgPool2d(7, stride=1) and flattened into a 512-dimensional vector. The vector is then concatenated with f1 and L1, L2, L3 in (5) to obtain a new V1. Similarly, V2, V3, V4, V5, and V6 are obtained. The vectors V1, V2, V3, V4, V5, and V6 are concatenated to obtain the total feature vector V. The vectors f1, f2, f3, f4, f5, and f6 in (4) are concatenated to obtain the total feature map f.
[0047] The features of the viewport map in the fusion step (6) are reduced in dimension by two-dimensional convolution Conv2d (3072, 112, kernel_size=1 (convolution kernel size is 1), padding=0 (no padding)) and fully connected layer Linear (1344, 224) on f and V in (6) respectively, to obtain features f of size 112x7x7 and vector V of size 224x1 respectively for easy calculation;
[0048] The feature map f of step (7) is input into five, the size of 112x7x7 feature f is subjected to two-dimensional convolution Conv2d(112, 512, kernel_size=3 (convolution kernel size is 3), stride=1 (step length is 1), padding=1 (no padding)), and Reshape into the size of 112x224x1x1 weight Weight_1, the feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 112) to generate the size of 112x1 bias Bais_1; The size of 112x7x7 feature f is subjected to two-dimensional convolution Conv2d(112, 128, kernel_size=3 (convolution kernel size is 3), stride=1 (step length is 1), padding=1 (no padding)), and Reshape into the size of 56x112x1x1 weight Weight_2, the feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 56) to generate the size of 56x1 bias Bais_2; The size of 112x7x7 feature f is subjected to two-dimensional convolution Conv2d(112, 32, kernel_size=3 (convolution kernel size is 3), stride=1 (step length is 1), padding=1 (no padding)), and Reshape into the size of 28x56x1x1 weight Weight_3, the feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 28) to generate the size of 28x1 bias Bais_3; The size of 112x7x7 feature f is subjected to two-dimensional convolution Conv2d(112, 8, kernel_size=3 (convolution kernel size is 3), stride=1 (step length is 1), padding=1 (no padding)), and Reshape into the size of 14x28x1x1 weight Weight_4, the feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 14) to generate the size of 14x1 bias Bais_4; The size of 112x7x7 feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 14) and Reshape into the size of 1x14x1x1 weight Weight_5, the feature f is subjected to average pooling AvgPool2d(7, stride=7 (step length is 7)) and full connection layer Linear(112, 1) to generate the size of 1 bias Bais_5.
[0049] The V in step (7) is operated by a full connection layer composed of five 1x1 convolution kernel convolution layer. The conv2d(input=V, weight=weight_1, bias=bias_1) obtains q1; q1 is operated by conv2d(input=q1, weight=weight_2, bias=bias_2) to obtain q2; q2 is operated by conv2d(input=q2, weight=weight_3, bias=bias_3) to obtain q3; q3 is operated by conv2d(input=q3, weight=weight_4, bias=bias_4) to obtain q4; q4 is operated by conv2d(input=q4, weight=weight_5, bias=bias_5) to obtain score. The input is an input tensor with a size of (minibatch (batch), in_channels (input channel), H (height), W (width)), the weight is a convolution kernel with a size of (out_channels (output channel), H, W), wherein the input is an input vector, the weight is equivalent to the weight of the full connection layer, and the bias is equivalent to the bias of the full connection layer. The convolution kernel size of the conv2d is 1x1.
[0050] The loss is calculated between the prediction score score and the subjective quality score score groud_truth of the corresponding panoramic image, and the network is trained according to the loss function Loss=(score-score ground_truth ) 2 , so that the loss gradually decreases, and after training, a 360° panoramic image objective quality evaluation algorithm with better robust performance is finally obtained.
[0051] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A panoramic image quality evaluation method characterized by comprising: The method comprises the following steps: obtaining a panoramic image, converting the panoramic image into a corresponding viewport image, and extracting a saliency map corresponding to the viewport image; extracting features of the viewport image to obtain corresponding multi-dimensional feature vectors and high-dimensional features; extracting features of the saliency map to obtain salient content region features; obtaining second high-dimensional features based on the high-dimensional features and the salient content region features, obtaining a second high-dimensional feature vector after average pooling of the second high-dimensional features, and obtaining a viewport total feature based on the second high-dimensional features; wherein the high-dimensional features and the salient content region features are multiplied element by element to obtain the second high-dimensional features; the second high-dimensional features are spliced to obtain the viewport total feature; splicing the second high-dimensional feature vector and the multi-dimensional feature vector to obtain a viewport total feature vector; obtaining weight values and bias values based on the viewport total feature through a weight generation network; obtaining a perceptual quality score based on the viewport total feature vector, the weight values and the bias values; The process of obtaining the weight values and bias values comprises: inputting the viewport total feature into five different weight generation networks to obtain five weight values and bias values respectively, and each of the five weight generation networks comprises a two-dimensional convolution layer, an average pooling layer and a fully connected layer; the fifth weight value and the bias value are obtained by sequentially passing the viewport total feature through the average pooling layer and the fully connected layer; wherein the two-dimensional convolution layer and the fully connected layer of each weight generation network are designed differently; constructing five fully connected layers composed of five convolution layers with a kernel size of 1x1 based on the weight values and the bias values, and obtaining the perceptual quality score after the viewport total feature vector sequentially passes through the five fully connected layers, wherein the weight values and the bias values are used as the weights and the bias of the fully connected layers respectively.
2. The panoramic image quality evaluation method according to claim 1, wherein the panoramic image is converted into six viewport images under a spherical structure, and a plurality of groups of viewport images are extracted based on a fixed angle value.
3. The panoramic image quality evaluation method according to claim 2, wherein the process of obtaining the viewport total feature vector comprises: splicing the second high-dimensional feature vectors with the multi-dimensional feature vectors respectively to obtain third high-dimensional feature vectors, the number of the second high-dimensional feature vectors being consistent with the number of the viewport images, the number of the third high-dimensional features being consistent with the number of the viewport images, and splicing the third high-dimensional feature vectors to obtain the viewport total feature vector.
4. The panoramic image quality evaluation method according to claim 1, wherein the process of obtaining the multi-dimensional feature vector comprises: constructing a multi-dimensional feature extraction network, obtaining multi-layer feature maps through a backbone network, inputting the multi-layer feature maps into the multi-dimensional feature extraction network, converting the multi-layer feature maps into vectors through two-dimensional convolution and average pooling dimension reduction, and processing the vectors through a fully connected layer to obtain the multi-dimensional feature vector; wherein the multi-dimensional feature extraction network comprises three structures, and outputs three multi-dimensional feature vectors; and the backbone network is Resnet34.
5. The panoramic image quality assessment method according to claim 1, characterized in that, The saliency map is also subjected to image dimension adjustment prior to feature extraction with the viewport image.
Citation Information
Patent Citations
Method of reasoning panoramic video quality according to field of view video quality
CN111696081A
Light field image angle reconstruction method
CN114066777A