A Blind Quality Assessment Method and System for Non-Uniformly Distorted Panoramic Images
Through the panoramic image quality evaluation method based on eye movement data and multi-layer multi-axis self-attention module, the accurate evaluation problem of non-uniform distorted panoramic images is solved, and higher prediction accuracy and feature sensitivity are achieved, which is suitable for the quality evaluation of non-uniform distorted panoramic images.
Patent Information
- Application Number
- CN202311013187.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-08-11
AI Technical Summary
The existing panoramic image quality evaluation method based on viewports is poor in the case of non-uniform distortion and cannot accurately reflect user visual perception. The traditional method ignores the deformation problem at the poles of the panoramic image.
Viewport sequence image preprocessing based on eye movement data, combined with multi-layer multi-axis self-attention module and deep semantic guidance, multi-scale quality perception features are extracted, and time information is captured through the gated loop unit to perform blind quality evaluation of non-uniform distorted panoramic images.
It improves the prediction accuracy of panoramic image quality of non-uniform distortion, enhances the sensitivity and abstract perception of feature representation, expands the application scope of the model, and can effectively deal with different types of non-uniform distortions.
Smart Images

Figure CN117237279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and multimedia digital image processing, and particularly relates to a method and system for blind quality evaluation of non-uniformly distorted panoramic images. Background Art
[0002] Panoramic images (OIs) are one of the important media for virtual reality. Through the characteristics of head-mounted displays (HMDs), immersive experiences can be provided for users. However, during processes such as acquisition, transmission, processing, and storage, OIs may suffer from distortion, so their quality is far from satisfactory, which significantly reduces the quality of the user experience. Therefore, accurately estimating the quality of OIs is very important for both system optimization and algorithm optimization. Generally, panoramic image quality assessment (OIQA) models can be divided into three categories, including full-reference OIQA (FR-IQA), reduced-reference OIQA (RR-OIQA), and no-reference / blind OIQA (NR- / BOIQA). The first two models require the use of complete and partial reference information during deployment, while the last model can evaluate image quality without reference information, so it is more practical than the first two models.
[0003] Some previous patch-based panoramic image quality assessment methods were achieved by dividing panoramic images in equirectangular projection (ERP) format into patches of equal size, and using a convolutional neural network (CNN) to extract quality features and regress them to scores. However, these methods ignore the distortion at the poles of panoramic images caused by projection stretching, resulting in inconsistent representations of patches and differences from users' visual perception. In contrast, viewport-based methods can solve the above problems because they provide a more reasonable representation that conforms to human perception. Recently, Vision Transformer has achieved great success in the field of computer vision and has been successfully applied to the OIQA task. Although these methods have shown good results on uniformly distorted OI quality databases such as OIQA and CVIQ, their performance on non-uniformly distorted OIs is not excellent enough. Summary of the Invention
[0004] In view of the above situation, the main purpose of the present invention is to propose a method and system for blind quality evaluation of non-uniformly distorted panoramic images to solve the above technical problems.
[0005] The present invention provides a method for blind quality evaluation of non-uniformly distorted panoramic images, and the method includes the following steps:
[0006] Step 1: Based on eye movement data, obtain a sequence of non-uniformly distorted viewport images from a non-uniformly distorted panoramic image, and perform image preprocessing on the sequence of viewport images;
[0007] Step 2: Based on the multi-layer and multi-axis self-attention module, extract the multi-level quality perception features of the non-uniformly distorted viewport sequence images by using the output of the previous layer as the input of the next layer;
[0008] Step 3: Aggregate the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to image quality at different perception scales, and obtain the multi-scale quality perception information
[0009] Step 4: Obtain the deep features of the multi-axis self-attention module, and in the way of deep semantic guidance, use the deep features to enhance the abstract perception degree of the multi-scale quality perception information, and obtain the final multi-scale quality perception information;
[0010] Judge whether the final multi-scale quality perception information contains time information. If so, use several gated recurrent units to capture the time features within each viewport sequence in the final multi-scale quality perception information, and model the time information of the viewport sequence to solve the time delay effect;
[0011] Step 5: Input the final multi-scale quality perception information into the fully connected layer to obtain the predicted evaluation score of the non-uniformly distorted panoramic image.
[0012] By extracting features based on the characteristics of multi-axis self-attention, aggregating multi-scale features, and enhancing the feature perception ability with deep semantics, the present invention can effectively predict the image quality of non-uniformly distorted panoramic images.
[0013] The present invention also provides a blind quality evaluation system for non-uniformly distorted panoramic images, and the system includes:
[0014] An extraction module, which obtains the non-uniformly distorted viewport sequence images from the non-uniformly distorted panoramic images based on eye movement data, and performs image preprocessing on the viewport sequence images;
[0015] And based on the multi-layer and multi-axis self-attention module, extract the multi-level quality perception features of the non-uniformly distorted viewport sequence images by using the output of the previous layer as the input of the next layer;
[0016] A multi-scale fusion module, which is used to aggregate the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to image quality at different perception scales, and obtain the multi-scale quality perception information;
[0017] A deep semantic guidance module, which is used to obtain the deep features of the multi-axis self-attention module, and in the way of deep semantic guidance, use the deep features to enhance the abstract perception degree of the multi-scale quality perception information, and obtain the final multi-scale quality perception information;
[0018] A quality regression module, which is used to input the final multi-scale quality perception information into a fully-connected layer to obtain the predicted evaluation score of the non-uniform distortion panoramic image.
[0019] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of a blind quality evaluation method for non-uniform distortion panoramic images proposed by the present invention;
[0021] Figure 2 It is a framework diagram of a blind quality evaluation system for non-uniform distortion panoramic images proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0023] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0024] Please refer to Figure 1 , the embodiments of the present invention provide a blind quality evaluation method for non-uniform distortion panoramic images, and the method includes the following steps:
[0025] Step 1: Based on eye movement data, obtain a sequence of viewport images with non-uniform distortion from the non-uniform distortion panoramic image, and perform image preprocessing on the sequence of viewport images;
[0026] Generally, the quality evaluation data of panoramic images belongs to uniformly distorted images. Compared with the images commonly used for quality evaluation, the storage format of panoramic images is different from that of 2D images. The viewport-based images conform to the human perception representation, and the sequence of viewport images using the viewing paths of subjects under different conditions is more in line with the actual viewing behavior.
[0027] The experimental images in the embodiments of the present invention are from the JUFE database, and a total of 258 original images taken by Insta360 pro2 are included, with a resolution of 8192×4096. The original images are reduced in quality to 3 distortion levels through four distortion types, which are respectively represented as BD, GB, GN, and ST, corresponding to the image distortion types of brightness discontinuity, Gaussian blur, Gaussian noise, and stitching distortion. The distortion levels are respectively represented as 1, 2, and 3, corresponding to low, medium, and high image distortion degrees.
[0028] Furthermore, the following relational expressions exist for the obtained non-uniform distortion viewport sequence images:
[0029]
[0030] Among them, D represents the non-uniform distortion viewport sequence image, c represents the starting condition, and the starting condition is four good starting points g5 including 5 seconds, bad starting point b5 of 5 seconds, good starting point g15 of 15 seconds, and bad starting point b15 of 15 seconds;
[0031] VS n represents the nth viewport sequence, and the following relational expressions exist for the viewport sequence:
[0032]
[0033] Among them, represents the mth viewport in the nth viewport sequence, and M represents the number of viewports in the sequence.
[0034] Therefore, according to the eye movement data, the non-uniform distortion viewport sequence images are obtained by adopting the relational expressions of the non-uniform distortion viewport sequence images.
[0035] Furthermore, for the convenience of input and training, the viewport sequence images are scaled to a unified size for easy input into the model;
[0036] The scaled viewport sequence images are normalized to make them easier to process and analyze
[0037] 80% of the normalized viewport sequence images are used as the training set, and 20% are used as the test set to prevent overfitting.
[0038] Step 2: Based on the multi-layer multi-axis self-attention module, the multi-level quality perception features of the non-uniform distortion viewport sequence images are extracted by using the output of the previous layer as the input of the next layer;
[0039] In the above solution, the non-uniform distortion viewport sequence images are initially extracted to obtain initial features. In this embodiment, the initial extraction uses the Stem module, and the Stem module includes two convolutional layers with a convolutional kernel of 3×3.
[0040] A multi - layer and multi - axis self - attention module is adopted, and the multi - layer and multi - axis self - attention modules are connected in series to form a network backbone;
[0041] Using the preliminary features as input features, and taking the output of the previous layer as the input of the next layer, the network backbone is used for feature extraction to obtain multi - level quality - perception features of different scales corresponding to the output of the multi - layer and multi - axis self - attention module.
[0042] In the above scheme, the extraction process of each layer of the multi - axis self - attention module is as follows:
[0043] Use the MBCony module to extract richer feature representations, then use the Block attention module to obtain local quality - perception information; finally, obtain global quality - perception information through the Grid attention module to get quality - perception features.
[0044] Then, using the quality - perception features as the input of the next - layer multi - axis self - attention module for cyclic extraction, and further obtaining multi - layer quality - perception features of different scales corresponding to different layers of the multi - axis self - attention module.
[0045] The relational expression of the extraction process of the multi - axis self - attention module is as follows:
[0046] F i =(GridAtten(BlockAtten(MBConv(x i ))));
[0047] Where, x i represents the input features of the multi - axis self - attention module of the current layer, F i represents the output features of the multi - axis self - attention module of the current layer, MBConv represents the mobile inverted residual bottleneck convolution operation, BlockAtten represents the block self - attention calculation operation, and GridAtten represents the grid self - attention calculation operation.
[0048] Among them, the specific operation steps of the MBConv operation are as follows: First, perform batch normalization on the features, then use pointwise convolution for channel up - dimensioning, perform depth convolution after up - dimensioning, then use the Squeeze - and - Excitation (SE) operation to enhance the representation of important channels, and finally use pointwise convolution again to restore the dimension and apply the residual connection. The relationship is as follows:
[0049] x i =x i +Conv(SE(DWConv(Conv(BN(x i )))));
[0050] Among them, Conv represents the point convolution operation (the convolution kernel is 1×1), DWConv represents the mobile inverted residual bottleneck convolution operation, and the SE operation includes a pooling layer and two fully connected layers.
[0051] Among them, the specific operation steps of the BlockAtten operation are as follows: The input feature map x i is windowed and divided into non-overlapping windows, and the operation is expressed as The size of each window is P×P. Self-attention calculation is performed on each window, and then residual connection is carried out. The relationship is as follows:
[0052] x i = x i + Unblock(RelAtten(Block(LN(x i )))) ;
[0053] The feature after residual connection is mapped through a fully connected layer operation and then another residual connection is carried out. The relationship is as follows:
[0054] x i = x i +(MLP(LN(x i ))) ;
[0055] Among them, LN represents layer normalization, the Block operation represents windowing, the RelAtten operation represents relative self-attention, the Unblock operation is the inverse process of the Block operation, and the fully connected layer mapping operation is through two linear layers.
[0056] Among them, the specific operation steps of the GridAtten operation are as follows: The output feature of the BlockAtten operation is grid-divided, and the operation is expressed as:
[0057]
[0058] Each grid has an adaptive size. Relative self-attention calculation is performed on each grid, and residual connection is carried out. The relationship is as follows:
[0059] x i = x i + Ungrid(RelAtten(Grid(LN(x i )))) ;
[0060] The feature after residual connection is then passed through a fully connected layer operation and residual connection is carried out. The relationship is as follows:
[0061] x i = x i+(MLP(LN(x i )));
[0062] Among them, the Grid operation represents grid segmentation, and the Ungrid operation represents the inverse process of the Grid operation.
[0063] Step 3: Aggregate the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to image quality at different perception scales, and obtain multi-scale quality perception information;
[0064] Perform a Generalized-Mean (GeM) pooling operation on the multi-level quality perception features to obtain dense information of the quality perception features at each scale. There is the following relationship for performing the GeM pooling operation on the multi-level quality perception features:
[0065] f i = GeM i (F i );
[0066] Among them, F i , i ∈ 1, 2, 3, 4 represents the input multi-level quality perception features at four different scales. GeM i represents performing the generalized mean pooling operation, and f i represents the multi-level quality perception features after pooling; In order to reduce the computational amount, four generalized mean pooling layers are adopted in this embodiment to perform compression processing on the features at four different scales respectively, and obtain four quality perception vectors representing features at different scales.
[0067] Concatenate the multi-level quality perception features at different scales and pass them through a Fully Connected (FC) layer, thereby generating a multi-scale quality perception vector v q , and further obtain multi-scale quality perception information. There is the following relationship for the multi-scale quality perception vector:
[0068] v q = FC(f1 ∪ f2 ∪ f3 ∪ f4);
[0069] Among them, v q represents the multi-scale quality perception vector, f1, f2, f3, and f4 respectively represent the multi-level quality perception features at four different scales after pooling, ∪ represents the concatenation operation, and FC represents the fully connected layer.
[0070] Step 4: Obtain the deep features of the multi-axis self-attention module, and adopt the method guided by deep semantics to enhance the abstract perception degree of the multi-scale quality perception information by using the deep features, and obtain the final multi-scale quality perception information;
[0071] Determine whether the final multi-scale quality perception information contains time information. If so, use a number of gated recurrent units to capture the time features within each viewport sequence in the final multi-scale quality perception information, and model the time information of the viewport sequence to address the latency effect;
[0072] Furthermore, there is the following relational expression for the final multi-scale quality perception information:
[0073]
[0074] Among them, represents the quality perception vector with both multi-scale information representation and deep semantic representation, that is, the final multi-scale quality perception information, v g represents the deep semantic guidance vector, and there is the following relational expression for the deep semantic guidance vector:
[0075] v g = GeM(F4);
[0076] Among them, F4 represents the deep features of the multi-axis self-attention module, that is, the multi-level quality perception features output by the last multi-axis self-attention module.
[0077] Use the generalized mean pooling layer to obtain the deep semantic guidance vector from the deep layer of the multi-axis self-attention layer, and without passing through the fully connected layer to retain the abstract perception information to the greatest extent. Concatenate the deep semantic guidance vector with the multi-scale quality perception information to obtain the final multi-scale quality perception information.
[0078] Step 5: Input the final multi-scale quality perception information into the fully connected layer to obtain the predicted evaluation score of the non-uniform distortion panoramic image, and its relational expression is as follows:
[0079]
[0080]
[0081] Among them, M represents the number of viewports in the viewport sequence, s m represents the predicted score of the m-th viewport, represents the predicted evaluation score.
[0082] Finally, to ensure the accuracy of the prediction, it needs to be trained. In the execution of the above steps 1 to 5, the corresponding training method includes the following training steps:
[0083] Obtain the average subjective score of all data in the non-uniform distortion panoramic image database;
[0084] Use the mean absolute error as the loss defining the average subjective score and the predicted evaluation score, and construct a loss function. The expression of the loss function is as follows:
[0085]
[0086] Among them, and respectively represent the model prediction quality score and the average subjective score quality score of the i-th panoramic image in the training data, and N represents the number of people participating in the experiment to evaluate the quality of non-uniformly distorted panoramic images;
[0087] Crop the non-uniformly distorted viewport sequence images and input them into the Adam optimizer for optimization. Set the weight decay strategy and learning parameters of the Adam optimizer. In this embodiment, the learning rate is set to 0.0001; the weight decay strategy, the decay rate is 0.0005;
[0088] Minimize the loss by updating the weights and learning parameters to improve the accuracy of the prediction evaluation score.
[0089] Furthermore, there is the following relationship for obtaining the average subjective score (MOS) of all data in the non-uniformly distorted panoramic image database:
[0090]
[0091] Among them, m i is the experience quality opinion score given by the i-th subject to the non-uniformly distorted panoramic picture, and MOS represents the average subjective score.
[0092] After the training of the present invention is completed, by comparing and calculating the finally obtained prediction evaluation score with the average subjective score, various test indexes are obtained to evaluate the performance of the model after training.
[0093] The test indexes include Spearman rank correlation coefficient (SRCC), Pearson correlation coefficient (PLCC) and root mean square error (RMSE):
[0094] The prediction monotonicity index, including the Spearman rank correlation coefficient (SRCC), is specifically expressed as:
[0095]
[0096] Among them, N represents the number of distorted images in the data, and d i represents the difference between the subjective score and the objective prediction score of the i-th image.
[0097] The prediction accuracy index, including the Pearson correlation coefficient (PLCC), is specifically expressed as:
[0098]
[0099] Among them, sub i and prei respectively represent the subjective score and the objective prediction score of the i-th image, and are the average subjective score and the average objective prediction score respectively.
[0100] The prediction error degree index includes root mean square error (RMSE)
[0101]
[0102] where sub i and pre i respectively represent the subjective score and the objective prediction score of the i-th image, and N represents the number of distorted images in the data.
[0103] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0104] 1. Based on the characteristics of the attention mechanism to capture quality perception features, obtain information content sensitive to non-uniform distortion features, guide the model to distinguish different types of non-uniform distortion, and enhance the discriminability of features;
[0105] 2. Aggregate multi-scale non-uniform distortion information to enhance feature representation and improve the sensitivity of features to image quality at different perception scales, and adopt deep semantics to enhance feature representation, enhance the abstract perception degree of multi-scale quality perception information, and improve the accuracy of subsequent prediction;
[0106] 3. Use the generalized mean pooling operation multiple times to dynamically compress feature information in a learnable manner. Compared with traditional pooling operations such as average pooling and max pooling, the generalized mean pooling can effectively reduce the loss of feature information during the pooling process;
[0107] 4. Through the flexible use of gated recurrent units, effective quality evaluation can be performed on both viewport sequence images containing time information and viewport sequence images not containing time information, expanding the application scope of the model.
[0108] Please refer to Figure 2 , the present invention also provides a blind quality evaluation system for non-uniform distortion panoramic images, and the system includes:
[0109] An extraction module, based on eye movement data, obtains a viewport sequence image of non-uniform distortion from a non-uniform distortion panoramic image, and performs image preprocessing on the viewport sequence image;
[0110] and based on a multi-layer multi-axis self-attention module, extracts multi-level quality perception features of the viewport sequence image of non-uniform distortion in a manner that uses the output of the previous layer as the input of the next layer;
[0111] A multi-scale fusion module, which is used to aggregate multi-level quality perception features to enhance feature representation and improve the sensitivity of features to image quality at different perception scales, so as to obtain multi-scale quality perception information;
[0112] A deep semantic guidance module, which is used to obtain the deep features of the multi-axis self-attention module, and in the way of deep semantic guidance, use the deep features to enhance the abstract perception degree of the multi-scale quality perception information, so as to obtain the final multi-scale quality perception information;
[0113] A quality regression module, which is used to input the final multi-scale quality perception information into a fully connected layer to obtain the predicted evaluation score of the non-uniform distortion panoramic image.
[0114] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0115] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A blind quality assessment method for non-uniformly distorted panoramic images, characterized in that, The method includes the following steps: Step 1: Based on eye movement data, obtain a sequence of non-uniformly distorted viewport images from a non-uniformly distorted panoramic image, and perform image preprocessing on the sequence of viewport images; Step 2: Based on a multi-layer multi-axis self-attention module, extract multi-level quality perception features of the sequence of non-uniformly distorted viewport images in a way that uses the output of the previous layer as the input of the next layer; Step 3: Aggregate the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to image quality at different perception scales, obtaining multi-scale quality perception information; Step 4: Obtain the deep features of the multi-axis self-attention module, and in a way guided by deep semantics, use the deep features to enhance the abstract perception degree of the multi-scale quality perception information, obtaining the final multi-scale quality perception information; Judge whether the final multi-scale quality perception information contains time information. If so, use several gated recurrent units to capture the time features within each viewport sequence in the final multi-scale quality perception information, and model the time information of the viewport sequence to solve the time delay effect; Step 5: Input the final multi-scale quality perception information into a fully connected layer to obtain the predicted evaluation score of the non-uniformly distorted panoramic image; In step 2, the method of using the output of the previous layer as the input of the next layer to extract information content sensitive to the features of the non-uniformly distorted viewport sequence image specifically includes: Extract the non-uniformly distorted sequence of viewport images to obtain preliminary features; Use a multi-layer multi-axis self-attention module, and form a network backbone by connecting the multi-layer multi-axis self-attention modules in series; Using the preliminary features as input features, extract features using the network backbone in a way that uses the output of the previous layer as the input of the next layer, in order to obtain multi-level quality perception features of different scales corresponding to the output of the multi-layer multi-axis self-attention module; In step 3, the method of aggregating the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to image quality at different perception scales, obtaining multi-scale quality perception information specifically includes: Perform a generalized mean pooling operation on the multi-level quality perception features to obtain dense information of the quality perception features at each scale. There is the following relationship for performing the generalized mean pooling operation on the multi-level quality perception features: ; Among them, represents the multi-level quality perception features input at four different scales, represents performing a generalized mean pooling operation, represents the multi-level quality perception features after pooling; Concatenate multi-level quality perception features at different scales and pass them through a fully connected layer to generate a multi-scale quality perception vector , and then obtain multi-scale quality perception information. The multi-scale quality perception vector has the following relationship: ; Among them, represents the multi-scale quality perception vector, respectively represent the multi-level quality perception features of four different scales after pooling, represents the concatenation operation, represents the fully connected layer; In step 4, there is the following relationship for the final multi-scale quality perception information: ; Among them, represents a quality perception vector with both multi-scale information representation and deep semantic representation, that is, the final multi-scale quality perception information, represents a deep semantic guidance vector, and there is the following relational expression for the deep semantic guidance vector: ; Among them, represents the deep features of the multi-axis self-attention module, that is, the multi-level quality perception features output by the last multi-axis self-attention module; The relationship formula for the extraction process of the multi-axis self-attention module is as follows: ; Among them, represents the input feature of the multi-axis self-attention module of the current layer, represents the output feature of the multi-axis self-attention module of the current layer, represents the mobile inverted residual bottleneck convolution operation, represents the block self-attention calculation operation, represents the grid self-attention calculation operation.
2. The blind quality evaluation method for non-uniformly distorted panoramic images according to claim 1, wherein In step 1, there is the following relationship for the non-uniformly distorted viewport sequence image: ; Among them, represents a non-uniform distortion viewport sequence image, represents the starting conditions, which are four good starting points g5 including 5 seconds, bad starting point b5 of 5 seconds, good starting point g15 of 15 seconds, and bad starting point b15 of 15 seconds; Indicates the th viewport sequence, and there is the following relationship for the viewport sequence: ; Among them, represents the th viewport in the th viewport sequence, indicating the number of viewports in the sequence.
3. A blind quality evaluation method for non-uniformly distorted panoramic images according to claim 1, characterized in that In step 1, the method of performing image preprocessing on the sequence of viewport images specifically includes: Resize the sequence of viewport images to a unified size; Normalize the resized sequence of viewport images; Use 80% of the normalized sequence of viewport images as the training set and 20% as the test set.
4. A blind quality evaluation method for non-uniformly distorted panoramic images according to claim 1, characterized in that There is the following relationship for the predicted evaluation score of the non-uniformly distorted panoramic image: ; Among them, represents the number of viewports in the viewport sequence, represents the th predicted score of the viewport, represents the predicted evaluation score.
5. A blind quality assessment method for non-uniformly distorted panoramic images according to claim 1, characterized in that, When performing the above steps 1 to 5, the corresponding training method includes the following training steps: Obtain the average subjective score of all data in the non-uniformly distorted panoramic image database; The mean absolute error is used as the loss for defining the mean subjective score and the predicted evaluation score, and a loss function is constructed. The expression of the loss function is as follows: ; Among them, and respectively represent the model prediction quality score and the average subjective score quality score of the th panoramic image in the training data, represents the number of people participating in the experiment to evaluate the quality of non-uniformly distorted panoramic images; The non-uniformly distorted viewport sequence images are cropped and input into the Adam optimizer for optimization. The weight decay strategy and learning parameters of the Adam optimizer are set; The loss is minimized by updating the weights and learning parameters to improve the accuracy of the predicted evaluation score.
6. A blind quality evaluation method for non-uniformly distorted panoramic images according to claim 5, characterized in that, In step 2, there is the following relational expression for obtaining the mean subjective score of all data in the non-uniformly distorted panoramic image database: ; Among them, is the score of the quality of experience given by the th subject for the non-uniform distortion panoramic picture, and MOS represents the average subjective score.
7. A blind quality evaluation system for non-uniformly distorted panoramic images, characterized in that, Applying the non-uniformly distorted panoramic image blind quality evaluation method according to any one of claims 1 to 6 above, the system includes: An extraction module, which, based on eye movement data, obtains non-uniformly distorted viewport sequence images from non-uniformly distorted panoramic images and performs image preprocessing on the viewport sequence images; And based on the multi-layer multi-axis self-attention module, the multi-level quality perception features of the non-uniformly distorted viewport sequence images are extracted in the way that the output of the previous layer is used as the input of the next layer; A multi-scale fusion module, which is used to aggregate the multi-level quality perception features to enhance the feature representation and improve the sensitivity of the features to the image quality at different perception scales, and obtain multi-scale quality perception information; A deep semantic guidance module, which is used to obtain the deep features of the multi-axis self-attention module and, in the way of deep semantic guidance, uses the deep features to enhance the abstract perception degree of the multi-scale quality perception information to obtain the final multi-scale quality perception information; A quality regression module, which is used to input the final multi-scale quality perception information into a fully connected layer to obtain the predicted evaluation score of the non-uniformly distorted panoramic image.
Citation Information
Patent Citations
No-reference image quality evaluation method based on spatial attention mechanism
CN114066812A
Incremental learning no-reference image quality evaluation method based on image semantic information
CN114972282A