Panoramic image quality evaluation method and system fusing multi-level distortion perception
By integrating multi-level distortion perception methods and utilizing backbone networks and feature map fusion modules, the problem of difficulty in capturing complex distortion features in panoramic image quality assessment is solved, achieving more efficient and accurate panoramic image quality assessment.
Patent Information
- Application Number
- CN202511097030.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing panoramic image quality assessment methods have difficulty in effectively capturing the various complex distortion features in panoramic images, especially in accurate assessment without relying on viewport priors.
A method of fusing multi-level distortion perception is adopted. Feature maps are extracted through the backbone network. The local-global distortion perception module, the deformable adaptive distortion perception module, the feature map fusion module, and the channel sense means phrase: attention fusion module are combined. Bilinear interpolation upsampling and spatial distortion perception attention processing are used to enhance the perception ability of complex distortion.
Without relying on viewport priors, it can automatically identify multi-scale distortion features, improving the accuracy and efficiency of panoramic image quality assessment.
Smart Images

Figure CN120598948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and multimedia digital image processing, and in particular to a panoramic image quality evaluation method and system integrating multi-level distortion perception. Background Art
[0002] Panoramic images (OI) have been widely used in virtual reality, autonomous driving, surveillance, and other fields due to their ability to provide 360-degree scene information. However, during the acquisition, transmission, and processing of panoramic images, various distortions, such as compression distortion, noise interference, and blur, are inevitably introduced. These distortions can seriously affect the quality of panoramic images and the user's viewing experience. Therefore, accurately evaluating the quality of panoramic images has become a pressing issue.
[0003] Existing panoramic image quality assessment methods are mainly divided into three categories: full-reference OIQA (FR-IQA), semi-reference OIQA (RR-OIQA), and no-reference / blind OIQA (NR- / BOIQA). FR-IQA and RR-OIQA require a reference image as a standard and assess image quality by comparing the differences between the reference image and the image to be assessed. However, in practical applications, reference images are often difficult to obtain. NR-IQA does not require a reference image and directly assesses the quality of the image to be assessed, making it more practical.
[0004] However, existing NR-IQA methods have limitations when processing panoramic images. Panoramic images are typically represented using the equirectangular projection (ERP), which can cause severe geometric distortion in polar regions. Panoramic images also exhibit multiple types of distortion, including global uniform distortion and local non-uniform distortion. Traditional NR-IQA methods struggle to effectively capture these complex distortion characteristics, resulting in inaccurate quality assessment results.
[0005] Current panoramic image quality assessment techniques often require explicit modeling of viewport information, such as locating key areas through user annotation or viewpoint prediction. This dependency not only introduces additional computational overhead but can also lead to subjective bias in the assessment results. To address this issue, we propose a method and system for panoramic image quality assessment that integrates multi-level distortion perception. This method and system can automatically identify multi-scale distortion features (such as projection distortion, compression artifacts, and local blur) in images without relying on viewport priors, thus providing a more efficient and versatile solution for panoramic image quality assessment. Summary of the Invention
[0006] In view of the above situation, the main purpose of the present invention is to propose a panoramic image quality evaluation method and system integrating multi-level distortion perception to solve the above technical problems.
[0007] The present invention proposes a panoramic image quality assessment method integrating multi-level distortion perception, the method comprising the following steps:
[0008] Step 1: Obtain a distorted panoramic image, process the distorted panoramic image to obtain a processed panoramic image; obtain a stage output feature map through the backbone network based on the processed panoramic image; and obtain a stage distortion-aware refinement feature map based on the stage output feature map;
[0009] Step 2: In the local branch of the local global distortion perception module, the stage-local distortion feature map is obtained based on the stage-distortion perception refinement feature map;
[0010] In the global branch, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module based on the more robust feature map and the stage local distortion feature map;
[0011] Step 3: Input the distortion-aware feature map processed by the deformable adaptive distortion-aware module into the feature map fusion module, fuse the distortion-aware feature map processed by the deformable adaptive distortion-aware module to obtain a fused output feature map, and obtain an output feature map processed by spatial distortion-aware attention based on the fused output feature map;
[0012] Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing;
[0013] Step 4: Input the stage output feature map into the channel perception enhancement module, optimize the stage output feature map to obtain an optimized feature map, and obtain the connection feature vector for quality regression based on the optimized feature map and the feature attention fusion feature map;
[0014] Based on the connected feature vector used for quality regression, the predicted quality score of the panoramic image is obtained after mapping through a fully connected layer.
[0015] The present invention also proposes a panoramic image quality assessment system integrating multi-level distortion perception, the system comprising:
[0016] Feature map extraction module, used for:
[0017] The distorted panoramic image is processed to obtain a processed panoramic image; based on the processed panoramic image, a stage output feature map is obtained through a backbone network; based on the stage output feature map, a stage distortion-aware refinement feature map is obtained;
[0018] Local-global distortion perception module, used to:
[0019] In the local branch, the feature map is refined based on the stage distortion perception to obtain the stage local distortion feature map;
[0020] In the global branch, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module based on the more robust feature map and the stage local distortion feature map;
[0021] Feature map fusion module, used for:
[0022] The distortion-aware feature maps processed by the deformable adaptive distortion-aware module are fused to obtain a fused output feature map, and an output feature map processed by the spatial distortion-aware attention is obtained based on the fused output feature map;
[0023] Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing;
[0024] Channel awareness enhancement module for:
[0025] The stage output feature map is optimized to obtain an optimized feature map, and a connection feature vector for quality regression is obtained based on the optimized feature map and the feature attention fusion feature map;
[0026] Quality score prediction module, used to:
[0027] Based on the connected feature vector used for quality regression, the predicted quality score of the panoramic image is obtained after mapping through a fully connected layer.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. The present invention uses a dual-branch local-global distortion perception module to mitigate the inherent geometric distortion of the global image while capturing global uniform distortion and local non-uniform distortion.
[0030] 2. The present invention uses a feature attention fusion module to upsample the deep feature map using bilinear interpolation, and after adding it to the shallow feature map, uses the spatial distortion perception attention module to adaptively highlight the areas related to the distortion, thereby enhancing the model's perception of complex distortions and obtaining the final fusion feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flow chart of the panoramic image quality assessment method integrating multi-level distortion perception proposed in the present invention;
[0032] Figure 2 Schematic diagram of the overall framework of the panoramic image quality assessment method integrating multi-level distortion perception proposed in the present invention;
[0033] Figure 3 This is a schematic diagram of the overall framework of the panoramic image quality assessment system integrating multi-level distortion perception proposed in the present invention. DETAILED DESCRIPTION
[0034] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0035] These and other aspects of the embodiments of the present invention will become clear with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0036] See also Figure 1 and Figure 2 The embodiment of the present invention proposes a panoramic image quality assessment method integrating multi-level distortion perception, which includes the following steps:
[0037] Step 1: Obtain a distorted panoramic image, process the distorted panoramic image to obtain a processed panoramic image; obtain a stage output feature map through the backbone network based on the processed panoramic image; and obtain a stage distortion-aware refinement feature map based on the stage output feature map;
[0038] In this step, based on the processed panoramic image, the stage output feature map is obtained through the backbone network. The relationship between the corresponding process is:
[0039] ;
[0040] in, Indicates the The stage output feature map, Indicates that the network is processed by the backbone network pre-training, represents the processed panoramic image, Represents the parameters of the backbone network pre-training network;
[0041] Based on the stage output feature map, the stage distortion perception refinement feature map is obtained, and the relationship between the corresponding process is:
[0042] ;
[0043] in, Indicates the Stage distortion-aware refinement feature map, Indicates that it has been processed by Gaussian error linear unit, Indicates that after layer normalization operation, Indicates that it has undergone a deformable convolution operation.
[0044] Furthermore, in this step, SwinTransformerV2 is used as the backbone network for hierarchical feature extraction. Its self-attention based on offset windows effectively models spatial dependencies and distortion patterns, making it well-suited for 2D and panoramic image content. For each processed panoramic image, multi-scale feature map representations are extracted from the four stages of the backbone to simultaneously capture low-level perceptual representations and high-level semantics.
[0045] Step 2: In the local branch of the local global distortion perception module, the stage-local distortion feature map is obtained based on the stage-distortion perception refinement feature map;
[0046] In the global branch, a more robust feature vector is obtained based on the stage distortion-aware feature map. Based on the more robust feature map and the stage local distortion feature map, the distortion-aware feature map of the stage is obtained after being processed by the deformable adaptive distortion-aware module.
[0047] In step 2, the stage local distortion feature map is obtained by refining the feature map based on the stage distortion perception. The specific steps are as follows:
[0048] The stage-wise distortion-aware refined feature map is subjected to 1×1 convolution and 3×3 depthwise convolution to expand the local receptive field. The feature map is evenly divided into three independent parts in the channel dimension, corresponding to querying key distortion information in the local area, identifying the key position or pattern of distortion features, and containing local distortion feature information that needs to be aggregated.
[0049] Based on querying key distortion information in the local area, identifying the key position or pattern of the distortion feature, and containing the local distortion feature information that needs to be aggregated, the stage local distortion feature map is obtained through the self-attention operation. The relationship between the corresponding process is:
[0050] ;
[0051] in, Indicates the Stage local distortion feature map, Indicates that after 1×1 convolution processing, Indicates that it has been processed by the Softmax function. Indicates in The key distortion information found in the local area of the stage, Indicates the Identify key locations or patterns of distortion features during the phase, Indicates the The stage contains the local distortion feature information that needs to be aggregated. represents the transpose symbol, express or Dimensions;
[0052] A more robust feature vector is obtained based on the stage distortion perception refinement feature map. Based on the more robust feature vector and the stage local distortion feature map, a distortion perception feature map of the stage after being processed by the deformable adaptive distortion perception module is obtained. The specific steps are as follows:
[0053] The stage-distortion-aware refined feature map is normalized, and the normalized stage-distortion-aware refined feature map is adaptively fused with the stage-distortion-aware refined feature map using learnable parameters to obtain a more robust feature map. The corresponding process has the following relationship:
[0054] ;
[0055] in, Indicates the The feature map of the stage is more robust, represents a small constant, The mean of the feature channel in the spatial dimension, The variance of the feature channel in the spatial dimension, Represents the scaling factor used to control the scale adjustment after feature normalization; Represents the global weighting factor, which acts on the normalized and offset feature map to control the contribution of this part of the feature to the final output; Represents the original feature weighting factor, which directly acts on the output feature map to control the degree of retention of the original features in the output; Represents the offset factor, which is used to adjust the center position of the feature distribution;
[0056] Use linear projection on the more robust feature map to reduce the number of channels to one-fourth of the original number of channels to obtain the feature vector after linear projection;
[0057] The feature vector after linear projection is processed by parallel depth convolution with receptive fields of 5×5, 7×7 and 9×9 to obtain the feature map output after multi-scale convolution processing. The relationship between the corresponding process is:
[0058] ;
[0059] in, Indicates the The feature map output after the multi-scale convolution processing of the stage, It means that the convolution kernel size is k×k, Indicates the The eigenvector of the stage after linear projection;
[0060] The feature map output after multi-scale convolution processing is restored to the original number of channels through 1×1 convolution, and then connected with the stage distortion-aware refined feature map through residual connection, and then projected to 512 dimensions to obtain the output feature map of the global branch;
[0061] Based on the output feature map of the global branch and the local distortion feature map of the stage, the distortion perception feature map of the stage after processing by the deformable adaptive distortion perception module is obtained. The relationship between the corresponding process is:
[0062] ;
[0063] in, Indicates the The distortion perception feature map after the deformable adaptive distortion perception module is processed in the stage. No. Output features of the global branch of the stage.
[0064] Furthermore, in this step, since there is an inherent two-level geometric distortion in the processed panoramic image, which greatly affects the accuracy of quality evaluation, deep convolution is combined with a hybrid local-global attention mechanism to achieve efficient complex local-global distortion perception.
[0065] Furthermore, in this step, the spatial dimension size is the width and height of the stage distortion-aware refined feature map.
[0066] Step 3: Input the distortion-aware feature map processed by the deformable adaptive distortion-aware module into the feature map fusion module, fuse the distortion-aware feature map processed by the deformable adaptive distortion-aware module to obtain a fused output feature map, and obtain an output feature map processed by spatial distortion-aware attention based on the fused output feature map;
[0067] Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing;
[0068] In step 3, the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the first stage is input into the feature map fusion module, and the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the first stage is fused to obtain a fused output feature map. Based on the fused output feature map, an output feature map processed by spatial distortion-aware attention is obtained. The specific steps are as follows:
[0069] A top-down progressive feature map fusion strategy is adopted for the distortion-aware feature map processed by the deformable adaptive distortion perception module in the first stage. The distortion-aware feature map processed by the deformable adaptive distortion perception module in the third stage and the distortion-aware feature map processed by the deformable adaptive distortion perception module in the fourth stage are fused to obtain the output feature map of the fusion of the third and fourth stages. The relationship between the corresponding processes is:
[0070] ;
[0071] in, Represents the output feature map of the fusion third and fourth stages, It represents the distortion perception feature map after the fourth stage is processed by the deformable adaptive distortion perception module. Represents the distortion perception feature map after processing by the deformable adaptive distortion perception module in the third stage;
[0072] Project the output feature maps of the fusion third and fourth stages to obtain the intermediate feature map;
[0073] The intermediate feature map is input into the spatial gating unit to generate the spatial attention map. The relationship between the corresponding process is:
[0074] ;
[0075] in, represents the spatial attention map, represents the intermediate feature map, Indicates a 3×3 depth convolution operation;
[0076] The spatial attention map is projected by 1×1 convolution and then element-wise multiplied with the intermediate features to obtain the output feature map of the spatial gating unit;
[0077] Based on the output feature map of the spatial gating unit and the output feature map of the fusion third and fourth stages, the residual connection is introduced to obtain the output feature map after spatial distortion-aware attention processing. The relationship between the corresponding processes is:
[0078] ;
[0079] in, represents the output feature map after spatial distortion-aware attention processing, Represents the output feature map of the spatial gating unit;
[0080] Based on the output feature map after spatial distortion-aware attention processing, the feature attention fusion feature map is obtained through bilinear interpolation upsampling and spatial distortion-aware attention processing. The specific steps are as follows:
[0081] The output feature map after spatial distortion-aware attention processing is fused with the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the second stage through bilinear interpolation upsampling, and then processed by spatial distortion-aware attention to obtain the feature map after spatial distortion-aware attention processing. The relationship between the corresponding processes is:
[0082] ;
[0083] in, represents the feature map after spatial distortion-aware attention processing, It represents the distortion perception feature map after the second stage is processed by the deformable adaptive distortion perception module. It means that after spatial distortion perception attention processing, Indicates that bilinear interpolation upsampling operation has been performed;
[0084] After the feature map processed by spatial distortion-aware attention is subjected to bilinear interpolation upsampling, it is fused with the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the first stage, and then processed by spatial distortion-aware attention to obtain the feature-attention fusion feature map. The relationship between the corresponding processes is:
[0085] ;
[0086] in, Represents the feature attention fusion feature map, which integrates the distortion perception features of all stages; Represents the distortion perception feature map after processing by the deformable adaptive distortion perception module in the first stage.
[0087] Step 4: Input the stage output feature map into the channel perception enhancement module, optimize the stage output feature map to obtain an optimized feature map, and obtain the connection feature vector for quality regression based on the optimized feature map and the feature attention fusion feature map;
[0088] Based on the connected feature vector used for quality regression, after mapping through a fully connected layer, the predicted quality score of the panoramic image is obtained;
[0089] In step 4, the stage output feature map is input into the channel perception enhancement module, and the stage output feature map is optimized to obtain an optimized feature map. Based on the optimized feature map and the feature attention fusion feature map, a connection feature vector for quality regression is obtained. The specific steps are as follows:
[0090] The fourth stage output feature map is subjected to a depthwise convolution operation with different receptive field channel projections to obtain an optimized feature map. The corresponding relationship is:
[0091] ;
[0092] in, represents the intermediate feature map obtained by the first stage enhancement, Indicates that after a 5×5 depth convolution operation, Represents the output feature map of the fourth stage, represents the optimized feature map, Indicates that it has been processed by the discard layer;
[0093] Layer normalization and residual connection are used on the optimized feature map to obtain channel-aware enhanced feature map;
[0094] The channel-aware enhanced feature map and the feature attention fusion feature map are flattened respectively, and then the flattened channel-aware enhanced feature map and the flattened feature attention fusion feature map are connected to obtain the connected feature vector for quality regression. The relationship between the corresponding process is:
[0095] ;
[0096] in, represents the concatenated feature vector for quality regression, represents the channel-aware enhanced feature map, Represents a flattening operation;
[0097] Based on the connection feature vector used for quality regression, after mapping through a fully connected layer, the predicted quality score of the panoramic image is obtained. The relationship between the corresponding process is:
[0098] ;
[0099] in, Indicates the The predicted quality score of the panoramic image, Indicates processing through the fully connected layer, represents the learnable parameters of the fully connected layer.
[0100] Furthermore, in this step, to enhance the model's ability to perceive and represent distortion-related features, the channel-aware enhancement module enhances each channel of the high-level feature map, improving the feature representation capabilities of important channels. Specifically, deep convolution operations with different receptive field channel projections are applied to the high-level output feature maps of the backbone network to adaptively enhance distortion-aware features while maintaining semantic integrity. The level of the output feature is determined by the features extracted at different levels of the backbone network, with higher-level features extracted at later network layers.
[0101] Furthermore, in this step, the purpose of connecting the flattened channel-aware enhanced feature map and the flattened feature attention fusion feature map is to construct a unified quality-aware representation.
[0102] Furthermore, in the present invention, during the execution of steps 1 to 4 above, the corresponding training method includes the following training steps:
[0103] When training the model, freeze the shallow parameters of the backbone network and only train the deep parameters to improve the efficiency and stability of feature map extraction;
[0104] The model is trained and optimized using the mean squared error loss function, using a set of predicted quality scores , and a given set of subjective quality scores Construct the loss function, the expression of the mean square error loss function is:
[0105] ;
[0106] in, represents the mean square error loss function, represents the total number of training samples, Indicates the The subjective quality score of the panoramic image;
[0107] Input the mean square error loss function into the Adam optimizer for optimization, and set the Adam optimizer weight decay strategy and learning parameters;
[0108] The loss is minimized by updating the weights and learning parameters to improve the prediction distortion range and quality.
[0109] Furthermore, the prediction results of the present invention are compared with the MOS score to calculate various indicators of the model, wherein the relationship formula for calculating the MOS score is:
[0110] ;
[0111] in, Indicates the The average score of the pictures; Indicates the number of evaluators, i.e. the total number of marks; Indicates the Picture courtesy of A mark indicates the quality of experience score (Quality of Experience QoE), Indicates the total number of pictures.
[0112] Furthermore, the present invention uses three IQA model evaluation metrics: the Pearson linear correlation coefficient (PLCC), the Spearman correlation coefficient (SRCC), and the root mean square error (RMSE). PLCC and RMSE are used to assess prediction accuracy, while SRCC is used to assess the consistency of prediction monotonicity. Higher PLCC and SRCC values, and lower RMSE values, indicate higher prediction accuracy and consistency. PLCC is calculated using a four-parameter nonlinear mapping function, with the specific mapping function being as follows:
[0113] ;
[0114] in, represents the objective score of the mapping, represents the quality score of the prediction, Both represent fitting parameters.
[0115] See also Figure 3 The embodiment of the present invention further provides a panoramic image quality assessment system integrating multi-level distortion perception, the system comprising:
[0116] Feature map extraction module, used for:
[0117] The distorted panoramic image is processed to obtain a processed panoramic image; based on the processed panoramic image, a stage output feature map is obtained through a backbone network; based on the stage output feature map, a stage distortion-aware refinement feature map is obtained;
[0118] Local-global distortion perception module, used to:
[0119] In the local branch, the feature map is refined based on the stage distortion perception to obtain the stage local distortion feature map;
[0120] In the global branch, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module based on the more robust feature map and the stage local distortion feature map;
[0121] Feature map fusion module, used for:
[0122] The distortion-aware feature maps processed by the deformable adaptive distortion-aware module are fused to obtain a fused output feature map, and an output feature map processed by the spatial distortion-aware attention is obtained based on the fused output feature map;
[0123] Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing;
[0124] Channel awareness enhancement module for:
[0125] The stage output feature map is optimized to obtain an optimized feature map, and a connection feature vector for quality regression is obtained based on the optimized feature map and the feature attention fusion feature map;
[0126] Quality score prediction module, used to:
[0127] Based on the connected feature vector used for quality regression, the predicted quality score of the panoramic image is obtained after mapping through a fully connected layer.
[0128] It should be understood that various components of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0129] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0130] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A panoramic image quality assessment method integrating multi-level distortion perception, characterized in that: The method comprises the following steps: Step 1: Obtain a distorted panoramic image, process the distorted panoramic image to obtain a processed panoramic image; obtain a stage output feature map through the backbone network based on the processed panoramic image; and obtain a stage distortion-aware refinement feature map based on the stage output feature map; Step 2: In the local branch of the local global distortion perception module, the stage-local distortion feature map is obtained based on the stage-distortion perception refinement feature map; In the global branch, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module based on the more robust feature map and the stage local distortion feature map; Step 3: Input the distortion-aware feature map processed by the deformable adaptive distortion-aware module into the feature map fusion module, fuse the distortion-aware feature map processed by the deformable adaptive distortion-aware module to obtain a fused output feature map, and obtain an output feature map processed by spatial distortion-aware attention based on the fused output feature map; Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing; Step 4: Input the stage output feature map into the channel perception enhancement module, optimize the stage output feature map to obtain an optimized feature map, and obtain the connection feature vector for quality regression based on the optimized feature map and the feature attention fusion feature map; Based on the connected feature vector used for quality regression, after mapping through a fully connected layer, the predicted quality score of the panoramic image is obtained; In step 2, the stage local distortion feature map is obtained based on the stage distortion perception refinement feature map, and the specific steps are as follows: The stage-wise distortion-aware refined feature map is subjected to 1×1 convolution and 3×3 depthwise convolution to expand the local receptive field. The resulting feature map is evenly divided into three independent parts in the channel dimension. These parts correspond to querying key distortion information in the local area, identifying the key position or pattern of distortion features, and containing local distortion feature information that needs to be aggregated. Based on querying key distortion information in the local area, identifying the key position or pattern of the distortion feature, and including the local distortion feature information that needs to be aggregated, a stage local distortion feature map is obtained through self-attention operation; In step 2, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage processed by the deformable adaptive distortion perception module is obtained based on the more robust feature map and the stage local distortion feature map. The specific steps are as follows: Normalize the stage-distortion-aware refined feature map, and adaptively fuse the normalized stage-distortion-aware refined feature map with the stage-distortion-aware refined feature map using learnable parameters to obtain a more robust feature map. Use linear projection on the more robust feature map to reduce the number of channels to one-fourth of the original number of channels to obtain the feature vector after linear projection; The linearly projected feature vector is processed by parallel depth convolution with receptive fields of 5×5, 7×7, and 9×9 to obtain the feature map output after multi-scale depth convolution processing; The feature map output after multi-scale convolution processing is restored to the original number of channels through 1×1 convolution, and then connected with the stage distortion-aware refined feature map through residual connection, and then projected to 512 dimensions to obtain the output feature map of the global branch; Based on the output feature map of the global branch and the local distortion feature map of the stage, the distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module.
2. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 1, characterized in that: In step 1, based on the processed panoramic image, the stage output feature map is obtained through the backbone network, and the relationship between the corresponding process is: ; in, Indicates the The stage output feature map, Indicates that the network is processed by the backbone network pre-training, represents the processed panoramic image, Represents the parameters of the backbone network pre-training network.
3. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 2, characterized in that: In step 1, the stage distortion perception refinement feature map is obtained based on the stage output feature map, and the relationship between the corresponding process is: ; in, Indicates the Stage distortion-aware refinement feature map, Indicates that it has been processed by Gaussian error linear unit, Indicates that after layer normalization operation, Indicates that it has undergone a deformable convolution operation.
4. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 3, characterized in that: Based on querying key distortion information in the local area, identifying the key position or pattern of the distortion feature, and containing the local distortion feature information that needs to be aggregated, the stage local distortion feature map is obtained through the self-attention operation. The relationship between the corresponding process is: ; in, Indicates the Stage local distortion feature map, Indicates that after 1×1 convolution processing, Indicates that it has been processed by the Softmax function. Indicates in The key distortion information found in the local area of the stage, Indicates the Identify key locations or patterns of distortion features during the phase, Indicates the The stage contains the local distortion feature information that needs to be aggregated. represents the transpose symbol, express or dimension.
5. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 4, characterized in that: The stage-distortion-aware refined feature map is normalized, and the normalized stage-distortion-aware refined feature map is adaptively fused with the stage-distortion-aware refined feature map using learnable parameters to obtain a more robust feature map. The corresponding process has the following relationship: ; in, Indicates the The feature map of the stage is more robust, represents a small constant, The mean of the feature channel in the spatial dimension, The variance of the feature channel in the spatial dimension, represents the scaling factor used to control the scale adjustment after feature normalization, represents the global weighting factor, represents the original feature weighting factor, represents the offset factor; Use linear projection on the more robust feature map to reduce the number of channels to one-fourth of the original number of channels to obtain the feature vector after linear projection; The feature vector after linear projection is processed by parallel depth convolution with receptive fields of 5×5, 7×7 and 9×9 to obtain the feature map output after multi-scale depth convolution processing. The relationship between the corresponding process is: ; in, Indicates the The feature map output after the stage-by-stage deep multi-scale convolution processing, It means that the convolution kernel size is k×k, Indicates the The eigenvector of the stage after linear projection; The feature map output after multi-scale convolution processing is restored to the original number of channels through 1×1 convolution, and then connected with the stage distortion-aware refined feature map through residual connection, and then projected to 512 dimensions to obtain the output feature map of the global branch; Based on the output feature map of the global branch and the local distortion feature map of the stage, the distortion perception feature map of the stage after processing by the deformable adaptive distortion perception module is obtained. The relationship between the corresponding process is: ; in, Indicates the The distortion perception feature map after the deformable adaptive distortion perception module is processed in the stage. No. Output features of the global branch of the stage.
6. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 5, characterized in that: In step 3, the distortion perception feature map processed by the deformable adaptive distortion perception module in the stage is input into the feature map fusion module, and the distortion perception feature map processed by the deformable adaptive distortion perception module in the stage is fused to obtain a fused output feature map. Based on the fused output feature map, an output feature map processed by spatial distortion perception attention is obtained. The specific steps are as follows: For the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the first stage, a top-down progressive feature map fusion strategy is adopted to fuse the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the third stage and the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the fourth stage to obtain the output feature map of the fusion of the third and fourth stages. The relationship between the corresponding processes is: ; in, Represents the output feature map of the fusion third and fourth stages, It represents the distortion perception feature map after the fourth stage is processed by the deformable adaptive distortion perception module. Represents the distortion perception feature map after processing by the deformable adaptive distortion perception module in the third stage; Project the output feature maps of the fusion third and fourth stages to obtain the intermediate feature map; The intermediate feature map is input into the spatial gating unit to generate the spatial attention map. The relationship between the corresponding process is: ; in, represents the spatial attention map, represents the intermediate feature map, Indicates a 3×3 depth convolution operation; The spatial attention map is projected by 1×1 convolution and then element-wise multiplied with the intermediate features to obtain the output feature map of the spatial gating unit; Based on the output feature map of the spatial gating unit and the output feature map of the fusion third and fourth stages, the residual connection is introduced to obtain the output feature map after spatial distortion-aware attention processing. The relationship between the corresponding processes is: ; in, represents the output feature map after spatial distortion-aware attention processing, Represents the output feature map of the spatial gating unit.
7. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 6, characterized in that: In step 3, based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing. The specific steps are as follows: The output feature map after spatial distortion-aware attention processing is fused with the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the second stage through bilinear interpolation upsampling, and then processed by spatial distortion-aware attention to obtain the feature map after spatial distortion-aware attention processing. The relationship between the corresponding processes is: ; in, represents the feature map after spatial distortion-aware attention processing, It represents the distortion perception feature map after the second stage is processed by the deformable adaptive distortion perception module. It means that after spatial distortion perception attention processing, Indicates that bilinear interpolation upsampling operation has been performed; After the feature map processed by spatial distortion-aware attention is subjected to bilinear interpolation upsampling, it is fused with the distortion-aware feature map processed by the deformable adaptive distortion-aware module in the first stage, and then processed by spatial distortion-aware attention to obtain the feature-attention fusion feature map. The relationship between the corresponding processes is: ; in, represents the feature attention fusion feature map, Represents the distortion perception feature map after processing by the deformable adaptive distortion perception module in the first stage.
8. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 7, characterized in that: In step 4, the stage output feature map is input into the channel perception enhancement module, the stage output feature map is optimized to obtain an optimized feature map, and a connection feature vector for quality regression is obtained based on the optimized feature map and the feature attention fusion feature map. The specific steps are as follows: The fourth stage output feature map is subjected to a depthwise convolution operation with different receptive field channel projections to obtain an optimized feature map. The corresponding relationship is: ; in, represents the intermediate feature map obtained by the first stage enhancement, Indicates that after a 5×5 depth convolution operation, Represents the output feature map of the fourth stage, represents the optimized feature map, Indicates that it has been processed by the discard layer; Layer normalization and residual connection are used on the optimized feature map to obtain channel-aware enhanced feature map; The channel-aware enhanced feature map and the feature attention fusion feature map are flattened respectively, and then the flattened channel-aware enhanced feature map and the flattened feature attention fusion feature map are connected to obtain the connected feature vector for quality regression. The relationship between the corresponding process is: ; in, represents the concatenated feature vector for quality regression, represents the channel-aware enhanced feature map, Represents a flatten operation.
9. The panoramic image quality assessment method integrating multi-level distortion perception according to claim 8, characterized in that: In step 4, based on the connection feature vector used for quality regression, after mapping through a fully connected layer, the predicted quality score of the panoramic image is obtained. The relationship between the corresponding process is: ; in, Indicates the The predicted quality score of the panoramic image, Indicates processing through the fully connected layer, represents the learnable parameters of the fully connected layer.
10. A panoramic image quality assessment system integrating multi-level distortion perception, characterized in that: The system applies the panoramic image quality assessment method integrating multi-level distortion perception according to any one of claims 1 to 9, and the system comprises: Feature map extraction module, used for: The distorted panoramic image is processed to obtain a processed panoramic image; based on the processed panoramic image, a stage output feature map is obtained through a backbone network; based on the stage output feature map, a stage distortion-aware refinement feature map is obtained; Local-global distortion perception module, used to: In the local branch, the feature map is refined based on the stage distortion perception to obtain the stage local distortion feature map; In the global branch, a more robust feature map is obtained based on the stage distortion perception refinement feature map, and a distortion perception feature map of the stage is obtained after being processed by the deformable adaptive distortion perception module based on the more robust feature vector and the stage local distortion feature map; Feature map fusion module, used for: The distortion-aware feature maps processed by the deformable adaptive distortion-aware module are fused to obtain a fused output feature map, and an output feature map processed by the spatial distortion-aware attention is obtained based on the fused output feature map; Based on the output feature map after spatial distortion-aware attention processing, a feature attention fusion feature map is obtained through bilinear interpolation upsampling processing and spatial distortion-aware attention processing; Channel awareness enhancement module for: The stage output feature map is optimized to obtain an optimized feature map, and a connection feature vector for quality regression is obtained based on the optimized feature map and the feature attention fusion feature map; Quality score prediction module, used to: Based on the connected feature vector used for quality regression, the predicted quality score of the panoramic image is obtained after mapping through a fully connected layer.
Citation Information
Patent Citations
Panoramic image blind quality evaluation method and system based on multi-collaborative network assistance
CN118196107A
Dynamic cross-level distortion information fusion reference-free panoramic image quality evaluation method
CN118823379A