Depth feature-based AI generated panoramic image quality evaluation method and system
Through the AI-generated panoramic image quality evaluation method based on depth features, using natural multivariate Gaussian model and Pap distance calculation, the problem of difficulty in evaluating the authenticity of AI-generated panoramic images and matching degree with text description in the prior art is solved, and higher evaluation accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510455312.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to accurately judge the authenticity of AI-generated panoramic images and lacks the ability to evaluate the degree of matching between the generated image and its corresponding text description.
A method of evaluating the quality of AI-generated panoramic image based on depth features is proposed. By processing the natural panoramic image in isometric columnar projection format, and using pre-trained convolutional neural network to extract multi-layer depth feature maps, combining spatial downsampling and stitching operations, a natural multivariate Gaussian model is constructed, and Papist distance is calculated to evaluate image quality.
It effectively reduces geometric distortion problems, improves the accuracy and reliability of panoramic images quality evaluation, and can more comprehensively and objectively evaluate the quality of AI-generated panoramic images.
Smart Images

Figure CN119963565A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and multimedia digital image processing, and in particular to a method and system for evaluating the quality of panoramic images generated by AI based on depth features. Background Art
[0002] With the rapid development of virtual reality technology, people can enjoy unprecedented immersive and interactive visual experiences in their daily lives. In particular, in the past few years, virtual reality content has not only attracted widespread attention from the industry, but has also become one of the hot areas of academic research. Panoramic images (also known as omnidirectional images or 360-degree images), as an important form of virtual reality content, are becoming increasingly popular because they can provide a full range of perspectives.
[0003] However, compared with panoramic images taken naturally, panoramic images generated by AI often have low-level technical distortion and high-level semantic distortion problems, which seriously interfere with the user's immersive experience. In view of this, it is particularly urgent to conduct in-depth research on the quality of panoramic images generated by AI and propose accurate evaluation methods. At present, panoramic image quality assessment mainly adopts manual feature-based methods and deep learning-based methods. Depending on the input data, the existing panoramic image quality evaluation models can be roughly divided into three categories: two-dimensional plane-based methods, sphere-based methods, and viewport-based methods.
[0004] Although these methods are able to assess the authenticity of generated images to a certain extent, they have difficulty in accurately judging the authenticity of a single generated image and generally lack the ability to assess the degree of match between a generated image and its corresponding text description. Summary of the invention
[0005] In view of the above situation, the main purpose of the present invention is to propose an AI-generated panoramic image quality evaluation method and system based on depth features to solve the above technical problems.
[0006] The present invention proposes a method for evaluating the quality of panoramic images generated by AI based on depth features, the method comprising the following steps: Step 1: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Step 2, converting the natural panoramic image in the equirectangular projection format into a viewport image; Step 3: Use the pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; Step 4: Use the intermediate feature map to calculate the local mean map and standard deviation map, and use the standard deviation map to construct the covariance matrix; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; Step 5, replace the natural panoramic image in equirectangular projection format with the panoramic image generated by the test AI, and obtain the natural multivariate Gaussian model of the test through steps 2 to 4; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; Step 6: Sum up the raw scores to get the final quality score.
[0007] The present invention also proposes an AI-generated panoramic image quality evaluation system based on depth features, the system comprising: Data preprocessing module, used to: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Image conversion module for: Convert natural panoramic images in equirectangular projection format to viewport images; Feature extraction module for: Use pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; Natural multivariate Gaussian model building blocks for: The local mean map and standard deviation map are obtained by calculation using the intermediate feature map, and the covariance matrix is constructed using the standard deviation map; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; The original score acquisition module is used to: Replace the natural panoramic image in equirectangular projection format with the test AI-generated panoramic image, and build the natural multivariate Gaussian model for the test; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; Aggregation module for: The raw scores are summed up to get the final quality score.
[0008] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention adopts an image conversion method from equirectangular projection format to cubic projection. In the traditional panoramic image processing process, geometric distortion often has an adverse effect on the final image quality. By converting equirectangular projection into cubic projection, the existence of such distortion can be effectively reduced, so that AI can provide a more accurate and clear visual experience when generating panoramic images, thereby improving the overall quality assessment level of panoramic images; 2. The present invention uses a deep neural network to perform multi-level feature extraction, which can not only effectively capture various distortion conditions in the image, but also provide a deeper understanding of the image content, providing a solid foundation for subsequent image quality evaluation; 3. The present invention adjusts the model parameters according to the specific type of viewport content to better simulate the human visual perception process, which can greatly improve the accuracy and reliability of panoramic image quality assessment; 4. The present invention fits the multivariate Gaussian model obtained by different conversion viewport images to calculate the quality score of each viewport respectively, and then determines the final panoramic image quality score by averaging these scores. This method not only takes into account the differences between different viewports, but also ensures the comprehensiveness and objectivity of the final evaluation results, greatly improving the performance of the model in practical applications.
[0009] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flow chart of the AI-generated panoramic image quality evaluation method based on depth features proposed in the present invention; Figure 2 This is the overall framework diagram of the AI-generated panoramic image quality evaluation method based on deep features proposed in the present invention; Figure 3 Schematic diagram of the overall framework of the AI-generated panoramic image quality evaluation system based on depth features proposed in the present invention. DETAILED DESCRIPTION
[0011] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0012] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0013] See also Figure 1 , an embodiment of the present invention proposes an AI-generated panoramic image quality evaluation method based on depth features, the method comprising the following steps: Step 1: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Specifically, in the rough stage of processing the data of natural panoramic images in this step, a set of original panoramic images without perceptible distortion is collected from the existing panoramic image database, and a total of 3903 panoramic images are obtained; In the refinement phase of processing the data of natural panoramic images in this step, panoramic images with similar contents were manually selected and removed, and ultimately 2,500 panoramic images were retained.
[0014] Step 2, converting the natural panoramic image in the equirectangular projection format into a viewport image; Specifically, in this step, a given set of natural panoramic images in equirectangular projection format is converted into a set of non-overlapping viewport images, i.e., six images converted by cubic projection; At the same time, in order to strike a balance between efficiency and effect, the present invention sets the starting viewing angle to 0° (longitude is 0 degrees) and obtains six cube projection conversion images: top, bottom, front, back, left and right; The two images transformed from the cubic projection are oriented toward the nadir and zenith, respectively, and the other four images are aimed at the horizon and rotated horizontally to include the entire band around the equator of the sphere. Therefore, for a particular panoramic image, we represent the projected viewport image as .
[0015] Step 3: Use the pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; In step 3, the viewport image is extracted using a pre-trained neural network to obtain a multi-layer depth feature map. The corresponding relationship is: ; in, represents a multi-layer deep feature map, Indicates that it has been processed by the backbone network. Indicates the parameters, represents the top part of the image, represents the bottom part of the image, represents the front part of the image, Indicates the back side of the image. represents the left part of the image, Represents the right part of the image; Through the layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of depth feature maps are gradually spliced to obtain the intermediate feature map. The relationship between the corresponding process is: ; in, represents the intermediate feature map, represents the spatial downsampling operation, Represents a splicing operation, represents the first layer of deep feature map, represents the second layer deep feature map, represents the third layer deep feature map, represents the fourth layer deep feature map, Represents the fifth layer deep feature map; Furthermore, in this step, the pre-trained neural network is based on deep neural networks such as VGG, residual network (ResNET) and efficient convolutional neural network (EfficientNet), and is constructed using the pre-trained efficient convolutional neural network (EfficientNet) seat backbone network for multi-layer deep feature representation, and the backbone network is replaceable, and different networks can be selected according to needs; At the same time, in order to unify the deep feature representations at different scales, the present invention uses a Gaussian filter to implement iterative sampling of the initial layer output. On this basis, the intermediate feature maps are obtained by connecting these spatially downsampled feature maps with the fifth layer feature maps along the channel dimension.
[0016] Step 4: Use the intermediate feature map to calculate the local mean map and standard deviation map, and use the standard deviation map to construct the covariance matrix; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; In step 4, the local mean map and the standard deviation map are obtained by calculation using the intermediate feature map, wherein the relationship in the process of obtaining the local mean map is: ; in, represents the local mean map, represents the spatial coordinates of the feature map, represents the height of the filter, represents the width of the filter, Indicates the vertical offset. Indicates the horizontal offset. represents a two-dimensional Gaussian filter, Indicates that the position on the intermediate feature map is The value of The local mean map and standard deviation map are obtained by calculation using the intermediate feature map. The standard sweat wiping map is obtained using the local mean map. The relationship between the corresponding process is: ; in, represents the standard deviation graph; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain the normalized local mean map. The corresponding process relationship is: ; in, represents the normalized local mean map, Indicates channel The corresponding local mean map is, Indicates that it has been processed by L2 norm; Using the mean vector and covariance matrix of the multivariate Gaussian model, a natural multivariate Gaussian model is constructed, and the relationship between the corresponding process is: ; in, represents the natural multivariate Gaussian model, represents the statistical eigenvector matrix extracted from the projected viewport image, represents the mean vector of the multivariate Gaussian model, represents the covariance matrix of the multivariate Gaussian model, The number of dimensions representing the statistical features, Represents pi.
[0017] Step 5, replace the natural panoramic image in equirectangular projection format with the panoramic image generated by the test AI, and obtain the natural multivariate Gaussian model of the test through steps 2 to 4; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; In step 5, the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model is calculated to obtain the original score. The relationship between the corresponding process is: ; in, represents the original score, It means that after the Bhattacharyya distance calculation, represents the mean vector of the natural multivariate Gaussian model under test, represents the covariance matrix of the natural multivariate Gaussian model of the test, represents the mean vector of the natural multivariate Gaussian model, Represents the covariance matrix of the natural multivariate Gaussian model.
[0018] Step 6: Summarize the original scores to get the final quality score; In step 6, the raw scores are summarized to obtain the final quality score. The relationship between the corresponding process is: ; in, represents the final quality score, Indicates that an aggregation operation has been performed.
[0019] Furthermore, the present invention uses two standard indicators to evaluate the performance of quality assessment, including Spearman rank correlation coefficient and Pearson linear correlation coefficient; Before calculating the Pearson linear correlation coefficient, a four-parameter logistic regression function was used to simulate the nonlinear relationship between the predicted score and the true score. The relationship is: ; in, represents the four-parameter logistic regression function, , , and All represent parameters. represents the original input score; The calculation formula of Pearson linear correlation coefficient is: ; in, represents the Pearson linear correlation coefficient, represents the number of observations, represents the sequence number of the observation value, and Represents the first Observations, and Respectively represent the average values of two variables; The calculation formula of Spearman's rank correlation coefficient is: ; in, represents the Spearman rank correlation coefficient, Indicates that for The rank difference between the observation points and the observation values in the two variables is Represents the number of paired observations.
[0020] See also Figure 3 , an embodiment of the present invention further provides an AI-generated panoramic image quality evaluation system based on depth features, the system comprising: Data preprocessing module, used to: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Image conversion module for: Convert natural panoramic images in equirectangular projection format to viewport images; Feature extraction module for: Use pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; Natural multivariate Gaussian model building blocks for: The local mean map and standard deviation map are obtained by calculation using the intermediate feature map, and the covariance matrix is constructed using the standard deviation map; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; The original score acquisition module is used to: Replace the natural panoramic image in equirectangular projection format with the test AI-generated panoramic image, and build the natural multivariate Gaussian model for the test; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; Aggregation module for: The raw scores are summed up to get the final quality score.
[0021] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0022] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0023] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for evaluating the quality of panoramic images generated by AI based on deep features, characterized in that: The method comprises the following steps: Step 1: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Step 2, converting the natural panoramic image in the equirectangular projection format into a viewport image; Step 3: Use the pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; Step 4: Use the intermediate feature map to calculate the local mean map and standard deviation map, and use the standard deviation map to construct the covariance matrix; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; Step 5, replace the natural panoramic image in equirectangular projection format with the panoramic image generated by the test AI, and obtain the natural multivariate Gaussian model of the test through steps 2 to 4; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; Step 6: Sum up the raw scores to get the final quality score.
2. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 1, characterized in that: In step 3, the viewport image is extracted using a pre-trained neural network to obtain a multi-layer depth feature map. The relationship between the corresponding process is: ; in, represents a multi-layer deep feature map, Indicates that it has been processed by the backbone network. Indicates the parameters, represents the top part of the image, represents the bottom part of the image, represents the front part of the image, Indicates the back side of the image. represents the left part of the image, Indicates the right part of the image.
3. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 2, characterized in that: In step 3, the first to fifth layers of depth feature maps are gradually spliced through layer-by-layer nested spatial downsampling and splicing operations to obtain intermediate feature maps. The relationship between the corresponding processes is: ; in, represents the intermediate feature map, represents the spatial downsampling operation, Represents a splicing operation, represents the first layer of deep feature map, represents the second layer deep feature map, represents the third layer deep feature map, represents the fourth layer deep feature map, Represents the fifth layer deep feature map.
4. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 3, characterized in that: In step 4, the local mean map and the standard deviation map are obtained by calculation using the intermediate feature map, wherein the relationship in the process of obtaining the local mean map is: ; in, represents the local mean map, represents the spatial coordinates of the feature map, represents the height of the filter, represents the width of the filter, Indicates the vertical offset. Indicates the horizontal offset. represents a two-dimensional Gaussian filter, Indicates that the position on the intermediate feature map is The value of .
5. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 4, characterized in that: In step 4, the local mean map and the standard deviation map are obtained by calculation using the intermediate feature map, wherein the standard sweat wiping map is obtained using the local mean map, and the relationship between the corresponding process is: ; in, Represents a standard deviation plot.
6. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 5, characterized in that: In step 4, the local mean map is normalized along the channel dimension using the L2 norm to obtain a normalized local mean map. The relationship between the corresponding process is: ; in, represents the normalized local mean map, Indicates channel The corresponding local mean map is, Indicates that it has been processed by L2 norm.
7. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 6, characterized in that: In step 4, the mean vector and covariance matrix of the multivariate Gaussian model are used to construct a natural multivariate Gaussian model, and the relationship between the corresponding process is: ; in, represents the natural multivariate Gaussian model, represents the statistical eigenvector matrix extracted from the projected viewport image, represents the mean vector of the multivariate Gaussian model, represents the covariance matrix of the multivariate Gaussian model, The number of dimensions representing the statistical features, Represents pi.
8. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 7, characterized in that: In step 5, the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model is calculated to obtain the original score. The relationship between the corresponding process is: ; in, represents the original score, It means that after the Bhattacharyya distance calculation, represents the mean vector of the natural multivariate Gaussian model under test, represents the covariance matrix of the natural multivariate Gaussian model of the test, represents the mean vector of the natural multivariate Gaussian model, Represents the covariance matrix of the natural multivariate Gaussian model.
9. The method for evaluating the quality of panoramic images generated by AI based on depth features according to claim 8, characterized in that: In step 6, the original scores are summarized to obtain the final quality score, and the relationship between the corresponding process is: ; in, represents the final quality score, Indicates that an aggregation operation has been performed.
10. An AI-generated panoramic image quality evaluation system based on depth features, characterized in that: The system applies any one of claims 1 to 9 of the AI-generated panoramic image quality assessment method based on depth features, and the system comprises: Data preprocessing module, used to: Processing the data of the natural panoramic image to obtain the natural panoramic image in an equirectangular projection format; Image conversion module for: Convert natural panoramic images in equirectangular projection format to viewport images; Feature extraction module for: Use pre-trained convolutional neural network as the backbone network to build a pre-trained neural network; Use the pre-trained neural network to extract the viewport image and obtain a multi-layer deep feature map; Through layer-by-layer nested spatial downsampling and splicing operations, the first to fifth layers of deep feature maps are gradually spliced to obtain intermediate feature maps; Natural multivariate Gaussian model building blocks for: The local mean map and standard deviation map are obtained by calculation using the intermediate feature map, and the covariance matrix is constructed using the standard deviation map; Along the channel dimension, the local mean map is normalized using the L2 norm to obtain a normalized local mean map; The multivariate Gaussian model was constructed using the normalized local mean map and covariance matrix; The mean vector and covariance matrix of the multivariate Gaussian model are obtained through maximum likelihood estimation; The natural multivariate Gaussian model is constructed by using the mean vector and covariance matrix of the multivariate Gaussian model; The original score acquisition module is used to: Replace the natural panoramic image in equirectangular projection format with the test AI-generated panoramic image, and build the natural multivariate Gaussian model for the test; Calculate the Bhattacharyya distance between the tested natural multivariate Gaussian model and the natural multivariate Gaussian model to get the raw score; Aggregation module for: The raw scores are summed up to get the final quality score.
Citation Information
Patent Citations
Non-reference image quality evaluation method based on high-quality natural image statistical magnitude model
CN103996192A
Virtual reality image quality evaluation method and system
CN115546162A
Virtual reality image quality evaluation method and system
CN116563210A
No-reference panoramic image quality evaluation method based on saliency weighting
CN116777886A
Radio foreground source modeling method based on deep learning and Gaussian fitting
CN117058524A