Multi-modal panoramic image blind quality evaluation method and system based on AI generation description

Through the multimodal panoramic image blind quality evaluation method based on AI generation description, the panoramic image features are extracted using visual language models and isometric deformable convolution blocks, which solves the problem of panoramic image quality degradation and achieves efficient and accurate quality evaluation.

CN120356071AActive Publication Date: 2025-07-22JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510852255.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively evaluate the quality decline caused by factors such as device performance and network bandwidth fluctuations during the acquisition, transmission and display of panoramic images, which affects the user experience.

Method used

The multimodal panoramic image blind quality evaluation method based on AI generation description is adopted, and image blocks are acquired through positioning and cropping, and features are extracted using visual language model pre-training and isometric deformable convolutional blocks to perform quality prediction.

Benefits of technology

It improves the accuracy and efficiency of panoramic image quality evaluation, reduces the computational burden, and enhances the richness of feature representation and the ability to integrate multimodal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356071A_ABST
    Figure CN120356071A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal panoramic image blind quality evaluation method and system based on AI generation description. The method comprises the following steps: generating a quality description text based on a distorted panoramic image; obtaining an image block containing rich visual semantic information from the distorted panoramic image; pre-training a visual language model by using the distorted panoramic image and the quality description text to obtain a pre-trained visual language model; extracting the quality description text and the image blocks containing the rich visual semantic information by using the pre-trained visual language model to obtain text features supervised by the image information; image block features are obtained through the equidistant deformable convolution blocks; a quality prediction score is obtained based on image block features and textual features supervised by image information. According to the method, the powerful visual representation capability of the multi-modal big language is explored, wide experiments are carried out in a panoramic image quality evaluation database, and the result shows that the method has remarkable advantages in the aspect of evaluating panoramic image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and multimedia digital image processing, and in particular to a method and system for blind quality assessment of multimodal panoramic images based on AI generated descriptions. Background Art

[0002] With the rapid development and popularization of virtual reality (VR) technology, panoramic images (OI), as an image form that can provide a 360-degree all-round visual experience, have shown their unique value in many application scenarios. Whether it is virtual tourism leading users to experience the beauty of the world without leaving home, or online education providing students with an immersive learning environment through panoramic classrooms, or real estate exhibitions allowing potential buyers to view properties online, OI plays a vital role.

[0003] However, OI is inevitably affected by a variety of factors throughout its life cycle of acquisition, transmission, and display. These factors include but are not limited to device performance limitations, bandwidth fluctuations in network transmission, and resolution differences in display devices. These factors work together to reduce image quality, thereby affecting the user's quality of experience (QoE). Summary of the invention

[0004] In view of the above situation, the main purpose of the present invention is to propose a multimodal panoramic image blind quality assessment method and system based on AI generated description to solve the above technical problems.

[0005] The present invention proposes a blind quality assessment method for multimodal panoramic images based on AI-generated descriptions, the method comprising the following steps: Step 1, generating quality description text based on the distorted panoramic image; By adopting the positioning and cropping method, image blocks containing rich visual semantic information are obtained from the distorted panoramic image; Step 2: Pre-training the visual language model using the distorted panoramic image and the quality description text to obtain a pre-trained visual language model; Step 3: Use the pre-trained visual language model to extract the quality description text and the image blocks containing rich visual semantic information to obtain text features supervised by the image information; Step 4: Use the isometric deformable convolution block to extract the image block to obtain the image block features; The image block features are spliced with the text features supervised by the image information to obtain the spliced and fused perceptual features; The concatenated and fused perceptual features are regressed for quality to obtain the quality prediction score.

[0006] The present invention also provides a multi-modal panoramic image blind quality evaluation system based on AI-generated descriptions, and the system includes: A data generation module, configured to: Generate a quality description text based on the distorted panoramic image; Obtain image patches containing rich visual semantic information from the distorted panoramic image by using a positioning and cropping method; A pre-training module, configured to: Pre-train a vision-language model using the distorted panoramic image and the quality description text to obtain a pre-trained vision-language model; A feature extraction module, configured to: Extract the quality description text and the image patches containing rich visual semantic information by using the pre-trained vision-language model to obtain text features supervised by image information; A quality regression module, configured to: Extract the image patches by using an equidistant deformable convolution block to obtain image patch features; Concatenate the image patch features and the text features supervised by image information to obtain perceptually fused features; Perform quality regression on the perceptually fused features to obtain a quality prediction score.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By leveraging the powerful capabilities of multi-modal large language models, the present invention develops a set of methods that can generate quality descriptions for single-modal panoramic image datasets, construct a multi-modal database, thereby enriching the information representation for panoramic image quality evaluation; 2. The present invention directly performs patching operations on the original ERP images without the need for viewport extraction, thereby avoiding additional computational burdens and improving the generality of the model; 3. The present invention pre-trains a vision-language model by using the constructed multi-modal panoramic image database, thereby achieving efficient cross-modal feature extraction and fusion in the panoramic image quality evaluation task; 4. By introducing an equidistant deformable convolution module, the present invention can more accurately capture the local features of image patches and solve the problem of bipolar distortion in panoramic images; 5. By constructing a multi-modal feature fusion module, the present invention not only enhances the richness of feature representation but also effectively aligns the feature spaces of images and texts through an attention mechanism, ensuring the effective integration of multi-modal information.

[0008] Additional aspects and advantages of the present invention will be partially given in the following description, will become apparent in part from the following description, or can be learned through the embodiments of the present invention. Description of the Drawings

[0009] Figure 1 Flow chart of the multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions proposed by the present invention; Figure 2 Quality description generation and pre-training Long-CLIP framework of the multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions proposed by the present invention; Figure 3 Long-CLIP and equidistant deformable convolution block double-branch prediction framework of the multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions proposed by the present invention; Figure 4 Schematic diagram of the overall framework of the multi-modal panoramic image blind quality evaluation system based on AI-generated descriptions proposed by the present invention. Detailed implementation manners

[0010] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0011] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited by this.

[0012] Please refer to Figure 1 and Figure 2 , the embodiments of the present invention propose a multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions. The method includes the following steps: Step 1, generating a quality description text based on the distorted panoramic image; By adopting the method of positioning and cropping, image blocks containing rich visual semantic information are obtained from the distorted panoramic image; In Step 1, when generating a quality description text based on the distorted panoramic image, the relational formula existing in the corresponding process is: ; Among them, represents the image-text multi-modal data set, represents being processed by the multi-modal large language model, represents the th image in the input data set, represents the prompt text input to the multi-modal large language model, represents the number of images in the data set, Represents the image data in the dataset, Represents the text data in the dataset.

[0013] Furthermore, in this step, the distorted panoramic image is obtained by ERP projection, and the number and size of the sampled image patches are set, and 10 image patches with a size of 0.1H×0.05W are extracted; to prevent overfitting, 80% of the images in the multi-modal panoramic image database are used for training, while 20% are used for testing.

[0014] Step 2: Use the distorted panoramic image and the quality description text to pre-train the vision-language model to obtain the pre-trained vision-language model; In Step 2, use the distorted panoramic image and the quality description text to pre-train the vision-language model to obtain the pre-trained vision-language model, and the process includes the following sub-steps: Pair each distorted panoramic image and the quality description text and input them into the vision-language model, and learn the feature representation by comparing the differences and similarities between different image-text pairs. The corresponding process has the following relational expression: ; Among them, Represents the feature representation of the quality description text of the th image, Represents the feature representation of the th image, Represents the learnable temperature parameter, Represents the contrastive learning loss function of image-text, Represents the contrastive learning loss function of text-image, Represents the transpose symbol, Represents the feature representation of the quality description text of the th other image, Represents the th feature representation of other images, Represents the number of samples of other images; The final contrastive learning loss expression is as follows: ; Among them, Represents the final contrastive learning loss function.

[0015] Learn the feature representation by comparing the differences and similarities between different image-text pairs, so that the corresponding image-text pairs are close in the feature space, while the non-corresponding image-text pairs are far apart in the feature space. Finally, optimize the model parameters of the vision-language model by minimizing the final contrastive learning loss function to obtain the pre-trained vision-language model Step 3: Use the pre-trained vision-language model to extract the quality description text and the image patches containing rich visual semantic information, and obtain the text features supervised by the image information; In Step 3, use the pre-trained vision-language model to extract the quality description text and the image patches containing rich visual semantic information, and obtain the text features supervised by the image information. The specific steps are as follows: Mark the quality description text through the pre-trained vision-language model, and then use the text encoder to extract it to obtain the text features extracted by the pre-trained vision-language model. The relational expression for the corresponding process is: ; where, represents the text features extracted by the pre-trained vision-language model, is processed by the text encoder representing the pre-trained vision-language model, represents being processed by the text tokenizer of the pre-trained vision-language model, represents the pre-trained vision-language model; Use the image encoder of the pre-trained vision-language model to extract the perceptual features of the image patches to obtain the intermediate image patch features. The relational expression for the corresponding process is: ; where, represents the intermediate image patch features, represents being processed by the image encoder of the pre-trained vision-language model, represents the th image patch, represents the number of image patches; Concatenate and fuse all the obtained intermediate image patch features to obtain the fused image features. The relational expression for the corresponding process is: ; where, represents the fused image features, all represent the intermediate image patch features; Based on the text features extracted by the pre-trained vision-language model, obtain the key sequence and value sequence, and obtain the query sequence through the fused image features. The relational expression for the corresponding process is: ; where, represents being processed by a convolutional layer with a size of 1×1, represents the query sequence, represents the key sequence, represents the value sequence; Perform attention fusion on the query sequence, key sequence, and value sequence to obtain text features supervised by image information. The relational expression for the corresponding process is as follows: ; Among them, represents the text features supervised by image information, represents the scaling factor in the attention mechanism, represents being processed by the attention function.

[0016] In this step, the visual language model used is Long-CLIP; The size of the text features extracted by the pre-trained visual language model is 1×1×512; In the step of obtaining the intermediate image patch features, first downsample the image patch to a size of 224×224, and then input it into the Long-CLIP network; The size of the intermediate image patch features is 1×1×512; Reduce the dimension of the fused image features through a 1×1 convolutional layer, use the downsampled image features as the query sequence, upsample the text features extracted by the pre-trained visual language model through a 1×1 convolutional layer, and use the upsampled text features as the key sequence and value sequence, where the size of the query sequence is 1×2048, and the size of the key sequence and value sequence is 1×1×2048; The size of the text features supervised by image information is 1×1×2048.

[0017] Step 4: Use the equidistant deformable convolutional block to extract the image patch to obtain the image patch features; Concatenate the image patch features with the text features supervised by image information to obtain the concatenated and fused perceptual features; Perform quality regression on the concatenated and fused perceptual features to obtain the quality prediction score; Please refer to Figure 3 , in Step 4, use the equidistant deformable convolutional block to extract the image patch to obtain the image patch features. The specific steps are as follows: Use the equidistant deformable convolutional block to extract the perceptual features of the image patch. The relational expression for the corresponding process is as follows: ; Among them, represents the perceptual features of the image patch, represents being processed by the equidistant deformable convolutional block; Concatenate and fuse the perceptual features of all image patches to obtain the image patch features. The relational expression for the corresponding process is as follows: ; Among them, represents the image patch feature, both represent the perceptual features of the image patch; The image patch feature is concatenated with the text feature supervised by the image information to obtain the perceptual feature after concatenation and fusion. The relational expression existing in the corresponding process is: ; Among them, represents the perceptual feature after concatenation and fusion, represents the processing by global average pooling; The perceptual feature after concatenation and fusion is subjected to quality regression to obtain the quality prediction score. The relational expression existing in the corresponding process is: ; Among them, represents the quality prediction score, represents the processing by a linear layer, represents the processing by a feature flattening operation.

[0018] In this step, in the process of obtaining the perceptual feature of the image patch, the image patch is first downsampled to a size of 224×224, and then an isometric variability convolution block is used to extract the image patch feature, and the size of the perceptual feature of the image patch is 28×28×512; The size of the image patch feature is 28×28×5120.

[0019] Furthermore, in performing the above steps 1 to 4, the corresponding training method includes the following training steps: Obtain training data. The training data includes several image patches of the panoramic image and the quality description corresponding to the panoramic image. In this embodiment, 10 image patches are used. Using the training data as the input, repeat steps 1 and 2 to obtain the parameters of the pre-trained Long-CLIP, and then freeze the Long-CLIP parameters and repeat steps 1, 3, and 4 to obtain the predicted quality score; Use the mean square error as the loss function for quality score prediction. Use the subjective score and the predicted quality score to construct the loss function. The expression of the loss function is: ; Among them, represents the MSE loss function, represents the number of images in the dataset, represents the subjective score of the th image in the training data, represents the th predicted quality score of the image in the training data; The MSE loss function is input into the Adam optimizer for optimization. The weight decay strategy and learning parameters of the Adam optimizer are set, where the learning rate is set to 0.00005; for the weight decay strategy, the decay rate is ; The loss is minimized by updating the weights and learning parameters to improve the model performance.

[0020] Furthermore, the present invention evaluates the model performance in the following manner: Obtain the average subjective score of all image data in the multi-modal panoramic image database. The relational expression existing in the corresponding process is: ; where, represents the average subjective score of the image data, represents the quality of experience opinion score given by the th subject to the non-uniformly distorted panoramic picture, represents the number of experiments participating in evaluating the quality of non-uniformly distorted panoramic images; Compare the perceived quality score of the obtained panoramic image with the average subjective score of the image data, and calculate various indicators of the model. The indicators include the following three types: The prediction monotonicity index, including the Spearman correlation coefficient, and the expression is: ; where, represents the Spearman correlation coefficient, represents the difference between the subjective score and the objective prediction score of the th image; The prediction accuracy index, including the Pearson correlation coefficient, and the expression is: ; where, represents the Pearson correlation coefficient, represents the average value of the subjective scores, represents the average value of the objective prediction scores; The prediction error degree index, including the root mean square error, and the expression is: ; where, represents the root mean square error.

[0021] Please refer to Figure 4 , an embodiment of the present invention also provides a multi-modal panoramic image blind quality evaluation system based on AI-generated descriptions. The system includes: A data generation module, used for: Generate quality description text based on the distorted panoramic image; By adopting the method of positioning and cropping, image patches containing rich visual semantic information are obtained from the distorted panoramic image; The pre-training module is used for: Pre-train the vision-language model using the distorted panoramic image and the quality description text to obtain the pre-trained vision-language model; The feature extraction module is used for: Use the pre-trained vision-language model to extract the quality description text and the image patches containing rich visual semantic information to obtain text features supervised by image information; The quality regression module is used for: Use the equidistant deformable convolution block to extract the image patches to obtain image patch features; Concatenate the image patch features with the text features supervised by image information to obtain the perceptually fused features after concatenation; Perform quality regression on the perceptually fused features after concatenation to obtain a quality prediction score.

[0022] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0023] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0024] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions, characterized in that The method includes the following steps: Step 1: Generate a quality description text based on the distorted panoramic image; By adopting a positioning and cropping method, obtain image patches containing rich visual semantic information from the distorted panoramic image; Step 2: Use the distorted panoramic image and the quality description text to pre-train a vision-language model to obtain a pre-trained vision-language model; Step 3: Use the pre-trained vision-language model to extract the quality description text and the image patches containing rich visual semantic information to obtain text features supervised by image information; Step 4: Use an isometric deformable convolution block to extract the image patches to obtain image patch features; Concatenate the image patch features and the text features supervised by image information to obtain a concatenated and fused perceptual feature; Perform quality regression on the concatenated and fused perceptual feature to obtain a quality prediction score.

2. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 1, characterized in that, In the said Step 1, when generating a quality description text based on the distorted panoramic image, the relational expression existing in the corresponding process is: ; Among them, represents an image-text multimodal dataset, represents being processed by a multimodal large language model, represents the th image in the input dataset, represents the prompt text input to the multimodal large language model, represents the number of images in the dataset, represents the image data in the dataset, represents the text data in the dataset.

3. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 2, wherein, In the said Step 2, when using the distorted panoramic image and the quality description text to pre-train a vision-language model to obtain a pre-trained vision-language model, the following sub-steps are included in the process: Input each distorted panoramic image and the quality description text into the vision-language model in pairs, and learn feature representations by comparing the differences and similarities between different image-text pairs. The relational expression existing in the corresponding process is as follows: ; Among them, represents the feature representation of the quality description text of the th image, represents the feature representation of the th image, represents the learnable temperature parameter, represents the contrastive learning loss function of image-text, represents the contrastive learning loss function of text-image, represents the transpose symbol, represents the feature representation of the quality description text of the th other image, represents the feature representation of the th other image, represents the number of samples of other images; The final contrastive learning loss expression is as follows: ; Among them, represents the final contrastive learning loss function; Optimize the model parameters of the vision-language model by minimizing the final contrastive learning loss function to obtain a pre-trained vision-language model.

4. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 3, wherein, In the said Step 3, when using the pre-trained vision-language model to extract the quality description text and the image patches containing rich visual semantic information to obtain text features supervised by image information, the specific steps are as follows: Mark the quality description text through the pre-trained vision-language model, and then use a text encoder to extract it to obtain text features extracted by the pre-trained vision-language model. The relational expression existing in the corresponding process is: ; Among them, represents the text features extracted by the pre-trained vision-language model, processed by the text encoder of the pre-trained vision-language model, represents being processed by the text tokenizer of the pre-trained vision-language model, represents the pre-trained vision-language model; Use the image encoder of the pre-trained vision-language model to extract the perceptual features of the image patches to obtain intermediate image patch features. The relational expression existing in the corresponding process is: ; Among them, represents the intermediate image patch feature, represents being processed by the image encoder of the pre-trained vision-language model, represents the th image patch, represents the number of image patches; Concatenate and fuse all the obtained intermediate image patch features to obtain a fused image feature. The relational expression existing in the corresponding process is: ; Among them, represents the fused image feature, both represent the intermediate image block feature; Fuse the text features extracted by the pre-trained vision-language model and the fused image feature by using an attention mechanism to obtain text features supervised by image information.

5. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 4, characterized in that, The specific steps of fusing the text features extracted by the pre-trained vision-language model and the fused image feature by using an attention mechanism include the following: Based on the text features extracted by the pre-trained vision-language model, obtain a key sequence and a value sequence, and obtain a query sequence through the fused image feature. The relational expression existing in the corresponding process is: ; Among them, represents being processed by a convolutional layer with a size of 1×1, represents the query sequence, represents the key sequence, represents the value sequence; Perform attention fusion on the query sequence, the key sequence, and the value sequence to obtain text features supervised by image information. The relational expression existing in the corresponding process is: ; Among them, represents the text features supervised by the image information, represents the scaling factor in the attention mechanism, represents being processed by the attention function.

6. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 5, characterized in that, In the said Step 4, when using an isometric deformable convolution block to extract the image patches to obtain image patch features, the specific steps are as follows: The perceptual features of the image patches are extracted using an equidistant deformable convolution block, and the relational expression in the corresponding process is as follows: ; Among them, represents the perceptual feature of the image block, indicating being processed by the isometric deformable convolutional block; The perceptual features of all the image patches are concatenated and fused to obtain the image patch features, and the relational expression in the corresponding process is as follows: ; Among them, represents the feature of an image block, both represent the perceptual features of the image block.

7. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 6, wherein In step 4, the image patch features are concatenated with the text features supervised by the image information to obtain the concatenated and fused perceptual features, and the relational expression in the corresponding process is as follows: ; Among them, represents the perceived features after splicing and fusion, represents the processing by global average pooling.

8. The multi-modal panoramic image blind quality evaluation method based on AI-generated descriptions according to claim 7, characterized in that, In step 4, the concatenated and fused perceptual features are subjected to quality regression to obtain the quality prediction scores, and the relational expression in the corresponding process is as follows: ; Among them, represents the quality prediction score, represents being processed by a linear layer, represents being processed by a feature flattening operation.

9. A multi-modal panoramic image blind quality evaluation system based on AI-generated descriptions, characterized in that, The system applies any one of the multi-modal panoramic image blind quality evaluation methods based on AI-generated descriptions in claims 1 to 8 above. The system includes: A data generation module for: Generating a quality description text based on the distorted panoramic image; Obtaining image patches containing rich visual semantic information from the distorted panoramic image by using a positioning and cropping method; A pre-training module for: Pre-training a vision-language model using the distorted panoramic image and the quality description text to obtain a pre-trained vision-language model; A feature extraction module for: Extracting the quality description text and the image patches containing rich visual semantic information using the pre-trained vision-language model to obtain text features supervised by the image information; A quality regression module for: Extracting the image patches using an equidistant deformable convolution block to obtain the image patch features; Concatenating the image patch features with the text features supervised by the image information to obtain the concatenated and fused perceptual features; Performing quality regression on the concatenated and fused perceptual features to obtain the quality prediction scores.

Citation Information

Patent Citations

  • Vision-text and self-supervised feature extraction-based quality evaluation method

    CN117876818A

  • Non-viewport-dependent distortion-resistant reference-free panoramic image quality evaluation method and system

    CN118096770A

  • Unified visual language model pre-training and adjusting method for image quality and aesthetic evaluation

    CN118607611A

  • Full-reference panoramic image quality evaluation method and system based on block sequence similarity

    CN119006982A

  • Visual language multi-modal large language model pre-training method based on image-text staggering

    CN119377678A