AI Generation Panoramic Image Quality Evaluation Method and System Based on Visual-Language Correspondence

Through the visual encoder and language encoder based on the Transformer architecture, combined with random block sampling and feature fusion methods, the problems of perception and alignment in AI-generated panoramic image quality evaluation are solved, and efficient and accurate quality prediction is achieved.

CN119919423BActive Publication Date: 2025-07-18JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510425337.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing AI-generated panoramic image quality evaluation methods lack comprehensive quality perception, fail to consider both perception and alignment, and rely on cues for specific tasks, limit their generalization capabilities and ignore the correlation and interaction between different modal features.

Method used

Using a visual encoder and language encoder based on the Transformer architecture, image features are extracted through random block sampling method, combined with L2 normalization and cosine similarity calculation, mass scores are predicted using a fully connected network and normalized function, and the model is optimized by the average absolute error loss and Pearson correlation induced loss.

Benefits of technology

The efficiency of visual feature extraction is improved, the alignment accuracy between text features and visual features is enhanced, and the accuracy and semantic correlation of AI-generated panoramic image quality evaluation is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919423B_ABST
    Figure CN119919423B_ABST
Patent Text Reader

Abstract

The present invention proposes an AI-generated panoramic image quality evaluation method and system based on visual-language correspondence. The method includes: obtaining an AI-generated panoramic image and sampling the AI-generated panoramic image; based on the set of image patches, using a visual encoder to represent the features of the image patches; using a language encoder to represent the features of the text description attached to the AI-generated panoramic image; performing L2 normalization processing and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence; using a fully connected network and a normalization function to process the fused feature vectors. The present invention uses visual-language correspondence analysis to perform joint analysis on the AI-generated panoramic image and its corresponding text description, and uses the learned visual-language correspondence relationship to efficiently and accurately predict the quality score of the AI-generated panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and multimedia digital image processing, and particularly relates to an AI-generated panoramic image quality evaluation method and system based on visual language correspondence. Background Art

[0002] The popularization of artificial intelligence (AI) technology, especially generative artificial intelligence and large language models (LLMs), has had a significant impact on all aspects of human life. Among these developments, AI-generated images (AGIs) are created by adopting text-to-image (T2I) generation models, including generative adversarial networks, variational autoencoders, and diffusion-based models. AI-generated omnidirectional images (AGOIs) are a specific category of AGIs and have great potential in virtual reality (VR) applications. Compared with two-dimensional images, AGOI has the attributes of generated content and omnidirectional images. However, there are currently some problems with AGOI, such as unrealistic structures, inappropriate combinations, etc. Therefore, there is an urgent need for an objective method to evaluate the quality of AGOI.

[0003] Recent research has utilized contrastive language-image pre-training models to enhance AGOI quality evaluation by simulating human preferences and evaluating visual-text correspondence. However, these methods have significant limitations. First, most methods for predicting human preferences for AGOI lack a comprehensive quality perception and fail to consider both perception and alignment simultaneously. Second, although measuring text-to-image alignment can indicate consistency, these methods rely heavily on task-specific prompting strategies (e.g., antonym prompt pairing, text words, and text templates), which limits their generalization ability and hinders unprompted image generation tasks. Current methods usually ignore the correlation and interaction between different modal features, resulting in unsatisfactory performance. Summary of the Invention

[0004] In view of the above situation, the main objective of the present invention is to propose an AI-generated panoramic image quality evaluation method and system based on visual language correspondence to solve the above technical problems.

[0005] The present invention proposes an AI-generated panoramic image quality evaluation method based on visual language correspondence, and the method includes the following steps:

[0006] Step 1: Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multi-modal feature fusion module based on the feature fusion mechanism. The language encoder, visual encoder, and multi-modal feature fusion module constitute a quality evaluation model;

[0007] Step 2: Obtain an AI-generated panoramic image, and sample the AI-generated panoramic image to obtain a set of image patches;

[0008] Step 3: Based on the set of image patches, use the visual encoder to represent the features of the image patches to obtain the visual features of the image patches;

[0009] Step 4: Use the language encoder to represent the features of the text description attached to the AI-generated panoramic image to obtain the text features of the text description;

[0010] Step 5: Perform L2 normalization processing and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain a fused feature vector;

[0011] Step 6: Use a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score;

[0012] Construct a mean absolute error loss and a Pearson correlation-induced loss based on the predicted quality score, and use the mean absolute error loss and the Pearson correlation-induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model;

[0013] Use the optimized quality evaluation model to obtain an evaluation result.

[0014] The present invention also proposes an AI-generated panoramic image quality evaluation system based on visual-language correspondence. The system includes:

[0015] A random block sampling module, which is used for:

[0016] Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multi-modal feature fusion module based on the feature fusion mechanism. The language encoder, visual encoder, and multi-modal feature fusion module constitute a quality evaluation model;

[0017] Obtain an AI-generated panoramic image, and sample the AI-generated panoramic image to obtain a set of image patches;

[0018] A visual feature extraction module, which is used for:

[0019] Based on the set of image patches, use the visual encoder to represent the features of the image patches to obtain the visual features of the image patches;

[0020] A text prompt embedding module:

[0021] Use a language encoder to perform feature representation on the text description attached to the AI-generated panoramic image to obtain the text features of the text description;

[0022] A multi-modal feature fusion module for:

[0023] Perform L2 normalization processing and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain a fused feature vector;

[0024] A quality regression module for:

[0025] Use a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score;

[0026] Construct an average absolute error loss and a Pearson correlation-induced loss based on the predicted quality score, and use the average absolute error loss and the Pearson correlation-induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model;

[0027] Obtain an evaluation result using the optimized quality evaluation model.

[0028] Compared with the prior art, the beneficial effects of the present invention are:

[0029] 1. Based on the random block sampling method independent of the viewport, the present invention efficiently extracts visual features while retaining local texture details and taking into account the overall spatial structure, thus reflecting the quality of the original image to the greatest extent and greatly improving the efficiency of visual feature extraction and calculation;

[0030] 2. By introducing a causal mask matrix and positional encoding in the language encoder, the present invention not only ensures the autoregressive generation characteristics of the text sequence but also accurately models the word order dependence relationship, thus significantly improving the alignment accuracy between text features and visual features and further enhancing the semantic relevance of quality evaluation;

[0031] 3. By using visual language correspondence analysis, the present invention performs a linkage analysis on the AI-generated panoramic image and its corresponding text description, and uses the learned visual language correspondence relationship to efficiently and accurately predict the quality score of the AI-generated panoramic image.

[0032] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flowchart of the method for evaluating the quality of AI-generated panoramic images based on visual language correspondence proposed by the present invention;

[0034] Figure 2 This is the algorithm framework diagram of the AI-generated panoramic image quality evaluation method corresponding to visual language proposed by the present invention;

[0035] Figure 3 This is the system framework diagram of the AI-generated panoramic image quality evaluation system corresponding to visual language proposed by the present invention. Detailed implementation manners

[0036] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0037] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0038] Please refer to Figure 1 , this embodiment provides an AI-generated panoramic image quality evaluation method corresponding to visual language, and the method includes the following steps:

[0039] Step 1: Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multi-modal feature fusion module based on a feature fusion mechanism. The language encoder, the visual encoder, and the multi-modal feature fusion module constitute a quality evaluation model.

[0040] Step 2: Obtain an AI-generated panoramic image, and sample the AI-generated panoramic image to obtain a set of image patches.

[0041] Please refer to Figure 2 , in Step 2, obtaining an AI-generated panoramic image and sampling the AI-generated panoramic image to obtain a set of image patches specifically includes the following sub-steps:

[0042] Perform random block sampling on the AI-generated panoramic image in ERP format to obtain a set of image patches , and there is the following relational expression in the corresponding process:

[0043] ;

[0044] Wherein, represents the th image patch and , represents the random sampling operation, represents the sampled AI-generated panoramic image and , represents the spatial coordinates of the upper left corner of the th image patch, represents the height of the image patch, represents the width of the image patch, represents the index of the image patch, represents the image height, represents the image width, represents the number of image channels, represents the total number of image patches.

[0045] It should be noted that the attachment Figure 2 in represents matrix multiplication.

[0046] Furthermore, the present invention samples the AI-generated panoramic image by a random block sampling method based on viewport-independent characteristics. To ensure that the sampling range remains within the image boundaries, the spatial coordinates of the upper left corner of each image patch

[0047] ;

[0048] wherein, represents a uniform distribution within the range of and .

[0049] This embodiment is carried out on the AIGCOIQA2024 database, which is the first AGOI (AI-generated panoramic image) subjective database, consisting of 300 AGOIs generated by five AIGC models according to 25 text prompts. Human visual preferences for each AGOI are collected from three perspectives: quality, comfort, and correspondence, and all preferences range from 0 to 100. The database is divided into 60% for training and 40% for testing according to the text prompts. To reduce the bias caused by random splitting, it is repeated ten times and the average effect is taken. The random sampling strategy adopted by the present invention can effectively capture the spatial and texture quality of the image, while significantly improving the model efficiency and minimizing the computational cost and model complexity.

[0050] Step 3: Based on the set of image patches, use a visual encoder to represent the features of the image patches to obtain the visual features of the image patches.

[0051] In step 3, based on the set of image patches, use a visual encoder to represent the features of the image patches to obtain the visual features of the image patches. The present invention uses a VisionTransformer feature encoder as the visual encoder. The visual encoder, when operating on the set of image patches In the image patch when performing feature extraction, it specifically includes the following sub-steps:

[0052] Flatten the image patch into a vector and map it to a high-dimensional embedding space to obtain the vector of the image patch mapped to the high-dimensional embedding space. The following relational expressions exist in the corresponding process:

[0053] ;

[0054] where represents the vector of the th image patch mapped to the high-dimensional embedding space, represents the linear projection matrix, represents the bias term.

[0055] Configure the learnable position encoding for the vector of the image patch mapped to the high-dimensional embedding space to construct a position encoding matrix, and use the position encoding matrix to perform an embedding operation on the vector of the image patch mapped to the high-dimensional embedding space to obtain an embedding vector fused with position information. The following relational expressions exist in the corresponding process:

[0056] ;

[0057] where represents the embedding vector corresponding to the th image patch fused with position information, represents the position encoding corresponding to the th image patch.

[0058] It should be noted that adding the learnable position encoding is to retain the position information of the image patch.

[0059] Send the embedding vector fused with position information into the Trnsformaer encoder for feature extraction to obtain the visual features of the image patch. The following relational expressions exist in the corresponding process:

[0060] ;

[0061] where represents the visual feature of the th image patch, represents the Transformer feature encoder with parameter .

[0062] Step 4: Use the language encoder to perform feature representation on the text description attached to the AI-generated panoramic image to obtain the text features of the text description.

[0063] In step 4, a language encoder is used to represent the text description attached to the AI-generated panoramic image in terms of features, so as to obtain the text features of the text description. In the present invention, GPT-2 is adopted as the language encoder, and its generative architecture is more suitable for open-ended text descriptions. In the process of the language encoder extracting features, the following sub-steps are specifically included:

[0064] Decompose the text description attached to the AI-generated panoramic image (such as "The panoramic view shows a dining area with a classic chandelier and a round table, next to a cozy living room with a fireplace and comfortable seats.") to obtain tokens, and the following relational expressions exist in the corresponding process:

[0065] ;

[0066] Among them, represents the text description attached to the AI-generated panoramic image, token represents a text processing unit, represents the total number of tokens, both represent tokens.

[0067] Map the tokens to a high-dimensional space by using an embedding matrix to obtain word vectors, and the following relational expressions exist in the corresponding process:

[0068] ;

[0069] Among them, represents the embedding matrix and , represents the embedding matrix operation, represents the th token, represents the th token corresponding word vector, represents the index of the token, represents the vocabulary size, represents the embedding dimension.

[0070] Configure the position encoding for the word vectors to construct a position encoding matrix, and use the position encoding matrix to perform an embedding operation on the word vectors to obtain word vectors fused with position encoding, and the following relational expressions exist in the corresponding process:

[0071] ;

[0072] Among them, represents the word vector corresponding to the th token of the fused position encoding, represents the th token corresponding position encoding and 。

[0073] Integrate the word vectors of the fused position encoding to obtain the input sequence of the fused position encoding. The following relational expressions exist in the corresponding process:

[0074] ;

[0075] Among them, represents the input sequence of the fused position encoding, both represent the word vectors of the fused position encoding.

[0076] Input the input sequence of the fused position encoding into the Transformer encoder for layer encoding processing to obtain the feature sequence. The following relational expressions exist in the corresponding process:

[0077] ;

[0078] Among them, represents the feature sequence after being processed by layers of the Transformer encoder.

[0079] Each layer of the Transformer encoder includes masked multi-head self-attention, a feed-forward neural network, and layer normalization;

[0080] When performing masked multi-head self-attention calculation, the following relational expressions exist in the corresponding process:

[0081] ;

[0082] Among them, represents the causal mask matrix, ensuring that the model can only access the current and previous tokens; represents the Query matrix, represents the Key matrix, represents the Value matrix.

[0083] Process the feature sequence using a trainable projection matrix to obtain the text features. The following relational expressions exist in the corresponding process:

[0084] ;

[0085] Among them, represents the text features and , represents the trainable projection matrix.

[0086] It should be noted that the trainable projection matrix is to ensure that the text features and the visual features Having the same feature dimension .

[0087] Step 5: Perform L2 normalization and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain the fused feature vector.

[0088] In Step 5, performing L2 normalization and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain the fused feature vector specifically includes the following sub-steps:

[0089] Perform cosine similarity calculation on the visual features and the text features to obtain the cosine similarity between the visual features and the text features. There are the following relational expressions in the corresponding process:

[0090] ;

[0091] where represents the cosine similarity between the visual features and the text features of the th image patch, represents the visual features of the th image patch, represents the text features, represents the dot product of vectors, represents the 2-norm, represents the total number of image patches.

[0092] Perform L2 normalization on the visual features to obtain the normalized visual features. There are the following relational expressions in the corresponding process:

[0093] ;

[0094] where represents the normalized visual features of the th image patch.

[0095] Perform L2 normalization on the text features to obtain the normalized text features. There are the following relational expressions in the corresponding process:

[0096] ;

[0097] where represents the normalized text features.

[0098] It should be noted that after performing L2 normalization on the visual features and the text features in sequence, and then performing cosine similarity calculation on the normalized visual features and the normalized text features, the influence of the feature amplitude difference on the cosine similarity calculation can be eliminated.

[0099] Calculate the cosine similarity between the normalized visual features and the normalized text features to obtain the cosine similarity between the enhanced visual features and the text features. The following relational expressions exist in the corresponding process:

[0100] ;

[0101] Among them, represents the cosine similarity between the enhanced visual features and the text features, represents the transpose of the normalized text features.

[0102] In the step of calculating the average value of the cosine similarity between the enhanced visual features and the text features to obtain the fused feature vector, the following relational expressions exist in the corresponding process:

[0103] ;

[0104] Among them, represents the fused feature vector.

[0105] Step 6: Process the fused feature vector using a fully connected network and a normalization function to obtain a predicted quality score;

[0106] Construct a mean absolute error loss and a Pearson correlation induced loss based on the predicted quality score, and use the mean absolute error loss and the Pearson correlation induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model;

[0107] Obtain an evaluation result using the optimized quality evaluation model.

[0108] In step 6, processing the fused feature vector using a fully connected network and a normalization function to obtain a predicted quality score specifically includes the following sub-steps:

[0109] Process the fused feature vector using a fully connected network to obtain an output score. The following relational expressions exist in the corresponding process:

[0110] ;

[0111] Among them, represents the output score after the fused feature vector passes through the fully connected layer, represents the weight matrix of the fully connected layer.

[0112] Normalize the output score after the fused feature vector passes through the fully connected layer using a normalization function to obtain a predicted quality score. The following relational expressions exist in the corresponding process:

[0113] ;

[0114] Among them, represents the predicted quality score, indicating that it has been processed by a general normalization method.

[0115] Furthermore, in order to optimize the parameters in the quality evaluation model, the present invention uses the mean absolute error loss and the Pearson correlation induced loss as loss functions to balance the coverage speed and prediction accuracy. The total loss is the weighted sum of the mean absolute error loss and the Pearson correlation induced loss. There is the following relational expression in the corresponding process:

[0116] ;

[0117] Among them, represents the total loss, represents the mean absolute error loss, represents the Pearson correlation induced loss, represents a hyperparameter.

[0118] It should be noted that is used to measure the difference between the predicted score and the true human preference score, is used to ensure the linearity and monotonicity of the predicted score, is used to balance and 's contributions.

[0119] The formula for the mean absolute error is as follows:

[0120] ;

[0121] Among them, represents the number of samples, represents the true value of the th sample (from the labeled data), represents the predicted value of the th sample (from the model output);

[0122] The formula for the Pearson correlation induced loss is as follows:

[0123] ;

[0124] Among them, represents the Pearson correlation coefficient, represents the average value of the subjective scores, represents the average value of the objective predicted scores;

[0125] From the above description, the total loss can be expressed as the following relational expression:

[0126] ;

[0127] Among them, represents the average value of the true values, represents the average value of the predicted values;

[0128] By comparing and calculating the average prediction result with the subjective evaluation score, various indicators of the model can be obtained. Among them, the subjective evaluation score is given by professional scorers from the constructor of the AIGCOIQA2024 database. The test indicators include the following three types:

[0129] The prediction monotonicity index, including the Spearman correlation coefficient (SRCC), is specifically expressed as:

[0130] ;

[0131] Among them, represents the Spearman correlation coefficient, represents the number of distorted images in the data, represents the th difference between the subjective score and the objective prediction score of the

[0132] The prediction accuracy index, including the Pearson correlation coefficient (PLCC), is specifically expressed as:

[0133] ;

[0134] Among them, represents the th subjective score of the image, represents the th objective prediction score of the image;

[0135] The prediction ranking correlation index, including the Kendall rank correlation coefficient (KRCC), is specifically expressed as:

[0136] ;

[0137] Among them, represents the number of pairs with inconsistent ranks in the paired data, represents the number of pairs with consistent ranks in the paired data, represents the total number of data.

[0138] Please refer to Figure 3 , this embodiment also provides an AI-generated panoramic image quality evaluation system based on vision-language correspondence. The system applies the above-mentioned AI-generated panoramic image quality evaluation method based on vision-language correspondence. The system includes:

[0139] A random block sampling module for:

[0140] Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multi-modal feature fusion module based on a feature fusion mechanism. The language encoder, the visual encoder, and the multi-modal feature fusion module constitute a quality evaluation model;

[0141] Obtain an AI-generated panoramic image and sample the AI-generated panoramic image to obtain a set of image patches;

[0142] A visual feature extraction module, which is used for:

[0143] Based on the set of image patches, use the visual encoder to perform feature representation on the image patches to obtain the visual features of the image patches;

[0144] A text prompt embedding module:

[0145] Use the language encoder to perform feature representation on the text description attached to the AI-generated panoramic image to obtain the text features of the text description;

[0146] A multi-modal feature fusion module, which is used for:

[0147] Perform L2 normalization processing and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain a fused feature vector;

[0148] A quality regression module, which is used for:

[0149] Use a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score;

[0150] Construct a mean absolute error loss and a Pearson correlation-induced loss based on the predicted quality score, and use the mean absolute error loss and the Pearson correlation-induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model;

[0151] Use the optimized quality evaluation model to obtain an evaluation result.

[0152] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0153] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. An AI-generated panoramic image quality evaluation method based on vision-language correspondence, characterized in that, The method includes the following steps: Step 1: Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multimodal feature fusion module based on a feature fusion mechanism. The language encoder, the visual encoder, and the multimodal feature fusion module form a quality evaluation model; Step 2: Obtain an AI-generated panoramic image, and sample the AI-generated panoramic image to obtain a set of image patches; Step 3: Based on the set of image patches, use the visual encoder to represent the features of the image patches to obtain the visual features of the image patches; Step 4: Use the language encoder to represent the features of the text description attached to the AI-generated panoramic image to obtain the text features of the text description; Step 5: Perform L2 normalization processing and cosine similarity calculation on the visual features of the image patches and the text features of the text description in sequence to obtain a fused feature vector; Step 6: Use a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score; Construct a mean absolute error loss and a Pearson correlation-induced loss based on the predicted quality score, and use the mean absolute error loss and the Pearson correlation-induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model; Use the optimized quality evaluation model to obtain an evaluation result; Wherein, in the said Step 4, using the language encoder to represent the features of the text description attached to the AI-generated panoramic image to obtain the text features of the text description specifically includes the following sub-steps: Decompose the text description attached to the AI-generated panoramic image to obtain tokens; Use an embedding matrix to map tokens to a high-dimensional space to obtain word vectors; Configure positional encoding for the word vectors to construct a positional encoding matrix, and use the positional encoding matrix to perform an embedding operation on the word vectors to obtain word vectors with fused positional encoding; Integrate the word vectors with fused positional encoding to obtain an input sequence with fused positional encoding; Input the input sequence with fused positional encoding into the Transformer encoder for layer encoding processing to obtain a feature sequence; Use a trainable projection matrix to process the feature sequence to obtain text features; Decompose the text description attached to the AI-generated panoramic image to obtain tokens. There is the following relational expression during the corresponding process ; Among them, represents the text description attached to the AI-generated panoramic image, and token represents the text processing unit, represents the total number of tokens, both represent tokens; In the step of using an embedding matrix to map tokens to a high-dimensional space to obtain word vectors, there is the following relational expression in the corresponding process: ; Among them, represents the embedding matrix, represents the embedding matrix operation, represents the th token, represents the th word vector corresponding to the token, represents the index of the token; In the step of configuring positional encoding for the word vectors to construct a positional encoding matrix and using the positional encoding matrix to perform an embedding operation on the word vectors to obtain word vectors with fused positional encoding, there is the following relational expression in the corresponding process: ; Among them, represents the word vector corresponding to the -th token of the fused position encoding, and represents the position encoding corresponding to the -th token; In the step of integrating the word vectors with fused positional encoding to obtain an input sequence with fused positional encoding, there is the following relational expression in the corresponding process; ; Among them, represents the input sequence of the fused positional encoding, both represent the word vectors of the fused positional encoding; After inputting the input sequence with fused positional encoding into the Transformer encoder for layer encoding processing to obtain a feature sequence, the following relational expressions exist in the corresponding process: ; Among them, represents the feature sequence processed by layers of Transformer encoders, represents the Transformer feature encoder with parameters ; In the step of using a trainable projection matrix to process the feature sequence to obtain text features, there is the following relational expression in the corresponding process: ; Among them, represents text features, represents a causal mask matrix, represents a trainable projection matrix; Wherein, in the said Step 6, using a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score specifically includes the following sub-steps: Use a fully connected network to process the fused feature vector to obtain an output score, and there is the following relational expression in the corresponding process: ; Among them, represents the output score after the fused feature vector passes through the fully connected layer, represents the weight matrix of the fully connected layer, represents the fused feature vector, represents the bias term; Use a normalization function to perform normalization processing on the output score after the fused feature vector passes through the fully connected layer to obtain a predicted quality score, and there is the following relational expression in the corresponding process: ; Among them, represents the predicted mass fraction, represents being processed by a general normalization method.

2. The AI-generated panoramic image quality evaluation method based on visual language correspondence according to claim 1, characterized in that, In step 2, obtain the AI-generated panoramic image, sample the AI-generated panoramic image to obtain a set of image patches, and there are the following relational expressions in the corresponding process: ; Among them, represents the th image patch, represents the random sampling operation, represents the AI-generated panoramic image to be sampled, represents the spatial coordinates of the upper left corner of the th image patch, represents the height of the image patch, represents the width of the image patch, represents the index of the image patch.

3. The AI-generated panoramic image quality evaluation method based on visual language correspondence according to claim 2, characterized in that, In step 3, based on the set of image patches, use a visual encoder to represent the features of the image patches to obtain the visual features of the image patches, which specifically includes the following sub-steps: Flatten the image patch into a vector and map it to a high-dimensional embedding space to obtain the vector of the image patch mapped to the high-dimensional embedding space; Configure learnable positional encoding for the vector of the image patch mapped to the high-dimensional embedding space to construct a positional encoding matrix, and use the positional encoding matrix to perform an embedding operation on the vector of the image patch mapped to the high-dimensional embedding space to obtain an embedded vector fused with positional information; Send the embedded vector fused with positional information into a Transformer encoder for feature extraction to obtain the visual features of the image patch.

4. The AI-generated panoramic image quality evaluation method based on visual language correspondence according to claim 3, characterized in that, Flatten the image patch into a vector and map it to a high-dimensional embedding space to obtain the vector of the image patch mapped to the high-dimensional embedding space, and there are the following relational expressions in the corresponding process: ; Among them, represents the vector of the $i$-th image patch mapped to the high-dimensional embedding space, represents the linear projection matrix; In the step of configuring learnable positional encoding for the vector of the image patch mapped to the high-dimensional embedding space to construct a positional encoding matrix, and using the positional encoding matrix to perform an embedding operation on the vector of the image patch mapped to the high-dimensional embedding space to obtain an embedded vector fused with positional information, there are the following relational expressions in the corresponding process: ; Among them, represents the embedding vector corresponding to the th image patch indicating the fusion position information, represents the positional encoding corresponding to the th image patch; In the step of sending the embedded vector fused with positional information into a Transformer encoder for feature extraction to obtain the visual features of the image patch, there are the following relational expressions in the corresponding process: ; Among them, represents the visual feature of the 5. The AI-generated panoramic image quality evaluation method corresponding to visual language according to claim 1, wherein In step 5, perform L2 normalization processing and cosine similarity calculation on the visual features of the image patch and the text features of the text description in sequence to obtain a fused feature vector, which specifically includes the following sub-steps: Calculate the cosine similarity between the visual features and the text features to obtain the cosine similarity between the visual features and the text features; Perform L2 normalization processing on the visual features to obtain the normalized visual features; Perform L2 normalization processing on the text features to obtain the normalized text features; Calculate the cosine similarity between the normalized visual features and the normalized text features to obtain the enhanced cosine similarity between the visual features and the text features; Calculate the average value of the enhanced cosine similarity between the visual features and the text features to obtain a fused feature vector.

6. The AI-generated panoramic image quality evaluation method corresponding to visual language according to claim 5, wherein Calculate the cosine similarity between the visual features and the text features to obtain the cosine similarity between the visual features and the text features, and there are the following relational expressions in the corresponding process: ; Among them, represents the cosine similarity between the visual features and text features of the th image patch, represents the visual features of the th image patch, represents the text features, represents the dot product of vectors, represents the L2 norm, represents the total number of image patches; In the step of performing L2 normalization processing on the visual features to obtain the normalized visual features, there are the following relational expressions in the corresponding process: ; Among them, represents the visual feature of the th image patch after normalization; In the step of performing L2 normalization processing on the text features to obtain the normalized text features, there are the following relational expressions in the corresponding process: ; Among them, represents the text features after normalization; In the step of calculating the cosine similarity between the normalized visual features and the normalized text features to obtain the enhanced cosine similarity between the visual features and the text features, there are the following relational expressions in the corresponding process: ; Among them, represents the cosine similarity between the enhanced visual feature and the text feature, represents the transpose of the normalized text feature; In the step of calculating the average value of the cosine similarity between the enhanced visual features and the text features to obtain the fused feature vector, the following relational expressions exist in the corresponding process: 。 7. An AI-generated panoramic image quality evaluation system based on visual language correspondence, characterized in that, The system applies the AI-generated panoramic image quality evaluation method based on visual language correspondence described in any one of claims 1 to 6. The system includes: A random block sampling module, configured to: Construct a visual encoder and a language encoder based on the Transformer architecture, and construct a multimodal feature fusion module based on the feature fusion mechanism. The language encoder, the visual encoder, and the multimodal feature fusion module constitute a quality evaluation model; Obtain an AI-generated panoramic image, and sample the AI-generated panoramic image to obtain a set of image blocks; A visual feature extraction module, configured to: Based on the set of image blocks, use the visual encoder to represent the features of the image blocks to obtain the visual features of the image blocks; A text prompt embedding module: Use the language encoder to represent the features of the text description attached to the AI-generated panoramic image to obtain the text features of the text description; A multimodal feature fusion module, configured to: Perform L2 normalization processing and cosine similarity calculation on the visual features of the image blocks and the text features of the text description in sequence to obtain a fused feature vector; A quality regression module, configured to: Use a fully connected network and a normalization function to process the fused feature vector to obtain a predicted quality score; Construct an average absolute error loss and a Pearson correlation-induced loss based on the predicted quality score, and use the average absolute error loss and the Pearson correlation-induced loss to optimize the quality evaluation model to obtain an optimized quality evaluation model; Use the optimized quality evaluation model to obtain an evaluation result.

Citation Information

Patent Citations

  • No-reference image quality evaluation method based on image features and semantic description

    CN118608467A

  • No-reference quality evaluation method based on visual compensation perception

    CN118784826A

  • Structured data modeling analysis method based on Transform

    CN119577402A