Generated image quality evaluation method based on pre-training cross-modal feature alignment embedding

By using a pre-trained cross-modal feature alignment embedding method, combined with multi-granularity decomposition and dynamic quality prototyping, the consistency and local region evaluation problems in generated image quality assessment are solved, achieving higher accuracy and wider applicability of image quality assessment.

CN120894650APending Publication Date: 2025-11-04ANHUI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511057271.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of generated images fail to adequately consider the consistency between the image and the text prompts, struggle to identify visual distortions in local areas, and are unable to adapt to quality evaluation in various generated image scenarios, resulting in limited evaluation accuracy.

Method used

We employ a pre-trained cross-modal feature alignment embedding method, which integrates image and text information through multi-granularity decomposition and dynamic quality prototype construction to build a multi-granularity image representation. We then utilize the BLIP model for visual question answering and quality assessment to form a dynamic quality prototype, which is finally mapped to a quality assessment score through a dual fully connected layer.

Benefits of technology

It significantly improves the accuracy and generalization ability of generated image quality assessment, can accurately assess the consistency between images and text, adapts to various generated image scenarios, and expands the practical application boundaries of image quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894650A_ABST
    Figure CN120894650A_ABST
Patent Text Reader

Abstract

The invention relates to a generated image quality evaluation method based on pre-training cross-modal feature alignment embedding, and compared with the prior art, the method solves the defect of limited quality evaluation precision caused by neglecting generation prompt information and neglecting local view quality degradation in the traditional image quality evaluation process. The method comprises the following steps: acquiring and generating an image quality evaluation data set; performing image multi-granularity decomposition; constructing a dynamic quality prototype; constructing a quality evaluation model; training a quality evaluation model; and obtaining a quality evaluation result. In a cross-modal feature alignment space constructed by a pre-training large model, object display index region division of an image is completed, dynamic quality prototype fitting data quality manifold distribution is constructed, an enhanced expression statement is constructed in combination with text prompt information when the image is generated, the consistency inspection of image content and text prompt is completed, and the image content and text prompt consistency is improved. And meanwhile, global view quality evaluation is considered, and the image quality evaluation task precision is remarkably improved in multiple scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of generating image quality evaluation, in particular to a method for generating image quality evaluation based on pre-training cross-modal feature alignment embedding. BACKGROUND

[0002] With the continuous development of general artificial intelligence technology, artificial intelligence generated content has gradually become a key bridge for new generation human-computer interaction. The current AI generated image technology is still significantly different in generation effect due to the influence of model architecture, inference strategy, etc. The volatility of such image quality not only affects the user experience, but also may cause misjudgment or mislead in practical application. Therefore, it is necessary to establish a scientific, automatic and scalable AI generated image quality evaluation framework.

[0003] Image quality evaluation has always been an important topic in the field of computer vision. In the past few decades, researchers have proposed many image quality evaluation methods. Generally, they can be divided into methods based on shallow statistical features and methods based on deep visual features. Methods based on shallow statistical features, such as SSIM and MSCN, fit the human perception of image visual quality indicators through manually designed natural statistical laws. However, the manually designed method is sometimes too complex and single in effect, resulting in its lack of wide application in image evaluation tasks.

[0004] With the development of deep learning technology, image quality evaluation methods based on CNN and Transformer have gradually become a hot topic in the field, with important practical value. More and more researchers have gradually entered this field. The typical approach is to capture high-level semantic information of images through deep networks, build rich global representations, and map visual quality features to image quality ratings through mapping relationships, so that the network has a certain visual quality perception ability consistent with humans.

[0005] However, the current research method generally faces the following key problems. First, the generated image is different from the natural scene image. The generated image needs to use the user to provide the prompt information to the AI tool, and then the AI processing is generated. Therefore, the content consistency between the prompt information and the generated image is a key factor for quality evaluation. The existing mainstream quality evaluation method often only considers the visual attributes of the image itself, and ignores the consistency of the prompt information, resulting in deviation of the image quality evaluation result. Second, with the development of the generated model, the existing technology generated image often has rich information and complex image content. The existing image quality evaluation method based on deep network maps the visual quality features to the quality score, and cannot sensitively pay attention to the local area visual distortion in the image, such as the local area may have extremely serious distortion, blur and meaningless content, which will greatly affect the overall visual quality. In order to adapt to the requirements of the network input window, some methods cut the image through sliding window to guide the model to pay attention to the local area difference, but this method may damage the semantic integrity of the image, causing the model evaluation error. Third, unlike the image quality evaluation in a single scene, the AI generated image is various in content, form and distortion, and a robust and robust model is needed to adapt to different distortion conditions, accurately evaluate the quality of the generated image, and be generalized to multiple scenes, which has practical application value. SUMMARY

[0006] In order to solve the problem of not considering the prompt information in the existing technology of generated image quality evaluation and the problem of local area quality insensitivity, which limits the precision of image quality evaluation, the present application proposes a generated image quality evaluation method based on pre-training cross-modal feature alignment embedding.

[0007] In order to achieve the above purpose, the technical scheme of the present application is as follows:

[0008] A generated image quality evaluation method based on pre-training cross-modal feature alignment embedding, characterized in that the method comprises the following steps:

[0009] S1, obtaining a generated image quality evaluation data set; the data set includes text prompt information, generated image and image quality score;

[0010] S2, using a pre-trained multi-modal large model to perform image multi-granularity decomposition, constructing a multi-granularity image representation containing global view granularity and local object subgraph granularity, and forming an image element matrix;

[0011] S3, using the pre-trained multi-modal large model to perform visual question answering on the generated images to obtain the visual quality of each image, to obtain the quality distribution of the generated images, and to extract hidden layer features in the visual question answering process for clustering to construct quality prototype vocabulary for evaluating image quality and form dynamic quality prototypes;

[0012] S4, constructing a quality evaluation model based on the image element matrix and the dynamic quality prototype to convert the quality evaluation problem into a classification problem to obtain the quality grade division probability of each element in the image element matrix, and mapping the quality evaluation score through a double fully connected layer;

[0013] S5, training the quality evaluation model to obtain a model weight file, loading the model weight file into the trained quality evaluation model, and inputting the image to be evaluated and the text prompt information corresponding to the image into the model, and the model outputs the quality evaluation result of the image to be evaluated.

[0014] Further, the step S1 specifically comprises the following steps:

[0015] S11, constructing text prompt information;

[0016] S12, generating corresponding images based on the text prompt information using a diffusion generation model; the diffusion generation model is Dalle-E or Midjourney;

[0017] S13, scoring the quality of all generated images;

[0018] S14, forming a data set containing text prompt information, generated images, and image quality scores as a generated image quality evaluation data set.

[0019] Further, the multi-modal large model adopts a BILP model.

[0020] Further, the step S2 specifically comprises the following steps:

[0021] S21, constructing an image multi-granularity decomposition module based on the pre-trained multi-modal large model;

[0022] The input of the image multi-granularity decomposition module is the text prompt information and the original image generated according to the text prompt information , the output is a global view granularity and a local object subgraph granularity set , both of which are constructed as a multi-granularity subgraph set corresponding to the original image , wherein the subscript represents the image sample index, and the superscript represents the global level superscript Represents object level ;

[0023] S22. Constructing a set of local object subgraph granularities Using a dependency parser to analyze text prompts Perform structured parsing to extract a set of object-level lexical items with clear semantic meaning. , where subscript Represents an object-level vocabulary index, with values... , For the number of words at the object level; Includes object words with clear semantic references, such as man, dog, house, etc., and discards words without clear entities, such as a, the, main, like, etc.

[0024] S23. Utilize a pre-trained multimodal alignment model to align the object-level vocabulary set. With the original image Embedded into the same semantic space, and cross-modal similarity calculated using the following formula, a saliency response map of the image is generated:

[0025] ;

[0026] in, Indicates the similarity results. The cosine function is used to calculate the cosine value of two variables as a representation of similarity in the feature space. This represents the encoding and embedding process for an image. This indicates the text encoding embedding process. Indicates the image sample index. Represents an object-level lexical index;

[0027] S24. Set the threshold filter coefficient. And the object-level sub-map region is obtained from the saliency response map of the image using the following formula:

[0028] ; ;

[0029] in, Represents the original image In regional location The mask value at the location; when the cosine similarity Greater than or equal to the threshold filter coefficient When the cosine similarity is 1, the mask value is equal to 1; when the cosine similarity is 1... Less than threshold filter coefficient When the mask value is 0, the mask value is equal to 0. Indicates the first the original image the first the local object subgraph, wherein the subscript represents an image sample index, and the superscript represents an object level ;

[0030] S25, repeating steps S22-S24 to obtain all object level vocabulary sets in the original image corresponding to the local object subgraph granularity , and merging each local object subgraph granularity to obtain a local object subgraph granularity set ;

[0031] S26, constructing a global view granularity : scaling the original image using the following formula to obtain a global view granularity :

[0032] ;

[0033] wherein represents the global view granularity, the subscript represents an image sample index, and the superscript represents a global level Global, represents an image size scaling function;

[0034] S27, integrating the global view granularity with the local object subgraph granularity set to form a multi-granularity subgraph set , and constructing an image element matrix based on the multi-granularity subgraph set.

[0035] Further, the step S3 specifically comprises the following steps:

[0036] S31, constructing a dynamic quality prototype construction module based on a pre-trained multi-modal large model;

[0037] The input of the dynamic quality prototype construction module is a fixed question and an original image , and the output is a quality prototype vocabulary ; wherein represents a quality prototype vocabulary, represents a clustering sub-cluster index, the value range of , is the number of quality prototype vocabularies; in the present application, The purpose of setting it to 5 is to construct cluster centers with a clear quality stratification, where the quality cluster centers are categorized as: Very Bad - Bad - Moderate - Good - Very Good.

[0038] S32. Setting fixed questions Using a multimodal large model to analyze the original image Perform visual question answering to obtain information about the original image. Image quality description ;

[0039] The visual question answering is implemented using the following formula:

[0040] ;

[0041] in, This indicates a description of image quality; it is an answer to the question about image quality. The subscript indicates the image quality. For image sample index, and Subscript The meaning is consistent; BLIP (Visual Question Answering) is used to develop visual question answering functionality for pre-trained multimodal large models. Original image;

[0042] S33, Define quality semantic embedding as The hidden state vector output by the last layer of the encoder in the execution response process;

[0043] The hidden state vector is shown in the following equation:

[0044] ;

[0045] in, For the first The hidden state vector of each sample; It is a vector space; For dimension, here, ; The BLIPVQA model incorporates an image attention-guided text encoder.

[0046] S34, based on Clustering algorithms use the following formula to... The hidden state vector output by the encoder during the response process Perform clustering operations to obtain semantic subclusters with similar quality levels. :

[0047] ;

[0048] in, For semantic subclusters, the index of the cluster sub-cluster (consistent with the quality prototype the value range of , the number of quality prototype words (consistent with the meaning of the number of quality prototype words), that is, the preset cluster number; clustering algorithm; the hidden state vector of the th sample, the number of samples in the data set;

[0049] S35, construct a dynamic quality prototype, and use the following formula to count the keyword with the highest frequency in each semantic sub-cluster , and use it as the dynamic semantic quality anchor point of the sub-cluster; the dynamic semantic quality anchor point is used to replace the traditional fixed label;

[0050] ;

[0051] wherein, the th dynamic quality prototype, the index of the cluster sub-cluster, the value range of , the set value of is 5; is a function of the independent variable corresponding to the maximum value, that is, among all in the sub-cluster, find the with the highest frequency; is a frequency statistics function, used to count the word with the highest frequency in ;

[0052] indicates the image quality description.

[0053] Further, the step S4 specifically comprises the following steps: S41, build a quality evaluation model, and set its network main body as a BLIP pre-trained multi-modal feature alignment embedding architecture, which includes an image encoder and a text encoder; the input of the image encoder is pixel image data for extracting image visual features; the input of the text encoder is an English text sequence for extracting text semantic features;

[0054] S42, perform multi-granularity decomposition on the original image, output a multi-granularity sub-graph set, and output the th dynamic quality prototype using a dynamic quality prototype construction module;

[0055] The multi-granularity decomposition of the original image outputs a multi-granularity sub-image set, which is achieved by the following formula: ;

[0056] Wherein, is the original image The multi-granularity sub-image set constructed by the multi-granularity decomposition module, is the global view granularity, is the local object-level sub-image granularity set, denotes the multi-granularity decomposition module;

[0057] The dynamic quality prototype output by the dynamic quality prototype construction module is achieved by the following formula:

[0058] ;

[0059] Wherein, denotes the i-th dynamic quality prototype, is the clustering sub-cluster index, the value range of is , the set value of is 5; is a function of taking the independent variable corresponding to the maximum value; is a frequency statistics function for counting the word with the highest frequency in the independent variable; denotes the clustering algorithm; denotes the hidden state vector of the i-th sample, denotes the number of samples in the data set, denotes the image sample index;

[0060] S43, considering the image quality evaluation, the generated content needs to be consistent with the image visual quality, therefore, the enhanced text representation is designed and constructed, which combines the text prompt information with the dynamic quality prototype;

[0061] The enhanced text representation is constructed as follows:

[0062] ;

[0063] Wherein, denotes the enhanced text representation, the subscript is the clustering sub-cluster index, which is consistent with the subscript in ; denotes an English text sequence, the internal is filled with dynamic quality prototype words, and the internal is filled with actual data sample content;​ represents the first dynamic mass prototype; represents the text prompt information for generating images; any one of the data samples , an enhanced text representation

[0064] S44, the enhanced text representation and the local object subgraph granularity set in the multi-granularity subgraph set and the global view granularity are converted into a unified vector space feature, and a multi-modal encoder cross-modal joint modeling is performed:

[0065]

[0066]

[0067]

[0068] wherein, represents the enhanced text representation encoding vector, represents the pre-training text encoder of the multi-modal large model BLIP, represents the enhanced text representation, is a vector space, is a dimension, and here represents the local object subgraph granularity set image encoding vector set, represents the pre-training image encoder of the multi-modal large model BLIP, represents the local object subgraph granularity set, represents the global object subgraph granularity image encoding vector, represents the global view granularity;

[0069] S45, the cross-modal semantic similarity is calculated by the following formula, and the semantic matching degree of the image feature and the text feature is measured by the feature cosine similarity:

[0070]

[0071]

[0072] wherein, is a local similarity set, and represents the local object subgraph granularity corresponding to the first object-level vocabulary and the​​​​​​​​ quality prototype vocabulary the result of cosine similarity calculation; representing the enhanced text representation encoding vector, representing the local object subgraph granularity set image encoding vector set, representing vector dot product; for global view similarity, representing global view granularity and the first quality prototype vocabulary the result of cosine similarity calculation; representing the enhanced text representation encoding vector, representing global view granularity image encoding vector;

[0073] S46, using the following formula to each local similarity and global similarity perform exponential normalization to convert the similarity into the classification probability of each image corresponding to each quality prototype:

[0074] ;

[0075] ;

[0076] wherein, local classification probability, representing the local object subgraph granularity corresponding to the first object-level vocabulary and the first quality prototype vocabulary enhanced text representation constructed consistent probability; representing local similarity, representing object-level vocabulary index, , the number of object-level vocabularies, representing clustering sub-cluster index; global view classification probability, representing global view granularity and the first quality prototype vocabulary enhanced text representation constructed consistent probability, representing global view similarity; the exponential normalization process is realized by using softmax function.

[0077] S47, set the double full connection layer structure to the local classification probability set and global view classification probability fusion is performed to obtain a multi-granularity quality score of the original image as the final output result of the quality assessment network;

[0078] The multi-granularity quality score of the original image is:

[0079]

[0080] wherein, is the quality assessment score, which is the final output result of the model; is a fully connected layer, and the subscript indicates different layer indexes, and have the same structure, only the internal parameters are different; is a nonlinear activation function; is a set of local classification probabilities; is a global view classification probability;

[0081] S48, set the quality assessment network loss function as a mean square error function, and use the loss function to measure the deviation between the output result of the model and the subjective quality score:

[0082] ;

[0083] wherein, is the loss function value, which is used to judge the deviation between the model prediction result and the true value; is the number of samples in the data set; indicates the quality assessment score of the i-th sample; is the subjective quality score, i.e. the true value. Further, the step S5 specifically comprises the following steps:

[0084] S51, set the number of iterations for training the quality assessment model

[0085] , and sample a pair of image texts from the data set as the to-be-evaluated data at each iteration, use the to-be-evaluated data as the model input, train the quality assessment model, and fit the subjective quality score; S52, configure parameters for the multi-granularity decomposition module and the dynamic quality prototype construction module, and output the multi-granularity subgraphs and the dynamic quality prototype;

[0086]

[0087] ​The multi-granularity grading module is set to a parameter of 8, meaning it acquires 8 object-level granularity sub-images from the original image. If there are fewer than 8 entities in the image, the acquisition can be repeated to ensure the input quantity for multi-granularity analysis. The dynamic quality prototype construction module is set to a parameter of 5, meaning it obtains 5 quality ratings from the overall dataset, with fixed semantics of poor, poor, moderate, good, and good, serving as semantic anchors for quality assessment.

[0088] S53. Based on the multi-granularity subgraph and dynamic quality prototype output in step S52, construct an enhanced text representation. Through multimodal feature alignment and similarity calculation, output the predicted probability of the image to various quality levels, form a probability prediction matrix, and provide a probabilistic basis for the final quality score.

[0089] S54. Use a two-layer fully connected layer to map the probability prediction matrix to the image quality assessment score to obtain the final result of the image quality assessment.

[0090] In a second aspect of the invention, a generated image quality assessment device based on pre-trained cross-modal feature alignment embedding is also included. The device includes: a dataset acquisition module, an image multi-granularity decomposition module, a dynamic quality model construction module, a quality assessment model construction module, and a quality assessment module.

[0091] The dataset acquisition module is used to acquire the generated image quality assessment dataset; the dataset includes text prompts, generated images, and image quality scores;

[0092] The image multi-granularity decomposition module is used to perform image multi-granularity decomposition using a pre-trained multimodal large model, construct a multi-granularity image representation that includes global view granularity and local object subgraph granularity, and form an image element matrix;

[0093] The dynamic quality model construction module is used to perform visual question answering on the generated images using a pre-trained multimodal large model, obtain the visual quality of each image, obtain the quality distribution of the generated images, extract the hidden layer features in the visual question answering process for clustering, construct quality prototype vocabulary for evaluating image quality, and form a dynamic quality prototype.

[0094] The quality assessment model construction module is used to transform the quality assessment problem into a classification problem based on the image element matrix and dynamic quality prototype, obtain the quality level classification probability of each element in the image element matrix, and map it into a quality assessment score through a double fully connected layer.

[0095] The quality assessment module is used to train the quality assessment model to obtain a model weight file, load the model weight file into the trained quality assessment model, input the image to be assessed and the corresponding text prompt information into the model, and the model outputs the quality assessment result of the image to be assessed.

[0096] In a third aspect of the present application, an electronic device is also included, comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the above-mentioned method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding.

[0097] In a third aspect of the present application, a machine-readable storage medium is also included, which stores executable instructions that, when executed, cause the machine to perform the above-mentioned method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding.

[0098] Advantages

[0099] The present application proposes a method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding. Compared with existing methods, this method fully integrates the text prompt information relied on by the generated images in the evaluation process, and introduces a multi-granularity image decomposition strategy to construct image sub-views and global view representations covering local details and overall structure, thereby realizing effective discrimination of the quality of generated images and significantly improving the evaluation accuracy. The present application uses the cross-modal pre-training model BLIP to construct a unified image-text embedding space, and jointly introduces multi-granularity feature modeling and dynamic quality prototype construction mechanisms, comprehensively considers the quality performance of generated images from the two dimensions of image visual quality and text semantic consistency. This method effectively overcomes the limitations of existing image quality evaluation methods based on deep neural networks when processing generated images, such as ignoring text prompt information and difficulty in identifying deviations between images and semantics. Through the above improvements, the present application significantly improves the evaluation accuracy and semantic alignment ability of the model, has stronger generalization performance and applicability, and can be widely applied to various generated image quality evaluation scenarios, expanding the practical application boundary of image quality evaluation technology. BRIEF DESCRIPTION OF DRAWINGS

[0100] Figure 1 A method flowchart of the method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding in the present application;

[0101] Figure 2 A general architecture diagram of the pre-trained multi-modal large model in the present application;

[0102] Figure 3 A processing flowchart of the image multi-granularity decomposition module in the present application;

[0103] Figure 4 A structure diagram of the dynamic quality prototype in the present application;

[0104] Figure 5This is a general visual diagram illustrating the overall accuracy of the quality assessment model in this invention.

[0105] Figure 6 This is a schematic diagram illustrating an example of the quality assessment model in this invention. Detailed Implementation

[0106] To provide a better understanding of the structural features and effects achieved by the present invention, a detailed description is provided below, accompanied by preferred embodiments and accompanying drawings:

[0107] Example 1

[0108] like Figure 1 The method shown is a generated image quality assessment method based on pre-trained cross-modal feature alignment embedding, which includes the following steps:

[0109] S1. Obtain the generated image quality assessment dataset; the dataset includes text prompts, generated images, and image quality scores.

[0110] Step S1 is used to obtain the generated image quality assessment dataset. This step involves constructing a variety of text prompts, using multiple diffusion generation models such as Dalle-E and Midjourney to generate a large number of images based on the prompts, and conducting quality scoring experiments on all the image data to form a dataset that includes text prompts, generated images, and image quality scores.

[0111] S2. Use a pre-trained multimodal large model to perform multi-granularity decomposition of the image, construct a multi-granularity image representation that includes global view granularity and local object subgraph granularity, and form an image element matrix.

[0112] Step S2 utilizes a pre-trained multimodal large model to perform multi-granularity image decomposition, taking the text prompt information and the corresponding generated image as model input. Within the multimodal aligned embedding space, each element in the text prompt information is explicitly indexed to the content region of the generated image, constructing a local granularity image representation. The original image is then scaled to preserve the overall view of the complete content, thereby obtaining a global granularity image representation without destroying the complete view. Combining the local and global granularities, a multi-granularity image representation is constructed, thus grasping the content parsing of the image from both local and global granularities, forming an image element matrix.

[0113] In traditional approaches, images are typically input directly into the network, which then crops and compresses them. The direct scaling of the image during the processing of the size may cause the image quality to degrade, which seriously affects the quality assessment of the generated image. The explicit modeling of the relationship between the local area of the image and the text description is beneficial to the image quality assessment. In the traditional approach, the image is usually processed by using a Transformer network, but the Transformer architecture lacks an inductive bias for local structure. If there is no explicit region-level guidance, the attention structure may be scattered and allocated to various regions of the image, which may fail to grasp the key points of the generated image, thereby affecting the evaluation of the image consistency.

[0114] Step S2 is a key step of feature extraction and structured processing of the generated image, which specifically includes the following contents: first, the generated image is decomposed in combination with local granularity and global granularity. The local granularity focuses on different content regions in the generated image, and these regions are described by corresponding text prompt information. The features of each region are recognized and extracted by a pre-trained multi-modal large model, and the region index representation is formed in combination with the text prompt information; the global granularity retains the overall information of the original image, and the image is scaled to a preset size as a global feature representation. Secondly, the region index representation of the local granularity and the scaled image representation of the global granularity are fused to form a multi-granularity image representation containing local details and global information, which comprehensively captures the visual features of the generated image. Then, the multi-granularity image representation is structured to be converted into a matrix form, i.e., an image element matrix. Each element in the matrix corresponds to the feature information of a specific granularity in the image, providing a standardized feature carrier for the image end input of the subsequent quality assessment model.

[0115] The role of step S2 is to convert the original generated image into a multi-scale image element matrix through multi-granularity decomposition and structured processing, so as to describe the content information of the image from multiple granularities, lay a foundation for the image end input of the quality assessment model in step S4, and at the same time, form a dynamic quality prototype with step S3, which supports the cross-modal quality assessment of the model from the image and text dimensions respectively.

[0116] S3, using a pre-trained multi-modal large model to perform visual question answering on the generated image, obtaining the visual quality of each image, obtaining the quality distribution of the generated image, and extracting the hidden layer features in the visual question answering process for clustering, constructing a quality prototype vocabulary for evaluating the image quality, and forming a dynamic quality prototype.

[0117] In the prior art, the quality classification anchor points (bad, poor, fair, good, perfect) are usually fixed by artificial setting, which has problems such as inability to accurately describe the image quality and difficulty for the model to distinguish the subtle differences between good and perfect. The dynamic quality prototype construction strategy introduced in the present application can construct a set of quality classification anchor points closely fitting the data quality manifold by controlling the quality manifold of the data samples, which is not only suitable for quality evaluation tasks, but also ensures that the constructed quality prototype is still in the alignment embedding space of the multi-modal large model, avoiding the situation that the model is not sensitive to the differentiation of synonyms.

[0118] Step S3 is used to construct a dynamic quality prototype, a multi-modal large model is used for visual question answering, and for each image, the question "how is the visual quality of the image" is asked to obtain the quality distribution of the generated images in the data set, and the hidden layer features in the visual question answering process are extracted for clustering, the quality description text of each class is counted, and a quality prototype vocabulary capable of evaluating the quality of the image is constructed, and finally a dynamic quality prototype is obtained.

[0119] S4, construct a quality evaluation model, based on the image element matrix and the dynamic quality prototype, convert the quality evaluation problem into a classification problem, obtain the quality level division probability of each element in the image element matrix, and map it to the quality evaluation score through a double fully connected layer.

[0120] The purpose of step S4 is to build a quality evaluation model. Specifically, the quality evaluation model is built on the Python platform, and the BLIP pre-trained multi-modal large model is used as the main framework of the quality evaluation model. During the pre-training process, the BLIP network is designed as a network architecture that combines image-text consistency understanding, image translation, and image visual question answering for joint training, which can effectively encode images and texts into the same semantic space. The text input of the quality evaluation model is a sentence with quality classification meaning composed of the text prompt information of the generated image and the dynamic quality prototype. The image input of the quality evaluation model is the image element matrix constructed by multi-granularity decomposition. The quality evaluation model converts the quality evaluation problem into a classification problem, obtains the quality level division probability of each element in the image element matrix, and maps it to the quality evaluation score through a double fully connected layer.

[0121] S5, train the quality evaluation model to obtain the model weight file, load the model weight file into the trained quality evaluation model, input the image to be evaluated and the text prompt information corresponding to the image into the model, and output the quality evaluation result of the image to be evaluated.

[0122] On the basis of the quality evaluation model constructed in step S4, the quality evaluation model is trained and optimized, the quality evaluation model can be more accurately evaluated by fine-tuning parameters and adjusting training strategy, and the model weight file obtained by training is the key basis for subsequent practical application of the quality evaluation model. Specifically, a full-amount fine-tuning framework is used to fine-tune the BLIP-ITC model. A plurality of iteration times are set, a suitable learning rate is adjusted, and the learning rate is periodically adjusted in combination with a cosine annealing strategy. The training process adopts mixed precision training. After the training is completed, the model weight file is obtained. On the basis of the model weight file obtained by training, the quality of the generated image is effectively evaluated by inputting the to-be-evaluated object, which is the link to obtain the final result in the whole method process. When applied to actual tasks, the model weight file is first loaded into the trained quality evaluation model, and then the to-be-evaluated image and the text prompt information corresponding to the image are input as the model, and the model outputs the quality evaluation result of the to-be-evaluated image.

[0123] Further, the step S1 specifically comprises the following steps:

[0124] S11, constructing text prompt information;

[0125] S12, generating a corresponding image based on the text prompt information by using a diffusion generation model; the diffusion generation model is Dalle-E or Midjourney;

[0126] S13, quality scoring of all generated images;

[0127] S14, forming a data set containing text prompt information, generated image and image quality score as a generated image quality evaluation data set.

[0128] Further, the step S2 specifically comprises the following steps:

[0129] S21, constructing an image multi-granularity decomposition module based on a pre-trained multi-modal large model.

[0130] The input of the image multi-granularity decomposition module is text prompt information and the original image generated according to the text prompt information , the output is a global view granularity and a local object subgraph granularity set , both of which are constructed as a multi-granularity subgraph set corresponding to the original image , wherein the subscript represents an image sample index, the superscript represents a global level , and the superscript represents an object level .

[0131] S22, construct a local object subgraph granularity set : use a dependency syntactic analyzer to analyze the text prompt information to perform structured analysis and extract an object-level vocabulary set with clear semantic direction , wherein the subscript represents the object-level vocabulary index, and the value range is , the number of object-level vocabularies; contains man, dog, house, and other such object words with clear semantic direction, and discards a, the, main, like, and other words without clear entities.

[0132] S23, using a pre-trained multi-modal alignment model, embedding the object-level vocabulary set into the same semantic space as the original image , and calculating the cross-modal similarity using the following formula to generate the image saliency response map:

[0133] ;

[0134] wherein represents the similarity result, represents the cosine function, which is used to calculate the cosine value of two variables as a feature space similarity representation, represents the encoding embedding process for the image, represents the encoding embedding process for the text, represents the image sample index, represents the object-level vocabulary index.

[0135] In step S23, the cross-modal similarity of each region of the image is first obtained, and then the cross-modal similarity of each region of the image is mapped to the corresponding position of the image in the form of a response value. The region with high similarity is given a high response (such as visualization methods such as highlighting, strong weight, etc.), and finally an image saliency response map is formed. The most relevant salient region in the image is intuitively presented through the map.

[0136] S24, set a threshold filtering coefficient , and use the following formula to obtain the object-level subgraph region from the image saliency response map:

[0137] ; ;

[0138] wherein represents the mask value of the original image at the region position ; when the cosine similarity greater than or equal to a threshold filtering coefficient , the mask value is equal to 1; when the cosine similarity is less than a threshold filtering coefficient , the mask value is equal to 0; denotes the first original image, the first local object subgraph, wherein the subscript denotes the image sample index, and the superscript denotes the object level .

[0139] S25, repeating steps S22-S24, obtaining all object level vocabulary sets in the original image corresponding to the local object subgraph granularity , and merging each local object subgraph granularity to obtain a local object subgraph granularity set .

[0140] S26, constructing a global view granularity : scaling the original image using the following formula to obtain a global view granularity :

[0141] ;

[0142] wherein, denotes the global view granularity, the subscript denotes the image sample index, and the superscript denotes the global level Global , and the image size scaling function.

[0143] S27, integrating the global view granularity with the local object subgraph granularity set to form a multi-granularity subgraph set , and constructing an image element matrix based on the multi-granularity subgraph set.

[0144] Further, the step S3 specifically comprises the following steps:

[0145] S31, constructing a dynamic quality prototype construction module based on a pre-trained multi-modal large model;

[0146] The input of the dynamic quality prototype construction module is a fixed question and an original image , and the output is a quality prototype vocabulary ; wherein, denotes the quality prototype vocabulary, and the subscript represents the index of the cluster sub-cluster, and takes a value of , is the number of quality prototype words; in the present application, is set to . The purpose of setting to 5 is to construct a cluster center with clear quality stratification, and the quality cluster center includes: very bad-bad-average-good-very good.

[0147] S32, set a fixed question , use a multi-modal large model to perform visual question and answer on the original image to obtain an image quality description about the original image ;

[0148] The visual question and answer is implemented as follows:

[0149] ;

[0150] wherein, represents the image quality description, which is an answer to the image quality, and the subscript is the image sample index, which has the same meaning as the subscript in the above formula; is the visual question and answer function BLIP-Visual-Question-Answer of the pre-trained multi-modal large model.

[0151] S33, the quality semantic embedding is defined as the hidden state vector output by the last layer of the encoder in the answer process;

[0152] The hidden state vector is shown in the following formula:

[0153] ;

[0154] wherein, is the hidden state vector of the i-th sample; is a vector space; is a dimension, and here, ; ; is a text encoder combined with image attention guidance inside the BLIPVQA model.

[0155] S34, based on the clustering algorithm, the hidden state vector output by the encoder in the answer process is clustered using the following formula to obtain semantic sub-clusters with similar quality levels :

[0156] ;​​​

[0157] wherein, is a semantic sub-cluster, is a clustering sub-cluster index (consistent with the quality prototype , subscript has the same meaning), the value range of , is a preset clustering number (consistent with the meaning of the number of quality prototype words); denotes a clustering algorithm; denotes the hidden state vector of the th sample, denotes the number of samples in the data set, denotes the image sample index.

[0158] S35, construct a dynamic quality prototype, and use the following formula to count the keywords with the highest frequency in each semantic sub-cluster , and take it as the dynamic semantic quality anchor point of the sub-cluster; the dynamic semantic quality anchor point is used to replace the traditional fixed label;

[0159] ;

[0160] wherein, denotes the th dynamic quality prototype, is a clustering sub-cluster index, the value range of , the set value of is 5; is a function of the independent variable corresponding to the maximum value, that is, among all in the sub-cluster, find the with the highest frequency; is a frequency statistics function, used to count the word with the highest frequency in denotes the image quality description.

[0161] Further, the step S4 specifically comprises the following steps:

[0162] S41, construct a quality evaluation model, and set its network main body as a BLIP pre-trained multi-modal feature alignment embedding architecture, which includes an image encoder and a text encoder; the input of the image encoder is image data of pixels, used to extract image visual features; the input of the text encoder is an English text sequence, used to extract text semantic features;

[0163] S42, perform multi-granularity decomposition on the original image, output a multi-granularity sub-graph set, and use the dynamic quality prototype construction module to output the a dynamic quality prototype;

[0164] The multi-granularity decomposition of the original image outputs a multi-granularity subgraph set, which is achieved by using the following formula: ;

[0165] wherein, is the original image The multi-granularity subgraph set constructed by the multi-granularity decomposition module, is the global view granularity, is the local object-level subgraph granularity set, denotes the multi-granularity decomposition module;

[0166] The dynamic quality prototype construction module outputs the first dynamic quality prototype, which is achieved by using the following formula:

[0167] ;

[0168] wherein, denotes the first dynamic quality prototype, is the clustering subcluster index, the value range of , the set value of is 5; is a function of the independent variable corresponding to the maximum value; is a frequency statistical function for counting the word with the highest frequency in the independent variable; denotes the clustering algorithm; denotes the hidden state vector of the first sample, denotes the number of samples in the data set, denotes the image sample index;

[0169] S43, considering the image quality evaluation, the generated content needs to be consistent with the image visual quality, therefore, an enhanced text representation is designed and constructed, which combines the text prompt information and the dynamic quality prototype together;

[0170] The enhanced text representation is constructed as follows:

[0171] ;

[0172] wherein, denotes the enhanced text representation, the subscript is the clustering subcluster index, which is consistent with the subscript in ; denotes an English text sequence, the internal Filling the vocabulary by dynamic quality prototypes, internally Filling by actual data sample content; Representing the first Dynamic quality prototype; Representing the text prompt information of the generated image; taking any one of the data samples , you can construct Enhanced text representation;

[0173] S44, the enhanced text representation And the local object subgraph granularity set In the multi-granularity subgraph set And the global view granularity Is converted into a unified vector space feature, and the multi-modal encoder cross-modal joint modeling is carried out:

[0174] ;

[0175] ;

[0176] ;

[0177] Among them, The enhanced text representation Encoding vector, Indicates the pre-training text encoder of the multi-modal large model BLIP, Indicates the enhanced text representation, Is a vector space, The dimension is ; Indicates the local object subgraph granularity set Image encoding vector set, Indicates the pre-training image encoder of the multi-modal large model BLIP, Indicates the local object subgraph granularity set, Indicates the global object subgraph granularity Image encoding vector, Indicates the global view granularity;

[0178] In step S44, three formulas are respectively used for text feature encoding, object-level subgraph feature encoding and global-level subgraph feature encoding.

[0179] S45, calculate the cross-modal semantic similarity by the following formula, measure the semantic matching degree of image features and text features by feature cosine similarity:

[0180] ;

[0181] ​ ;

[0182] wherein, is a set of local similarities, representing the local object subgraph granularity corresponding to the j-th object-level vocabulary and the i-th quality prototype vocabulary ; ; represents an enhanced text representation encoding vector, represents a set of local object subgraph granularity image encoding vectors, represents a vector dot product; is a global view similarity, representing the global view granularity corresponding to the j-th object-level vocabulary and the i-th quality prototype vocabulary ; represents an enhanced text representation encoding vector, represents a global view granularity image encoding vector;

[0183] S46, exponentially normalizing each local similarity and global similarity to transform the similarity into a classification probability of each image corresponding to each quality prototype using the following formula:

[0184] ;

[0185] ;

[0186] wherein, is a local classification probability, representing the local object subgraph granularity corresponding to the j-th object-level vocabulary and the i-th quality prototype vocabulary ; ; represents a probability of consistency of the enhanced text representation representing a local similarity, representing an object-level vocabulary index, , representing a clustering sub-cluster index; is a global view classification probability, representing the global view granularity corresponding to the j-th object-level vocabulary and the i-th quality prototype vocabulary ; ​This represents the global view similarity; the exponential normalization process is implemented using the softmax function.

[0187] S47. Define the local classification probability set for the dual fully connected layer structure. Global view classification probability By fusing the images, we can obtain information about the original images. The multi-granularity quality score is used as the final output of the quality assessment network;

[0188] The original image The multi-particle size quality score is:

[0189]

[0190] in, The quality assessment score is the final output of the model. It is a fully linked layer; the subscripts indicate different layer indices. and The structures are completely identical, only the internal parameters are different; It is a non-linear activation function; It is a set of local classification probabilities; Classify the probability of the global view;

[0191] S48. Set the loss function of the quality assessment network to the mean squared error function, and use the loss function to measure the deviation between the model's output and the subjective quality score:

[0192] ;

[0193] in, This is the loss function value, used to determine the deviation between the model's prediction and the actual value; This represents the number of samples in the dataset. Indicates the first The quality assessment score of each sample; For image sample indexing; This is a subjective quality score, i.e., the true value.

[0194] Furthermore, step S5 specifically includes the following steps:

[0195] S51. Set the number of iterations for training the quality assessment model. In each iteration, a pair of image and text is sampled from the dataset as the data to be evaluated. The data to be evaluated is used as the input to the model to train the quality evaluation model and fit the subjective quality score.

[0196] S52. Configure the parameters of the multi-granularity hierarchical module and the dynamic quality prototype construction module, and output the multi-granularity subgraph and the dynamic quality prototype.

[0197] The parameters of the multi-granularity classification module are set to 8, that is, 8 object-level granularity subgraphs are collected from the original image, and if there are less than 8 entities in the image, the collection can be repeated to ensure the input quantity of multi-granularity analysis. The parameters of the dynamic quality prototype construction module are set to 5, that is, 5 quality ratings are obtained from the overall data set, and the fixed semantics are poor, relatively poor, medium, relatively good and good, which are used as semantic anchors for quality evaluation.

[0198] S53, based on the multi-granularity subgraph and the dynamic quality prototype output in step S52, construct an enhanced text representation, output the prediction probability of the image to each quality level through multi-modal feature alignment and similarity calculation, and form a probability prediction matrix to provide probabilistic basis for the final quality score.

[0199] S54, use a double-layer fully connected layer to map the probability prediction matrix to an image quality evaluation score to obtain the final result of image quality evaluation.

[0200] In order to verify the effect of the present application, the data set used in the present application is derived from the generated image quality evaluation public data set AIGIQA-20K, which contains 20000 generated images and their corresponding generated text prompts and artificial subjective scores. 14000 images are selected as the training set, 4000 images are selected as the verification set, and 2000 images are selected as the test set. As shown in Table 1, the evaluation results (SRCC index, PLCC index) of the method proposed in the present application exceed the existing image quality evaluation algorithm, and the quality evaluation result is improved to 0.9024, 0.9229.

[0201] Table 1 Comparison of the present application and the prior art on AIGIQA-20K data set

[0202] ResultMethods SRCC PLCC StairIQA 0.7895 0.9425 MUSIQ 0.8329 0.8646 RE-IQA 0.8219 0.8333 CLIPIQA 0.7863 0.7117 CLIPAGIQA 0.8715 0.8836 The present invention 0.9024 0.9229

[0203] Figure 2 The overall architecture schematic diagram of the pre-trained multi-modal large model in the present application is shown in the figure, which details the network framework of the generated image quality evaluation. As shown in Figure 2As shown, the entire system consists of three core components: a BLIP-based multi-modal encoder for constructing a text-image deep alignment embedding space; a dynamic quality prototype construction module for adaptively generating quality-graded semantic anchors from training data; and a multi-granularity semantic-guided image decomposition module for extracting local and global granularity representations of images to improve the modeling ability of detail quality defects. Specifically, a set of prompt text prompts generates a set of AI-generated images through an AI tool. After the images are decomposed into sub-image sets by the MGD module, image features about the sub-image sets are obtained through the image encoder. At the same time, the set of AI-generated images is processed by the DQPC dynamic quality prototype construction module to generate five quality gradient words (such as bad, poor, fair, good, perfect, etc.). The five quality gradient words are combined according to the fixed sentence pattern A photo that Q matches P to construct five text sentences, and text features are obtained through the text encoder. The cosine similarity between the image and text features is calculated and exponentially normalized to obtain the classification of each sub-image with respect to the quality gradient. The complete sub-image classification probabilities are concatenated through a double-layer fully connected layer to obtain the quality evaluation result of the generated image. Through this framework, the consistency of image content and text prompt content and the visual quality attributes of the image itself can be ensured.

[0204] Figure 3 The figure illustrates the processing flow of the image multi-granularity decomposition module in the present application. The multi-granularity decomposition module takes the generated image and its corresponding text prompt as input, indexes the image content using a multi-modal semantic embedding space, extracts image regions that explicitly correspond to key semantic elements in the text prompt, and constructs multi-granularity quality representations. As shown, Figure 3 First, the prompt text prompt is analyzed by dependency syntax to obtain an object-level vocabulary set with explicit semantic pointing . Next, based on the MLLM multi-modal large model, each region is indexed and divided according to the object-level vocabulary in the original image. Finally, the object-level sub-image is obtained. At the same time, the original image is directly scaled to obtain the global-level sub-image , and the multi-granularity sub-image set is obtained by merging.

[0205] Figure 4 The figure illustrates the structure of the dynamic quality prototype in the present application. As shown, Figure 4As shown, the quality description and its semantic embedding are generated by BLIP-VQA, and representative quality prototype words are obtained by clustering as adaptive semantic anchors in subsequent quality evaluation. Specifically, for a set of images, the BLIP-VQA module is used to align the visual question and answer "What is the visual quality of this image?", and the BLIP-VQA answer about the image is obtained. The hidden layer information is clustered, and the number of clusters is set to 5. The most frequent words in each class are counted to obtain quality prototype words at different levels of visual quality description (cluster result example: bad, poor, fair, good, perfect).

[0206] Figure 5 The overall visualization diagram of the quality evaluation accuracy of the quality evaluation model in the application is shown in the performance radar chart of the method described in the application. As shown in the figure, Figure 5 all schemes start from the center, and different angles correspond to different indicators. The figure shows the results of AIGCIQA2023, PKU-I2IQA, and PKU-AIGIQA-4K three data sets on the quality Q-SRCC, quality Q-PLCC, consistency C-SRCC, consistency C-PLCC, authenticity A-SRCC, and authenticity A-PLCC, a total of 6 indicators. As can be seen from Figure 5 , the area surrounded by the performance effect of the image quality evaluation method described in the application covers the largest area, indicating that the effect on the six indicators has reached the best. As a comparison, 14 schemes of the same type from 2022 to 2025 are selected for comparison, and the right side is the identification line of the comparison scheme.

[0207] Figure 6 The example evaluation diagram of the quality evaluation model in the application is shown in the figure. The classification probability of each image with respect to five different quality prototype words is shown. Specifically, each group of examples is composed of the left original image and the right classification probability. In the right figure, the histogram corresponds to the probability of the original image and the five quality prototype classifications. The vertical axis represents the value from 0 to 1. The higher the value, the higher the classification probability. The upper right corner of the histogram is marked with the original image corresponding score, model prediction score, and difference between the two. As can be seen from Figure 6 , the performance effect of the quality evaluation model described in the application is very close to score, and the classification probability also basically follows a Gaussian distribution. For images with poor quality, the model gives a lower score, and for images with good quality, the model gives a higher score, which can be considered to be consistent with human subjective quality perception.

[0208] In summary, the application designs a generated image quality evaluation method based on pre-training cross-modal feature alignment embedding, which uses the cross-modal feature alignment space constructed in the pre-training large model to complete the object display index region division of the image, constructs the dynamic quality prototype fitting data quality manifold distribution, combines the enhanced expression sentence constructed by the text prompt information during generation to complete the consistency examination of the image content and the text prompt, and simultaneously considers the global view quality evaluation. The image quality evaluation task precision is significantly improved in multiple scenes. The defects of limited quality evaluation precision caused by ignoring the generated prompt information in the traditional image quality evaluation process and ignoring the local view quality degradation in the image quality evaluation process are solved. It is suitable for a wide range of generated image quality evaluation scenes.

[0209] Embodiment two

[0210] The application also includes a generated image quality evaluation device based on pre-training cross-modal feature alignment embedding, which comprises a data set acquisition module, an image multi-granularity decomposition module, a dynamic quality model construction module, a quality evaluation model construction module and a quality evaluation module.

[0211] The data set acquisition module is used to acquire the generated image quality evaluation data set; the data set comprises text prompt information, generated image and image quality score;

[0212] The image multi-granularity decomposition module is used to perform image multi-granularity decomposition by using a pre-trained multi-modal large model, construct a multi-granularity image representation containing global view granularity and local object subgraph granularity, and form an image element matrix;

[0213] The dynamic quality model construction module is used to perform visual question answering on the generated image by using a pre-trained multi-modal large model, acquire the visual quality of each image, obtain the quality distribution of the generated image, extract the hidden layer features in the visual question answering process for clustering, construct quality prototype words for evaluating image quality, and form a dynamic quality prototype;

[0214] The quality evaluation model construction module is used to convert the quality evaluation problem into a classification problem based on the image element matrix and the dynamic quality prototype, obtain the quality level division probability of each element in the image element matrix, and map the quality evaluation score through a double full connection layer;

[0215] The quality evaluation module is used to train the quality evaluation model to obtain a model weight file, load the model weight file into the trained quality evaluation model, input the image to be evaluated and the text prompt information corresponding to the image into the model, and output the quality evaluation result of the image to be evaluated by the model.

[0216] Embodiment three

[0217] The present application also includes an electronic device comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the above method for evaluating quality of generated images based on pre-trained cross-modal feature alignment embedding.

[0218] In this embodiment, the electronic device can include, but is not limited to, a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile computing device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable computing device, a consumer electronic device, and the like.

[0219] Embodiment Four

[0220] The present application also includes a machine-readable storage medium storing executable instructions that, when executed, cause the machine to perform the above method for evaluating quality of generated images based on pre-trained cross-modal feature alignment embedding.

[0221] In particular, a system or apparatus equipped with a readable storage medium on which a software program code implementing the functions of any of the above embodiments is stored, and causing the computer or processor of the system or apparatus to read and execute the instructions stored in the readable storage medium, can be provided.

[0222] In this case, the program code read from the readable medium itself can implement the functions of any of the above embodiments, and thus the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present specification.

[0223] Embodiments of the readable storage medium include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as a CD-ROM, a CD-R, a CD-RW, a DVD-ROM, a DVD-RAM, a DVD- RW, a DVD-RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer or the cloud over a communication network.

[0224] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) embodying computer usable program code.

[0225] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the

[0226] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the

[0227] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the

[0228] The essential features of the application, the main characteristics and the advantages of the application have been shown and described above. It should be understood by those skilled in the art that the application is not limited to the above-described embodiments, and that the above-described embodiments and descriptions in the specification are only the principles of the application. Without departing from the spirit and scope of the application, various changes and improvements can be made to the application, and these changes and improvements all fall within the scope of the claimed application. The scope of protection of the application is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding, characterized in that, The method includes the following steps: S1. Obtain the generated image quality assessment dataset; the dataset includes text prompts, generated images, and image quality scores. S2. Use a pre-trained multimodal large model to perform multi-granularity decomposition of the image, construct a multi-granularity image representation that includes global view granularity and local object subgraph granularity, and form an image element matrix; S3. Use a pre-trained multimodal large model to perform visual question answering on the generated images, obtain the visual quality of each image, obtain the quality distribution of the generated images, extract the hidden layer features in the visual question answering process for clustering, construct quality prototype vocabulary for evaluating image quality, and form a dynamic quality prototype. S4. Construct a quality assessment model. Based on the image element matrix and dynamic quality prototype, the quality assessment problem is transformed into a classification problem. The quality level classification probability of each element in the image element matrix is ​​obtained and mapped to the quality assessment score through a double fully connected layer. S5. Train the quality assessment model to obtain the model weight file. Load the model weight file into the trained quality assessment model, and input the image to be assessed and the corresponding text prompt information into the model. The model outputs the quality assessment result of the image to be assessed.

2. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11. Construct text prompt information; S12. Based on the text prompt information, generate the corresponding image using a diffusion generation model; the diffusion generation model is Dalle-E or Midjourney. S13. Score the quality of all generated images; S14. Create a dataset containing text prompts, generated images, and image quality scores, as the generated image quality assessment dataset.

3. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 1, characterized in that, The multimodal large model adopts the BILP model.

4. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 1, characterized in that, Step S2 specifically includes the following steps: S21. Based on the pre-trained multimodal large model, construct an image multi-granularity decomposition module; The input to the image multi-granularity decomposition module is a text prompt message. Compared to the original image generated based on the text prompt information The output is at the global view granularity. Set of local object subgraph granularity ,in, Indicates the image sample index. Indicates global level , Represents object level ; S22. Constructing a set of local object subgraph granularities Using a dependency parser to analyze text prompts Perform structured parsing to extract a set of object-level lexical items with clear semantic meaning. ,in, This represents an object-level vocabulary index, with values ​​ranging from 1 to 2. , For the number of words at the object level; S23. Utilize a pre-trained multimodal alignment model to align the object-level vocabulary set. With the original image Embedded into the same semantic space, and cross-modal similarity calculated using the following formula, a saliency response map of the image is generated: ; in, Indicates the similarity results. Represents the cosine function. This represents the encoding and embedding process for an image. This indicates the encoding and embedding process for text. S24. Set the threshold filter coefficient. The object-level sub-map region can be obtained from the saliency response map of an image using the following formula: ; ; in, Represents the original image In regional location The mask value at that location; Indicates the first Zhang Original Image The Zhang's local object subgraph; S25. Repeat steps S22-S24 to obtain the original image. The set of all object-level vocabularies Corresponding local object subgraph granularity The granularity of each local object subgraph is then merged to obtain a set of local object subgraph granularities. ; S26. Use the following formula to process the original image. Scaling is performed to obtain the global view granularity. : ; in, Indicates the granularity of the global view. This represents the image scaling function; S27. Adjust the global view granularity Set of local object subgraph granularity Integration to form multi-granularity sub-atlases An image element matrix is ​​constructed based on multi-granularity sub-map sets.

5. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 4, characterized in that, Step S3 specifically includes the following steps: S31. Construct a dynamic quality prototype building module based on a pre-trained multimodal large model; The input to the dynamic quality prototype construction module is a fixed question. and the original image The output is a vocabulary of quality prototypes. ;in, The prototypical word for quality. Indicates the index of the cluster sub-cluster. The range of values ​​is , For the number of prototype words in quality; S32. Setting fixed questions Using a pre-trained multimodal large model to process the original image Perform visual question answering to obtain information about the original image. Image quality description ; The visual question answering is implemented using the following formula: ; in, For image sample indexing; BLIP (Visual Question Answering) is used to develop visual question answering functionality for pre-trained multimodal large models. S33, Define quality semantic embedding as The hidden state vector output by the last layer of the encoder in the execution response process; The hidden state vector is shown in the following equation: ; in, For the first The hidden state vector of each sample; It is a vector space; For dimensions; The BLIPVQA model incorporates an image attention-guided text encoder. S34, based on Clustering algorithms use the following formula to... The hidden state vector output by the encoder during the response process Perform clustering operations to obtain semantic subclusters with similar quality levels. : ; in, It is a semantic sub-cluster; express Clustering algorithms; Indicates the number of samples in the dataset; S35. Construct a dynamic quality prototype and use the following formula to statistically analyze each semantic subcluster. The most frequently occurring keyword within the sub-cluster is used as the dynamic semantic quality anchor point for that sub-cluster. ; in, The function of the independent variable that takes the maximum value. It is a frequency statistics function. This indicates a description of image quality.

6. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 5, characterized in that, Step S4 specifically includes the following steps: S41. Construct a quality assessment model, setting its network core as a BLIP pre-trained multimodal feature alignment embedding architecture, which includes an image encoder and a text encoder. S42. Perform multi-granularity decomposition on the original image, output multi-granularity sub-image sets, and use the dynamic quality prototype construction module to output the first... A dynamic quality prototype; The process of decomposing the original image into multiple granularities and outputting multiple granularity sub-image sets is achieved using the following formula: ; in, Original image The multi-granularity sub-maps constructed by the multi-granularity decomposition module For global view granularity, For local object-level subgraph granularity sets, This indicates a multi-granularity decomposition module. For image sample indexing; The output of the dynamic quality prototype construction module is as follows: A dynamic quality prototype is implemented using the following formula: ; in, Indicates the first A dynamic quality prototype For cluster indexes, The range of values ​​is ; The function that takes the maximum value of the independent variable; It is a frequency statistics function; express Clustering algorithms; Indicates the first The hidden state vector of each sample; Indicates the number of samples in the dataset; S43. Design and construct enhanced text representation This combines text prompts with dynamic quality prototypes. The constructed enhanced text representation is as follows: ; in, This represents enhanced text representation; This represents a sequence of English text. Vocabulary filling is constructed from dynamic quality prototypes. Filled with actual data sample content; Text prompts indicating the generated image; S44. Use the following formula to enhance text representation. and multi-granularity sub-atlas Local object subgraph granularity set and global view granularity The features are converted into a unified vector space and then used for cross-modal joint modeling by a multimodal encoder. ; ; ; in, Enhanced text representation The encoded vector, This represents a pre-trained text encoder for the multimodal large model BLIP. This indicates enhanced text representation. For a vector space, For dimensions; Represents a set of granularity for local object subgraphs. Image encoding vector set, This represents a pre-trained image encoder for a multimodal large model BLIP. Represents a set of local object subgraph granularities. Indicates global view granularity The image encoding vector, Indicates the granularity of the global view; S45. Calculate cross-modal semantic similarity using the following formula, and measure the semantic matching degree between image features and text features using feature cosine similarity: ; ; in, Let be the set of local similarities, representing the th Local object subgraph granularity corresponding to each object-level vocabulary With the Quality prototype vocabulary The result of cosine similarity calculation; Represents the dot product of vectors; For global view similarity, indicating the granularity of the global view. With the Quality prototype vocabulary The result of cosine similarity calculation; Indicates global view granularity The image encoding vector; S46. Use the following formula to calculate the local similarity for each local similarity. and global similarity Exponential normalization is performed to convert the similarity into the classification probability of each image corresponding to each quality prototype: ; ; in, Let be the local classification probability, representing the probability of the first classification. Local object subgraph granularity corresponding to each object-level vocabulary With the Quality prototype vocabulary Constructed Enhanced Text Representation The probability of consistency; Indicates local similarity; , Represents an object-level lexical index. For the number of words at the object level; The classification probability of the global view represents the granularity of the global view. With the Quality prototype vocabulary Constructed Enhanced Text Representation The probability of consistency Indicates the similarity of global views; S47. Define the local classification probability set for the dual fully connected layer structure. Global view classification probability By fusing the images, we can obtain information about the original images. The multi-granularity quality score is used as the final output of the quality assessment network; The original image The multi-particle size quality score is: ; in, The quality assessment score is the final output of the model. It is a fully linked layer, and the subscripts represent different layer indices; It is a non-linear activation function; It is a set of local classification probabilities; S48. Set the loss function of the quality assessment network to the mean squared error function, and use the loss function to measure the deviation between the model's output and the subjective quality score: ; in, This is the loss function value, used to determine the deviation between the model's prediction and the actual value; The number of samples in the dataset; Indicates the first The quality assessment score of each sample; This is a subjective quality score, i.e., the true value.

7. The method for evaluating the quality of generated images based on pre-trained cross-modal feature alignment embedding according to claim 6, characterized in that, Step S5 specifically includes the following steps: S51. Set the number of iterations for training the quality assessment model. In each iteration, a pair of image and text is sampled from the dataset as the data to be evaluated. The data to be evaluated is used as the input to the model to train the quality evaluation model and fit the subjective quality score. S52. Configure the parameters of the multi-granularity decomposition module and the dynamic quality prototype construction module, and output the multi-granularity subgraph and the dynamic quality prototype. S53. Based on the multi-granularity subgraph and dynamic quality prototype output in step S52, construct an enhanced text representation. Through multimodal feature alignment and similarity calculation, output the predicted probability of the image to various quality levels to form a probability prediction matrix. S54. Use a two-layer fully connected layer to map the probability prediction matrix to the image quality assessment score to obtain the final result of the image quality assessment.

8. A generative image quality assessment device based on pre-trained cross-modal feature alignment embedding, characterized in that, The device includes: a dataset acquisition module, an image multi-granularity decomposition module, a dynamic quality model construction module, a quality assessment model construction module, and a quality assessment module; The dataset acquisition module is used to acquire the generated image quality assessment dataset; the dataset includes text prompts, generated images, and image quality scores; The image multi-granularity decomposition module is used to perform image multi-granularity decomposition using a pre-trained multimodal large model, construct a multi-granularity image representation that includes global view granularity and local object subgraph granularity, and form an image element matrix; The dynamic quality model construction module is used to perform visual question answering on the generated images using a pre-trained multimodal large model, obtain the visual quality of each image, obtain the quality distribution of the generated images, extract the hidden layer features in the visual question answering process for clustering, construct quality prototype vocabulary for evaluating image quality, and form a dynamic quality prototype. The quality assessment model construction module is used to transform the quality assessment problem into a classification problem based on the image element matrix and dynamic quality prototype, obtain the quality level classification probability of each element in the image element matrix, and map it into a quality assessment score through a double fully connected layer. The quality assessment module is used to train the quality assessment model to obtain a model weight file, load the model weight file into the trained quality assessment model, input the image to be assessed and the corresponding text prompt information into the model, and the model outputs the quality assessment result of the image to be assessed.

9. An electronic device, characterized in that, include: At least one processor; And a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the generated image quality assessment method based on pre-trained cross-modal feature alignment embedding as described in any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that, It stores executable instructions that, when executed, cause the machine to perform the generated image quality assessment method based on pre-trained cross-modal feature alignment embedding as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Quality evaluation system and method for generated image

    CN121527018A

  • Image generation quality evaluation method for semantic evidence learning

    CN121937462A