Image Aesthetic Quality Evaluation Method, Device, Electronic Device and Storage Medium

By obtaining text and image feature information, combining global and local similarity calculations, and using preset weight rules to generate aesthetic quality evaluation results, the problem of inaccurate aesthetic evaluation of AI generated images in the prior art is solved, and the accuracy of evaluation is improved.

CN119206261BActive Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411345887.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-07-29
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

The existing image aesthetic quality evaluation method is poor in evaluating AI-generated images, lacks consideration of text information, and ignores the correlation between AI-generated images and the text prompt words used.

Method used

By obtaining text feature information and image feature information, using preset evaluation keywords to extract multiple text features of matching degree of prompt word text, combining global image features and local image features to calculate similarity, determining the matching level weight based on the preset weight assignment rules, and generating aesthetic quality evaluation results.

Benefits of technology

It improves the accuracy of image aesthetic quality evaluation, comprehensively considers the local characteristics and complete semantic information of the image, and enhances the aesthetic evaluation effect of AI-generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206261B_ABST
    Figure CN119206261B_ABST
Patent Text Reader

Abstract

The present application provides an image aesthetic quality evaluation method, apparatus, electronic device and storage medium. The method includes: obtaining text feature information and image feature information; determining a global similarity between the global image feature and the target text feature, and a local similarity between the local image feature and the target text feature according to the target text feature in the text feature information and the image feature information; determining an average similarity between the target text feature and the image feature information according to the global similarity and the local similarity; determining a matching degree weight of the target text feature in the text feature information and a weight value of the matching degree weight according to the average similarity between the target text feature and the image feature information; and obtaining an aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight. The method of the present application improves the accuracy of picture aesthetic quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an image aesthetic quality evaluation method, apparatus, electronic device, and storage medium. Background Art

[0002] Aesthetic quality is a measure of the beauty perceived visually, which measures the visual attractiveness of an image in the eyes of humans. The research of computational aesthetics explores how to use computable technologies to predict the emotional responses of humans to visual stimuli, so that the computer can simulate the human aesthetic process, and thus automatically predict the aesthetic quality of an image through computable methods.

[0003] Traditional aesthetic evaluation methods usually adopt methods designed based on professional knowledge or general features, and evaluate images by extracting a fixed feature set and following certain rules. In the field of deep learning, the basic method of image aesthetic evaluation is to input a single image, and after a series of processes of a deep learning network, output the aesthetic evaluation score of the image. By guiding the training of the deep learning network, it is enabled to have a strong ability to perceive and extract image aesthetic features, so as to perceive the aesthetic quality during the image encoding and decoding process and realize image aesthetic evaluation.

[0004] However, the existing image aesthetic quality evaluation methods have the problem of poor evaluation effect. Summary of the Invention

[0005] This application provides an image aesthetic quality evaluation method, apparatus, electronic device, and storage medium, which are used to solve the problem of poor effect when evaluating the aesthetic quality of an image generated by evaluation keywords.

[0006] In a first aspect, this application provides an image aesthetic quality evaluation method, including:

[0007] Obtain text feature information and image feature information, where the text feature information is the text features obtained by extracting features from the prompt text according to preset evaluation keywords, the evaluation keywords represent the matching degree between the prompt text and the image to be evaluated, and the image feature information includes the global image features of the image to be evaluated and the local image features of the target local image in the image to be evaluated;

[0008] Determine the global similarity between the global image features and the target text features, and the local similarity between the local image features and the target text features according to the target text features in the text feature information and the image feature information;

[0009] Determine the average similarity between the target text features and the image feature information according to the global similarity and the local similarity;

[0010] Determining, based on the average similarity between the target text feature and the image feature information, a matching degree weight of the target text feature in the text feature information and a weight of the matching degree weight, wherein the weight of the matching degree weight is determined based on a preset weight assignment rule and the size of the matching degree weight of the target text feature in the text feature information;

[0011] According to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight, the aesthetic quality evaluation result of the image to be evaluated is obtained.

[0012] In the embodiment of the present application, obtaining text feature information and image feature information includes:

[0013] Obtain prompt word text and image feature information;

[0014] Obtaining a prompt word evaluation text based on the prompt word text and preset evaluation keywords, wherein the number of evaluation keywords is not less than 2;

[0015] The prompt word evaluation text is input into the text encoder of the visual language model to obtain text feature information.

[0016] In the embodiment of the present application, obtaining text feature information and image feature information includes:

[0017] Obtain text feature information and the image to be evaluated;

[0018] According to the preset size requirements, the image to be evaluated is scaled to obtain a global image of the image to be evaluated;

[0019] According to the preset size requirements, the image to be evaluated is intercepted and processed to obtain the target local image;

[0020] The global image and the target local image of the image to be evaluated are respectively input into the image encoder of the visual language model to obtain the global image features and the local image features;

[0021] Image feature information is obtained based on global image features and local image features.

[0022] In the embodiment of the present application, the average similarity between the target text feature and the image feature information is determined based on the global similarity and the local similarity, including:

[0023] Perform weighted averaging on the local similarities to obtain the initial average similarity;

[0024] According to the initial average similarity and the global similarity, the average similarity of the target text feature and image feature information is obtained.

[0025] In an embodiment of the present application, determining the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight according to the average similarity between the target text feature and the image feature information includes:

[0026] Perform a weighted summation process on the average similarity between the target text feature and the image feature information to obtain the global average similarity;

[0027] Determine the matching degree weight of the target text feature in the text feature information according to the global average similarity and the average similarity;

[0028] Obtain the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight assignment rule.

[0029] In an embodiment of the present application, obtaining the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight assignment rule includes:

[0030] Perform a sorting process on the matching degree weight of the target text feature in the text feature information to obtain a matching degree weight sequence;

[0031] According to the weight assignment rule, perform an assignment process on the matching degree weights in the matching degree weight sequence to obtain the weight value of the matching degree weight, where the earlier the position of the matching degree weight in the matching degree weight sequence, the greater the weight value of the matching degree weight.

[0032] In an embodiment of the present application, obtaining the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight includes:

[0033] Obtain a matching degree weight value according to the matching degree weight and the weight value of the matching degree weight;

[0034] Perform a weighted summation process on the matching degree weight value to obtain a matching degree weight value;

[0035] Adjust the matching degree weight value according to the preset weight distribution range to obtain the aesthetic quality evaluation result of the image to be evaluated.

[0036] In a second aspect, the present application provides an image aesthetic quality evaluation device, including:

[0037] An acquisition module is configured to acquire text feature information and image feature information. The text feature information is obtained by extracting features from the prompt text based on preset evaluation keywords. The evaluation keywords represent the degree of match between the prompt text and the image to be evaluated. The image feature information includes global image features of the image to be evaluated and local image features of the target local image in the image to be evaluated.

[0038] A first determining module is configured to determine a global similarity between a global image feature and a target text feature, and a local similarity between a local image feature and a target text feature based on the target text feature and the image feature information in the text feature information;

[0039] A second determination module is used to determine the average similarity between the target text feature and the image feature information based on the global similarity and the local similarity;

[0040] A third determining module is configured to determine, based on the average similarity between the target text feature and the image feature information, a matching degree weight of the target text feature in the text feature information and a weight of the matching degree weight, wherein the weight of the matching degree weight is determined based on a preset weight assignment rule and the size of the matching degree weight of the target text feature in the text feature information;

[0041] The obtaining module is used to obtain the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight of the matching degree weight.

[0042] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0043] Memory stores computer-executable instructions;

[0044] The processor executes the computer-executable instructions stored in the memory to implement the image aesthetic quality assessment method of the embodiment of the present application.

[0045] In a fourth aspect, a computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the image aesthetic quality assessment method of an embodiment of the present application.

[0046] The image aesthetic quality evaluation method, device, electronic device, and storage medium provided by this application generate an image to be evaluated through a prompt text, then generate a global image and a local image based on the image to be evaluated, obtain global image features and local image features, calculate the similarity between the global image features and the local image features and the target text features obtained according to the prompt text and evaluation keywords respectively, perform an average process on the obtained similarities to obtain the average similarity of the target text features, determine the matching degree weight of the target text features in the text feature information according to the average similarity of the target text features and the image feature information, and then determine the aesthetic quality evaluation result of the image to be evaluated according to the preset weight assignment rule and the magnitude of the matching degree weight of the target text features in the text feature information, thereby improving the effect of evaluating the aesthetic quality of the image generated according to keywords. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0048] Figure 1 It is a schematic flowchart of the image aesthetic quality evaluation method provided by an embodiment of this application;

[0049] Figure 2 It is a schematic flowchart of another image aesthetic quality evaluation method provided by an embodiment of this application;

[0050] Figure 3 It is a schematic structural diagram of the image aesthetic quality evaluation device provided by an embodiment of this application;

[0051] Figure 4 It is a schematic structural diagram of the electronic device provided by an embodiment of this application.

[0052] Through the above-mentioned accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and the written description are not intended to limit the scope of the concept of this application in any way, but to explain the concept of this application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0054] Glossary:

[0055] AI-generated models: Leveraging deep learning technology, these models learn and train on massive datasets to identify and understand the meaning of keywords and generate images that meet the requirements. Users can enter a set of keywords to describe a scene, object, etc., and the model generates corresponding images based on these keywords. Examples include GauGAN, GauGAN2, BigGAN, and DF-GAN.

[0056] Visual Language Model: A technology that combines image and natural language processing to understand and interpret the relationship between images and text, converting image and text data into representations in the same space, enabling cross-modal interaction and joint learning. This model can be applied to various tasks such as visual question answering, image description generation, and converting text and images into vectors.

[0057] In the existing technology, traditional image aesthetics assessment methods are limited to image content and mostly rely on human subjective evaluation or expert opinions. The evaluation process is relatively cumbersome and lacks consideration of textual information. It ignores the correlation between AI (Artificial Intelligence) generated images and the text prompt words used. Therefore, it is limited in effectiveness when evaluating the aesthetics of AI-generated images.

[0058] The embodiment of the present application can extract text features of multiple matching degrees for the prompt word text based on preset evaluation keywords. In order to comprehensively consider the local features and complete semantic information of the image to be evaluated, the text feature information is respectively calculated with the global image feature information and the local image feature information. Then, the similarity between the image to be evaluated and the text features of each matching degree and the matching degree weight of each text feature in all text features are obtained, thereby obtaining the aesthetic quality evaluation result according to the preset weight assignment rules and matching degree weights. As a result, the image aesthetic evaluation is performed according to the similarity between the image and the prompt word of the generated image, thereby improving the evaluation accuracy.

[0059] The embodiments of the present application provide a method, device, electronic device, and storage medium for evaluating image aesthetic quality.

[0060] Figure 1 This is a flow chart of the image aesthetic quality assessment method provided in the embodiment of the present application. The execution subject of this method can be a server or other server, and this embodiment is not particularly limited here. Figure 1 As shown, the method may include:

[0061] S101. Obtain text feature information and image feature information. The text feature information is the text features obtained by extracting features from the prompt text according to preset evaluation keywords. The evaluation keywords characterize the matching degree between the prompt text and the image to be evaluated. The image feature information includes the global image features of the image to be evaluated and the local image features of the target local image in the image to be evaluated.

[0062] Among them, the text feature information can refer to multiple landmark data extracted from text data that can describe the content or nature of the text. In the embodiments of the present application, the text feature information can refer to the vector representation processed by the text encoder. They can capture the semantic and context information in the text. One text feature information corresponds to one evaluation keyword, and there are feature information of multiple evaluation keywords with different matching degrees for the text feature information.

[0063] The preset evaluation keywords can refer to words that describe different matching degrees between the image and the prompt text, and can be words with different degrees from light to heavy, such as "first level", "second level", "third level", etc.

[0064] The prompt text can refer to the text input by the user for generating the image to be evaluated, such as "sunset", "ocean", "cat", etc.

[0065] The image feature information can refer to the numerical representation extracted from the image that can describe attributes such as the content, structure, texture, and style of the image. These feature information are usually extracted by a pre-trained image encoder (such as ResNet, VGG, EfficientNet, etc.) and transformed into points in a high-dimensional vector space.

[0066] The global image features of the image to be evaluated can refer to the sub-images that can represent the complete semantic information of the image, and the local image features of the image to be evaluated can refer to the sub-images that can represent the local features and basic quality of the image.

[0067] The image to be evaluated can be an image generated by an AI generation model according to the prompt text.

[0068] Among them, in the embodiments of the present application, the method for obtaining text feature information and image feature information may include:

[0069] Obtain the prompt text and image feature information;

[0070] According to the prompt text and the preset evaluation keywords, obtain the prompt evaluation text, where the number of evaluation keywords is not less than 2;

[0071] Input the prompt evaluation text into the text encoder of the vision-language model to obtain the text feature information.

[0072] Among them, the prompt evaluation text can refer to the text composed of the prompt text and the evaluation keyword. The expression form of the prompt evaluation text can be templates such as "This is an image where the {evaluation keyword} matches the {prompt text}". By embedding the same prompt text and different evaluation keywords into the text template, multiple text features can be extracted. For example, when the evaluation keywords are "excellent", "average", and "very poor", and the prompt text is "sunset", the prompt evaluation texts include "excellent sunset", "average sunset", and "very poor sunset".

[0073] Input the prompt evaluation text into the text encoder of the vision-language model. The text encoder will output multiple text feature vectors, and each vector corresponds to a sentence describing different degrees of relevance, and multiple text feature vectors describing different degrees of relevance between the image and the prompt can be obtained.

[0074] Among them, in the embodiments of the present application, the methods for obtaining text feature information and image feature information may include:

[0075] Obtain text feature information and the image to be evaluated;

[0076] According to the preset size requirements, perform scaling processing on the image to be evaluated to obtain the global image of the image to be evaluated;

[0077] According to the preset size requirements, perform cropping processing on the image to be evaluated to obtain the target local image;

[0078] Input the global image and the target local image of the image to be evaluated into the image encoder of the vision-language model respectively to obtain the global image feature and the local image feature;

[0079] According to the global image feature and the local image feature, obtain the image feature information.

[0080] Among them, the preset size requirements can be set according to the size requirements of the image encoder.

[0081] The method of performing cropping processing on the image to be evaluated according to the preset size requirements to obtain the target local image may include, by means of a sliding window, taking the image obtained each time as the local image, and then randomly selecting (or selecting at intervals) a fixed number of images from them as the target local image, or automatically determining the position of the target local image to be selected according to the algorithm, and then performing targeted cropping.

[0082] Input the global image and the target local image of the image to be evaluated into the image encoder of the vision-language model respectively. The obtained global image feature can represent the information of all regions in the global image and can be used to describe the content, style, and layout of the entire image, etc. The local image feature can represent the local attributes and information of this region.

[0083] S102 : Determine the global similarity between the global image feature and the target text feature, and the local similarity between the local image feature and the target text feature, based on the target text feature and the image feature information in the text feature information.

[0084] The target text feature may refer to one of multiple text feature information describing different degrees of relevance between the image and the prompt word. For example, when the evaluation keywords are "excellent," "average," and "very poor," and the prompt word text is "book," the target text feature may be feature information for "excellent book." In this embodiment of the present application, by continuously repeating step S102, the global and local similarities of all text features and image feature information in the text feature information can be obtained.

[0085] Similarity may refer to the degree of matching between image features and text features, global similarity may refer to the degree of matching between features of the complete semantic information of the image and text features, and local similarity may refer to the degree of matching between local features of the image and text features. In some embodiments, methods for determining similarity may include Euclidean distance, Manhattan distance, and Pearson correlation coefficient, etc. In an embodiment of the present application, global similarity may be obtained by determining the cosine similarity of global image features and target text features, and local similarity may be obtained by determining the cosine similarity of local image features and target text features. Cosine similarity may refer to the similarity of two vectors obtained by evaluating the cosine value of the angle between them.

[0086] S103: Determine the average similarity between the target text feature and the image feature information based on the global similarity and the local similarity.

[0087] Among them, the average similarity can refer to the similarity of the target text features obtained by calculating the average value of the global similarity and the local similarity. By evaluating the calculation of the similarity, the information of the two dimensions of the original quality of the image and the complete semantics can be comprehensively considered.

[0088] In the embodiment of the present application, the method for determining the average similarity between the target text feature and the image feature information based on the global similarity and the local similarity may include:

[0089] Perform weighted averaging on the local similarities to obtain the initial average similarity;

[0090] According to the initial average similarity and the global similarity, the average similarity of the target text feature and image feature information is obtained.

[0091] Among them, the initial average similarity may refer to the average of all local similarities of the target text feature. For example, if the local similarities of the target text feature include three values: 23, 36, and 67, the calculation method of the initial average similarity is (23 + 36 + 67) / 3, and the calculation result is 42.

[0092] The average similarity can be obtained based on the initial average similarity and the global similarity. For example, when the initial average similarity is 42 and the global similarity is 64, the average similarity is 53.

[0093] In some embodiments, the method for determining the local similarity may further include:

[0094] According to the number of local images and a preset cleaning rule, determine the similarity cleaning threshold of the local images, then calculate the initial local similarity between the target text feature and the local images, and screen the initial local similarity according to the cleaning threshold to obtain the local similarity. For example, when the number of remaining local images is 64, since the number of local images is large and the difference in the relevance of each image to the prompt text is large, the similarity cleaning threshold is set to a relatively high 60%. When the similarity of a local image is lower than 60%, delete the image; when the similarity of a local image is higher than 60%, retain the image, eliminate the untrustworthy data with large deviation values, and leave the reasonable data with relatively high credibility to improve the accuracy of image aesthetics evaluation.

[0095] S104. Determine the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight according to the average similarity between the target text feature and the image feature information, where the weight value of the matching degree weight is determined according to a preset weight value assignment rule and the size of the matching degree weight of the target text feature in the text feature information.

[0096] Among them, the matching degree weight may refer to the proportion of the average similarity of the target text feature in the average similarities of all text features.

[0097] The weight value of the matching degree weight may refer to different weight values assigned according to different rankings of the proportion of the matching degree weight.

[0098] The method for determining the weight value of the matching degree weight according to a preset weight value assignment rule and the size of the matching degree weight of the target text feature in the text feature information may include sorting the weights from largest to smallest, from smallest to largest, etc., and then assigning values according to the weight value assignment rule, assigning the largest value to the largest weight, or assigning the largest value to the smallest weight. It is also possible to judge the value range of the weight. For example, when the weight is between 0 and 0.5, the weight value is 1; when the weight is between 0.5 and 1, the weight value is 1.5.

[0099] Among them, in the embodiments of the present application, the method for determining the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight according to the average similarity between the target text feature and the image feature information may include:

[0100] Perform a weighted summation process on the average similarity between the target text feature and the image feature information to obtain the global average similarity;

[0101] Determine the matching degree weight of the target text feature in the text feature information according to the global average similarity and the average similarity;

[0102] Obtain the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight value assignment rule.

[0103] Among them, the global average similarity may refer to the weighted sum of the average similarities between all text features in the text feature information and the image feature information. The global average similarity can be obtained by performing a weighted summation process on all average similarities. For example, when the average similarity between the first text feature and the image feature information is 0.3, the average similarity between the second text feature and the image feature information is 0.2, and the average similarity between the third text feature and the image feature information is 0.7, the calculation method of the global average similarity is 0.3 + 0.2 + 0.7, and the result is 1.2.

[0104] The method for determining the matching degree weight of the target text feature in the text feature information according to the global average similarity and the average similarity may include calculating the proportion of the similarity of each target text feature in the global average similarity. This proportion is the matching degree weight of each target text feature. For example, when the average similarity between the first text feature and the image feature information is 0.3, the average similarity between the second text feature and the image feature information is 0.2, the average similarity between the third text feature and the image feature information is 0.7, and the global average similarity is 1.2, the matching degree weight of the average similarity between the second text feature and the image feature information in the text feature information is 0.17.

[0105] Among them, in the embodiments of the present application, the method for obtaining the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight value assignment rule may include:

[0106] Perform a sorting process on the matching degree weights of the target text features in the text feature information to obtain a matching degree weight sequence;

[0107] According to the weight assignment rule, the matching degree weights in the matching degree weight sequence are assigned to obtain the weights of the matching degree weights. Among them, the earlier the position of the matching degree weight in the matching degree weight sequence, the greater the weight of the matching degree weight.

[0108] Among them, the matching degree weight sequence can refer to the sorting of each matching degree weight among all weights. For example, when the matching degree weights are 0.2, 0.3, and 0.5, the matching degree weight sequence of the weight 0.2 is 3, the matching degree weight sequence of the weight 0.3 is 2, and the matching degree weight sequence of the weight 0.5 is 1.

[0109] The weight assignment rule can be determined according to the evaluation keywords. For example, when there are five evaluation keywords, there can be five weights for assignment. When there are 5 evaluation keywords, but there are 3 levels of evaluation keywords, there can be 3 weights for assignment. When the evaluation keywords at the first level are "excellent" and "very good", the evaluation keywords at the second level are "average", and the evaluation keywords at the third level are "very poor" and "bad", the set weights are 5, 3, and 1 respectively.

[0110] The weight assignment rule can also be determined according to the sorting of each matching degree weight. For example, when the weight sequence is 1, the weight is 5; when the weight sequence is 2, the weight is 3; when the weight sequence is 3, the weight is 1. It is also possible to multiply the above two weight assignment rules to comprehensively obtain the weight of the matching degree weight.

[0111] S105. Obtain the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight of the matching degree weight.

[0112] Among them, the aesthetic quality evaluation result of the image to be evaluated can refer to the aesthetic evaluation obtained according to the similarity between the prompt text and the image generated by the prompt text. The higher the evaluation, the better the image quality is characterized. The aesthetic quality evaluation result is notified to the user by means of text messages, emails, etc.

[0113] Among them, in the embodiments of the present application, the method for obtaining the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight of the matching degree weight may include:

[0114] Obtain the matching degree weight value according to the matching degree weight and the weight of the matching degree weight;

[0115] Perform a weighted summation process on the matching degree weight value to obtain the matching degree weight;

[0116] Adjust the matching degree weight according to the preset weight distribution range to obtain the aesthetic quality evaluation result of the image to be evaluated.

[0117] The matching degree weight value may refer to a result obtained by combining the weight and the weight of the weight. For example, when the matching degree weight is 0.3 and the weight of the matching degree weight is 5, the matching degree weight value is 1.5.

[0118] The matching degree weight may refer to the result obtained by weighted summing the matching degree weights of all text features. For example, when the matching degree weight of the first text feature is 2 and the matching degree weight of the second text feature is 4, the matching degree weight is 6.

[0119] In order to ensure that all features are on the same scale numerically, adjusting the matching degree weight may refer to standardizing the matching degree weight according to a preset weight distribution range. In an embodiment of the present application, the weight distribution range may be [0, +∞]. The maximum and minimum values of the matching degree weight are first calculated, and the matching degree weight is scaled according to the maximum and minimum values to ensure that the matching degree weight is standardized to a numerical range starting from 0. For example, when the distribution of the matching degree weight is [1,5], the matching degree weight is calculated as (matching degree weight-1)*5 / 4, where 1 is the minimum value of the matching degree weight, 5 is the maximum value of the matching degree weight, and 4 is the result of the maximum value-minimum value of the matching degree weight to standardize the matching degree weight.

[0120] The image aesthetic quality assessment method provided in the embodiment of the present application can extract text features of multiple matching degrees for the prompt word text based on preset evaluation keywords. In order to comprehensively consider the local features and complete semantic information of the image to be evaluated, the text feature information is respectively calculated with the global image feature information and the local image feature information. Then, the similarity between the image to be evaluated and the text features of each matching degree and the matching degree weight of each text feature in all text features are obtained, thereby obtaining the aesthetic quality assessment result according to the preset weight assignment rules and matching degree weights. As a result, the image aesthetic assessment is performed based on the similarity between the image and the prompt word of the generated image, thereby improving the accuracy of the assessment.

[0121] Figure 2 A flow chart of another method for evaluating image aesthetic quality provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes:

[0122] S201: Acquire an image to be evaluated and a prompt word text for generating the image to be evaluated.

[0123] The image to be evaluated may be an image generated by an AI generation model based on a prompt word text.

[0124] S202. Obtain the evaluation text information and the characteristics of the evaluation text information according to the prompt text and the preset evaluation words, where the preset evaluation words include "very good", "average", and "very poor".

[0125] Among them, the evaluation text information can refer to the text that restricts the prompt word to different degrees of similarity.

[0126] S203. Obtain the global image, local image, global image characteristics, and local image characteristics according to the image to be evaluated.

[0127] Among them, the global image can refer to the image that contains all regions of the image, and the global image characteristics can refer to the vector information that characterizes all regions in the global image.

[0128] The local image can refer to the image of some regions in the image, and the local image characteristics can refer to the vector information of the image of some regions.

[0129] S204. Obtain the global similarity and local similarity according to the global image characteristics, local image characteristics, and the characteristics of the evaluation text information.

[0130] Among them, the global similarity includes the global similarity of "very good 'prompt word'", the global similarity of "average 'prompt word'", and the global similarity of "very poor 'prompt word'", and the local similarity includes the local similarity of "very good prompt word", the local similarity of "average prompt word", and the local similarity of "very poor prompt word".

[0131] S205. Obtain the average similarity of each evaluation text information characteristic according to the global similarity and local similarity.

[0132] Among them, the method of obtaining the average similarity of each evaluation text information characteristic can include taking the mean of the global similarity of "very good 'prompt word'" and the local similarity of "very good 'prompt word'" to obtain the average similarity of "very good 'prompt word'", taking the mean of the global similarity of "average 'prompt word'" and the local similarity of "average 'prompt word'" to obtain the average similarity of "average 'prompt word'", and taking the mean of the global similarity of "very poor 'prompt word'" and the local similarity of "very poor 'prompt word'" to obtain the average similarity of "very poor 'prompt word'".

[0133] S206. Obtain the aesthetic quality evaluation result of the image to be evaluated according to the average similarity of each evaluation text information characteristic.

[0134] Among them, the method for obtaining the aesthetic quality evaluation result of the image to be evaluated may include summing the average similarities of each evaluation text information feature to obtain the global similarity between the evaluation text information feature and the image to be evaluated as the aesthetic quality evaluation result of the image to be evaluated.

[0135] Another image aesthetic quality assessment method provided in an embodiment of the present application can limit the prompt word text according to the evaluation words "very good", "average", and "very poor". By comparing the local image features and global image features with the feature information of the prompt word text under evaluation words of different degrees, the average similarity is obtained, which enriches the dimensions of aesthetic assessment and improves the accuracy of image aesthetic assessment.

[0136] Figure 3 This is a schematic diagram of the structure of the image aesthetic quality assessment device provided in the embodiment of the present application. Figure 3 As shown, the image aesthetic quality assessment device 30 includes: an acquisition module 301, a first determination module 302, a second determination module 303, a third determination module 304, and a obtaining module 305.

[0137] Acquisition module 301 is used to acquire text feature information and image feature information. The text feature information is obtained by extracting features from the prompt text based on preset evaluation keywords. The evaluation keywords represent the degree of match between the prompt text and the image to be evaluated. The image feature information includes global image features of the image to be evaluated and local image features of the target local image in the image to be evaluated.

[0138] A first determining module 302 is configured to determine a global similarity between a global image feature and a target text feature, and a local similarity between a local image feature and a target text feature, based on the target text feature and the image feature information in the text feature information;

[0139] A second determination module 303 is used to determine the average similarity between the target text feature and the image feature information based on the global similarity and the local similarity;

[0140] A third determining module 304 is configured to determine, based on the average similarity between the target text feature and the image feature information, a matching degree weight of the target text feature in the text feature information and a value of the matching degree weight, wherein the value of the matching degree weight is determined based on a preset weight assignment rule and the size of the matching degree weight of the target text feature in the text feature information;

[0141] The obtaining module 305 is used to obtain the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight.

[0142] In the embodiment of the present application, the acquisition module 301 may also be used to:

[0143] Obtain prompt word text and image feature information;

[0144] Obtaining a prompt word evaluation text based on the prompt word text and preset evaluation keywords, wherein the number of evaluation keywords is not less than 2;

[0145] The prompt word evaluation text is input into the text encoder of the visual language model to obtain text feature information.

[0146] In the embodiment of the present application, the acquisition module 301 may also be used to:

[0147] Obtain text feature information and the image to be evaluated;

[0148] According to the preset size requirements, the image to be evaluated is scaled to obtain a global image of the image to be evaluated;

[0149] According to the preset size requirements, the image to be evaluated is intercepted and processed to obtain the target local image;

[0150] The global image and the target local image of the image to be evaluated are respectively input into the image encoder of the visual language model to obtain the global image features and the local image features;

[0151] Image feature information is obtained based on global image features and local image features.

[0152] In the embodiment of the present application, the second determining module 303 may also be used to:

[0153] Perform weighted averaging on the local similarities to obtain the initial average similarity;

[0154] According to the initial average similarity and the global similarity, the average similarity of the target text feature and image feature information is obtained.

[0155] In the embodiment of the present application, the third determining module 304 may also be used to:

[0156] Perform weighted summation on the average similarity of target text features and image feature information to obtain the global average similarity;

[0157] Determine the matching degree weight of the target text feature in the text feature information based on the global average similarity and the average similarity;

[0158] According to the matching degree weight of the target text feature in the text feature information and the preset weight assignment rule, the weight of the matching degree weight is obtained.

[0159] In the embodiment of the present application, the third determining module 304 may also be used to:

[0160] Sort the matching degree weights of the target text features in the text feature information to obtain a sequence of matching degree weights;

[0161] According to the weight assignment rule, assign values to the matching degree weights in the sequence of matching degree weights to obtain the weights of the matching degree weights. Among them, the earlier the position of the matching degree weight in the sequence of matching degree weights, the greater the weight of the matching degree weight.

[0162] In the embodiment of the present application, the obtaining module 305 can also be used for:

[0163] Obtain the matching degree weight value according to the matching degree weight and the weight of the matching degree weight;

[0164] Perform a weighted summation process on the matching degree weight values to obtain the matching degree weight value;

[0165] Adjust the matching degree weight value according to the preset weight distribution range to obtain the aesthetic quality evaluation result of the image to be evaluated.

[0166] As can be seen from the above, the image aesthetic quality assessment device of the embodiment of the present application is composed of an acquisition module 301, which is used to obtain text feature information and image feature information. The text feature information is the text feature obtained after feature extraction of the prompt word text according to the preset evaluation keywords. The evaluation keywords represent the degree of matching between the prompt word text and the image to be evaluated. The image feature information includes the global image features of the image to be evaluated and the local image features of the target local image in the image to be evaluated; the first determination module 302 is used to determine the global similarity between the global image features and the target text features, and the local similarity between the local image features and the target text features according to the target text features and image feature information in the text feature information. similarity; a second determining module 303 is used to determine the average similarity between the target text feature and the image feature information based on the global similarity and the local similarity; a third determining module 304 is used to determine the matching degree weight of the target text feature in the text feature information and the weight of the matching degree weight based on the average similarity between the target text feature and the image feature information, wherein the weight of the matching degree weight is determined according to a preset weight assignment rule and the size of the matching degree weight of the target text feature in the text feature information; an obtaining module 305 is used to obtain the aesthetic quality evaluation result of the image to be evaluated based on the matching degree weight of the target text feature in the text feature information and the weight of the matching degree weight. Therefore, the embodiment of the present application can extract text features of multiple matching degrees for the prompt word text based on the preset evaluation keywords. In order to comprehensively consider the local features and complete semantic information of the image to be evaluated, the text feature information is respectively calculated with the global image feature information and the local image feature information. Then, the similarity between the image to be evaluated and the text features of each matching degree and the matching degree weight of each text feature in all text features are obtained, so as to obtain the aesthetic quality evaluation result according to the preset weight assignment rules and matching degree weights, and produce an image aesthetic evaluation based on the similarity between the image and the prompt word of the generated image, thereby improving the evaluation accuracy.

[0167] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 4 As shown, the electronic device 40 includes:

[0168] The electronic device 40 may include one or more processors 401 , one or more computer-readable storage media memories 402 , a communication component 403 , and other components. The processor 401 , the memory 402 , and the communication component 403 are connected via a bus 404 .

[0169] In a specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402 , so that the at least one processor 401 performs the above-mentioned image aesthetic quality assessment method.

[0170] For the specific implementation process of the processor 401, reference may be made to the above method embodiments. Their implementation principles and technical effects are similar, and thus will not be elaborated herein.

[0171] In the above Figure 4 In the illustrated embodiment, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention may be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0172] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0173] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0174] In some embodiments, a computer program product is also proposed, including a computer program or instruction, which when executed by a processor implements the steps in any of the above image aesthetic quality assessment methods.

[0175] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.

[0176] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware with instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0177] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the image aesthetic quality evaluation methods provided by the embodiments of the present application.

[0178] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0179] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium.

[0180] Since the instructions stored in the storage medium can execute the steps in any one of the image aesthetic quality evaluation methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the image aesthetic quality evaluation methods provided by the embodiments of the present application can be realized. For details, please refer to the previous embodiments and will not be elaborated here.

[0181] Those skilled in the art will readily think of other implementation manners of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0182] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An image aesthetic quality assessment method, characterized in that, Applied to an image aesthetic quality calculation system, including: Obtain text feature information and image feature information. The text feature information is the text feature obtained by extracting features from the prompt text according to preset evaluation keywords, and the evaluation keywords represent the matching degree between the prompt text and the image to be evaluated. The image feature information includes the global image feature of the image to be evaluated and the local image feature of the target local image in the image to be evaluated; According to the target text feature in the text feature information and the image feature information, determine the global similarity between the global image feature and the target text feature, and the local similarity between the local image feature and the target text feature; the target text feature refers to one of the text feature information describing the different degrees of relevance between the image and the prompt; According to the global similarity and the local similarity, determine the average similarity between the target text feature and the image feature information; According to the average similarity between the target text feature and the image feature information, determine the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight, where the weight value of the matching degree weight is determined according to the preset weight assignment rule and the size of the matching degree weight of the target text feature in the text feature information; among them, the preset weight assignment rule is that the earlier the position of the matching degree weight in the matching degree weight sequence, the greater the weight value of the matching degree weight, and the matching degree weight sequence is the sequence obtained by sorting the matching degree weights; According to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight, obtain the aesthetic quality evaluation result of the image to be evaluated; The obtaining the aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight includes: According to the matching degree weight and the weight value of the matching degree weight, obtain the matching degree weight value; Perform weighted summation processing on the matching degree weight value to obtain the matching degree weight value; According to the preset weight distribution range, adjust the matching degree weight value to obtain the aesthetic quality evaluation result of the image to be evaluated; The obtaining the text feature information and the image feature information includes: Obtain the prompt text and the image feature information; According to the prompt text and the preset evaluation keywords, obtain the prompt evaluation text, where the number of the evaluation keywords is not less than 2; Input the prompt evaluation text into the text encoder of the vision-language model to obtain the text feature information.

2. The method according to claim 1, characterized in that, The obtaining the text feature information and the image feature information includes: Obtain the text feature information and the image to be evaluated; According to the preset size requirement, perform scaling processing on the image to be evaluated to obtain the global image of the image to be evaluated; Intercept the to-be-evaluated image according to the preset size requirement to obtain the target local image; Input the global image of the to-be-evaluated image and the target local image into the image encoder of the vision-language model respectively to obtain the global image feature and the local image feature; Obtain the image feature information according to the global image feature and the local image feature.

3. The method according to claim 1, wherein The determining the average similarity between the target text feature and the image feature information according to the global similarity and the local similarity includes: Perform weighted averaging on the local similarity to obtain an initial average similarity; Obtain the average similarity between the target text feature and the image feature information according to the initial average similarity and the global similarity.

4. The method according to claim 1, wherein The determining the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight according to the average similarity between the target text feature and the image feature information includes: Perform weighted summation on the average similarity between the target text feature and the image feature information to obtain a global average similarity; Determine the matching degree weight of the target text feature in the text feature information according to the global average similarity and the average similarity; Obtain the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight value assignment rule.

5. The method according to claim 4, characterized in that, The obtaining the weight value of the matching degree weight according to the matching degree weight of the target text feature in the text feature information and the preset weight value assignment rule includes: Sort the matching degree weights of the target text feature in the text feature information to obtain a matching degree weight sequence; Perform assignment processing on the matching degree weights in the matching degree weight sequence according to the weight value assignment rule to obtain the weight value of the matching degree weight.

6. An image aesthetic quality evaluation device, characterized in that, including: An acquisition module, configured to acquire text feature information and image feature information, where the text feature information is text features obtained by performing feature extraction on a prompt text according to a preset evaluation keyword, the evaluation keyword represents the matching degree between the prompt text and the to-be-evaluated image, and the image feature information includes the global image feature of the to-be-evaluated image and the local image feature of the target local image in the to-be-evaluated image; A first determination module, configured to determine the global similarity between the global image feature and the target text feature and the local similarity between the local image feature and the target text feature according to the target text feature in the text feature information and the image feature information; the target text feature refers to one of multiple text feature information describing different degrees of relevance between the image and the prompt; A second determination module, configured to determine the average similarity between the target text feature and the image feature information according to the global similarity and the local similarity; A third determination module, configured to determine a matching degree weight of the target text feature in the text feature information and a weight value of the matching degree weight according to an average similarity between the target text feature and the image feature information, where the weight value of the matching degree weight is determined according to a preset weight assignment rule and a magnitude of the matching degree weight of the target text feature in the text feature information; where the preset weight assignment rule is that the more forward the position of the matching degree weight in the matching degree weight sequence, the greater the weight value of the matching degree weight, and the matching degree weight sequence is a sequence obtained by sorting the matching degree weights; An obtaining module, configured to obtain an aesthetic quality evaluation result of the image to be evaluated according to the matching degree weight of the target text feature in the text feature information and the weight value of the matching degree weight; The obtaining module is configured to obtain a matching degree weight value according to the matching degree weight and the weight value of the matching degree weight; Perform a weighted summation process on the matching degree weight value to obtain a matching degree weight value; Adjust the matching degree weight value according to a preset weight distribution range to obtain the aesthetic quality evaluation result of the image to be evaluated; The obtaining module is specifically configured to obtain the prompt text and the image feature information; Obtain a prompt evaluation text according to the prompt text and the preset evaluation keywords, where the number of the evaluation keywords is not less than 2; Input the prompt evaluation text into a text encoder of a vision-language model to obtain the text feature information.

7. An electronic device, characterized in that, Includes: A processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the image aesthetic quality evaluation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image text matching model training method, bidirectional search method and related device

    CN108288067A

  • AI evaluation method and system based on AI generation content matching degree evaluation

    CN118245760A

  • Unified visual language model pre-training and adjusting method for image quality and aesthetic evaluation

    CN118607611A