Image evaluation method and electronic equipment
By extracting key image information to generate question information and using a visual question-answering model for evaluation, combined with unsupervised quality evaluation, the problem of evaluating the consistency between image content and input text description in generative artificial intelligence image generation technology is solved, and efficient and accurate screening of generated images is achieved.
Patent Information
- Application Number
- CN202510897985.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-03
AI Technical Summary
Existing generative AI image generation technology faces challenges in the stability and controllability of output results, especially the semantic consistency between image content and input text description is difficult to effectively evaluate.
By acquiring the target image and extracting key information, generating question information corresponding to the key information, using the visual question answering model to conduct question answering, generating evaluation results, and combining unsupervised image quality evaluation methods for comprehensive evaluation.
It realizes the automated and refined evaluation of the content details of the generated image, can accurately judge whether the image is faithful to the input information, and improves the accuracy and reliability of the quality screening of the generated image.
Smart Images

Figure CN120747010A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to an image evaluation method and electronic device. Background Art
[0002] Generative AI technologies, exemplified by large, multimodal models, are rapidly developing. Text-to-image generation, in particular, can generate relevant, visually rich images based on user-provided text descriptions. This technology demonstrates enormous potential for application in a variety of fields, including content creation, art design, virtual scene construction, and data augmentation.
[0003] However, despite the growing power of generative models, the stability and controllability of their outputs still face challenges. In practical applications, the quality of images generated by the model is often uncertain, and not all generated images can meet specific application requirements. Currently, for the screening and evaluation of generated images, some automated evaluation methods focus on the visual fidelity of the image itself, such as evaluating the image's clarity, noise level, or the presence of artifacts. However, such methods may ignore the semantic consistency between the image content and the input text description, resulting in the screened images having good quality but possibly biased content. Summary of the Invention
[0004] In view of this, the present disclosure provides an image evaluation method and an electronic device.
[0005] One aspect of the present disclosure provides an image evaluation method, including: acquiring a target image, where the target image is an image generated based on target input information, the target input information including multiple key information; generating multiple question information based on the target input information, where the question information corresponds to at least one key information; using the target image and the question information as inputs of a visual question answering model, and generating answer information for each question information based on the visual question answering model; and generating a first evaluation result based on the multiple answer information.
[0006] According to an embodiment of the present disclosure, multiple question information is generated based on the target input information, including: generating corresponding question information based on each key information; and / or obtaining at least one associated information based on the target input information, the associated information representing the association relationship between multiple first key information in the multiple key information; generating at least one question information corresponding to the multiple key information based on each associated information.
[0007] According to an embodiment of the present disclosure, at least one question information corresponding to multiple key information is generated, including: generating at least one question information based on each associated information and multiple first key information corresponding to the associated information, and the question information is used to verify whether the contents corresponding to the multiple first key information in the target image satisfy the association relationship.
[0008] According to an embodiment of the present disclosure, question information corresponding to multiple key information is generated, including: selecting at least one second key information different from the first key information from the multiple key information, and no association relationship exists between any second key information and the first key information; generating at least one question information based on the association information, at least one first key information, and at least one second key information, and the question information is used to verify whether the content corresponding to the first key information and the content corresponding to the second key information in the target image do not satisfy the association relationship.
[0009] According to an embodiment of the present disclosure, the number of question information is greater than the number of key information.
[0010] According to an embodiment of the present disclosure, the question information is a question about the existence of key information and / or the relationship between multiple key information.
[0011] According to an embodiment of the present disclosure, generating a first evaluation result includes: generating the first evaluation result according to the proportion of answer information that is a positive answer in all answer information.
[0012] According to an embodiment of the present disclosure, the image evaluation method further includes: evaluating the target image according to an unsupervised image quality evaluation method to generate a second evaluation result; and generating a target evaluation result for the target image by combining the first evaluation result and the second evaluation result.
[0013] According to an embodiment of the present disclosure, the image evaluation method further includes: in response to the target evaluation result being greater than a preset threshold, treating the target image as an image that meets preset quality requirements.
[0014] Another aspect of the present disclosure provides an image evaluation device, including: a first acquisition module, used to acquire a target image, where the target image is an image generated based on target input information, and the target input information includes multiple key information; a first generation module, used to generate multiple question information based on the target input information, where the question information corresponds to at least one key information; a second generation module, used to use the target image and question information as input of a visual question answering model, and generate answer information for each question information based on the visual question answering model; and a third generation module, used to generate a first evaluation result based on the multiple answer information.
[0015] Another aspect of the present disclosure provides an electronic device comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image evaluation method of any one of the aforementioned embodiments.
[0016] Another aspect of the present disclosure provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the image evaluation method according to any one of the aforementioned embodiments.
[0017] Another aspect of the present disclosure provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the operation of the image evaluation method of any of the aforementioned embodiments is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0019] Figure 1 The following schematically shows a flow chart of an image evaluation method according to an embodiment of the present disclosure;
[0020] Figure 2 Schematically shows a flow chart of generating question information in an image evaluation method according to an embodiment of the present disclosure;
[0021] Figure 3 Schematically shows another flow chart of generating question information in the image evaluation method according to an embodiment of the present disclosure;
[0022] Figure 4 Schematically shows another flow chart of generating question information in the image evaluation method according to an embodiment of the present disclosure;
[0023] Figure 5 Another flow chart of the image evaluation method according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 6 Another flow chart of the image evaluation method according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 7 Another flow chart of the image evaluation method according to an embodiment of the present disclosure is schematically shown;
[0026] Figure 8 Schematically shows an overall flow chart according to an embodiment of the present disclosure;
[0027] Figure 9A block diagram schematically illustrates an image evaluation apparatus according to an embodiment of the present disclosure; and
[0028] Figure 10 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0030] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0032] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0033] In the embodiments of this disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.
[0034] An embodiment of the present disclosure provides an image evaluation method, including: acquiring a target image, where the target image is an image generated based on target input information, and the target input information includes multiple key information; generating multiple question information based on the target input information, and the question information corresponds to at least one key information; using the target image and the question information as inputs of a visual question answering model, and generating answer information for each question information based on the visual question answering model; and generating a first evaluation result based on the multiple answer information.
[0035] Figure 1 The flowchart of the image evaluation method according to the embodiment of the present disclosure is schematically shown.
[0036] like Figure 1 As shown, the image evaluation method may include at least operations S110 to S140.
[0037] In operation S110, a target image is acquired. The target image is an image generated based on target input information, and the target input information includes multiple key information. In this embodiment, the target input information may refer to one or more paragraphs of text used to guide the image generation model to generate an image, such as a descriptive text prompt. Key information may refer to core elements extracted from the target input information that are expected to be reflected in the target image, such as specific objects, object attributes, scene style, etc. The target image is the image file output by the image generation model (for example, a text-based image model based on a diffusion model) after receiving the target input information.
[0038] Specifically, obtaining the target image can include reading an image pre-generated by the image generation model from an image library, or invoking the image generation model in real time and inputting the target input information to generate the target image. For example, a folder containing hundreds of images to be evaluated can be used to read one image as the target image, and the target input information associated with the image can also be read simultaneously.
[0039] For example, if the target input is "a photo of a red apple on a wooden table," the target image is the image generated by the image generation model based on the text. The words "wooden table," "red," and "apple" can all be used as key information extracted from the target input.
[0040] In operation S120, multiple question messages are generated based on the target input information. Each question message corresponds to at least one key information. A question message can be a data structure used to determine whether the target image contains specific content or features. Typically, it can be expressed as a question in natural language, with its content corresponding to the key information. Each question message is used to inquire about the presence or expression of one or more key information in the image.
[0041] Specifically, a preset language model can be used to convert one or more key information contained in the target input information into one or more forms of question information. This process can be configured to be executed automatically, that is, when the target input information is input, the language model can output multiple question information corresponding to the key information contained therein.
[0042] For example, continuing with the above example, for the key information "apple", the question information "Is there an apple in the image?" can be generated; for the key information "wooden table", the question information "Is there a table in the image?" can be generated; for the key information "red", the question information "Does the image contain red elements?" can be generated.
[0043] In operation S130, the target image and question information are used as input to the visual question answering model, which then generates answers to each question. The visual question answering model is a multimodal artificial intelligence model that can understand image content and answer related questions. The answers are the responses given by the visual question answering model to each input question after analyzing the content of the target image.
[0044] Operation S130 may include simultaneously inputting target image data (e.g., an image file encoded in JPEG or PNG format) and a question (e.g., a text string) into an interface of a deployed visual question answering model. After processing, the model returns a text string or an identifier as the answer. This process is repeated for each question until all questions have corresponding answers.
[0045] For example, if the target image generated above, "There is a red apple on a wooden table," is fed into a visual question answering model along with the question "Is there an apple in the image?", the model might output "yes" or "yes." If the target image is then fed into another question, "Is there a car in the image?", the model might output "no" or "not found."
[0046] In operation S140, a first evaluation result is generated based on the multiple answer information. The first evaluation result can be an indicator for characterizing the degree of consistency between the target image and the target input information at the content level. For example, all the answer information output by the visual question-answering model can be aggregated and processed, and this set of discrete answer information sets can be used as input according to a preset evaluation logic or evaluation model, and processed to generate a comprehensive evaluation score or evaluation grade as the first evaluation result. For example, assuming that for a target image, a set of three answer information (for example, "yes", "no", and "yes") is obtained, after obtaining the answer set, a first evaluation result that can characterize the overall evaluation situation is output according to the established processing rules, such as a value of "0.75" or a label of "basically in line".
[0047] According to the disclosed embodiments, by extracting key information from the target input and generating corresponding question information, and then using a visual question-answering model to answer questions about the target image one by one, it is possible to achieve an automated and refined assessment of the details of the generated image content. This assessment of the generated image is deepened from a macroscopic, global feature similarity comparison to an alignment test of the local, specific elements of the image, thereby more accurately determining whether the generated image faithfully reflects the specific requirements of the input information.
[0048] Figure 2 The flowchart of generating question information in the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0049] like Figure 2 As shown, based on the above embodiment, operation S120 may include operation S210 and / or operations S220 to S230.
[0050] In operation S210, corresponding question information is generated based on each key information. Operation S210 performs an existence check on each independent, core component element decomposed from the target input information, with each key information being considered as an independent evaluation point. Specifically, operation S210 may include traversing all extracted key information and applying a preset question generation rule or template to each key information, thereby generating one or more question information for each key information.
[0051] For example, if the target input information is "a cat is sleeping on a blue sofa," the key information extracted may include "cat," "sofa," and "blue." Then, questions such as "Is there a cat in the image?" will be generated for "cat," "Is there a sofa in the image?" will be generated for "sofa," and "Does the image contain blue?" will be generated for "blue."
[0052] And / or operations S230~S240.
[0053] In operation S220, at least one piece of association information is obtained based on the target input information. The association information represents the association relationship between multiple pieces of first key information in the multiple key information. Association information may be information used to describe the logical, spatial, attribute, or state relationship between two or more pieces of key information. The first key information refers to the key information linked together by the association information.
[0054] For example, natural language understanding (NLU) analysis can be performed on the target input information, such as syntactic analysis or semantic role labeling, to identify the relationships between different key information (words or phrases). For example, it can identify which object performs which action, or the spatial relationship between one object and another.
[0055] For example, continuing with the example "A cat is sleeping on a blue sofa," we can analyze and obtain a piece of association information representing a spatial relationship, linking the two first key pieces of information, "cat" and "sofa." At the same time, we can also obtain a piece of association information representing an attribute relationship, linking the two first key pieces of information, "sofa" and "blue."
[0056] In operation S230, at least one question information corresponding to the plurality of key information is generated based on the associated information. The single question information generated in operation S230 may simultaneously involve the plurality of key information and be used to verify the embodiment of the key information as a whole combination in the image.
[0057] For example, based on the previously identified association information linking "cat" and "sofa," a question corresponding to these two first key information items can be generated, such as: "Do both a cat and a sofa appear in the image?" Similarly, based on the association information linking "sofa" and "blue," a question can be generated, such as: "Is there a sofa in the image, and is the sofa associated with the color blue?"
[0058] According to an embodiment of the present disclosure, by providing two question generation paths, a more comprehensive and in-depth evaluation of the generated image content can be achieved. Operation S210 ensures the existence of basic elements in the image, while operations S220 to S230 further explore whether the correct combination relationship is formed between these elements. This progressive evaluation from "element existence" to "element combination" can more effectively determine whether the generated image accurately understands and reproduces the scene structure and complex semantics contained in the target input information, compared to evaluating only a single element.
[0059] Figure 3Another flowchart of generating question information in the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0060] like Figure 3 As shown, based on the above embodiment, operation S230 may include operation S310.
[0061] In operation S310, at least one question information is generated according to each piece of association information and a plurality of first key information corresponding to the association information. The question information is used to verify whether the contents corresponding to the plurality of first key information in the target image satisfy an association relationship.
[0062] Operation S310 is used to generate a more refined question information, the purpose of which is not only to confirm whether multiple key information exist at the same time, but also to further explore whether the contents of these key information presented in the image conform to the specific relationship described in the target input information.
[0063] For example, by analyzing the specific semantics of the associated information (e.g., whether it represents spatial location, a subordinate relationship, or dynamic behavior) and combining it with the multiple first key information it connects to, a question sentence can be constructed to directly inquire about the specific relationship. For example, when the associated information is a verb, the question information can be constructed to verify whether the action exists between the subject and the object; when the associated information is a preposition, the question information can be constructed to verify whether the spatial relationship exists between the two objects.
[0064] For example, if the target input information is "a boy is wearing a yellow hat," the associated information "wearing" and its corresponding first key information "boy" and "hat" can be identified. Based on this, a question information is generated to directly verify the "wearing" relationship, such as: "Is the boy in the image wearing a hat?" This question directly verifies the wearing relationship between "boy" and "hat," rather than simply confirming whether the two appear at the same time. For another example, if the target input information is "a vase is on the table," the associated information "on..." and its corresponding first key information "vase" and "table" can be identified, and the question information is generated: "Is the vase on the table?" to verify the spatial position relationship.
[0065] According to the disclosed embodiments, by generating question information used to verify specific associations between key information, the semantic accuracy of the generated image can be deeply assessed. Compared to simply confirming the existence of multiple elements, the ability to determine whether these elements are correctly organized according to the requirements of the input information improves the precision and reliability of the assessment, effectively distinguishing low-quality images that are simply keyword-stacked but contain scene logic errors, thereby ensuring that the selected images are highly faithful to the original input information in both content and structure.
[0066] Figure 4 Another flowchart of generating question information in the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0067] like Figure 4 As shown, based on the above embodiment, operation S310 may include operations S410 to S420.
[0068] In operation S410, at least one second key information different from the first key information is selected from the plurality of key information, where no association relationship exists between any of the second key information and the first key information. The second key information may be key information that exists in the target input information but is not directly associated with the specific associated information currently being analyzed. The purpose of selecting this second key information is to construct a control or interference item for subsequent negative verification.
[0069] For example, first, a piece of associated information and its corresponding multiple first key information are determined, then these first key information are excluded from the entire key information set of the target input information, and one or more of the remaining key information are selected as the second key information.
[0070] For example, if the target input is "a glass placed on a wooden table with a book next to it," when analyzing the associated information "placed on...", the first key information corresponding to it is "glass" and "wooden table." In this case, "book" can be selected as the second key information from the remaining key information "book" because there is no direct association between "book" and "glass" in the original input.
[0071] In operation S420, at least one question is generated based on the association information, the at least one first key information, and the at least one second key information. The question is used to verify whether the content corresponding to the first key information and the content corresponding to the second key information in the target image do not satisfy the association relationship. Operation S420 is used to generate a reverse verification question, that is, by incorrectly applying the original association relationship to an unrelated object to verify whether the generated model has produced incorrect associations or feature "leakage."
[0072] For example, a first key information (such as a subject), a related information (such as a relationship), and a second key information (such as an incorrect object) are combined into a question sentence. The structure of the question sentence is used to verify whether such an incorrect association has not occurred in the image.
[0073] For example, continuing with the example above, using the first key information "glass," the associated information "placed on...", and the second key information "book," a negative verification question can be generated, such as: "Is the glass not placed on the book?" If the generated image is correct (i.e., the glass is on the table, and the book is next to it), the visual question answering model's expected answer to this question should be affirmative (e.g., "yes"). For another example, if the target input information is "a girl wearing red clothes and a boy wearing blue pants," when verifying the association between "girl" and "red clothes," "boy" can be selected as the second key information and the question "Is the girl not wearing blue pants?" can be generated to verify whether the color attribute has been incorrectly assigned.
[0074] According to the embodiment of the present disclosure, by introducing negative verification questions, the rigor and robustness of the evaluation can be significantly improved. It can not only confirm whether the correct association relationship exists (such as Figure 3 This combined "positive confirmation" and "negative exclusion" evaluation strategy can more effectively identify generated images that are simply random stacking of elements without accurately understanding and expressing the precise relationships between them, significantly improving the accuracy of selecting high-quality images.
[0075] According to an embodiment of the present disclosure, the number of question information is greater than the number of key information. For example, this quantitative transcendence relationship can be achieved in a variety of ways. For example, not only is a basic question generated for each key information to verify its existence, but additional questions can also be generated for some or all of the key information to explore its specific attributes, status, or relationship with other key information. Through this "one-to-many" or "many-to-many" question generation strategy, the total number of question information can naturally be greater than the total number of key information.
[0076] For example, if the target input is "a white cat on a red carpet," four key pieces of information can be extracted: "cat," "white," "carpet," and "red." For a more comprehensive assessment, five question pieces can be generated: "Is there a cat in the image?" (corresponding to the key piece "cat"), "Is there a carpet in the image?" (corresponding to the key piece "carpet"), "Is the cat in the image white?" (corresponding to the key pieces "cat" and "white"), "Is the carpet in the image red?" (corresponding to the key pieces "carpet" and "red"), and "Is the cat on the carpet?" (corresponding to multiple key pieces of information and their associated relationships).
[0077] According to the disclosed embodiments, by ensuring that the amount of problematic information exceeds the amount of critical information, a more comprehensive and detailed verification of the generated image can be performed. Rather than simply checking the existence of isolated elements, this can cover more local details, such as whether the attributes of objects are correct and whether the relationships between objects meet the description. This increased "resolution" of the assessment can more effectively identify images that contain all key elements but have deviations in their combination or detailed attributes, significantly improving the depth and accuracy of the assessment.
[0078] According to an embodiment of the present disclosure, the question information is a question about the existence of key information and / or the association between multiple key information. The question information is not a simple keyword or declarative sentence, but information in the form of a question that can guide the visual question-answering model to make judgments and answers. Specifically, it is divided into two cases. The first is a question about "existence", the purpose of which is to verify whether the entity or attribute corresponding to a single key information exists in the target image. The second is a question about "association", the purpose of which is to verify whether the entities or attributes corresponding to multiple key information present the specific relationship described in the target input information. In actual applications, the multiple question information generated may contain only one type of question, or it may be a combination of two types of question.
[0079] For example, if the target input information is "a man holding a black umbrella," a question about the key information "man" can be generated regarding its "presence," such as "Is there a man in the image?" A question about the key information "umbrella" and "black" can be generated regarding their "association" (attribute association), such as "Is the umbrella in the image black?" A question about the key information "man" and "umbrella" can be generated regarding their "association" (behavior association), such as "Is the man in the image holding an umbrella?"
[0080] According to the embodiment of the present disclosure, by clearly defining the question information as interrogative sentences about "existing situations" and "related situations", standardized and operational input is provided for subsequent evaluation using the visual question answering model. It can directly guide the visual question answering model to conduct targeted analysis and judgment on the specific content and structure of the image, making the evaluation process more focused and efficient. Compared with using vague instructions or declarative descriptions, the use of structured interrogative sentences can ensure that the focus of the evaluation is precisely aligned with the core elements and semantic relationships in the target input information.
[0081] In another embodiment, the question information can be configured in another format. The question information can be declarative information to be verified regarding the existence of key information and / or the relationship between multiple key information. "Declarative information to be verified" refers to a data structure in the form of an affirmative or negative statement. It does not pose a question itself, but rather a factual assertion, with the goal of allowing the visual question answering model to determine whether the assertion is consistent with the content of the target image.
[0082] For example, the process of generating question information is transformed into generating a series of propositions to be verified. For example, instead of generating a question sentence like "Is there a cat in the image?", a statement like "There is a cat in the image" is generated. The task of the visual question answering model has also changed from "answering questions" to "verifying statements." The answer information it outputs represents the confirmation (e.g., "yes" or "true") or denial (e.g., "no" or "false") of the statement, or the degree of match to the statement content (e.g., a match score).
[0083] For example, if the target input information is "a man holding a black umbrella," the presence of the key information "man" can generate the declarative information to be verified: "The image contains a man." The association between the key information "umbrella" and "black" can generate the declarative information to be verified: "The umbrella in the image is black." The association between the key information "man" and "umbrella" can generate the declarative information to be verified: "The man in the image is holding an umbrella."
[0084] Figure 5 Another flowchart of the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0085] like Figure 5 As shown, based on the above embodiment, operation S140 may include operation S510.
[0086] In operation S510, a first evaluation result is generated based on the proportion of answer information that is a positive answer in all answer information. A "positive answer" refers to an answer output by the visual question answering model that indicates an affirmative response to the content of the question information. A "positive answer" refers to answer information output by the visual question answering model that is used to affirm or confirm the existence of the inquiry content in the question information. For example, for the question "Is there a cat in the image?", the answer information "yes", "exists" or "has" can all be defined as a positive answer. Accordingly, all answer information is the totality of the answer information obtained for all question information.
[0087] For example, first, all the answer information returned by the visual question answering model is traversed and classified, and the total number of answer information that is positive answer is counted; then, the total number of all answer information is obtained; finally, the total number of positive answers is divided by the total number of all answer information, and the obtained quotient (i.e., ratio or percentage) is used as the first evaluation result.
[0088] For example, suppose 10 questions are generated for a target image and 10 corresponding answers are obtained. After analyzing these 10 answers, it is found that 7 are "yes" (defined as positive answers) and 3 are "no." The percentage is then calculated: 7 / 10 = 0.7. Therefore, the first evaluation result is 0.7. This value intuitively indicates that the target image meets the target input requirements in 70% of the details.
[0089] According to the disclosed embodiments, by calculating the proportion of positive responses to generate an evaluation result, qualitative judgments about image content can be converted into a standardized, objective quantitative score. This score intuitively reflects the extent to which the generated image is faithful to the detailed requirements of the input information, allowing for direct comparison of content alignment between different images.
[0090] In another embodiment, in combination with the aforementioned embodiment, the answer information represents the degree of match with the statement content (e.g., a match score). For example, the match score can be a floating point number between 0 and 1, where 1 indicates a perfect match and 0 indicates a complete mismatch. For each piece of statement information to be verified, the corresponding match score is obtained from the visual question answering model; then, a preset mathematical operation is performed on the entire set of obtained match scores to derive a comprehensive indicator that can represent the overall match level. This mathematical operation can be an arithmetic mean, weighted mean, or geometric mean of all scores. This comprehensive indicator serves as the first evaluation result.
[0091] For example, suppose the system generates three statements for verification for a target image: 1. "There is a car in the image," 2. "The car is red," and 3. "The car is moving on the road." After verification, the visual question answering model generates three matching scores as responses: 0.99 for statement 1, 0.95 for statement 2, and 0.80 for statement 3 (possibly indicating that the car is stationary, not moving). Taking the arithmetic average, (0.99 + 0.95 + 0.80) / 3 ≈ 0.913. Therefore, the first evaluation result is 0.913.
[0092] Figure 6Another flowchart of the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0093] like Figure 6 As shown, based on the above embodiment, the image evaluation method may further include operations S610 to S620.
[0094] In operation S610, the target image is evaluated according to an unsupervised image quality assessment method to generate a second assessment result. Unsupervised image quality assessment methods are algorithms that directly analyze the intrinsic characteristics of the target image (such as clarity, contrast, noise level, artifacts, etc.) without referencing an ideal, undistorted original image and output a quantitative score. The "second assessment result" is the score output by the algorithm that represents the visual quality of the target image itself.
[0095] For example, the target image is input into a pre-trained unsupervised image quality assessment model (e.g., BRISQUE, NIQE, or similar models). The model extracts and analyzes the statistical features of the natural scene of the image and outputs a numerical value. This numerical value is the second evaluation result and can be configured so that a higher score represents better image quality. For example, the target image generated above is input into an image quality assessment model. After analysis, the model outputs a score, such as "85.4" (assuming a percentage system). This score of "85.4" is the second evaluation result, which is independent of the image content and only reflects the technical quality of the image, such as clarity and color naturalness.
[0096] In operation S620, the first evaluation result and the second evaluation result are combined to generate a target evaluation result for the target image. The target evaluation result is a comprehensive final score that considers both the content accuracy (represented by the first evaluation result) and the visual quality (represented by the second evaluation result) of the image.
[0097] For example, this can be achieved by taking a weighted sum of the first and second evaluation results. Preset weighting coefficients are assigned to the first and second evaluation results, reflecting the relative importance of content accuracy and visual quality in the final evaluation. The target evaluation result is obtained by multiplying the respective scores by their corresponding weighting coefficients and then adding them together.
[0098] For example, assuming that the above operations have obtained: a first evaluation result (for example, according to Figure 7The calculated percentage of positive responses for the content of the illustrated embodiment is 0.9; the second evaluation result (image visual quality) is 85.4. For ease of calculation, the second evaluation result is first normalized to the range of 0-1, for example, 85.4 becomes 0.854. The weight of the first evaluation result is set to 0.6, and the weight of the second evaluation result is set to 0.4. The calculation process for the target evaluation result is: (0.9 × 0.6) + (0.854 × 0.4) = 0.54 + 0.3416 = 0.8816. Therefore, the generated target evaluation result is 0.8816.
[0099] According to the disclosed embodiments, by introducing unsupervised image quality assessment and combining it with content alignment assessment, a more comprehensive and multi-dimensional evaluation of generated images is achieved. This overcomes the limitations of single-dimensional evaluation and avoids screening out images with correct content but poor quality (e.g., blurry or noisy), or images with excellent quality but severely inconsistent content with the input information. This dual standard, which considers both the "overall" and "detailed" aspects, ensures that the final target evaluation results more realistically and completely reflect the comprehensive usability of the generated images.
[0100] Figure 7 Another flowchart of the image evaluation method according to an embodiment of the present disclosure is schematically shown.
[0101] like Figure 7 As shown, based on the above embodiment, the image evaluation method may further include operation S710.
[0102] In operation S710, in response to the target evaluation result being greater than a preset threshold, the target image is considered to meet the preset quality requirements. The "preset threshold" is a benchmark score used to determine whether an image is qualified. For example, the target evaluation result calculated in the above operation is compared with a preset value (i.e., the preset threshold). If the target evaluation result is greater than the threshold, the target image is labeled "qualified," "passed," or a similar label and is included in the list of qualified images. Conversely, if the target evaluation result is not greater than the threshold, it can be marked as "unqualified" and discarded or included in another list.
[0103] For example, assume that the preset threshold is set to 0.85. For the target image whose target evaluation result is 0.8816 obtained in the previous embodiment, since 0.8816 is greater than 0.85, the image will be judged to meet the preset quality requirements and will be retained. For another target image, if its calculated target evaluation result is 0.79, since 0.79 is not greater than 0.85, the image will be judged to not meet the requirements. It should be noted that the preset threshold can be static, that is, a fixed value (such as always 0.85); it can also be dynamic, for example, according to the total number of image batches to be evaluated, it is dynamically adjusted to only retain the images ranked in the top 10% of the evaluation results. At this time, the threshold is the image score at the 10% position.
[0104] Figure 8 The figure schematically shows an overall flow chart according to an embodiment of the present disclosure.
[0105] like Figure 8 As shown, input information is first received. The input information is fed into a multimodal large visual model (MVLM) to generate a target image. This process corresponds to operation S110, which involves obtaining a target image generated based on the target input information. Furthermore, the input information is processed by a generation rule module to generate multiple questions (as shown in the figure, Q1, Q2, Q3, etc.). This process corresponds to operation S120, which involves generating multiple questions based on the target input information. After obtaining the target image and question information, the system enters a two-branch evaluation phase: the target image and question information are jointly fed into a visual question answering (VQA) model. The VQA model answers the questions and generates a first evaluation result based on the answers. This process fully embodies operation S130 (generating answer information using the VQA model) and operation S140 (generating the first evaluation result based on the answer information). The first evaluation result indicates the consistency between the image content and the input information. In the upper evaluation branch, the target image is fed into an image quality assessment (IQA) model to generate a second evaluation result. This process corresponds to operation S610, where the target image is evaluated according to the unsupervised image quality assessment method to generate a second evaluation result representing its visual quality. Finally, the first and second evaluation results are processed together to generate a final target evaluation result. This process corresponds to operation S620, where the first and second evaluation results are combined to generate a comprehensive target evaluation result. The target evaluation result can be further used in subsequent screening steps, such as comparison with a preset threshold (corresponding to operation S710, not shown), to determine whether the target image meets the final quality requirements.
[0106] Figure 9 The figure schematically shows a block diagram of an image evaluation device according to an embodiment of the present disclosure.
[0107] like Figure 9 As shown, the image evaluation device 900 may include a first acquisition module 910 , a first generation module 920 , a second generation module 930 , and a third generation module 940 .
[0108] The first acquisition module 910 is used to acquire a target image, which is an image generated based on target input information. The target input information includes multiple key information. In some embodiments, the first acquisition module 910 can be used to perform operation S110 in the above-mentioned image evaluation method, which will not be described in detail here.
[0109] The first generating module 920 is used to generate multiple question information according to the target input information, and the question information corresponds to at least one key information. In some embodiments, the first generating module 920 can be used to perform operation S120 in the above-mentioned image evaluation method, which will not be described in detail here.
[0110] The second generation module 930 is configured to use the target image and question information as inputs to the visual question answering model and generate answer information for each question information based on the visual question answering model. In some embodiments, the second generation module 930 can be used to perform operation S130 in the above-mentioned image evaluation method, which will not be described in detail here.
[0111] The third generating module 940 is used to generate a first evaluation result according to the plurality of answer information. In some embodiments, the third generating module 940 can be used to perform operation S140 in the above-mentioned image evaluation method, which will not be described in detail here.
[0112] According to an embodiment of the present disclosure, the first generation module may include a first submodule, and / or a second submodule and a third submodule.
[0113] The first submodule is used to generate corresponding question information according to each key information. In some embodiments, the first submodule can be used to perform operation S210 in the above-mentioned image evaluation method, which will not be described in detail here.
[0114] The second submodule is used to obtain at least one associated information based on the target input information, where the associated information represents the associated relationship between the multiple first key information in the multiple key information. In some embodiments, the second submodule can be used to perform operation S220 in the above-mentioned image evaluation method, which will not be described in detail here.
[0115] The third submodule is used to generate at least one question information corresponding to the multiple key information based on the respective associated information. In some embodiments, the third submodule can be used to perform operation S230 in the above-mentioned image evaluation method, which will not be described in detail here.
[0116] According to an embodiment of the present disclosure, the third submodule may include an association satisfaction module.
[0117] The association satisfaction module is configured to generate at least one question based on each piece of association information and the multiple pieces of first key information corresponding to the association information. The question is used to verify whether the content corresponding to the multiple pieces of first key information in the target image satisfies the association relationship. In some embodiments, the association satisfaction module can be configured to perform operation S310 in the aforementioned image evaluation method, and further description thereof is omitted here.
[0118] According to an embodiment of the present disclosure, the third submodule may include an association unsatisfied module.
[0119] The association failure module is configured to select at least one second key information different from the first key information from the plurality of key information, wherein no association relationship exists between any second key information and the first key information, and to generate at least one question information based on the association information, the at least one first key information, and the at least one second key information. The question information is configured to verify whether the content corresponding to the first key information and the content corresponding to the second key information in the target image do not satisfy the association relationship. In some embodiments, the preset module 16 can be configured to perform operations S410-S420 in the above-described image evaluation method, which will not be described in detail herein.
[0120] According to an embodiment of the present disclosure, the third generating module may include a calculating module.
[0121] The calculation module is used to generate a first evaluation result based on the proportion of the number of positive answers in the total answer information. In some embodiments, the calculation module can be used to perform operation S510 in the above-mentioned image evaluation method, which will not be described in detail here.
[0122] According to an embodiment of the present disclosure, the image evaluation device may further include a second evaluation module and a third evaluation module.
[0123] The second evaluation module is used to evaluate the target image according to the unsupervised image quality evaluation method to generate a second evaluation result. In some embodiments, the second evaluation module can be used to perform operation S610 in the above-mentioned image evaluation method, which will not be described in detail here.
[0124] The third evaluation module is used to combine the first evaluation result and the second evaluation result to generate a target evaluation result for the target image. In some embodiments, the third evaluation module can be used to perform operation S620 in the above-mentioned image evaluation method, which will not be described in detail here.
[0125] According to an embodiment of the present disclosure, the image evaluation device may further include a screening module.
[0126] The screening module is used to treat the target image as an image that meets the preset quality requirements in response to the target evaluation result being greater than a preset threshold. In some embodiments, the screening module can be used to perform operation S710 in the above-mentioned image evaluation method, which will not be described in detail here.
[0127] According to the embodiments of the present invention, any number of modules, sub-modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.
[0128] For example, any multiple of the first acquisition module 910, the first generation module 920, the second generation module 930, and the third generation module 950 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first acquisition module 910, the first generation module 920, the second generation module 930, and the third generation module 950 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first acquisition module 910 , the first generation module 920 , the second generation module 930 , and the third generation module 950 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0129] It should be noted that the data processing system part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. The description of the data processing system part specifically refers to the data processing method part and will not be repeated here.
[0130] Figure 10 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure is schematically shown. Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0131] like Figure 10 As shown, the electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage portion 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0132] Various programs and data required for the operation of the electronic device 1000 are stored in the RAM 1003. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the programs may also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 may also perform various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0133] According to an embodiment of the present disclosure, electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to bus 1004. Electronic device 1000 may also include one or more of the following components connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or modem. Communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 1010 as needed, so that computer programs read from the removable media can be installed into storage section 1008 as needed.
[0134] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0135] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0136] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0137] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 1002 and / or the RAM 1003 described above and / or one or more memories other than the ROM 1002 and the RAM 1003 .
[0138] An embodiment of the present disclosure also includes a computer program product, which includes a computer program containing program code for executing the method provided by the embodiment of the present disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the image evaluation method provided by the embodiment of the present disclosure.
[0139] When the computer program is executed by the processor 1001, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0140] In one embodiment, the computer program may be based on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 1009, and / or installed from a removable medium 1011. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above. According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computer programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to languages such as Java, C++, Python, "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. Where a remote computing device is involved, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, and all of these combinations and / or couplings fall within the scope of the present disclosure.
[0142] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An image evaluation method, comprising: Acquire a target image, where the target image is an image generated based on target input information, where the target input information includes a plurality of key information; generating a plurality of question information according to the target input information, wherein the question information corresponds to at least one of the key information; Taking the target image and the question information as inputs of a visual question answering model, and generating answer information for each of the question information based on the visual question answering model; A first evaluation result is generated based on the plurality of answer information.
2. The method according to claim 1, wherein generating a plurality of question information according to the target input information comprises: Generate the corresponding problem information according to each key information; and / or, Obtaining at least one piece of association information according to the target input information, wherein the association information represents an association relationship between a plurality of first key information in the plurality of key information; At least one piece of question information corresponding to the plurality of key information is generated according to each piece of associated information.
3. The method according to claim 2, wherein generating at least one piece of question information corresponding to the plurality of key information comprises: At least one question information is generated based on each of the association information and the multiple first key information corresponding to the association information. The question information is used to verify whether the contents corresponding to the multiple first key information in the target image satisfy the association relationship.
4. The method according to claim 2, wherein generating the problem information corresponding to the plurality of key information comprises: Selecting at least one second key information from the plurality of key information, which is different from the first key information, wherein no association relationship exists between any second key information and the first key information; At least one question information is generated based on the association information, at least one of the first key information, and at least one of the second key information. The question information is used to verify whether the content corresponding to the first key information and the content corresponding to the second key information in the target image do not satisfy the association relationship. 5 . The method according to claim 1 , wherein the number of the problem information is greater than the number of the key information.
6. The method according to claim 1, wherein the question information is a question about the existence of the key information and / or the relationship between multiple key information.
7. The method according to claim 1, wherein generating the first evaluation result comprises: The first evaluation result is generated according to the proportion of the answer information that is a positive answer in all the answer information.
8. The method according to claim 1, further comprising: Evaluating the target image according to an unsupervised image quality evaluation method to generate a second evaluation result; The first evaluation result and the second evaluation result are combined to generate a target evaluation result for the target image.
9. The method according to claim 7, further comprising: In response to the target evaluation result being greater than a preset threshold, the target image is regarded as an image that meets preset quality requirements.
10. An electronic device comprising: at least one processor; as well as a memory connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the following operations: obtaining a target image, wherein the target image is an image generated based on target input information, and the target input information contains multiple key information; generating multiple question information based on the target input information, and the question information corresponds to at least one of the key information; using the target image and the question information as inputs of a visual question answering model, and generating answer information for each of the question information based on the visual question answering model; and generating a first evaluation result based on the multiple answer information.