An evaluation method, device and equipment of a handwritten graph and a storage medium
By combining intent analysis and visual analysis, natural language descriptions are decomposed into structured semantic information to generate visual question-answering tasks. This solves the problems of the single and subjective nature of existing text-based image evaluation methods and achieves efficient and accurate multi-dimensional evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-05
AI Technical Summary
Existing evaluation methods for text-based graph models are mostly limited to a single indicator and are highly subjective, failing to understand complex natural language requirements and thus failing to meet practical application needs.
Intent analysis is used to decompose natural language descriptions into multiple structured semantic information, and visual analysis is combined to generate visual question answering tasks. Quantitative evaluation results are obtained through multi-dimensional evaluation, and data parallel processing is performed using a distributed cluster system.
It enables a comprehensive, systematic, and quantifiable evaluation of the quality of generated text-based images, ensuring the objectivity and accuracy of the evaluation. It can understand user intent and deeply analyze the content of target images, thereby improving evaluation efficiency and accuracy.
Smart Images

Figure CN122156866A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to artificial intelligence, generative search, knowledge graphs, and text-to-graph. Background Technology
[0002] With the rapid development of generative artificial intelligence technology, text-to-image generation models have demonstrated increasingly powerful capabilities in image generation tasks. However, existing evaluation methods for text-to-image models are mostly limited to a single metric and are highly subjective, unable to understand complex "natural language requirements" instructions, resulting in unsatisfactory evaluation results and failing to meet the needs of practical applications. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for evaluating text images.
[0004] According to one aspect of this disclosure, a method for evaluating text-based images is provided, comprising: The system receives an evaluation request instruction described in natural language, and a target image to be evaluated; wherein the target image is an image generated based on the text corresponding to the natural language description. When the evaluation request instruction is parsed, the natural language description is decomposed into multiple structured semantic information based on intent analysis; Based on the multiple structured semantic information, combined with visual analysis of the target image, a set of visual question answering tasks for evaluation and verification are obtained. A multi-dimensional evaluation is performed based on the visual question-answering task to obtain quantitative evaluation results.
[0005] According to another aspect of this disclosure, an evaluation apparatus for textural images is provided, comprising: The instruction receiving module is used to receive an evaluation request instruction described in natural language, and a target image to be evaluated; wherein, the target image is an image generated based on the text corresponding to the natural language description; The instruction parsing module is used to decompose the natural language description into multiple structured semantic information based on intent analysis when parsing the evaluation request instruction; The question-answering task generation module is used to take the multiple structured semantic information as a basis and combine it with the visual analysis of the target image to obtain a set of visual question-answering tasks for evaluation and verification. The evaluation module is used to perform multi-dimensional evaluation based on the visual question answering task and obtain quantitative evaluation results.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method provided in any embodiment of this disclosure.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method provided according to any embodiment of this disclosure.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to any embodiment of this disclosure.
[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of a distributed cluster processing scenario according to an embodiment of the present disclosure; Figure 2 This is a schematic flowchart of a text-based image evaluation method according to an embodiment of the present disclosure. Figures 3-9 This is a flowchart illustrating another textual image evaluation method according to an embodiment of the present disclosure; Figure 10 This is a schematic diagram of a text image evaluation workflow as an application example according to an embodiment of this disclosure; Figure 11 This is a schematic diagram of the composition structure of the evaluation apparatus for textual analysis according to an embodiment of the present disclosure; Figure 12 This is a block diagram of an electronic device used to implement the text image evaluation method of the embodiments of this disclosure. Detailed Implementation
[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0012] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.
[0013] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0014] According to embodiments of this disclosure, Figure 1 This is a schematic diagram of a distributed cluster processing scenario according to an embodiment of the present disclosure. The distributed cluster system is an example of a cluster system, and an exemplary description is provided of a method for evaluating raw text images using this distributed cluster system. This disclosure is not limited to raw text image evaluation methods on a single machine or multiple machines; employing distributed processing can further improve the efficiency of raw text image evaluation methods. Figure 1 As shown, the distributed cluster system includes multiple nodes (such as server cluster 101, server 102, server cluster 103, server 104, and server 105; server 105 can also connect to electronic devices, such as mobile phone 1051 and desktop computer 1052). Multiple nodes, as well as multiple nodes and connected electronic devices, can jointly execute one or more textural graph evaluation tasks. Optionally, the multiple nodes in the distributed cluster system can adopt a data parallel processing method, allowing multiple nodes to execute textural graph evaluation tasks based on the same processing method. Furthermore, one or more processing logics from the processing method can be distributed across multiple nodes to collaboratively complete the textural graph evaluation task.
[0015] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 2This is a flowchart illustrating a textural mapping evaluation method according to an embodiment of the present disclosure. This method can be applied to a textural mapping evaluation device. For example, the device can be deployed on a terminal, server, or other processing device in a single-machine, multi-machine, or cluster system to perform the textural mapping evaluation task. The terminal can be a user equipment (UE), mobile device, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 2 As shown, this method is applied to Figure 1 In any node or electronic device in the cluster system shown, the following are included: S201. Receive an evaluation request instruction described in natural language, and a target image to be evaluated; wherein the target image is an image generated based on the text corresponding to the natural language description.
[0016] In some examples, the user inputs an evaluation request instruction described in natural language. This instruction could be: "Evaluate whether this image meets the requirement of the prompt 'An orange cat wearing a red bow is sitting on a sofa,' and simultaneously upload or specify a cat image generated by a text-based image model based on that prompt, using that cat image as the target image." By clearly defining the object to be evaluated (the target image) and the evaluation reference standard (meeting the requirements of the evaluation request instruction), an evaluation reference is provided for subsequent evaluations.
[0017] S202. Parse the evaluation request instruction. During the parsing process, the natural language description is decomposed into multiple structured semantic information based on intent analysis.
[0018] In some examples, the core of the intent analysis of the evaluation request instruction is to evaluate the consistency between the target image and the text. The core prompt word "an orange cat wearing a red bow sitting on the sofa" quoted in the evaluation request instruction is taken as the natural language description to be parsed and broken down into three structured pieces of information, such as: object (cat, sofa), attribute (cat: orange, wearing a red bow), and relationship (cat sitting on the sofa).
[0019] S203. Based on multiple structured semantic information, combined with visual analysis of the target image, a set of visual question answering tasks for evaluation and verification are obtained.
[0020] In some examples, based on the three structured semantic information obtained from S202, a judgment question requiring visual analysis to answer is generated for each structured semantic information. For example, for the attribute information "the cat is orange", a visual question-answering task is generated: "Is the cat in the picture orange?". To obtain a further answer, a visual model can be invoked to analyze and judge the color of the cat in the target image.
[0021] S204. Conduct multi-dimensional evaluation based on the visual question-answering task to obtain quantitative evaluation results.
[0022] In some examples, when all visual question-answering tasks are generated according to S203, for example, the total task is broken down into three sub-tasks, and the three questions described are: "Do you have a cat?", "Is the color orange?", and "Does it wear a bow?". The answers obtained are: "Yes" (there is a cat), "Yes" (the cat is orange), and "No" (it does not wear a bow).
[0023] In some examples, a comprehensive evaluation can be conducted based on these answers. For instance, a simple quantification method is to divide the number of correctly answered questions by the total number of questions, resulting in a quantification score of 2 / 3 ≈ 0.67 (or 67%). This score represents a quantitative evaluation result for the consistency between the text and images. Furthermore, the quantitative evaluation result can also be assessed from multiple dimensions, including objective and subjective dimensions. These will be elaborated in detail in subsequent quantitative evaluation examples and will not be repeated here.
[0024] This embodiment of the disclosure combines natural language processing and visual analysis technologies to achieve a comprehensive, systematic, and quantifiable evaluation of the quality of generated text-to-image (TPE) images. It not only accurately understands the user's intent but also performs in-depth analysis based on the content of the target image, thereby ensuring the objectivity and accuracy of the evaluation. Furthermore, the multi-dimensional evaluation method makes the evaluation results richer and more comprehensive.
[0025] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 3 This is a flowchart illustrating a text-based graph evaluation method according to an embodiment of the present disclosure. Based on intent analysis, the natural language description is decomposed into multiple structured semantic information, such as... Figure 3 As shown, it includes: S301. Based on intent analysis, extract multiple key pieces of information for evaluation from the natural language description.
[0026] In some examples, based on intent analysis, multiple key pieces of information can be automatically extracted from the prompt "An ancient stone bridge spans a misty river". For example, the extracted key information includes: "ancient" (attribute), "stone bridge" (object), "spans" (relationship), "misty" (attribute), and "river" (object).
[0027] S302. Decompose the natural language description based on multiple key information to obtain multiple structured semantic information.
[0028] In some examples, the extracted key information can be organized and structured to form multiple explicit structural semantic information pieces. For example, the resulting multiple structural semantic information pieces include: 1) Objects: stone bridge, river surface; 2) Attributes: the stone bridge is ancient, the river surface is shrouded in mist; 3) Relationship: the stone bridge spans the river. The evaluation checklist composed of this structured semantic information forms an explicit checklist for subsequent quantitative evaluation.
[0029] This embodiment of the present disclosure details how to decompose natural language descriptions into multiple structured semantic information based on intent analysis. Specifically, through intent analysis and the extraction of multiple key information, complex natural language descriptions can be efficiently transformed into multiple structured semantic information. This method not only improves the accuracy of semantic understanding but also provides a reliable foundation for subsequent visual question answering tasks, thereby enhancing overall evaluation efficiency and accuracy.
[0030] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 4 This is a schematic flowchart of the evaluation method for text-based images according to embodiments of the present disclosure, such as... Figure 4 As shown, it includes: S401. Based on intent analysis, extract at least two key pieces of information from the natural language description, including the generated object, attribute information, relationships between objects, text content requirements, style requirements, and constraints.
[0031] S402. Decompose the natural language description based on at least two key pieces of information to obtain multiple structured semantic information.
[0032] This embodiment of the disclosure details how to extract key information from natural language descriptions, specifically extracting at least two types of key information from "generated objects, attribute information, relationships between objects, text content requirements, style requirements, and constraints," comprehensively covering multiple key information dimensions of text-based graph evaluation. This method not only improves the comprehensiveness of the evaluation but also provides rich semantic basis for subsequent visual question-answering tasks, thereby ensuring the accuracy and reliability of the evaluation results.
[0033] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 5 This is a flowchart illustrating the text-based image evaluation method according to an embodiment of the present disclosure. It uses multiple structured semantic information sets as a basis, combined with visual analysis of the target image, to obtain a set of visual question-answering tasks for evaluation and verification, such as... Figure 5 As shown, it includes: S501. Use multiple structured semantic information as a unified semantic basis for question answer generation and visual analysis.
[0034] In some examples, extracting multiple key pieces of information from the prompts yields a range of structured semantic information (e.g., object - astronaut, attribute - wearing a silver spacesuit, relationship - standing on the lunar surface), which serves as a common benchmark for all subsequent steps (such as question-answering generation and visual analysis). In other words, both question generation and image analysis must strictly revolve around this information to ensure the relevance and consistency of the evaluation.
[0035] S502. Based on multiple structured semantic information, obtain a decision-type question for question-answer generation.
[0036] In some examples, based on the unified semantic basis of S501 above, a judgment question with an answer of "yes / no" can be automatically generated for each verifiable piece of information. For example: 1) "Is there an astronaut in the picture?"; 2) "Is the astronaut wearing a silver spacesuit?"; 3) "Is the astronaut standing on the surface of the moon?".
[0037] S503. Perform visual analysis on the target image based on the judgment question to obtain the visual question answering task.
[0038] In some examples, a "visual question-answering task" can consist of "question-answer pairs," which are explicitly composed of a decision question (a decision question is a question in the form of a decision question generated by S502 above) and a visual analysis action for that question. For example, for a "question-answer pair," the task for the question "Are astronauts wearing silver spacesuits?" is to call a visual model to analyze the target image to identify the color of the astronaut's suit and determine whether it is silver, thereby generating the corresponding answer to the question "Are astronauts wearing silver spacesuits?".
[0039] This embodiment of the present disclosure details how to use structured semantic information as a unified semantic basis and combine it with visual analysis to generate visual question-answering tasks. Specifically, through a unified semantic basis and a judgment-type question generation mechanism, semantic information can be efficiently combined with image content. This method not only improves the relevance of visual question-answering tasks but also provides high-quality question-answer pairs for subsequent evaluation, thereby improving the accuracy and efficiency of the overall evaluation.
[0040] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 6 This is a flowchart illustrating the evaluation method for text-based images according to embodiments of the present disclosure. Visual analysis of the target image is performed based on decision-type questions to obtain a visual question-answering task, such as... Figure 6 As shown, it includes: S601. Obtain image evaluation instructions based on the decision question, and use the decision question as the question for question and answer generation.
[0041] In some examples, the judgment question is "Is the apple in the picture placed on the wooden table?", which inherently contains an image evaluation instruction: "Please check whether the positional relationship between the apple and the wooden table is 'placed on the wooden table'." This judgment question can be directly used as a "question" to be answered.
[0042] S602. Based on the image evaluation command, a visual analysis of the target image is triggered to obtain the visual analysis results.
[0043] In some examples, in response to the image evaluation instruction of S601 above, visual models such as object detection or scene graph generation can be used to analyze the image. The visual analysis results obtained after the visual analysis can be the detected objects "apple" and "table", as well as the predicted relationship labels between them such as "on" (on top) or "near" (next to).
[0044] S603. The judgment result of whether the visual analysis result meets the image evaluation instruction is used as the answer to the question and answer generation.
[0045] In some examples, the visual analysis result (relation label "on") can be compared with the image evaluation instruction requirement (the relationship should be "place on the wooden table"). Since "on" matches the requirement, the judgment result is "satisfied", therefore, the answer to the "question" in S601 above is generated as "yes"; if the relation label is "near", the judgment result is "not satisfied", and the answer to the "question" in S601 above is generated as "no".
[0046] S604. Based on the questions and answers generated from the question-and-answer session, a visual question-and-answer task is obtained.
[0047] In some examples, the visual question-answering task is constructed based on the "question" in S601 above and the "answer" obtained to form a "question-answer pair". The visual question-answering task is: (Question: "Is the apple in the picture placed on the wooden table?", Answer: "Yes"). At this point, a complete and executed visual question-answering task is recorded. When all visual question-answering tasks have been recorded, multiple chains of evidence can be formed for subsequent quantifiable evaluation.
[0048] This disclosure describes in detail how to generate visual question-answering tasks based on decision-type questions. Specifically, by combining decision-type questions with image evaluation instructions, the generation of visual question-answering tasks can be completed efficiently. This method not only improves the quality of question-answering generation but also ensures the relevance of visual analysis results to the questions, thereby enhancing the accuracy and efficiency of the overall evaluation.
[0049] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 7 This is a flowchart illustrating the text-based image evaluation method according to an embodiment of the present disclosure. It performs multi-dimensional evaluation based on a visual question-answering task to obtain quantitative evaluation results, such as... Figure 7 As shown, it includes: S701. A comprehensive evaluation is conducted based on at least two indicators from the objective evaluation dimensions to obtain the first evaluation result.
[0050] In some examples, after all visual question answering tasks are completed, two objective metrics can be used for comprehensive evaluation. For example, (1) Visual Question Answering Accuracy (VQA Accuracy): calculate the proportion of correct answers to questions in all visual question answering tasks; (2) Image-Text Semantic Similarity Score (CLIP Score): use a Contrastive Language-Image Pre-training (CLIP) model to calculate the cosine similarity between the entire target image and the complete prompt word. Different weights can be assigned to the visual question answering accuracy and the image-text semantic similarity score, and a weighted average can be performed to obtain a comprehensive score, which can be used as the first evaluation result.
[0051] S702. The first evaluation result shall be used as the quantitative evaluation result.
[0052] This embodiment of the present disclosure details how to perform multi-dimensional objective evaluation based on a visual question-answering task. This embodiment, through a multi-dimensional objective evaluation method, can comprehensively reflect the quality of the generated text-based image. This method not only improves the comprehensiveness of the evaluation but also ensures the objectivity and quantifiability of the evaluation results.
[0053] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 8 This is a flowchart illustrating the text-based image evaluation method according to an embodiment of the present disclosure. It performs multi-dimensional evaluation based on a visual question-answering task to obtain quantitative evaluation results, such as... Figure 8 As shown, it includes: S801. Based on at least one indicator of the objective evaluation dimension, and by integrating at least one indicator of the subjective evaluation dimension, a comprehensive evaluation is conducted to obtain a second evaluation result.
[0054] In some examples, after all visual question-answering tasks are completed, a comprehensive evaluation can be conducted using one objective metric and one subjective metric. For example, the following can be used: (1) Visual question-answering accuracy: Calculate the proportion of correct answers to questions in all visual question-answering tasks to obtain an objective score; (2) Human preference standards derived from pre-set, structured criteria: such as the human preference standard of the "aesthetic quality" dimension. A large language model (LLM) can score subjective scores at multiple levels based on criteria (such as composition and color harmony), and then the objective score (normalized to 0-1) is fused with the subjective score, for example, by weighted averaging again, to obtain a second evaluation result.
[0055] S802. The second evaluation result shall be used as the quantitative evaluation result.
[0056] This embodiment of the disclosure details how to combine objective and subjective evaluation dimensions for comprehensive assessment. By combining objective and subjective evaluation dimensions, this embodiment can comprehensively reflect the quality of the generated text-based images. This method not only improves the comprehensiveness of the assessment but also ensures the accuracy and reliability of the assessment results.
[0057] According to embodiments of this disclosure, a method for evaluating text-based images is provided. Figure 9 This is a flowchart illustrating a text-based image evaluation method according to an embodiment of the present disclosure. The method involves performing a multi-dimensional evaluation based on the visual question-answering task to obtain a quantitative evaluation result, such as... Figure 9 As shown, it includes: S901. A comprehensive evaluation is conducted based on at least two indicators from the objective evaluation dimensions to obtain a third evaluation result.
[0058] In some examples, after all visual question answering tasks are completed, two objective metrics can be used for comprehensive evaluation. For example, (1) visual question answering accuracy: calculate the proportion of correct answers to questions in all visual question answering tasks; (2) image-text semantic similarity score: use the CLIP model to calculate the cosine similarity between the entire target image and the complete prompt word. Different weights can be assigned to the visual question answering accuracy and the image-text semantic similarity score, and a weighted average can be performed to obtain a comprehensive score, which can be used as the third evaluation result.
[0059] S902. Optimize the third evaluation result based on at least one indicator of the subjective evaluation dimension to obtain the fourth evaluation result.
[0060] In some examples, subjective metrics can be introduced to calibrate or correct the objective third evaluation result obtained in S901 above. This can be achieved through calibration coefficients. For example, LLM can determine whether an image has obvious inconsistencies (such as a distorted hand or a floating object) based on the "image realism / oddity" criterion. If LLM determines that "there are no obvious oddities and it conforms to common sense," it assigns a "positive calibration coefficient" (e.g., +0.05); if it determines that "there are obvious oddities," it assigns a "negative calibration coefficient" (e.g., -0.1). Finally, the calibration coefficients are applied to the third evaluation result to obtain the optimized fourth evaluation result.
[0061] S903. The fourth evaluation result shall be used as the quantitative evaluation result.
[0062] This embodiment of the disclosure details how to optimize evaluation results through multi-dimensional comprehensive evaluation. This embodiment, through multi-dimensional comprehensive evaluation and optimization, can comprehensively reflect the quality of the generated text-based images. This method not only improves the comprehensiveness of the evaluation but also ensures the accuracy and reliability of the evaluation results.
[0063] The evaluation method for text images provided in the above-described embodiments of this disclosure will be illustrated below.
[0064] With the rapid development of text-based image generation technology, model capabilities have evolved from simple image generation to understanding and stably executing complex instructions, finding wide application in real-world scenarios such as design, content production, product prototyping, education, and business. However, current text-based image evaluation systems are lagging behind: traditional evaluations often focus on low-level visual metrics such as resolution and image clarity, making it difficult to measure the model's true capabilities in complex semantic understanding, complex reasoning, constraint execution, and reusability; manual evaluation is costly, subjective, and inefficient, failing to meet the evaluation needs of rapid model iteration; and existing automated evaluation metrics are limited, lacking comprehensive coverage of key dimensions such as text rendering accuracy, instruction compliance completeness, and consistency with physical common sense.
[0065] The following application example of the embodiments of this disclosure is an efficient, comprehensive, and objective automated evaluation scheme for text-based graph models. The design objective is to understand and stably execute complex instructions. The technical essence is the collaboration of an LLM-based master control agent with multiple execution agents (specifically, a multi-modal execution agent architecture) and multi-dimensional quantitative indicators for coordinated evaluation, forming a complete automated execution technology chain. Specifically, it revolves around the evaluation main line of "input determination—preprocessing—instruction parsing—visual understanding-based question-answering verification—multi-indicator scoring result output," integrating Optical Character Recognition (OCR), Visual Question Answering (VQA), and Image Quality Assessment (IQA) modules. The Assessment module, CLIP module, and other modules work together based on various visual and semantic analysis technologies. This modular approach, with its decoupled design and collaborative operation, enables a systematic and quantifiable automatic evaluation of the text-to-image model's key capabilities, such as instruction compliance, text rendering, image-text consistency, image quality, and human preferences. This solves the technical problems of existing evaluation systems, such as misalignment between the model's actual capabilities, low evaluation efficiency, strong subjectivity, and limited dimensions.
[0066] Figure 10 This is a schematic diagram of a text-based image evaluation workflow, illustrating an application example according to an embodiment of this disclosure. Figure 10 As shown, the workflow for evaluating the raw image includes the following: (a) Input Decision: Data Input and Task Decision Stage The system receives the text query to be evaluated and its corresponding "target image" (AI-generated image) as input. First, it determines whether the text query contains a requirement for visible text content in the target image: if it does, the OCR evaluation branch implemented by OCR module 1001 is triggered; otherwise, it directly enters the instruction parsing process. This input determination mechanism avoids redundant calls to unrelated modules, thereby improving overall evaluation efficiency.
[0067] (II) Preprocessing: OCR preprocessing and text consistency assessment When the query involves text content in an image, the OCR module 1001 can be called to extract the text information in the image and evaluate the recognition results to obtain character-level confidence, text accuracy, and semantic similarity score with the original query. The OCR success rate, OCR accuracy, and semantic similarity score are output as independent evaluation results of text rendering capability.
[0068] (iii) Instruction parsing: performing structured parsing of evaluation request instructions in natural language requirement descriptions. The input text query can be deeply parsed. The evaluation request instruction is: Generate apples, place the apples in a bamboo basket, the basket is three green apples, the basket is labeled "fresh," and the style is like an oil painting. Note that "there are no red apples." This evaluation request instruction needs to be broken down into multiple structured semantic pieces of information, mainly including: 1. Main object (Object); 2. Sub-Objects; 3. Attribute information (such as color, material, quantity, etc.); 4. Relationships between objects (positional and logical relationships); 5. Text requirements; 6. Style description; 7. Negative constraints; These multiple structured semantic information are used as a unified semantic basis for subsequent question-answering generation and visual understanding.
[0069] (iv) Visual understanding-based question-answering verification: generation of judgment questions, and visual question-answering verification of "question-answer pairs" in visual question-answering tasks. Based on structured instruction information, a set of judgment questions corresponding one-to-one with each instruction is automatically generated, such as "Are there green apples?" or "Are there baskets?". The target image and the judgment questions are then input into the VQA module 1002 for visual understanding reasoning. The output of the VQA module 1002 is used to determine whether "the target image meets the requirements of each instruction in the evaluation request instruction." If the question-and-answer result can be clearly determined, the instruction compliance scoring process begins; if it cannot be determined or information is missing, it is marked as a bad case for anomaly analysis.
[0070] (v) Output of results from multi-indicator scoring: multi-dimensional scoring and result output Assuming the instruction compliance judgment is valid, the following multi-category, multi-dimensional evaluation indicators can be further output: 1. VQA Score: Used to measure the overall compliance capability of VQA module 1002 with structured instructions, and is an objective indicator; 2. CLIP score: Used to evaluate the semantic consistency between text and images; it is an objective indicator. 3. IS, FID, and IQA scores: These are used to evaluate the overall image quality and the rationality of the image distribution of the target image, and are objective indicators. 4. Human Preference (HP) score: This score measures the overall appeal and usability of a target image to human users in real-world usage scenarios. It focuses on subjective dimensions that are difficult to fully capture with automated metrics, such as aesthetic quality, compositional rationality, stylistic integrity, common sense and physical consistency, and creative expression. These are subjective metrics.
[0071] Finally, by combining the above multi-dimensional evaluation indicators, a complete and quantifiable evaluation result is formed, covering semantic consistency, image quality, and subjective preferences.
[0072] Based on the above-mentioned text-based image evaluation workflow, which consists of "input determination—preprocessing—instruction parsing—visual understanding-based question-answering verification—multi-indicator scoring result output," this approach integrates multiple modules working collaboratively based on various visual and semantic analysis technologies, including OCR, VQA, IQA, and CLIP modules. The specific descriptions of each module are as follows: (1) OCR module 1001: Extracts text information from images using high-precision OCR algorithms, and quantitatively evaluates the quality of text generation from three levels: character confidence, text accuracy and semantic similarity, providing direct evidence for the text rendering capability of the text-to-image model.
[0073] (2) VQA module 1002: Transforms complex instructions into a set of verifiable visual question-answering tasks. By judging whether the object, attribute, relationship and negation condition are satisfied item by item, it realizes a fine evaluation of the degree of instruction compliance.
[0074] (3) IQA module 1003: integrates IS, FID and no-reference image quality indicators to quantify the sharpness, diversity and overall quality of the generated image and perform unified normalization processing.
[0075] (4) Semantic Consistency Module 1004: Based on cross-modal models such as CLIP, it calculates the semantic matching degree between text and images, focusing on the alignment of objects, attributes and details.
[0076] (5) Human Preference Scoring Module 1005: Based on a real human evaluation process, and based on a predefined set of structured human preference scoring criteria (Rubrics), subjective evaluation is decomposed into several decidable high-level semantic dimensions, as shown in Table 1 below: Table 2 Each dimension employs a discrete-level scoring mechanism, accompanied by clear judgment descriptions, ensuring high consistency among different evaluators under the same standards. This scoring system can be derived from manual annotation experience, as shown in the evaluation criteria in Table 2 below: Table 2 During the operation of the text-based image evaluation, the human scoring criteria do not directly rely on real-time human participation. Instead, they are input into the LLM-based master control agent in the form of scoring rule text, example judgments, and historical scoring samples. This allows the large language model to learn and internalize the implicit rules of human aesthetic and preference judgments, thereby simulating the logic of human evaluation during the reasoning stage.
[0077] The judgment results of each preference dimension are numerically mapped and weighted to generate the final HP score. This HP score can be represented as a continuous value or mapped as a level label, and is used for model cross-comparison, version iteration monitoring, and result screening, as shown in the evaluation total score criteria in Table 3 below: Table 3 The text-based image evaluation workflow used in this application example, and its collaboration between the master control intelligence centered on LLM and multiple execution agents (specifically, a multi-modal execution agent architecture), specifically through the collaborative work of multiple modules based on various visual and semantic analysis technologies such as OCR, VQA, IQA, and CLIP, achieves at least the following technical effects: (1) Achieve comprehensive and systematic multi-dimensional evaluation: Through modular design, multiple core dimensions such as text quality, instruction compliance, image quality, text-image alignment, and common sense logic are integrated, covering the core capabilities of the text-to-image model and solving the problem of single evaluation dimensions in traditional methods.
[0078] (2) Achieving objective and quantifiable evaluation results: Each module adopts quantitative indicators or clear judgment rules, reducing subjective intervention; through the standardized judgment logic with LLM as the core, the consistency between common sense and logical evaluation is ensured, solving the problem of strong subjectivity in manual evaluation.
[0079] (3) Accuracy of complex task evaluation capability: Through the structured question-answering design of the VQA module and the logical reasoning capability of LLM, the effective evaluation of long text, multi-constraint, and complex reasoning generation tasks is realized, which solves the problem of insufficient evaluation of complex tasks by traditional automation indicators.
[0080] (4) Improved evaluation efficiency: The entire evaluation process is automated and requires no manual intervention. It supports rapid evaluation of batch samples, meets the evaluation requirements of high-frequency model iteration, and shortens the model iteration cycle by meeting the high-frequency testing requirements in the model development process.
[0081] (5) Reduced evaluation costs: It does not require a large number of professional evaluators or crowdsourcing personnel, significantly reducing the human and time costs in the evaluation process, and is especially suitable for large-scale model optimization and multi-model comparison scenarios.
[0082] (6) Improve the objectivity and consistency of the assessment: By using quantitative indicators and standardized judgment logic, the subjective bias of manual assessment is avoided, the consistency and reliability of the assessment results are ensured, and objective basis is provided for the assessment of model capabilities.
[0083] (7) Covering the core capability assessment requirements in real application scenarios: It not only focuses on basic visual indicators such as image clarity and resolution, but also covers key capabilities such as instruction compliance, text rendering accuracy, object and attribute consistency, physical and common sense rationality and human preferences, which can more realistically reflect the usability of text-based graph models in actual application scenarios such as design, content production and product prototyping.
[0084] In some examples of embodiments of this disclosure, without departing from the overall inventive concept of the embodiments of this disclosure described above, the following alternative methods can be used for some modules or processes, which can also achieve automated evaluation of the embodiments of this disclosure.
[0085] (1) VQA module, which can be replaced by a rule engine-based target detection, attribute recognition or relation extraction model to automatically detect the number of objects, attribute matching and spatial relationships; this method is suitable for evaluation tasks with relatively simple structure and clear rules.
[0086] (2) Semantic consistency module, CLIP can be replaced by other cross-modal alignment models, such as semantic matching models based on dual-tower structure or multimodal encoder, to calculate the semantic similarity between text and image.
[0087] (3) The human preference scoring module can be replaced by a dedicated scoring model (such as a reward model or preference prediction model) trained by human preference data in some scenarios to output the overall preference score, but its interpretability and adaptability to complex constraints are relatively limited.
[0088] According to embodiments of this disclosure, an evaluation apparatus for text-based images is provided. Figure 11 This is a schematic diagram of the composition structure of the evaluation device for textural images according to embodiments of the present disclosure, such as... Figure 11 As shown, the device includes: The instruction receiving module 1101 is used to receive an evaluation request instruction described in natural language and a target image to be evaluated; wherein the target image is an image generated based on the text corresponding to the natural language description. The instruction parsing module 1102 is used to decompose the natural language description into multiple structured semantic information based on intent analysis when parsing the evaluation request instruction; The question-answering task generation module 1103 is used to take multiple structured semantic information as a basis and combine them with visual analysis of the target image to obtain a set of visual question-answering tasks for evaluation and verification. Evaluation module 1104 is used to perform multi-dimensional evaluation based on the visual question answering task and obtain quantitative evaluation results.
[0089] In one embodiment of this disclosure, the instruction parsing module 1102 is used for: Based on intent analysis, several key pieces of information for evaluation are extracted from natural language descriptions; By deconstructing the natural language description based on multiple key pieces of information, multiple structured semantic information is obtained.
[0090] In one embodiment of this disclosure, several key pieces of information include at least two of the following: generated object, attribute information, relationships between objects, text content requirements, style requirements, and constraints.
[0091] In one embodiment of this disclosure, the question-and-answer task generation module 1103 is used for: Multiple structured semantic information are used as a unified semantic basis for question answering generation and visual analysis; Based on multiple structured semantic information, a decision-type question is obtained for question-answer generation; Visual analysis of the target image is performed based on judgment questions to obtain a visual question answering task.
[0092] In one embodiment of this disclosure, the question-and-answer task generation module 1103 is used for: Image evaluation instructions are obtained based on decision-type questions, and decision-type questions are used as questions in question-and-answer generation. The visual analysis of the target image is triggered by the image evaluation command to obtain the visual analysis results. The result of whether the visual analysis meets the image evaluation instructions is used as the answer to the question and answer generation. The visual question-answering task is derived from the questions and answers generated by the question-answering process.
[0093] In one embodiment of this disclosure, the evaluation module 1104 is used for: The first evaluation result is obtained by comprehensively evaluating at least two indicators from the objective evaluation dimensions. The first assessment result will be used as the quantitative assessment result.
[0094] In one embodiment of this disclosure, the evaluation module 1104 is used for: A second evaluation result is obtained by comprehensively evaluating at least one indicator from the objective evaluation dimension and integrating at least one indicator from the subjective evaluation dimension. The second assessment result will be used as the quantitative assessment result.
[0095] In one embodiment of this disclosure, the evaluation module 1104 is used for: A third evaluation result is obtained by comprehensively evaluating at least two indicators from the objective evaluation dimensions. The third evaluation result is optimized based on at least one indicator of the subjective evaluation dimension to obtain the fourth evaluation result; The fourth assessment result will be used as the quantitative assessment result.
[0096] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0097] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.
[0098] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0099] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0100] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1202 or a computer program loaded from storage unit 1208 into random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0101] Multiple components in device 1200 are connected to I / O interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of monitors, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0102] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the text-based image evaluation method. For example, in some embodiments, the text-based image evaluation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the text-based image evaluation method described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform a text-based graph evaluation method by any other suitable means (e.g., by means of firmware).
[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0107] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0108] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0109] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for evaluating text-based images, comprising: The system receives an evaluation request instruction described in natural language, and a target image to be evaluated; wherein the target image is an image generated based on the text corresponding to the natural language description. When the evaluation request instruction is parsed, the natural language description is decomposed into multiple structured semantic information based on intent analysis; Based on the multiple structured semantic information, combined with visual analysis of the target image, a set of visual question answering tasks for evaluation and verification are obtained. A multi-dimensional evaluation is performed based on the visual question-answering task to obtain quantitative evaluation results.
2. The method according to claim 1, wherein, The intent-based analysis decomposes the natural language description into multiple structured semantic information, including: Based on the intent analysis, several key pieces of information for evaluation are extracted from the natural language description; The natural language description is decomposed based on the multiple key information to obtain the multiple structured semantic information.
3. The method according to claim 2, wherein, The key information includes at least two of the following: generated object, attribute information, relationships between objects, text content requirements, style requirements, and constraints.
4. The method according to any one of claims 1-3, wherein, The step of using the multiple structured semantic information as a basis, combined with visual analysis of the target image, to obtain a set of visual question-answering tasks for evaluation and verification includes: The aforementioned structured semantic information is used as a unified semantic basis for question-answering generation and visual analysis; Based on the multiple structured semantic information, a decisional question is obtained for question-answer generation; The visual analysis of the target image is performed based on the judgment question to obtain the visual question answering task.
5. The method according to claim 4, wherein, The step of performing visual analysis on the target image based on the judgment question to obtain the visual question-answering task includes: Image evaluation instructions are obtained based on the judgment question, and the judgment question is used as the question for question and answer generation; The visual analysis of the target image is triggered according to the image evaluation instruction to obtain the visual analysis result; The visual analysis result is used as the answer to the question-and-answer generation, based on whether it satisfies the image evaluation instruction. The visual question-answering task is obtained based on the questions and answers generated by the question-answering process.
6. The method according to any one of claims 1-5, wherein, The multi-dimensional evaluation based on the visual question-answering task, to obtain quantitative evaluation results, includes: The first evaluation result is obtained by comprehensively evaluating at least two indicators from the objective evaluation dimensions. The first evaluation result is used as the quantitative evaluation result.
7. The method according to any one of claims 1-5, wherein, The multi-dimensional evaluation based on the visual question-answering task, to obtain quantitative evaluation results, includes: A second evaluation result is obtained by comprehensively evaluating at least one indicator from the objective evaluation dimension and integrating at least one indicator from the subjective evaluation dimension. The second evaluation result is used as the quantitative evaluation result.
8. The method according to any one of claims 1-5, wherein, The multi-dimensional evaluation based on the visual question-answering task, to obtain quantitative evaluation results, includes: A third evaluation result is obtained by comprehensively evaluating at least two indicators from the objective evaluation dimensions. The third evaluation result is optimized based on at least one indicator of the subjective evaluation dimension to obtain the fourth evaluation result; The fourth evaluation result is used as the quantitative evaluation result.
9. An evaluation device for textual images, comprising: The instruction receiving module is used to receive an evaluation request instruction described in natural language, and a target image to be evaluated; wherein, the target image is an image generated based on the text corresponding to the natural language description; The instruction parsing module is used to decompose the natural language description into multiple structured semantic information based on intent analysis when parsing the evaluation request instruction; The question-answering task generation module is used to take the multiple structured semantic information as a basis and combine it with the visual analysis of the target image to obtain a set of visual question-answering tasks for evaluation and verification. The evaluation module is used to perform multi-dimensional evaluation based on the visual question answering task and obtain quantitative evaluation results.
10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.