Evaluation method, device, storage medium and program product of a large model of a text-to-image generation model

By using an automated text-based image large model evaluation method, prompt text and discrimination conditions are obtained, and images are generated and evaluated. This solves the problems of slow evaluation speed and insufficient accuracy in existing technologies, and achieves efficient and accurate evaluation results.

CN122220554APending Publication Date: 2026-06-16BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2024-12-16
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In existing technologies, the evaluation methods for large-scale models of text images rely on manual evaluation or model evaluation, which results in a large workload, slow speed, and inaccurate results.

Method used

By acquiring the prompt text and its corresponding discrimination conditions, the system automatically generates the image to be judged in the large model of the text image to be evaluated, and evaluates it according to the discrimination conditions to determine the target evaluation result. The fully automated evaluation process avoids human error.

Benefits of technology

It improves the accuracy and objectivity of the large-scale model evaluation of textual graphs, reduces human intervention, increases evaluation speed, and lowers costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220554A_ABST
    Figure CN122220554A_ABST
Patent Text Reader

Abstract

The application discloses an evaluation method and device of a text-to-image large model, a storage medium and a program product; the method comprises the following steps: obtaining prompt text and corresponding discrimination conditions; for each prompt text, inputting the prompt text into a text-to-image large model to be evaluated, determining a to-be-discriminated image according to the output of the text-to-image large model to be evaluated, evaluating the to-be-discriminated image according to the discrimination conditions corresponding to the prompt text, and determining a target evaluation result, thereby solving the problem of inaccurate evaluation result of the text-to-image large model; the evaluation process is fully automatically realized, human intervention is not needed, human errors are avoided, the result is more accurate and objective, the evaluation speed is fast, and the cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an evaluation method, device, storage medium, and program product for large-scale text-to-image models. Background Technology

[0002] Current technologies for evaluating large-scale text-based image models typically involve either manual evaluation or model-based evaluation. Manual evaluation is labor-intensive, very slow, and prone to misjudgments. Model-based evaluation, which directly assesses the images generated by the large-scale text-based image model based on its prompts, yields inaccurate results. Summary of the Invention

[0003] This invention provides a method, device, storage medium, and program product for evaluating large text-based graph models, addressing the problem of inaccurate evaluation results for large text-based graph models.

[0004] According to one aspect of the present invention, an evaluation method for a large model of text-based graphs is provided, comprising:

[0005] Obtain the prompt text and its corresponding judgment conditions;

[0006] For each of the aforementioned prompt texts, the prompt text is input into the large-scale model of the text-to-image to be evaluated. The image to be judged is determined based on the output of the large-scale model of the text-to-image to be evaluated. The image to be judged is evaluated based on the discrimination conditions corresponding to the prompt text, and the target evaluation result is determined.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0008] At least one processor, and a memory communicatively connected to said at least one processor;

[0009] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the evaluation method for the large model of the text image according to any embodiment of the present invention.

[0010] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the evaluation method for large text-image models according to any embodiment of the present invention.

[0011] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the evaluation method for large-scale text-based models as described in any embodiment of the present invention.

[0012] The technical solution of this invention, through obtaining prompt text and its corresponding discrimination conditions; for each prompt text, inputting the prompt text into the large-scale image-to-text model to be evaluated, determining the image to be discriminated based on the output of the large-scale image-to-text model to be evaluated, evaluating the image to be discriminated based on the discrimination conditions corresponding to the prompt text, and determining the target evaluation result, solves the problem of inaccurate evaluation results of the large-scale image-to-text model; obtaining prompt text and discrimination conditions, the large-scale image-to-text model to be evaluated generates the image to be discriminated based on the prompt text, and automatically detects the image to be discriminated based on the discrimination conditions to obtain the target evaluation result. The evaluation process is fully automated, requiring no human intervention, avoiding human error, resulting in more accurate and objective results, fast evaluation speed, and low cost.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of an evaluation method for a large text-based image model according to Embodiment 1 of the present invention;

[0016] Figure 2 This is a flowchart of an evaluation method for a large text-based image model according to Embodiment 2 of the present invention;

[0017] Figure 3 This is a schematic diagram of the structure of an evaluation device for a large-scale text image model according to Embodiment 3 of the present invention;

[0018] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the evaluation method for large-scale text-based models according to embodiments of the present invention. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] Example 1

[0022] Figure 1 This is a flowchart of an evaluation method for a large-scale text-to-image model provided in Embodiment 1 of the present invention. This embodiment is applicable to evaluating the image generation effect of a large-scale text-to-image model. The method can be executed by an evaluation device for the large-scale text-to-image model, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0023] S101. Obtain the prompt text and its corresponding judgment conditions.

[0024] In this embodiment, the prompt text can be understood as text information used to prompt the large image model when generating an image from the raw image model. For example, the prompt text is "Draw a cat". The discrimination condition can be understood as the condition used to determine whether the generated image is correct. For example, the discrimination condition is "Subject: Cat". The discrimination condition can be generated according to the prompt file.

[0025] Pre-generate prompt texts and their corresponding judgment conditions. For example, pre-set different text-to-image tasks, and generate prompt texts and judgment conditions corresponding to each type of text-to-image task. For the same type of text-to-image task, one or more prompt texts and judgment conditions can be generated, or different prompt texts and their corresponding judgment conditions can be generated randomly.

[0026] S102. For each prompt text, input the prompt text into the large model of the text-to-image to be evaluated, determine the image to be judged based on the output of the large model of the text-to-image to be evaluated, evaluate the image to be judged based on the judgment conditions corresponding to the prompt text, and determine the target evaluation result.

[0027] In this embodiment, the text-based image model to be evaluated can be understood as a text-based image model with evaluation requirements, which can generate images based on text prompts. The image to be judged can be understood as the image generated by the text-based image model to be evaluated. The image to be judged needs to be detected to determine whether the image generated by the text-based image model to be evaluated is correct, thus completing the evaluation of the text-based image model to be evaluated. The target evaluation result can be understood as the result obtained by evaluating the image to be judged, such as whether the generated image to be judged is correct / incorrect, whether the generated image to be judged passes / fails, etc.

[0028] The large-scale image model to be evaluated is determined in advance. For example, different large-scale image models are pre-trained. Different large-scale image models may have different focuses when generating images. For example, some large-scale image models focus on aesthetics when generating images, while others focus on the accuracy of the number of subjects, etc. The method provided in this application embodiment can evaluate any type of large-scale image model, and the large-scale image model to be evaluated is selected as the large-scale image model to be evaluated.

[0029] The image generation model to be evaluated is tested for each prompt text to determine the corresponding target evaluation result. The testing method is as follows: The prompt text is input into the image generation model to be evaluated. The model generates an image based on the prompt text. For example, if the prompt text is "Draw a cat," the model generates an image containing a cat. The image output by the model is used as the image to be judged. The information to be detected is determined according to the judgment condition. For example, if the judgment condition is "Subject: Cat," the model checks whether the subject in the image to be judged is a cat. By detecting and recognizing the image to be judged, the result is compared with the judgment condition to determine whether the image to be judged is correct, and thus the target evaluation result is determined. The target evaluation result can be used to evaluate the accuracy of the image generation model. The target evaluation results corresponding to different prompt texts are statistically analyzed, and all target evaluation results are comprehensively analyzed. The image generation model to be evaluated is then evaluated based on the analysis results. For example, the accuracy of the images generated by the text-to-image model to be evaluated is calculated based on the comprehensive analysis of the target evaluation results, and the text-to-image model to be evaluated is evaluated based on the accuracy. For instance, if the accuracy of the images generated by the text-to-image model to be evaluated is greater than a preset accuracy threshold, the evaluation result of the text-to-image model to be evaluated is determined to be qualified; otherwise, the evaluation result of the text-to-image model to be evaluated is determined to be unqualified. The text-to-image model to be evaluated with an unqualified evaluation result can be retrained or applied to scenarios or business where the accuracy requirement of the generated results is lower.

[0030] This invention provides an evaluation method for a large-scale text-to-image model. The method involves acquiring prompt text and its corresponding discrimination conditions. For each prompt text, the prompt text is input into the large-scale text-to-image model to be evaluated. The output of the large-scale text-to-image model is used to determine the image to be judged. The image to be judged is then evaluated based on the discrimination conditions corresponding to the prompt text to determine the target evaluation result. This method solves the problem of inaccurate evaluation results for large-scale text-to-image models. The method acquires the prompt text and discrimination conditions, the large-scale text-to-image model to be evaluated generates the image to be judged based on the prompt text, and automatically detects the image to be judged using the discrimination conditions to obtain the target evaluation result. The evaluation process is fully automated, requiring no manual intervention, avoiding human error, resulting in more accurate and objective results, faster evaluation speed, and lower cost.

[0031] Example 2

[0032] Figure 2 This is a flowchart of an evaluation method for a large-scale text-based image model provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiments. Figure 2 As shown, the method includes:

[0033] S201. Obtain the prompt text and its corresponding judgment conditions.

[0034] Optionally, the prompt text and the corresponding discrimination conditions include subject information; or, the prompt text and the corresponding discrimination conditions include subject information and attribute information; or, the prompt text and the corresponding discrimination conditions include subject information and positional relationship information; or, the prompt text and the corresponding discrimination conditions include subject information, attribute information, and positional relationship information.

[0035] Among them, subject information can be understood as information used to describe the generated subject, such as subject type; attribute information can be understood as information used to describe the attributes of the subject, such as color, shape, quantity, etc., one or more; positional relationship information can be understood as information used to describe the positional relationship between subjects, such as above, below, left, etc.

[0036] Existing technologies for evaluating large-scale text-based image models typically use a single algorithm to assign a score, which cannot accurately represent the model's quality. For example, models that score aesthetics often exhibit significant data bias, largely depending on the distribution of the training set. Secondly, they lack objective quantitative metrics for some tasks. Large-scale text-based image models offer a high degree of freedom, allowing for numerous performance scenarios and optimization directions from an evaluation perspective. Some models focus on the accuracy of single-subject quantity, others on aesthetics, and still others on text generation, attribute binding (e.g., red object A and green object B). In short, existing technologies for evaluating large-scale text-based image models rely on a single evaluation dimension, resulting in inaccurate results. This application's embodiments evaluate models from one or more dimensions, such as subject information, attribute information, and positional relationship information, leading to more accurate results.

[0037] Optionally, the steps for generating the prompt text and its corresponding discrimination conditions include A1-A2:

[0038] A1. Select at least one type of text-to-image task from the task library as the task to be generated.

[0039] In this embodiment, the text-to-image task can be understood as a task for generating images from text. Each prompt text corresponds to at least one type of text-to-image task. The task information included in each type of text-to-image task is different, and the task information is used to describe the image generation requirements. Each type of text-to-image task describes the image generation requirements differently. For example, a text-to-image task is "single subject generation," where the task information is "single subject," and the described image generation requirement is that the generated image includes only one type of subject. The task to be generated can be understood as a task that requires generating evaluation information to evaluate the large-scale text-to-image model. Different types of text-to-image tasks are pre-generated and saved in a task library. Before evaluating the large-scale text-to-image model, at least one type of text-to-image task is selected from the task library as the task to be generated. When selecting a text-to-image task, at least two types of text-to-image tasks can be randomly selected, or appropriate at least two types of text-to-image tasks can be selected according to certain rules or requirements.

[0040] Optionally, the task library includes at least two of the following types of text-based image tasks:

[0041] Single-subject generation;

[0042] Single main color;

[0043] Number of single entities;

[0044] Single main body shape;

[0045] Dual-subject generation;

[0046] Dual-subject attribute binding;

[0047] The positional relationship between the two main entities;

[0048] Three entities are generated;

[0049] Three main entities generate attribute binding;

[0050] The positional relationship of the three main entities;

[0051] Difficult tasks are characterized by a greater number of subject types than the first preset value, a greater number of each attribute of a subject than the corresponding second preset value, and a greater number of positional relationships than the third preset value.

[0052] The attributes can be quantity, color, shape, etc. Single-subject generation means the generated image includes only one type of subject, for example, a cat. Single-subject color generation means the generated image includes only one type of subject, and the color of the subject is limited. Single-subject quantity generation means the generated image includes only one type of subject, and the number of subjects is limited. Single-subject shape generation means the generated image includes only one type of subject, and the shape of the subjects is limited. Dual-subject generation means the generated image includes two types of subjects, for example, a cat and a table. Dual-subject attribute binding means the generated image includes two types of subjects, and each subject has a corresponding attribute; for example, the cat's attribute is color, and the table's attribute is shape. Dual-subject positional relationship means the generated image includes two types of subjects, and the positional relationship between the two subjects needs to be limited. Three-subject generation, three-subject attribute binding, and three-subject positional relationship are similar to two-subject generation and will not be elaborated further here.

[0053] A challenging task refers to a task where the types, attributes (such as color, quantity, shape, etc.), and positional relationships of the subjects are all subject to corresponding quantity constraints. This typically involves multiple subjects + multiple colors + multiple shapes + multiple quantities + positional relationships. The specific quantities can be randomly selected, as long as they exceed the corresponding preset values. "Number of subject types greater than the first preset value" means that the number of subject types in the generated image must exceed the first preset value. For example, if the first preset value is 4, a challenging task means that the generated image must contain at least 4 types of subjects. "Number of each subject attribute greater than the corresponding second preset value" means that the number of subject attributes, including color, quantity, and shape, must exceed 2 colors, 5 quantities, and 1 shape. For example, generating vases with 3 colors would require more than 5 vases and more than 1 vase shape. "Number of positional relationships greater than the third preset value" means that the number of positional relationships that define the positional relationships between subjects in the generated image must exceed the third preset value. For example, if the third preset value is 1, then the generated image must define at least two positional relationships, such as defining the positional relationship between a table and a cat, or defining the positional relationship between a vase and a table.

[0054] This application embodiment sets different types of text-to-image tasks, which cover different image generation scenarios and ranges, and are applicable to different large text-to-image models, with a high degree of freedom.

[0055] The embodiments of this application can freely adjust the text-to-image tasks included in the task library according to needs, thereby expanding the scenarios and difficulty.

[0056] A2. For each task to be generated, determine the task information, extract data from the corresponding database based on the task information, and generate prompt text and judgment conditions.

[0057] For each task to be generated, corresponding prompt text and judgment conditions can be generated as follows: The task to be generated can be analyzed using a pre-trained neural network model, or through word segmentation, template matching, etc., to determine the task information. For example, if the task is for a single subject, the task information is that the subject type is one; if the task is for dual-subject attribute binding, the task information is that the subject type is two. Attribute elements are selected for each subject; these elements can be limited in the task or randomly selected. Data is extracted from the corresponding database based on the task information, such as the subject, color, shape, and quantity. Prompt text and judgment conditions for the task to be generated are generated based on the extracted data. The prompt text and judgment conditions can be pre-set with corresponding templates, and the extracted data is added to the templates to form the prompt text and judgment conditions. Multiple data points can be extracted to form multiple prompt texts for the task to be generated and corresponding judgment conditions for each prompt text. In this embodiment, each prompt text has its corresponding judgment condition.

[0058] Each task to be generated can have one or more corresponding evaluation information. In practical applications, to improve the accuracy of evaluation results, a large amount of data is needed for evaluation. Therefore, multiple evaluation information is generated for each type of task to be generated. This application embodiment pre-stores at least one type of text-to-image task in the task library. Different types of text-to-image tasks guide the generation of prompt text and discrimination conditions, accurately generating different types of prompt text and discrimination conditions, enriching the types of prompt text, and improving the accuracy of test results.

[0059] Optionally, the database includes: a subject database, an attribute database, and a location relationship database; the task information includes at least one of the following: the number of types of subjects to be generated, the attribute status information of each type of target subject, and the location relationship status of different target subjects.

[0060] In this embodiment, the subject database stores different types of subjects, such as cats, dogs, vases, tables, flowers, etc. The attribute database stores different attributes, such as yellow, green, blue, 1, 10, square, circle, ellipse, rhombus, etc. The attributes described above are different types of attributes, and different types of attributes can be stored in different attribute databases. For example, the attribute database includes a color database, a quantity database, and a shape database, where the color database stores different colors, the quantity database stores different quantities, and the shape database stores different shapes. The positional relationship database stores different positional relationships, such as above, below, left, upper right, etc. The subject database, attribute database, and positional relationship database are pre-generated.

[0061] In this embodiment, the target subject can be understood as the subject generated by the large-scale model of the text image to be evaluated; the number of target subject types describes how many types of target subjects there are, for example, two. Attribute status information can be understood as information describing whether an attribute exists. For example, attribute status information indicating whether an attribute exists or does not exist indicates whether the attribute of the target subject is limited. When an attribute exists, the attribute status information can include the specific type of the attribute; for example, empty attribute status information indicates that the attribute does not exist, attribute status information indicating color indicates that the attribute exists and the attribute is color, thus limiting the color of the subject; or, attribute status information indicating the existence of an attribute, but without limiting the specific type of the attribute, color, quantity, and shape can be randomly selected. The positional relationship status of different target subjects can be understood as information describing whether a positional relationship exists, for example, positional relationship status indicating that a positional relationship exists.

[0062] Optionally, data can be extracted from the corresponding database based on the task information, and prompt text and judgment conditions for the task to be generated can be generated, including steps B1-B4:

[0063] B1. Based on the type and number of target subjects, query the subject database to determine at least one primary target subject.

[0064] In this embodiment, the first target subject can be understood as the subject that needs to be drawn in the image. Based on the number of types of the first target subject, a pre-generated subject database is queried, and at least one first target subject with a number equal to the number of types is selected from the subject database. For example, if the number of types is 1, a cat is selected as the first target subject from the subject database; if the number of types is 2, both cats and dogs are selected as the first target subjects from the subject database.

[0065] B2. If the attribute status information indicates that the attribute exists, query the attribute database to determine the first target attribute; if the attribute status information indicates that the attribute does not exist, determine that the first target attribute is empty.

[0066] In this embodiment, the first target attribute can be understood as the relevant attribute of the subject to be drawn in the image, such as color, shape, quantity, etc. If the attribute status information indicates that the attribute exists, the attribute type is determined based on the attribute status information, or one or more attribute types are randomly selected. The corresponding attribute database is queried based on the attribute type, and the attribute of the corresponding type is selected from the attribute database as the first target attribute of the first target subject. For example, if the attribute type is color, the corresponding first target attribute is red; if the attribute type is shape, the corresponding first target attribute is square; if the attribute type is quantity, the corresponding first target attribute is two. The first target attribute determined in this embodiment can be multiple, and the attribute type can be of various types. Each type of attribute can correspond to multiple first target attributes. When the attribute status information indicates that the attribute exists, the specific first target attribute can be determined by querying the attribute database, and the query result is used as the first target attribute. If the attribute status information indicates that the attribute does not exist, the first target attribute is determined to be empty. This step can determine the first target attribute corresponding to each type of first target subject.

[0067] B3. If the positional relationship status is "a first positional relationship exists", query the positional relationship database to determine the first target positional relationship; if the positional relationship status is "no positional relationship exists", determine that the first target positional relationship is empty.

[0068] In this embodiment, the first target positional relationship can be understood as the relevant positional relationship of the subject to be drawn in the image, such as above, below, etc. If the positional relationship status is "existing," the positional relationship is selected from the positional relationship database as the first target positional relationship of the first target subject. The first target positional relationship determined in this embodiment can be one or more, which can be carried in the attribute status information or determined by analyzing the number of subjects to be generated. When the positional relationship status is "existing," the specific first target positional relationship can be determined by querying the positional relationship database, that is, the query result is used as the first target positional relationship. If the positional relationship status information is "no positional relationship," the first target positional relationship is determined to be empty. This step can determine the first target position relationship corresponding to the first target subject. The first target position relationship obtained at this time can be the position relationship between different first target subjects. The different first target subjects can be two different first target subjects of the same type, or two different first target subjects of different types. In the embodiments of this application, when determining the target position relationship of different first target subjects, it can directly associate it with two first target subjects when generating the first target position relationship, that is, directly determine the first target position relationship of two first target subjects, or it can not be associated, and then randomly select two first target subjects to make their position relationship the first target position relationship.

[0069] B4. Generate prompt text and judgment conditions based on the first target subject, the first target attributes, and the positional relationship of the first target.

[0070] The system combines all the obtained primary target subjects, primary target attributes, and primary target positional relationships. This combination can be random or based on existing relationships to generate prompt text and judgment conditions. For example, if primary target subject 1 is a cat, primary target subject 2 is a table, primary target subject 1's primary target attribute is black, primary target subject 2's primary target attribute is round, and primary target subject 1 is above primary target subject 2, the generated prompt text would be: "Draw a black cat on a round table." The judgment conditions would be: Subject 1: Cat, Color: Black, Position: Above Subject 2; Subject 2: Table, Shape: Round, Position: Below Subject 1.

[0071] For example, for a text-based image task generating a single subject, the prompt text is "Draw a cat," and the discrimination condition is "Subject: cat, Missing subject: 0." For a color-bound text-based image task, taking the number of subjects as equal to one as an example, the prompt text is "Draw a yellow cat," and the discrimination condition is "Subject: cat, color: yellow." For a quantity-bound text-based image task, taking the type of subjects as equal to one as an example, the prompt text is "Draw three cats," and the discrimination condition is "Subject: cat; Quantity: 3." For a text-based image task with multiple subjects, the prompt text is "Draw a cat and a table," and the discrimination condition is "Subject: cat, table; Missing subject: 0." The difficulty of the task can be controlled by manually setting the discrimination conditions and the range of quantities and colors. For example, a prompt text containing five subjects can be set with a number for "Missing subject," so that even if one subject is missing from the image, it is still considered correct. It also allows setting text-to-image tasks with multiple subjects and attributes. For example, the prompt text could be "Draw a yellow cat on a round red table." The criteria for this prompt text would be "Subject 1: Cat, Color: Yellow, Position: Subject 2, Above; Subject 2: Table, Color: Red, Shape: Round, Position: Subject 1, Below." The prompt text can be expanded according to different text-to-image tasks and their difficulty levels. The more criteria set, the more difficult the task becomes. For example, "Three red dogs sitting on four green dogs" requires generating both cats and dogs. The cats must be red, the dogs must be green, there must be three cats and four dogs, and the cats must be on the dogs. The corresponding criteria are generated simultaneously with the prompt text. For example, if there are 100 subjects in the subject library, 100 prompt texts and 100 criteria for each subject can be generated.

[0072] This application embodiment extracts relevant information from at least one of the following databases—a subject database, an attribute database, and a location relationship database—by analyzing task information. It accurately generates prompt text and judgment conditions, and the generated prompt text and judgment conditions are diverse, covering different scenarios and meeting various evaluation needs, thus improving the accuracy of the evaluation results. The database includes a subject database; or, a subject database and an attribute database; or, a subject database and a location relationship database; or, a subject database, an attribute database, and a location relationship database. If the database does not include an attribute database, the attribute status information can be considered as non-existent; if the database does not include a location relationship database, the location relationship status can be considered as non-existent.

[0073] Optionally, the steps for generating the prompt text include: extracting data from the database, generating the prompt text and discrimination conditions; the database includes a subject database; or, the database includes a subject database and an attribute database; or, the database includes a subject database and a location relationship database; or, the database includes a subject database, an attribute database, and a location relationship database.

[0074] The system randomly extracts data from a database to generate prompt text and judgment conditions. The database may include a subject database; or, it may include a subject database and an attribute database; or, it may include a subject database and a location relationship database; or, it may include a subject database, an attribute database, and a location relationship database. In other words, the database may only include a subject database, and further include at least one of the attribute database and a location relationship database. For example, N1 subjects are randomly extracted from the subject database, N2 attributes are extracted from the attribute database, and N3 location relationships are extracted from the location relationship database. This embodiment of the application randomly extracts corresponding data from the database and randomly generates different prompt texts and judgment conditions to cover different scenarios and improve the accuracy of model evaluation.

[0075] Optionally, data can be extracted from the database to generate prompt text and judgment conditions, including:

[0076] C1. Extract at least one second target subject from the subject database;

[0077] In this embodiment, the second target subject can be understood as the subject that needs to be drawn in the image. A pre-generated subject database is queried, and one or more second target subjects are randomly selected from the database. For example, a cat can be selected as the second target subject from the subject database, or a cat and a dog can be selected as the second target subjects.

[0078] C2. For each second target subject, generate prompt text and discrimination conditions based on the second target subject; or, extract at least one second target attribute from the attribute database, and generate prompt text and discrimination conditions based on the second target subject and the second target attribute; or, extract at least one second target positional relationship from the positional relationship database, and generate prompt text and discrimination conditions based on the second target subject and the second target positional relationship; or, extract at least one second target attribute from the attribute database and at least one second target positional relationship from the positional relationship database, and generate prompt text and discrimination conditions based on the second target subject, the second target attribute, and the second target positional relationship.

[0079] In this embodiment, the second target attribute can be understood as the relevant attributes of the subject to be drawn in the image, such as color, shape, quantity, etc. The second target positional relationship can be understood as the relevant positional relationship of the subject to be drawn in the image, such as above, below, etc.

[0080] One or more attributes are randomly selected from the attribute database as the second target attribute of the second target subject. For example, red is selected as the second target attribute from colors, square is selected as the second target attribute from shapes, and two are selected as the second target attribute from quantities. The second target attribute determined in this application embodiment can be multiple, and the attribute types can be various, with each type of attribute corresponding to one or more second target attributes. Prompt text and judgment conditions are generated based on the second target subject and the second target attributes.

[0081] The location relationships are selected from the location relationship database as the second target location relationships for the second target subject. The second target location relationships determined in this application embodiment can be one or more. This step determines the second target location relationships corresponding to the second target subject. The obtained second target location relationships can be location relationships between different second target subjects. These different second target subjects can be two different second target subjects of the same type, or two different second target subjects of different types. In determining the target location relationships of different second target subjects, this application embodiment can directly associate them with two second target subjects when generating the second target location relationships, i.e., directly determine the second target location relationships of two second target subjects, or it can choose not to associate them and subsequently randomly select two second target subjects to make their location relationships the second target location relationships. Prompt text and judgment conditions are generated based on the second target subject and the second target location relationships.

[0082] One or more attributes are randomly selected from the attribute database as the second target attribute of the second target subject, and positional relationships are selected from the positional relationship database as the second target positional relationships of the second target subject. Based on the second target subject, the second target attribute, and the second target positional relationships, prompt text and judgment conditions are generated.

[0083] In the embodiments of this application, when generating prompt text and discrimination conditions, the prompt text and discrimination conditions can be generated solely based on the second target subject, or the second target subject can be extracted first, and then at least one of the second target attributes and the positional relationship between the second target and the second target can be extracted to generate prompt text and discrimination conditions.

[0084] In this embodiment, relevant information is randomly extracted from at least one of the subject database, attribute database, and location relationship database, and different prompt texts and discrimination conditions are randomly generated. The generated prompt texts and discrimination conditions are diverse, covering different scenarios, meeting different evaluation needs, and improving the accuracy of evaluation results.

[0085] S202. For each prompt text, input the prompt text into the large model of text-to-image to be evaluated, and determine the image to be judged based on the output of the large model of text-to-image to be evaluated.

[0086] S203. Perform target detection and segmentation on the image to be judged, and determine the image detection result. The image detection result includes at least one detection box corresponding to the subject to be detected, the detection box attributes, and the mask.

[0087] In this embodiment, the image detection result can be understood as the result obtained by performing target detection on the image to be judged. The subject to be detected can be understood as the subject detected from the image to be judged; the detection box attribute is the type of subject selected by the detection box. In image processing, a mask is a special type of image used to specify the area to be operated on in the original image. The mask is usually a binary image (that is, each pixel in the image has only two possible values, usually 0 and 255, representing black and white respectively), but it can also be a grayscale image or a multi-channel image.

[0088] Target detection and segmentation of an image to be judged can be achieved through a network model or algorithm. For example, an image segmentation network model is predetermined, and the image to be judged is input into the image segmentation network model for detection and segmentation to determine the image detection result. The image segmentation network model is used to identify the targets in the image to be judged, for example, identifying all targets in the image to be judged and segmenting them, with each target being taken as a subject to be detected. Based on the segmentation result, the detection box, detection box attributes, and mask corresponding to each subject to be detected are obtained. Alternatively, in this embodiment of the application, when performing target detection and segmentation on the image to be judged, some targets can be identified. For example, the subjects to be detected are determined in advance based on prompt text and / or discrimination conditions. The image segmentation network model performs target detection and segmentation on the image to be judged according to the type of subject, obtaining the detection box, detection box attributes, and mask corresponding to the subject to be detected. For example, the subject to be detected is a cat. In addition to the cat, there may be other subjects (such as dogs, tables, etc.) in the image to be judged. The cat in the image to be judged is detected, and its corresponding detection box, detection box attributes, and mask are determined.

[0089] The image segmentation network model in this application embodiment can be Grounded-Segment-Anything (GSAM). GSAM is a project based on Grounding DINO and Segment-Anything that performs image detection and segmentation tasks using a single sentence. Grounding DINO usually refers to the Dialogue-Information-Integration module (DIM) in the DINO algorithm. Its main purpose is to provide relevant information for the detection task based on a single sentence. The DINO algorithm includes several modules such as Dialogue Policy (DP), Question Selection Network (QSN), and the aforementioned DIM. In the entire DINO model, each dialogue action (such as speaking a sentence) is generated after weighted processing by each module. Segment-Anything is generally used to refer to image and semantic segmentation technology, which can segment any image into several meaningful segments or regions according to semantics. This technology has wide applications, such as scene understanding, medical image diagnosis, and object recognition. Its main purpose is to segment target objects. The specific method typically employs a convolutional neural network structure for image feature extraction, followed by a semantic segmentation network for pixel-level classification. Finally, contextual understanding and processing techniques are used to connect related regions into a meaningful paragraph.

[0090] The GSAM model enables object detection and segmentation, recognizing different types of targets with results independent of the training set. Traditional techniques use CNNs for image object detection and recognition, but evaluating a large model with millions of parameters using tens of thousands of CNN parameters is challenging and fails to cover all possible capabilities. For example, detection and segmentation models struggle to accurately detect and segment objects not present in the training set, impacting the evaluation task. This application utilizes the GSAM model for object detection and segmentation, achieving more accurate results.

[0091] S204. Perform attribute detection on the mask to determine the attribute detection results of the subject to be detected.

[0092] In this embodiment, the attribute detection result can be understood as the result obtained by detecting the attributes of the mask; for example, if the attribute is color, the attribute detection result can be a specific color such as red or yellow; or, the probability or confidence level of a certain color, such as the confidence level of belonging to red. Based on a pre-set model and algorithm, the attributes of all masks are detected, such as color and shape, to obtain the attribute detection result of the subject to be detected. By performing attribute detection on the mask of each subject to be detected, the attribute detection result of each subject to be detected is obtained. Combining the image detection result and the attribute detection result, the detection box, detection box attributes, and attribute detection result of all subjects to be detected can be obtained, such as color and shape.

[0093] Optional attributes include color and shape; attribute detection is performed on the mask to determine the attribute detection results of the subject to be detected, including C1-C3:

[0094] C1. Determine the target color and target shape corresponding to the mask based on the discrimination conditions.

[0095] In this embodiment, the target color can be understood as the color of the target subject to be generated; the target shape can be understood as the shape of the target subject to be generated.

[0096] The discrimination criteria are analyzed to determine the target color and shape corresponding to the target subject. The target subject is then matched with a mask to determine the target color and shape corresponding to the mask. Since multiple subjects may exist in the image to be discriminated against, and these subjects may be of the same or different types, the discrimination criteria determine the target color and shape corresponding to each subject, which needs to be associated with a mask. Therefore, when determining the target color and shape corresponding to the mask, the type and positional relationship of the subject to be detected can be analyzed. Based on the analysis results, the target subject is matched with the subject to be detected, and then the target color and shape corresponding to the mask are determined based on the correspondence between the subject to be detected and the mask.

[0097] C2. If the target color exists, perform color detection on the mask based on the first visual language pre-trained model and the target color, determine the first confidence level output by the first visual language pre-trained model, and use the first confidence level as the attribute detection result of the subject to be detected.

[0098] In this embodiment, the first visual language pre-training model is a Contrastive Language-Image Pre-Training (CLIP) model. The first confidence score is the confidence score output by the first visual language pre-training model, which is used to describe the reliability of the color detection result.

[0099] If the target color exists, the target color and mask are input into the first visual language pre-trained model, instructing the model to detect and recognize the target color in the mask. The model uses its learned knowledge to determine whether the mask's color matches the target color, determines and outputs a first confidence score, and uses this first confidence score as the attribute detection result for the subject. If the target color does not exist, the first confidence score can be set to empty or a specific flag value so that subsequent processing can determine that the target color does not exist, eliminating the need to judge whether the generated subject's color is correct.

[0100] C3. If the target shape exists, perform shape detection on the mask based on the second visual language pre-trained model and the target shape, determine the second confidence level output by the second visual language pre-trained model, and use the second confidence level as the attribute detection result of the subject to be detected.

[0101] In this embodiment, the second visual language pre-training model is a Contrastive Language-Image Pre-Training (CLIP) model; the first and second visual language pre-training models can be the same model or different models. The second confidence score is the confidence score output by the second visual language pre-training model, used to describe the reliability of the shape detection result.

[0102] If the target shape exists, the target shape and mask are input into the second visual language pre-trained model, instructing the model to detect and recognize the target shape from the mask. The model uses its learned knowledge to determine whether the mask's shape matches the target shape, determines and outputs a second confidence score, and uses this score as the attribute detection result for the subject. If the target shape does not exist, the second confidence score can be set to empty or a specific flag value so that subsequent processing can determine that the target shape does not exist, eliminating the need to judge the correctness of the generated subject's shape.

[0103] This application embodiment uses CLIP to perform shape and color detection on the mask, which can be extended to target detection of thousands of categories, balancing accuracy and breadth, and avoiding secondary training in downstream tasks. Furthermore, by decoupling shape and color detection, this application embodiment can accurately determine the result of each detection, and accurately pinpoint whether the generated shape or color is incorrect when errors occur.

[0104] S205. Determine the target evaluation result based on the discrimination conditions, image detection results, and attribute detection results.

[0105] Using the discrimination criteria as the detection standard, the image detection results and attribute detection results are compared with the discrimination criteria, and the target detection result is determined based on the comparison results. For example, the type, quantity, and positional relationship of the generated subject are determined based on the image detection results and attribute detection results, and then compared with the discrimination criteria to determine whether the type, quantity, and positional relationship of the generated subject are correct. If correct, the target detection result is determined to be passed; otherwise, the target detection result is determined to be failed.

[0106] Optionally, the target evaluation result, including D1-D3, is determined based on the discrimination criteria, image detection results, and attribute detection results:

[0107] D1. Determine the discrimination information of the target subject based on the discrimination conditions. The discrimination information includes at least one of the following: type, color, quantity, shape, and positional relationship.

[0108] In this embodiment, the discrimination information can be understood as information used as discrimination criteria. The discrimination information includes at least one of the following: type, color, quantity, shape, and positional relationship. Positional relationship refers to the positional relationship between two target subjects. The two target subjects can be subjects of the same type or subjects of different types. The discrimination conditions are analyzed to determine all subjects included in the discrimination conditions. Each subject or each subject in the discrimination conditions is taken as a target subject, and the discrimination information for each subject or each target subject is determined. For example, if the discrimination conditions are "Subject: Cat; Quantity: 3", the target subject is a cat, and the discrimination information includes: quantity is 3; if the discrimination conditions are "Subject 1: Cat, Color: Yellow; Subject 2: Table, Shape: Round", the target subject 1 is a cat, and the discrimination information includes: color is yellow; the target subject 2 is a table, and the discrimination information includes: shape is round; if the discrimination conditions are "Subject 1: Cat, Color: Yellow; Subject 2: Cat, Color: White", the target subject 1 is a cat, and the discrimination information includes: color is yellow; the target subject 2 is a cat, and the discrimination information includes: color is white.

[0109] D2. Determine the discriminant information of the subject to be detected based on the attribute detection results and image detection results. The discriminant information includes at least one of the following: type, color, quantity, shape, and positional relationship.

[0110] In this embodiment, the information to be discriminated can be understood as the information of the subject to be detected that needs to be discriminated. The information to be discriminated includes at least one of the following: type, color, quantity, shape, and positional relationship; wherein, the positional relationship refers to the positional relationship between two subjects to be detected. The two subjects to be detected can be subjects of the same type or subjects of different types. The attribute detection results and image detection results are analyzed. Based on the detection box attributes and detection boxes of the image detection results, the type, quantity, and position of the subjects to be detected are determined. Based on the attribute detection results, the color and shape of the subjects to be detected are determined. Based on the above analysis results, the discriminant information for each type of subject to be detected is determined through comprehensive analysis.

[0111] D3. Compare the discriminant information with the information to be discriminated to determine the target evaluation result.

[0112] The discrimination information of all target subjects is compared with the discrimination information of the subject to be detected. Using the discrimination information of the target subjects as a standard, it is determined whether the discrimination information of the subject to be detected is consistent with the discrimination information of the target subjects; that is, whether the subject to be detected is generated according to the relevant requirements of the target subjects. In this embodiment, when comparing discrimination information and the information to be detected, an error range can be determined based on discrimination conditions, or an error range can be preset, with corresponding error ranges set according to different text-based image tasks. For example, "missing subject: 1" indicates that the allowed error range is that one subject can be missing, meaning the number of subjects to be detected can be one less than the number of target subjects. Finally, the comparison results are analyzed to determine the target evaluation result.

[0113] The embodiments of this application evaluate the subject from one or more dimensions, such as the subject's type, color, quantity, shape, and positional relationship information, resulting in more accurate results. Furthermore, it can accurately locate errors when the image generated by the Wenshengtu large model is inaccurate, such as errors in subject type generation, subject attribute generation, or subject positional relationship generation.

[0114] Optionally, the attribute detection results include a first confidence level and a second confidence level; the discriminant information of the subject to be detected is determined based on the attribute detection results and the image detection results, including at least one of the following:

[0115] E1. Determine the type of the subject to be detected based on the detection frame attributes of the subject to be detected.

[0116] The detection box attribute indicates the type of the subject to be detected, so the detection box attribute of the subject to be detected can be directly used as the type of the subject to be detected.

[0117] E2. Count the number of detection frames of the same type of subject to be detected, and determine the number of subjects to be detected of each type.

[0118] Based on the detection box attributes, determine the detection boxes of the same type of subject to be detected, count the number of detection boxes of the same type of subject to be detected, and obtain the number of subjects to be detected of this type; count all subjects to be detected by type, and obtain the number of subjects to be detected of each type. For example, there are 3 kittens and 2 puppies.

[0119] E3. Determine the color of the subject to be detected based on the first confidence level and the first confidence threshold.

[0120] In this embodiment, the first confidence threshold can be preset, for example, set according to accuracy requirements; for example, the first confidence threshold is 0.75. The first confidence level is compared with the first confidence threshold. If the first confidence level is not less than the first confidence threshold, it indicates that the generated target color of the subject to be detected is correct, and the color of the subject to be detected is determined to be the corresponding target color. That is, in this embodiment, when performing color detection on the mask based on the first visual language pre-trained model and the target color, if the first confidence level is not less than the first confidence threshold, the color of the subject to be detected is the target color. If the first confidence level is less than the first confidence threshold, the color of the subject to be detected is not the target color.

[0121] E4. Determine the shape of the subject to be detected based on the second confidence level and the second confidence threshold.

[0122] In this embodiment, the second confidence threshold can be preset, for example, set according to accuracy requirements; for example, the second confidence threshold is 0.75. The second confidence level is compared with the second confidence threshold. If the second confidence level is not less than the second confidence threshold, it indicates that the generated target shape of the subject to be detected is correct, and the shape of the subject to be detected is determined to be the corresponding target shape. That is, in this embodiment, when performing shape detection on the mask based on the second visual language pre-trained model and the target shape, if the second confidence level is not less than the second confidence threshold, the shape of the subject to be detected is the target shape. If the second confidence level is less than the second confidence threshold, the shape of the subject to be detected is not the target shape.

[0123] E5. Determine the positional relationship between the subjects to be detected based on the positions of the detection frames of at least two subjects to be detected.

[0124] Detection boxes are typically represented by coordinates, and their positions can be determined based on the coordinates of the detection boxes of the subjects to be detected. When determining positional relationships, any two subjects to be detected can be selected, and their positional relationship can be determined based on the position of each subject. In this embodiment, when determining positional relationships, pairwise selection can be performed sequentially to determine the positional relationships between all subjects to be detected, or the subjects whose positional relationships need to be determined can be identified based on discrimination conditions, and the corresponding two subjects can be selected to determine the positional relationship between them.

[0125] In this embodiment of the application, when detecting a generated subject or a type of subject to be detected, the detection box and its attributes are used to determine whether the target subject to be generated has been generated; the confidence level of CLIP output is 0 to 1 to determine whether the generated color and shape are correct, with a confidence level greater than 0.75 indicating correctness and otherwise incorrectness, thus determining whether the generated color and shape are the target color and shape; the number of subjects to be detected is determined by counting the number of detection boxes; and the positional relationships such as up, down, left, and right are returned by comparing the positions of two detection boxes.

[0126] The target evaluation results of the images generated for each evaluation prompt text are compiled and analyzed. For each image, the algorithm outputs whether the generated subject exists, and whether the generated color / quantity, etc., are correct. Python tools are used to read the judgment result file or Excel spreadsheet to organize the target detection results of all images, facilitating the analysis of the model's performance across different dimensions. The evaluation information corresponding to the same type of text-based image task can be analyzed to determine the model's performance in that specific text-based image task scenario. In different text-based image task scenarios, the target evaluation result for each image will have pass and fail. If the target evaluation result is pass, it means the model generated correctly.

[0127] This application's embodiments evaluate from multiple dimensions, determining the subject type, positional relationship, and quantity based on the location and attributes of the detection box. The detection of subject type, positional relationship, and quantity is converted into the detection of detection boxes, making the process simpler and the results more accurate. The detection of color and shape is converted into a comparison of confidence levels, and the detection accuracy can be adjusted by setting thresholds to meet the needs of different business scenarios.

[0128] The method provided in this application can also evaluate multiple large-scale text-based image models, and through a multi-task evaluation method, provide the detection results of different large-scale text-based image models under different text-based image tasks. The best model is selected through comparison.

[0129] This invention provides an evaluation method for large-scale text-generated image models, solving the problem of inaccurate evaluation results. It covers different image generation scenarios and ranges through various text-generated image tasks, and is applicable to different large-scale text-generated image models, offering a high degree of flexibility. The large-scale text-generated image model to be evaluated generates an image to be judged based on prompt text, and then automatically detects the image to be judged according to discrimination conditions to obtain the target evaluation result. Finally, the large-scale text-generated image model to be evaluated is evaluated based on the target evaluation results corresponding to different evaluation information. The evaluation process is fully automated, requiring no human intervention, greatly replacing manual judgment, avoiding human error, and resulting in more accurate and objective results. The indicators are accurately quantified, the evaluation speed is fast, and costs are effectively reduced. Furthermore, arbitrary scenarios and difficulties can be added and expanded by adjusting the prompt text and discrimination conditions. Simultaneously, using GSAM and CLIP large-scale models to evaluate large-scale text-generated image models yields accurate results with high coverage.

[0130] Example 3

[0131] Figure 3 This is a schematic diagram of the structure of an evaluation device for a large-scale text-based image model provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes a text acquisition module 31 and an evaluation module 32.

[0132] Text acquisition module 31 is used to acquire the prompt text and its corresponding discrimination conditions;

[0133] The evaluation module 32 is used to input the prompt text into the large model of the text-to-image to be evaluated for each prompt text, determine the image to be judged based on the output of the large model of the text-to-image to be evaluated, evaluate the image to be judged based on the discrimination conditions corresponding to the prompt text, and determine the target evaluation result.

[0134] This invention provides an evaluation device for large-scale text-based image models, which solves the problem of inaccurate evaluation results. It acquires prompt text and discrimination conditions, generates a discrimination image based on the prompt text, and automatically detects the discrimination image using the discrimination conditions to obtain the target evaluation result. The evaluation process is fully automated, requiring no manual intervention, avoiding human error, resulting in more accurate and objective results, faster evaluation speed, and lower cost.

[0135] Optionally, the prompt text and the corresponding discrimination condition include subject information; or, the prompt text and the corresponding discrimination condition include subject information and attribute information; or, the prompt text and the corresponding discrimination condition include subject information and positional relationship information; or, the prompt text and the corresponding discrimination condition include subject information, attribute information, and positional relationship information.

[0136] Optionally, the device may also include:

[0137] The task selection module is used to select at least two types of text-to-image tasks from the task library as tasks to be generated.

[0138] The first data extraction module is used to analyze each task to be generated, determine the task information of the task to be generated, extract data from the corresponding database based on the task information, and generate prompt text and discrimination conditions.

[0139] Optionally, the database includes: a subject database, an attribute database, and a location relationship database; the task information includes at least one of the following: the number of target subject types, attribute status information of each type of target subject, and location relationship status of different target subjects;

[0140] Optional, the data extraction module includes:

[0141] The target subject determination unit is used to query the subject database based on the type and number of the target subjects to determine at least one first target subject;

[0142] The target attribute determination unit is configured to query the attribute database and determine the first target attribute if the attribute status information indicates that the attribute exists; and to determine that the first target attribute is empty if the attribute status information indicates that the attribute does not exist.

[0143] A positional relationship determination unit is configured to, if the positional relationship status is that a positional relationship exists, query a positional relationship database to determine a first target positional relationship; and if the positional relationship status is that a positional relationship does not exist, determine that the first target positional relationship is empty.

[0144] The prompt text generation unit is used to generate prompt text and discrimination conditions based on the first target subject, the first target attribute and the first target position relationship.

[0145] Optionally, the task library includes at least two types of text-based image tasks:

[0146] Single-subject generation;

[0147] Single main color;

[0148] Number of single entities;

[0149] Single main body shape;

[0150] Dual-subject generation;

[0151] Dual-subject attribute binding;

[0152] The positional relationship between the two main entities;

[0153] Three entities are generated;

[0154] Three main entities generate attribute binding;

[0155] The positional relationship of the three main entities;

[0156] The difficult task is defined as follows: the number of subject types in the difficult task is greater than a first preset value, the number of each attribute of the subject is greater than a corresponding second preset value, and the number of positional relationships is greater than a third preset value.

[0157] Optionally, the device may also include:

[0158] The second data extraction module is used to extract data from the database and generate prompt text and its corresponding judgment conditions.

[0159] The database includes a subject database; or, the database includes a subject database and an attribute database; or, the database includes a subject database and a location relationship database; or, the database includes a subject database, an attribute database, and a location relationship database.

[0160] Optionally, the second data extraction module is specifically configured to: extract at least one second target subject from the subject database; for each second target subject, generate prompt text and discrimination conditions based on the second target subject; or, extract at least one second target attribute from the attribute database, and generate prompt text and discrimination conditions based on the second target subject and the second target attribute; or, extract at least one second target positional relationship from the positional relationship database, and generate prompt text and discrimination conditions based on the second target subject and the second target positional relationship; or, extract at least one second target attribute from the attribute database and at least one second target positional relationship from the positional relationship database, and generate prompt text and discrimination conditions based on the second target subject, the second target attribute, and the second target positional relationship.

[0161] Optional, evaluation module 32 includes:

[0162] The target detection unit is used to perform target detection and segmentation on the image to be judged, and determine the image detection result. The image detection result includes at least one detection box, detection box attributes and mask corresponding to the subject to be detected.

[0163] An attribute detection unit is used to perform attribute detection on the mask and determine the attribute detection result of the subject to be detected.

[0164] The target evaluation result determination unit is used to determine the target evaluation result based on the discrimination conditions, the image detection result, and the attribute detection result.

[0165] Optionally, the attributes include color and shape;

[0166] Optionally, the attribute detection unit is specifically configured to: determine the target color and target shape corresponding to the mask based on the discrimination conditions; if the target color exists, perform color detection on the mask according to the first visual language pre-trained model and the target color, determine the first confidence level output by the first visual language pre-trained model, and use the first confidence level as the attribute detection result of the subject to be detected; if the target shape exists, perform shape detection on the mask according to the second visual language pre-trained model and the target shape, determine the second confidence level output by the second visual language pre-trained model, and use the second confidence level as the attribute detection result of the subject to be detected.

[0167] Optionally, the target evaluation result determination unit is specifically used for: determining the discrimination information of the target subject according to the discrimination conditions, wherein the discrimination information includes at least one of the following: type, color, quantity, shape, and positional relationship; determining the discrimination information to be determined of the subject to be detected according to the attribute detection result and the image detection result, wherein the discrimination information to be determined includes at least one of the following: type, color, quantity, shape, and positional relationship; comparing the discrimination information with the discrimination information to be determined, and determining the target evaluation result.

[0168] Optionally, the attribute detection result includes a first confidence level and a second confidence level; determining the discriminant information of the subject to be detected based on the attribute detection result and the image detection result includes at least one of the following:

[0169] The type of the subject to be detected is determined based on the detection frame attributes of the subject to be detected;

[0170] Count the number of detection frames of the same type of subject to be detected, and determine the number of subjects to be detected of each type;

[0171] The color of the subject to be detected is determined based on the first confidence level and the first confidence threshold.

[0172] The shape of the subject to be detected is determined based on the second confidence level and the second confidence threshold.

[0173] The positional relationship between the subjects to be detected is determined based on the positions of the detection frames of at least two subjects to be detected.

[0174] The evaluation device for large-scale text-based images provided in this embodiment of the invention can execute the evaluation method for large-scale text-based images provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0175] Example 4

[0176] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0177] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded into the RAM 43 from storage unit 48. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0178] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0179] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as the evaluation methods for large text-based graph models.

[0180] In some embodiments, the evaluation method for the Wensheng large model can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the Wensheng large model evaluation method described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to perform the Wensheng large model evaluation method by any other suitable means (e.g., by means of firmware).

[0181] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0182] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0183] This invention provides a computer program product, which includes a computer program that, when executed by a processor, implements the evaluation method for large-scale text-based graph models as described in any embodiment of this invention.

[0184] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0185] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0186] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0187] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0188] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0189] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An evaluation method for a large-scale text-to-image model, characterized in that, include: Obtain the prompt text and its corresponding judgment conditions; For each of the aforementioned prompt texts, the prompt text is input into the large-scale model of the text-to-image to be evaluated. The image to be judged is determined based on the output of the large-scale model of the text-to-image to be evaluated. The image to be judged is evaluated based on the discrimination conditions corresponding to the prompt text, and the target evaluation result is determined.

2. The method according to claim 1, characterized in that, The prompt text and the corresponding discrimination conditions include subject information; or, the prompt text and the corresponding discrimination conditions include subject information and attribute information; or, the prompt text and the corresponding discrimination conditions include subject information and positional relationship information; or, the prompt text and the corresponding discrimination conditions include subject information, attribute information, and positional relationship information.

3. The method according to claim 1 or 2, characterized in that, Also includes: Select at least one type of text-based image task from the task library as the task to be generated; For each task to be generated, the task information of the task to be generated is determined, and data is extracted from the corresponding database based on the task information to generate a prompt text and the discrimination conditions corresponding to the prompt text.

4. The method according to claim 3, characterized in that, The database includes: a subject database, an attribute database, and a location relationship database; the task information includes at least one of the following: the number of target subject types, the attribute status information of each type of target subject, and the location relationship status of different target subjects; The step of extracting data from the corresponding database based on the task information and generating the prompt text and discrimination conditions for the task to be generated includes: Based on the number of target subjects, query the subject database to determine at least one first target subject; If the attribute status information indicates that the attribute exists, query the attribute database to determine the first target attribute; if the attribute status information indicates that the attribute does not exist, determine that the first target attribute is empty. If the positional relationship status indicates that a positional relationship exists, query the positional relationship database to determine the first target positional relationship; if the positional relationship status indicates that a positional relationship does not exist, determine that the first target positional relationship is empty. The prompt text and discrimination conditions are generated based on the first target subject, the first target attribute, and the first target position relationship.

5. The method according to any one of claims 3 or 4, characterized in that, The task library includes at least two types of text-based image tasks: Single-subject generation; Single main color; Number of single entities; Single main body shape; Dual-subject generation; Dual-subject attribute binding; The positional relationship between the two main entities; Three entities are generated; Three main entities generate attribute binding; The positional relationship of the three main entities; The difficult task is defined as follows: the number of subject types in the difficult task is greater than a first preset value, the number of each attribute of the subject is greater than a corresponding second preset value, and the number of positional relationships is greater than a third preset value.

6. The method according to any one of claims 1 or 2, characterized in that, The steps for generating the prompt text and its corresponding discrimination conditions include: Extract data from the database to generate prompt text and its corresponding judgment conditions; The database includes a subject database; or, the database includes a subject database and an attribute database; or, the database includes a subject database and a location relationship database; or, the database includes a subject database, an attribute database, and a location relationship database.

7. The method according to claim 6, characterized in that, Data is extracted from the database to generate prompt text and its corresponding judgment conditions, including: Extract at least one second target subject from the subject database; For each second target subject, generate prompt text and discrimination conditions based on the second target subject; or, Extract at least one second target attribute from the attribute database, and generate prompt text and discrimination conditions based on the second target subject and the second target attribute; or... Extract at least one second target location relationship from the location relationship database, and generate prompt text and discrimination conditions based on the second target subject and the second target location relationship; or... At least one second target attribute is extracted from the attribute database, and at least one second target positional relationship is extracted from the positional relationship database. Based on the second target subject, the second target attribute, and the second target positional relationship, a prompt text and a discrimination condition corresponding to the prompt text are generated.

8. The method according to any one of claims 1-7, characterized in that, The step of evaluating the image to be judged according to the discrimination criteria and determining the target evaluation result includes: The image to be judged is subjected to target detection and segmentation to determine the image detection result, which includes at least one detection box, detection box attributes and mask corresponding to the subject to be detected; Attribute detection is performed on the mask to determine the attribute detection result of the subject to be detected; The target evaluation result is determined based on the discrimination criteria, the image detection result, and the attribute detection result.

9. The method according to claim 8, characterized in that, The attributes include color and shape; the attribute detection of the mask to determine the attribute detection result of the subject to be detected includes: The target color and target shape corresponding to the mask are determined based on the discrimination conditions; If the target color exists, perform color detection on the mask based on the first visual language pre-trained model and the target color, determine the first confidence level output by the first visual language pre-trained model, and use the first confidence level as the attribute detection result of the subject to be detected; If the target shape exists, shape detection is performed on the mask based on the second visual language pre-trained model and the target shape to determine the second confidence level output by the second visual language pre-trained model, and the second confidence level is used as the attribute detection result of the subject to be detected.

10. The method according to claim 8, characterized in that, The determination of the target evaluation result based on the discrimination criteria, the image detection result, and the attribute detection result includes: The discrimination information of the target subject is determined according to the discrimination conditions, and the discrimination information includes at least one of the following: type, color, quantity, shape and positional relationship; The discriminant information of the subject to be detected is determined based on the attribute detection results and the image detection results. The discriminant information includes at least one of the following: type, color, quantity, shape, and positional relationship. The discrimination information is compared with the information to be discriminated to determine the target evaluation result.

11. The method according to claim 10, characterized in that, The attribute detection result includes a first confidence level and a second confidence level; the step of determining the discriminant information of the subject to be detected based on the attribute detection result and the image detection result includes at least one of the following: The type of the subject to be detected is determined based on the detection frame attributes of the subject to be detected; Count the number of detection frames of the same type of subject to be detected, and determine the number of subjects to be detected of each type; The color of the subject to be detected is determined based on the first confidence level and the first confidence threshold. The shape of the subject to be detected is determined based on the second confidence level and the second confidence threshold. The positional relationship between the subjects to be detected is determined based on the positions of the detection frames of at least two subjects to be detected.

12. An electronic device, characterized in that, The electronic device includes: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the evaluation method of the Wensheng large model according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the evaluation method for the large model of the text image according to any one of claims 1-11.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the evaluation method for the large model of the text image according to any one of claims 1-11.