Multi-object text generation image semantic evaluation method and system based on scene graph
By constructing text and visual scene diagrams, multimodal feature coding and relational coding, the problem of difficult to evaluate text-generated images in complex scenarios in the prior art is solved, and accurate evaluation is achieved in multi-object scenarios.
Patent Information
- Application Number
- CN202411945550.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
Existing text-generating image evaluation methods are difficult to conduct objective and quantitative semantic consistency evaluation in complex scenarios containing multiple objects, especially in terms of object existence, attributes and relationships.
By constructing text scene diagrams and visual scene diagrams, using multimodal object feature encoding and relational encoding, the similarity between text object features and image object features, as well as the similarity between text relationship features and visual relationship features, comprehensively evaluate the semantic consistency of text generated images.
It realizes objective and accurate evaluation of the object existence, attributes and relationships of text-generated images in multi-object scenarios, and improves the accuracy and consistency of the evaluation.
Smart Images

Figure CN120014418A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation, and in particular relates to a method and system for semantic evaluation of multi-object text generation images based on scene graphs. Background Art
[0002] In recent years, with the rapid development of deep learning, the text-to-image technology that generates corresponding images based on user input text prompts has shown explosive growth. However, unlike visual tasks such as image classification, text-to-image generation has a variety of generated results and there is no single correct result. Therefore, how to objectively and quantitatively evaluate the image generation results is a difficult research problem in the text-to-image task.
[0003] Existing text-generated image evaluation methods are mainly divided into two aspects: image quality evaluation and semantic consistency evaluation. Among them, image quality evaluation focuses on the quality of generated images, including clarity, authenticity, rationality, etc. The representative evaluation indicator is FID (Fréchet Inception Distance). The pre-trained visual backbone model InceptionNetv3 extracts visual features for the real image set and the generated image set respectively, and then calculates the distance between the distributions based on the mean and variance of the feature distributions of the two image sets. The smaller the distance, the higher the quality of the generated image. However, the FID indicator does not take into account whether the generated image conforms to the user input text prompt, that is, it lacks the evaluation of semantic consistency.
[0004] In terms of semantic consistency evaluation, the representative evaluation index is CLIPSIM. The CLIP model is used to extract image features for the generated image and text features for the text prompts input by the user through the image-text pre-training model CLIP. Since the CLIP model is pre-trained on large-scale image-text pairs, it can align image and text features in the semantic space. Therefore, by calculating the cosine similarity between image features and text features, the semantic consistency of the image and text can be obtained, thereby evaluating the semantic consistency of the text-generated image. However, when the user wants to generate a complex scene containing multiple objects, the CLIP model is directly used for evaluation, and the results obtained are often inconsistent with objective facts. Specifically, there are the following situations: "object existence" problem: when the user input includes multiple objects, one of the objects is often omitted in the generated image; "object attribute" problem: when the user specifies attributes such as color and shape for multiple objects, the attributes of the objects in the generated image are often mixed or wrong; "object relationship problem": when the user specifies the spatial position, action subject and object and other relationships between multiple objects, the relationship between the corresponding objects in the generated image is often wrong. For the above situations, the results of semantic consistency evaluation using the CLIP model often do not match the objective facts.
[0005] In response to the above problems, how to conduct semantic consistency evaluation on the text-to-image model and obtain objective, quantitative and correct evaluation results in complex scenes containing multiple objects has become a difficult problem of great significance. Summary of the invention
[0006] In response to the above difficulties, the present invention proposes a method and system for semantic evaluation of multi-object text-generated images based on scene graphs. By parsing the text prompts input by the user, objects in text form and the relationships between objects are obtained, a text scene graph is constructed, and a visual scene graph is constructed according to the generated image through a target detection model. According to the similarity between the text scene graph and the visual scene graph, semantic evaluation of the text-generated image is achieved.
[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0008] A method for semantic evaluation of multi-object text generation images based on scene graphs, comprising the following steps:
[0009] Perform multimodal object feature encoding on text prompts and generated images to obtain text object features and image object features;
[0010] Calculate the similarity between text object features and image object features to obtain the semantic consistency of the objects;
[0011] Encode the object relations between text prompts and generated images to obtain textual relation features and visual relation features;
[0012] Calculate the similarity between textual relationship features and visual relationship features to obtain the semantic consistency of the relationship;
[0013] The semantic consistency of objects and the semantic consistency of relations are integrated to obtain the final semantic consistency evaluation result of text-generated image.
[0014] Furthermore, the multimodal object feature encoding of the text prompt and the generated image is used to parse the text and the image and extract features for each text object and image object, including the following steps:
[0015] (1) Text parsing: Parse the text input by the user to obtain each object in the text and the relationship between objects;
[0016] (2) Text encoding: Encode the objects in the text obtained in step (1) to obtain text object features;
[0017] (3) Image parsing: The generated image is divided into candidate regions to generate multiple candidate regions. Since the number of candidate regions is greater than the number of image objects in the image, each candidate region contains an image object or does not contain an image object.
[0018] (4) Image coding: Encode the candidate region obtained in step (3) to obtain image object features.
[0019] Furthermore, in the above step (1), the dependency relationship between the elements in the text is found through rule-based text parsing, thereby obtaining each relationship in the text and its subject and object.
[0020] Furthermore, in the above step (2), the pre-trained CLIP model is fine-tuned on the target detection and positioning dataset, and the fine-tuned CLIP model is used as a text encoder to encode the subject and object texts obtained in step (1) to obtain text object features;
[0021] Furthermore, in the above step (3), a plurality of candidate bounding boxes are generated as image candidate regions by using a region proposal network trained on a target detection dataset.
[0022] Furthermore, in the above step (4), the fine-tuned CLIP model is used as an image encoder to encode the generated image, and the feature map is cropped according to each bounding box through the RoIAlign method to obtain the image object features of each candidate region.
[0023] Furthermore, the step of calculating the similarity between the text object features and the image object features includes a text and image node matching method for finding the image object corresponding to each text object in the image, thereby obtaining the semantic consistency of the generated image in terms of the object. The method specifically includes two optimization objectives, which are respectively intended to minimize the semantic distance between the text object and the image object, thereby finding the closest match, and to minimize the overlapping areas between the selected image areas, thereby avoiding repeatedly assigning the same image object to different text objects. By minimizing the two optimization objectives, the image object corresponding to each text object can be found from multiple candidate areas in the generated image.
[0024] Furthermore, the step of encoding the relationship between the text prompt and the object in the generated image includes a method for aligning the relationship between the text and the image object, the method comprising the following steps:
[0025] (1) Encode the relationship between two text objects and their text forms to obtain text relationship features;
[0026] (2) Encode the two candidate regions assigned to the two text objects in the generated image to obtain visual relationship features;
[0027] (3) The text and image relationship encoding model is trained through the contrastive learning loss function to align the text relationship features and visual relationship features.
[0028] Furthermore, in the step of calculating the similarity between the textual relationship features and the visual relationship features, the semantic consistency of the relationship is obtained, including: obtaining the semantic consistency of the generated image in terms of the relationship according to the cosine similarity between the textual relationship features and the visual relationship features.
[0029] Furthermore, the visual relationship feature is obtained by using a three-stage parallel visual relationship encoding network, comprising the following steps:
[0030] Text stage: The subject and object of the relationship to be encoded are input into the encoder in text form to obtain the first visual feature vector;
[0031] Position stage: The bounding boxes of the subject and object in the image are input into the encoding network composed of multi-layer perceptrons in the form of coordinates to obtain the second visual feature vector, which contains the absolute and relative position relationship information of the subject and object;
[0032] Visual semantic stage: extract the feature map of the image, extract three features on the feature map based on the subject bounding box, object bounding box, and the minimum bounding box that contains both the subject and the object, and then obtain the third visual feature vector containing the visual features of the subject and the object after splicing;
[0033] The first visual feature vector, the second visual feature vector and the third visual feature vector are concatenated and input into the fully connected layer to obtain the final visual relationship feature.
[0034] Corresponding to the above method, the present invention also provides a scene graph-based multi-object text generation image semantic evaluation system, which includes:
[0035] The object semantic consistency calculation module is used to perform multimodal object feature encoding on the text prompt and the generated image, obtain text object features and image object features, calculate the similarity between the text object features and the image object features, and obtain the semantic consistency of the object;
[0036] The module for calculating the semantic consistency of the relationship is used to encode the object relationship in the text prompt and the generated image, obtain the text relationship features and the visual relationship features, calculate the similarity between the text relationship features and the visual relationship features, and obtain the semantic consistency of the relationship;
[0037] The comprehensive module is used to integrate the semantic consistency of objects and the semantic consistency of relations to obtain the final text-generated image semantic consistency evaluation results.
[0038] The effect of the present invention is that compared with the existing text-generated image evaluation method, the text-generated image semantic evaluation method of the present invention can take into account the consistency between the generated image and the text prompt given by the user, and in a complex scene containing multiple objects, it can make objective and correct evaluation results on the object existence, object attributes, object relationships and other aspects of the generated image.
[0039] The reason why the present invention has the above-mentioned inventive effects is that: by constructing scene graphs for text prompts and generated images respectively, the characteristics of each text object, image object, text relationship, and visual relationship can be obtained. By aligning in the feature space, the semantic consistency of text and image can be measured at both the object and object relationship levels, thereby obtaining the semantic consistency evaluation results of text-generated images in multi-object scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flow chart of a method for semantic evaluation of multi-object text generation images based on scene graphs of the present invention.
[0041] Figure 2 Schematic diagram of text object encoding in the embodiment.
[0042] Figure 3 Schematic diagram of image object encoding in the embodiment.
[0043] Figure 4 Schematic diagram of text and image node matching in the embodiment.
[0044] Figure 5 Schematic diagram of text and image relationship coding in the embodiment. DETAILED DESCRIPTION
[0045] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] The present embodiment provides a method for semantic evaluation of multi-object text generation images based on scene graphs, and its process is as follows: Figure 1 As shown, the following steps are included:
[0047] (1) Text object encoding
[0048] The rule-based text parser is used to parse the text prompts input by the user to obtain the objects contained in the text and the relationship between the objects, all in the form of text. Then, the parsed text objects are encoded by the text encoder of the pre-trained CLIP model to obtain the features of each text object in the text prompt. Figure 2 It is a diagram of text object encoding.
[0049] (2) Image object coding
[0050] The Region Proposal Network trained on the target detection dataset generates multiple candidate regions for the generated image. Since the number of candidate regions is greater than the number of image objects, each region may contain an image object or not. The Region Proposal Network first extracts the feature map of the input image, and then performs a sliding window operation on the feature map. Each sliding window position will generate a series of anchor points, each of which has a different size and aspect ratio, covering a variety of possible object sizes. Finally, each anchor point is classified to determine whether it contains the target object, and the bounding box regression is performed to correct the position of the anchor point to generate the final candidate region.
[0051] After obtaining the candidate regions, the fine-tuned CLIP model is used as the image encoder to extract features for each region. First, the CLIP image encoder is fine-tuned on the object detection dataset, and the image features of the region are obtained by performing feature pooling on the feature map through the RoIAlign operation. Then, the CLIP text encoder is used to obtain the text features of the region category, and the image encoder is fine-tuned through the contrastive learning loss of the alignment of the regional image features and text features. Figure 3 It is a schematic diagram of candidate region generation and region feature encoding for a generated image.
[0052] (3) Matching text and image nodes
[0053] After encoding the text object in step (1) and encoding the image region in step (2), the cosine similarity between a text object feature and an image region feature can be calculated to obtain their semantic similarity, that is, the semantic consistency of the object. However, since there are currently multiple candidate regions, it is necessary to find the image region referred to by the text object. Therefore, the present invention proposes a text and image node matching method. By optimizing the following optimization target, the corresponding image region can be found for each text object. The optimization target can be expressed as the following formula:
[0054]
[0055] in, is a set of selected text objects and corresponding image areas. is the selected image area set, Dis is the text object feature t i With the image region feature v j The cosine distance, r i 、r j It represents two image regions in the selected image region set, IoU is the intersection over union ratio between the two image regions, and λ is a hyperparameter, which can be set to 0.1 in actual use.
[0056] The first optimization goal is to maximize the similarity between the text object and the corresponding image region, so as to find the most suitable match. The second optimization goal is to avoid repeatedly selecting the same image region for multiple different text objects, so it is necessary to minimize the intersection-over-union ratio between the selected image regions, that is, to avoid too large an overlap between image regions. Figure 4 It is a schematic diagram for matching text and image nodes.
[0057] (4) Text and image relationship encoding
[0058] After node matching in step (3), one-to-one correspondence between text objects and image objects has been obtained. It is then necessary to measure whether the relationship between the two objects in the text and the image is consistent.
[0059] First, the object relationship in the text is encoded. The encoder with Transformer structure is used to set 3 special tokens to represent the subject, object and relationship features respectively. Because the relationship encoding only hopes to reflect the information of the relationship itself, it is necessary to avoid the relationship features obtained by encoding with subject and object information. Only the textual relationship is input after tokenization, and the names of the relationship subject and object are replaced by universal subject tokens and object tokens. After Transformer encoding, the textual relationship features are obtained.
[0060] Then, the object relationship in the image is encoded. Since the generated image may contain distortion noise, it is necessary to avoid the interference of the generated noise on the object relationship in the image as much as possible. The present invention designs a three-stage parallel visual relationship encoding network. (1) In the text stage, the subject and object of the relationship to be encoded are input into the encoder in the form of text to obtain the first visual feature vector. At this time, the subject and object in the image are known, and the characteristics of each subject and object can be accurately described by the text; (2) In the position stage, the bounding boxes of the subject and object in the image are input into the encoding network composed of a multi-layer perceptron (MLP) in the form of coordinates to obtain the second visual feature vector, which contains the absolute and relative position relationship information of the subject and object; (3) In the visual semantic stage, the feature map of the image is extracted by the Resnet convolutional neural network, and then the RoIAlign operation is performed on the feature map to extract three features according to the subject bounding box, the object bounding box, and the minimum bounding box containing the subject and object at the same time. After splicing, the third visual feature vector containing the visual features of the subject and object is obtained. The above three visual feature vectors are spliced and input into the final fully connected layer to obtain the final visual relationship feature. Figure 5 It is a schematic diagram for encoding the relationship between text and image.
[0061] By calculating the cosine similarity between textual relationship features and visual relationship features, the semantic consistency of a relationship in the generated image can be obtained. Combining the semantic consistency of each object and the semantic consistency of each relationship, the final semantic consistency evaluation result can be obtained.
[0062] The comprehensive semantic consistency of each object and the semantic consistency of each relationship may be a weighted sum of the two types of semantic consistency scores. Considering the different actual ranges of the two types of scores, in order to balance the impact of the two types of scores on the comprehensive evaluation results, the weight of the semantic consistency score of the object may be set to 1, and the weight of the semantic consistency score of the relationship may be set to 0.1.
[0063] The following experimental results show that compared with the existing text-to-image evaluation method, the multi-object text-to-image semantic evaluation method based on scene graph of the present invention can obtain more accurate and objective evaluation results. On the one hand, it is more consistent with the general tendency of human subjects, and on the other hand, it is more sensitive to erroneous results of inconsistent text-to-image semantics.
[0064] This embodiment is based on the MS-COCO image and text dataset for experiments. The dataset is proposed by the document "Microsoft COCO: Common Objects in Context" (authors Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick and Piotr Dollár, published in 2015), which includes 80 common object categories, and each image has 5 corresponding text descriptions. The present invention is compared with the following three existing text generation image evaluation methods as experiments:
[0065] Existing method 1: The method in the paper "Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis" (authors Minfeng Zhu, Pingbo Pan, Wei Chen and Yi Yang, published in the 2019 IEEE Conference on Computer Vision and Pattern Recognition), which provides a text description for the generated image through an image description generation model, calculates the similarity between the generated text description and the input text description, and obtains the semantic evaluation result of the text-generated image.
[0066] Existing method 2: The method in the document "Attngan: Fine-grained text to image generation with attentional generative adversarial networks" (authors Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang and Xiaodong He, published in the 2018 IEEE Conference on Computer Vision and Pattern Recognition). This method uses text to search in multiple images, calculates the recall rate of the retrieved generated images, and obtains the semantic evaluation results of the text-generated images.
[0067] Existing method three: The method in the document "Nuwa: Visual synthesis pre-training for neural visual world creation" (authors Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang and Nan Duan, published in the 2022 European Conference on Computer Vision), which encodes images and texts through the pre-trained CLIP model and calculates the similarity to obtain the semantic evaluation results of the text-generated image.
[0068] The present invention: the method of this embodiment.
[0069] The experiment uses the accuracy metric to measure the accuracy of different evaluation methods. Specifically, 10 images are generated for each of the 200 texts, and 20 human subjects are invited to score the 10 images to measure whether the generated images are consistent with the semantics of the text, and select the images with the highest and lowest scores corresponding to each text. At the same time, these two images are evaluated using an objective evaluation method, and the ratio of the number of better generated images correctly distinguished by the evaluation method to the number of all generated images is taken as the accuracy. The larger the accuracy value, the more reliable the evaluation method.
[0070] In addition, in order to verify that the present invention is more sensitive to generating wrong images, image-text pairs containing object existence and attributes, and object relationship information were selected from the MS-COCO dataset, including 2,300 pairs containing quantifiers (reflecting object existence), 1,901 pairs containing adjectives (reflecting object attributes), 720 pairs containing prepositions (reflecting object spatial relationships), and 1,144 pairs containing verbs (reflecting object subject-object relationships). By replacing the above key information in the text, incorrect image-text pairs are constructed, and statistics are performed to determine whether the text-generated image evaluation method can accurately distinguish between correct and incorrect images.
[0071] Table 1. Experimental results of human evaluation consistency with existing text generation image evaluation methods
[0072] method Accuracy Existing method 1 0.535 Existing method 2 0.124 Existing method three 0.700 The present invention 0.710
[0073] Table 2. Experimental results on error image sensitivity compared with existing text generation image evaluation methods
[0074]
[0075]
[0076] As can be seen from Table 1, the present invention has achieved text-generated image evaluation results that are more consistent with human evaluation. As can be seen from Table 2, the present invention also has a higher recognition ability for generated images with incorrect key semantic information. Existing text-generated image evaluation methods do not explicitly model these key semantic information, so the ability to distinguish is insufficient; while the present invention explicitly parses and encodes the semantic consistency of objects and relationships between objects, so that the semantic consistency of text-generated images can be more accurately evaluated, and more accurate evaluation results can be obtained.
[0077] Corresponding to the above method, another embodiment of the present invention provides a scene graph-based multi-object text generation image semantic evaluation system, which includes:
[0078] The object semantic consistency calculation module is used to perform multimodal object feature encoding on the text prompt and the generated image, obtain text object features and image object features, calculate the similarity between the text object features and the image object features, and obtain the semantic consistency of the object;
[0079] The module for calculating the semantic consistency of the relationship is used to encode the object relationship in the text prompt and the generated image, obtain the text relationship features and the visual relationship features, calculate the similarity between the text relationship features and the visual relationship features, and obtain the semantic consistency of the relationship;
[0080] The comprehensive module is used to integrate the semantic consistency of objects and the semantic consistency of relations to obtain the final text-generated image semantic consistency evaluation results.
[0081] It should be understood that the methods and systems disclosed in the above embodiments provided by the present invention can be implemented in other ways. For example, the division of the above modules can have other division methods in actual implementation, and multiple modules can be combined or integrated into another system. The specific working process of each of the above modules can refer to the corresponding process in the above method embodiment, and will not be repeated here.
[0082] Each module in the present invention can be implemented in the form of a software functional unit, which can be stored in a computer-readable storage medium, including several instructions for enabling a computer device to perform some or all of the steps of the method of the present invention. For example, one embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention. For example, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, CD, etc.), the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the various steps of the method of the present invention are implemented.
[0083] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for semantic evaluation of multi-object text generation images based on scene graphs, characterized in that: The following steps are involved: Perform multimodal object feature encoding on text prompts and generated images to obtain text object features and image object features; Calculate the similarity between text object features and image object features to obtain the semantic consistency of the objects; Encode the object relations between text prompts and generated images to obtain textual relation features and visual relation features; Calculate the similarity between textual relationship features and visual relationship features to obtain the semantic consistency of the relationship; The semantic consistency of objects and the semantic consistency of relations are integrated to obtain the final semantic consistency evaluation result of text-generated image.
2. The method according to claim 1, characterized in that The multimodal object feature encoding of the text prompt and the generated image includes: Parse the text input by the user to obtain each object in the text and the relationship between objects; Encode the objects in the text to obtain the text object features; Divide the generated image into candidate regions, each candidate region contains an image object or does not contain an image object; Encode the candidate region to obtain image object features.
3. The method according to claim 2, characterized in that The parsing of the text input by the user is to find the dependency relationship between the elements in the text through rule-based text parsing, so as to obtain each relationship in the text and its subject and object; the encoding of the objects in the text is to fine-tune the pre-trained CLIP model on the target detection and positioning dataset, and use the fine-tuned CLIP model as a text encoder to encode the text of the subject and object to obtain text object features.
4. The method according to claim 2, characterized in that: The candidate region division of the generated image is to generate multiple candidate bounding boxes through a candidate region generation model trained on the target detection dataset; the encoding of the candidate regions is to use the fine-tuned CLIP model as an image encoder to encode the generated image, and to crop the feature map according to each bounding box through the RoIAlign method to obtain the image object features of each candidate region.
5. The method according to claim 1, characterized in that The calculating of the similarity between the text object features and the image object features includes matching the text and the image nodes. The matching of the text and the image nodes includes two optimization objectives, which are respectively aimed at minimizing the semantic distance between the text object and the image object, and minimizing the overlapping area between the selected image areas. The corresponding image object is found for each text object from the multiple candidate areas in the generated image through the two optimization objectives.
6. The method according to claim 1, characterized in that The encoding of the object relationship between the text prompt and the generated image includes: Encode the relationship between two text objects and their text forms to obtain text relationship features; Encode the two candidate regions assigned to the two text objects in the generated image to obtain visual relationship features; The text and image relationship encoding model is trained through the contrastive learning loss function to align the text relationship features and visual relationship features.
7. The method according to claim 6, characterized in that The visual relationship feature is obtained by using a three-stage parallel visual relationship encoding network, including the following steps: Text stage: The subject and object of the relationship to be encoded are input into the encoder in text form to obtain the first visual feature vector; Position stage: The bounding boxes of the subject and object in the image are input into the encoding network composed of multi-layer perceptrons in the form of coordinates to obtain the second visual feature vector, which contains the absolute and relative position relationship information of the subject and object; Visual semantic stage: Extract the feature map of the image, extract three features on the feature map based on the subject bounding box, object bounding box, and the minimum bounding box that contains both the subject and the object, and then obtain the third visual feature vector containing the visual features of the subject and the object after splicing; The first visual feature vector, the second visual feature vector and the third visual feature vector are concatenated and input into the fully connected layer to obtain the final visual relationship feature.
8. A scene graph-based multi-object text generation image semantic evaluation system, characterized in that: include: The object semantic consistency calculation module is used to perform multimodal object feature encoding on the text prompt and the generated image, obtain text object features and image object features, calculate the similarity between the text object features and the image object features, and obtain the semantic consistency of the object; The module for calculating the semantic consistency of the relationship is used to encode the object relationship in the text prompt and the generated image, obtain the text relationship features and the visual relationship features, calculate the similarity between the text relationship features and the visual relationship features, and obtain the semantic consistency of the relationship; The comprehensive module is used to integrate the semantic consistency of objects and the semantic consistency of relations to obtain the final text-generated image semantic consistency evaluation results.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.