Automated assessment of spatial relationships in images

The spatial relationship of objects in images is evaluated through VISOR indicators, and the problem of inaccurate evaluation of spatial relationship between images and text in the prior art is solved, and more efficient image generation and search results are achieved.

CN120303709APending Publication Date: 2025-07-11MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083003.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-17
Filing Date
2023-11-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When evaluating the spatial relationship between objects in images, the prior art cannot effectively reflect the corresponding spatial relationship expressed by text, resulting in inaccurate automation evaluation.

Method used

Using VISOR indicators, by detecting the object position and spatial relationship in the image, we calculate whether the spatial relationship between objects matches the corresponding relationship of text expressions, and output the corresponding evaluation value.

Benefits of technology

It improves the accuracy of the evaluation of the spatial relationship between images and text, and can effectively filter out images that meet text descriptions, save computing resources, and improve image generation and search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303709A_ABST
    Figure CN120303709A_ABST
Patent Text Reader

Abstract

This document relates to automated analysis of images. One example method involves acquiring an image and text associated with the image, detecting two or more objects in the image, and determining respective positions of the detected two or more objects in the image. The example method also involves determining whether a spatial relationship between the detected two or more objects matches a corresponding spatial relationship expressed by the text based at least on the detected respective positions of the two or more objects. The example method also involves outputting a value reflecting whether the detected spatial relationship between the two or more objects matches a corresponding spatial relationship expressed by the text.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In many computing scenarios, automated metrics are beneficial for evaluating data. For example, an automated mean opinion score rating can replace a human rating of speech quality. Similarly, an automated score can replace a human rating to evaluate image quality. However, the prior art for automated evaluation of image quality has certain deficiencies in the spatial relationships between objects in a given image, which will be discussed in detail below. Summary of the Invention

[0002] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0003] This specification generally relates to techniques for automated evaluation of spatial relationships in images. One example includes a method or technique that can be executed on a computing device. The method or technique can include obtaining an image and text associated with the image and detecting two or more objects in the image. The method or technique can also include determining the respective positions of the two or more detected objects in the image. The method or technique can also include determining whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text, at least based on the respective positions of the two or more detected objects. The method or technique can also include outputting a value reflecting whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text.

[0004] Another example includes a system having a hardware processing unit and a storage resource storing computer-readable instructions. When executed by the hardware processing unit, the computer-readable instructions can cause the system to determine the respective positions of two or more objects in an image. The computer-readable instructions can also cause the system to: determine whether the spatial relationship between the two or more objects matches the corresponding spatial relationship expressed by the text associated with the image, at least based on the respective positions of the two or more objects. The computer-readable instructions can also cause the system to output a reflection of whether the spatial relationship between the two or more objects matches the corresponding relationship expressed by the text associated with the image.

[0005] Another example includes a computer-readable storage medium. The computer-readable storage medium can store instructions that, when executed by a computing device, cause the computing device to perform actions. The actions can include obtaining an image and text associated with the image, detecting two or more objects in the image, and determining the respective positions of the two or more detected objects in the image. The actions can also include determining whether a spatial relationship between the two or more detected objects matches a corresponding spatial relationship expressed by the text, at least based on the respective positions of the two or more detected objects. The actions can further include outputting a value reflecting whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text.

[0006] The examples listed above are intended to provide a quick reference to assist the reader and are not intended to limit the scope of the concepts described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The detailed description is described with reference to the drawings. In the drawings, the leftmost digit of a reference number identifies the drawing in which the reference number first appears. The use of like reference numbers in different instances in the specification and drawings may indicate like or identical items.

[0008] Figure 1 An example system consistent with some embodiments of the present concept is shown.

[0009] Figure 2 An example method or technique consistent with some embodiments of the present concept is shown.

[0010] Figure 3 An example image, text string, and corresponding metrics that can be computed for the image, consistent with some embodiments of the present concept, are shown.

[0011] Figure 4 An example text string that can be used to generate a corresponding image, consistent with some embodiments of the present concept, is shown.

[0012] Figure 5 An example of how an object detector can determine a bounding box and centroid from an object in an image, consistent with some embodiments of the present concept, is shown.

[0013] Figure 6 An example of how the centroid of a bounding box can be converted to a predicate expressing a spatial relationship, consistent with some embodiments of the present concept, is shown.

[0014] Figures 7 to 11 An example of experimental results consistent with some embodiments of the present concept is shown. DETAILED DESCRIPTION OVERVIEW

[0015] There are many computational scenarios where images are associated with text. For example, a user can submit a text query to a search engine to search for images, a user can add captions to images they post on social media, or a user can input text into a text-to-image synthesis model that synthesizes images based on the text. There are various metrics for evaluating the quality of a given image independent of the text, and there are other metrics for evaluating the degree of match between a given image and the associated text.

[0016] However, in some cases, the text can convey spatial relationships between objects, and existing metrics cannot effectively characterize the extent to which the corresponding image reflects the spatial relationships conveyed by the text. The disclosed embodiments provide several metrics, collectively referred to herein as VISOR metrics, that can be used to evaluate whether the spatial relationships between objects in an image match the corresponding spatial relationships expressed by the text. Thus, the disclosed metrics can be used in a wide range of applications, such as ranking text-to-image synthesis models, filtering search results, or evaluating image captions. Definitions

[0017] As used herein, the term "text-to-image synthesis model" refers to a model that receives text as input and generates an image as output. For example, a text-to-image synthesis model can receive a phrase or sentence identifying one or more object categories (e.g., dogs and cats) and output an image that includes instances of the one or more object categories (e.g., an image of a dog standing next to a cat). The term "image-to-text model" refers to a model that receives an image as input and generates text as output. For example, an image-to-text model can receive an image of a German shepherd standing next to a robin and output the phrase "a dog standing next to a bird".

[0018] The term "model" is generally used herein to refer to a series of processing techniques and includes models trained using machine learning as well as manually coded (e.g., heuristic-based) models. For example, a machine learning model can be a neural network, a support vector machine, a decision tree, etc. The term "image" as used herein refers to both still images (e.g., pictures) and videos. The term "text" as used herein refers to expressions of natural language, such as letters, special characters, and / or combinations thereof - words, phrases, complete sentences, paragraphs, etc. Overview of Machine Learning

[0019] There are various types of machine learning frameworks that can be trained to perform a given task. Support vector machines, decision trees, and neural networks are just a few examples of machine learning frameworks that have been widely used in a variety of applications (such as image processing and natural language processing). Some machine learning frameworks (such as neural networks) use layers of nodes that perform specific operations.

[0020] In a neural network, nodes are connected to each other via one or more edges. A neural network can include an input layer, an output layer, and one or more intermediate layers. Individual nodes can process their respective inputs according to a predefined function and provide an output to a subsequent layer, or, in some cases, to a previous layer. The inputs to a given node can be multiplied by corresponding weight values for the edges between that input and the node. Additionally, nodes can have independent bias values, which are also used to produce the output. Various training processes can be used to learn the edge weights and / or bias values. When the term "parameter" is used without a modifier, it refers in this document to learnable values, such as edge weights and bias values that can be learned by training a machine learning model, such as a neural network.

[0021] Neural network architectures can have different layers that perform different specific functions. For example, one or more layers of nodes can together perform a specific operation, such as pooling, encoding, or convolutional operations. In this document, the term "layer" refers to a group of nodes that share inputs and outputs, such as sharing inputs and outputs with an external source or other layers in the network. The term "operation" refers to a function that can be performed by one or more layers of nodes. The term "model architecture" refers to the overall architecture of a layered model, including the number of layers, the connectivity of the layers, and the type of operations performed by individual layers. The term "neural network architecture" refers to the model architecture of a neural network. The terms "trained model" and / or "tuned model" refer to a model architecture and the values that have been trained or tuned for the model architecture. Note that two trained models may share the same model architecture but have different parameter values, for example, if the two models are trained with different training data or there is an underlying random process during training. Example System

[0022] This implementation can be executed in various scenarios and on various devices. As discussed in more detail below, Figure 1 FIG. 100 shows an example system in which this embodiment can be employed.

[0023] As Figure 1 shown, system 100 includes client device 110, server 120, server 130, and server 140 connected via one or more networks 150. Note that the client device can be embodied as a mobile device such as a smart phone or tablet, or as a fixed device such as a desktop computer, server device, etc. Similarly, the server can be implemented using various types of computing devices. In some cases, especially in addition to the servers, Figure 1 any of the devices shown can be implemented in a data center, server farm, etc.

[0024] Figure 1Certain components of the devices shown herein may be referred to herein by reference numerals in parentheses. In the following description, the parentheses (1) indicate the occurrence of a given component on the client device 110, (2) indicate the occurrence of a given component on the server 120, (3) indicate the occurrence on the server 130, and (4) indicate the occurrence on the server 140. Unless a particular instance of a given component is identified, this document will generally refer to the components without parentheses.

[0025] Generally, the devices 110, 120, 130, and / or 140 may have corresponding processing resources 101 and storage resources 102, which will be discussed in more detail below. The devices may also have various modules that utilize the processing and storage resources to perform the techniques discussed herein. The storage resources may include both persistent storage resources (such as magnetic or solid-state drives) and volatile storage (such as one or more random access memory devices). In some cases, the modules may be provided as executable instructions that are stored on a persistent storage device, loaded into a random access memory device, and read by the processing resources from the random access memory to execute.

[0026] The server 120 may include a text-to-image synthesis model 121 that receives text and automatically generates an image. The image may be uploaded to the server 140 for processing by the image evaluation module 141. The server 130 may include an image library 131 having searchable images. Images from the image library may also be uploaded to the server 140 for processing by the image evaluation module 141. The image library may provide a search function, for example, as part of a general web search engine and / or as part of an image hosting service, where each user can search for their own photos in the library.

[0027] The image evaluation module 141 may evaluate the images received from the server 130 and / or 140 to produce values that characterize the relationships between the objects in the images. For example, these values may characterize whether these relationships match the corresponding relationships expressed in text, such as the text used by the text-to-image synthesis model to generate a particular image, a query submitted by a user of the client application 111, and / or a caption generated manually or automatically for an image in the image library on the server 130.

[0028] To evaluate a given image, an image evaluation module 141 can input the image into an object detection module 142. The object detection module can automatically detect objects in the image and output the category and bounding box for each detected object. A relationship evaluation module 143 can determine whether the spatial relationships between the detected objects match the text, which is associated with the image, such as the text used to generate the image, the query used to search for the image, and / or the caption associated with the image. The image evaluation module can calculate a value of a metric for the image, where the value indicates whether the detected objects match the corresponding object categories from the text and the corresponding spatial relationships expressed by the text. Example method

[0029] Figure 2 An example method 200 consistent with the present concept is shown. The method 200 can be implemented on many different types of devices, for example, by one or more cloud servers, by a client device such as a laptop, a tablet, or a smart phone, or by a combination of one or more servers, client devices, etc.

[0030] The method 200 begins at block 202, where an image and text associated with the image are obtained. The text can be a query for searching the image, an automatically or manually generated caption for the image, and / or text input into a text-to-image synthesis model to generate the image.

[0031] The method 200 continues at block 204, where two or more objects in the image are detected. For example, an object detector can automatically identify the corresponding categories of the detected objects.

[0032] The method 200 continues at block 206, where the corresponding positions of the two or more detected objects are determined. For example, an object detector can identify the bounding boxes for the detected objects.

[0033] The method 200 continues at block 208, where it is determined whether the spatial relationships between the two or more detected objects match the corresponding spatial relationships expressed by the text. For example, the centroid of the bounding boxes can be used to determine whether one object is above, below, to the left, or to the right of another object.

[0034] The method 200 continues at block 210, where a value is output that reflects whether the spatial relationships between the two or more detected objects match the corresponding spatial relationships expressed by the text. For example, the value can be calculated as described below for the VISOR metric.

[0035] Note that method 200 can be executed multiple times for one or more text-to-image synthesis models. Multiple instances of text can be input into a given text-to-image synthesis model to generate multiple images. The same or different text instances can be input into one or more other text-to-image synthesis models to evaluate, rank, or otherwise compare different text-to-image synthesis models. Specific algorithm

[0036] The following describes a metric method collectively referred to as "VISOR" for quantifying spatial reasoning performance. To evaluate the performance of VISOR, a dataset called SR2D is created. SR2D contains sentences describing the spatial relationship (left / right / above / below) between a pair of objects. Spatial relation dataset

[0037] Let be a set of object categories. Let be a set of spatial relations between objects. Consider the following two-dimensional relations, namely, and 80 object categories derived from the MS-COCO dataset. Lin et al., (2014). Microsoft coco: Common objects in context. In *Computer Vision - ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6 - 12, 2014, Proceedings, Part V*, Volume 13 (pp. 740 - 755), Springer International Publishing. Then, for each and let the predicate R(A,B) indicate that there is a spatial relation R between object A and object B. For example, (cat, dog) describes a scenario where the cat is to the left of the dog. For each pair, 8 types of spatial relations can be constructed: left(A, B), right(A, B), above(A, B), below(A, B) left(B, A), right(B, A), above(B, A), below(B, A).

[0038] Each predicate R(A, B) can be converted into a template <r> , and rewrite it in natural language. Add appropriate articles "a" or "an" before object names A and B respectively to obtain four templates: a < / r> In one on the left side a In one on the right side a In one above a In one below

[0039] This template-based process has been able to mitigate linguistic ambiguities, subjectivity, and grammatical errors. Additionally, the template-based approach can be extended to new object categories and additional spatial relations. Although the following discussion focuses on two-dimensional relations, the templates can be extended to generate test inputs for evaluating more complex spatial relations and geometric features of objects.

[0040] Given object categories from MS-COCO and two objects per image, there are 3160 unique object pair (A, B) combinations from the binomial coefficient. For each pair of objects, the 8 types of spatial relations listed above can be constructed, resulting in a total of 3160 × 8 = 25280 predicates. The SR2D dataset contains 25280 text examples that are evenly distributed across 80 COCO object categories, with each object appearing in 632 images. Table 1 below shows a few examples: VISOR metric calculation

[0041] Let h be an oracle function that returns a set of detected objects from the set in image x. Then, the object accuracy (OA) for image x generated by a sentence containing objects A and B is: OA(x, A, B) = 1 h(x) (A ∩ B). (1)

[0042] Note that the oracle function h here can be an automated model or a human who detects the presence of objects mentioned in the sentence. The object accuracy is independent of the relation R, whose presence is captured by the VISOR metric.

[0043] When R gen is the true relation mentioned in the text, let R be the generated spatial relation. Then, for each image x,

[0044] For artists and designers, a useful feature of text-to-image synthesis models is the ability to generate multiple images for each input text prompt. This allows the creator to pick an appropriate image from the N generated images. The metric VISOR is defined in this paper n , to reflect the ability of the text-to-image synthesis model to generate at least n spatially correct images when given a text input that mentions spatial relationships. From the perspective of usability, this more lenient version of VISOR is useful for measuring whether at least n images that satisfy the input sentence (e.g., a threshold number of images) can be found, in which task the creator can select images from the set of output images.

[0045] VISOR n is the probability of generating an image such that for each text prompt t, VISOR = 1 for at least n images out of N images:

[0046] The relationship between VISOR and VISOR is given below n is as follows.

[0047] The following discussion uses N = 4 images for each text prompt. Therefore, the following metrics will be discussed: VISOR1, VISOR2, VISOR3, and VISOR4. Note that VISOR = 1 only when both objects are generated in the image, i.e., OA = 1. However, the text-to-image synthesis model may not be able to generate multiple objects in a large subset of images. Therefore, it is useful to distinguish between two capabilities of the model: (1) generating multiple objects, and (2) generating them according to the spatial relationships described in the prompt text. To this end, a conditional VISOR metric is defined, which is the conditional probability that the correct spatial relationship is generated given that two objects are correctly generated:

[0048] The summary of the VISOR calculation process is in Figure 3 In Figure 3 shows an example of the intuition behind OA, VISOR, VISOR cond and VISOR 1 / 2 / 3 / 4 Four text strings 302 are shown in the figure, as well as four images 304, 306, 308, and 310, which are, for example, generated by a text-to-image synthesis model based on those text strings. The legend 312 shows that alternating long dashed and short dashed boundaries are used to express the absence of one or two objects in the generated image, solid boundaries are used to express images in which two objects are generated but the spatial relationship is incorrect, and short dashed boundaries are used to express successfully generated images in which the objects and their spatial relationships are correctly generated. Each generated image is surrounded by a corresponding boundary to convey whether that particular image has the correct objects and / or correct spatial relationships expressed by the corresponding text string.

[0049] The VISOR percentage score 314 is shown for each group of four images. The first row of images includes two images with correct objects and spatial relationships, resulting in a correct percentage of 50% for the VISOR score. The second and third rows of images do not have any images with correct objects and spatial relationships, so the correct percentage of the VISOR score is 0%. The fourth row of images includes three images with correct objects and spatial relationships, so the correct percentage of the VISOR score is 75%.

[0050] In addition, Figure 3 shows the following additional VISOR metrics. VISOR n Percentage 316 is shown for each row of four images. For example, since the first row includes two correct images, the percentage values of VISOR1, VISOR2, VISOR3, and VISOR4 are 100, 100, 0, and 0, respectively. The total score 320 can be calculated for all 16 images. The total score includes VISOR cond 322, average VISOR 324, and average VISOR n 326. Here, VISOR cond has a value of 5 / 11 because out of a total of 11 images with the correct object classes, 5 images have the correct spatial relationship (short dashed boundary), and 6 images have an incorrect spatial relationship (solid boundary). The average VISOR 324 is calculated for all 16 images, and 5 of them have the correct spatial relationship. The average VISOR n 326 is also calculated for all 16 images. For example, the average percentage value of VISOR1 is 50% because VISOR1 is 100% in the first and last rows of images and 0% in the second and third rows of images.

[0051] To compute the previously described VISOR metric, the following process can be performed. Given any text prompt t and a text-to-image synthesis model g, first generate an image x = g(t), and use an object detector to localize the objects in x. The object accuracy OA can be computed as described above. The centroid coordinates of objects A and B can be obtained from the bounding boxes of the detected objects.

[0052] For example, Figure 4 An example is shown of using the text string 402 to generate the image 404. Figure 5 An example is shown in which the image 404 is processed using an object detector to obtain the bounding boxes and centroids of each detected object. The bounding box 502 around the elephant has a centroid 504, the bounding box 506 around the motorcycle has a centroid 508, and the bounding box 510 around the tree has a centroid 512.

[0053] Figure 6 An example is shown of how to localize object centroids and convert them into predicates expressing the spatial relationships between them. These predicates are compared with the ground truth relationship R to obtain the VISOR score. The centroid 508 of the motorcycle (object A) is shown to have positions xa, ya, and the centroid 504 of the elephant (object B) is shown to have positions xb, yb.

[0054] Based on the centroids, the spatial relationship R between them can be derived using the rules shown in the predicate transformer 406 (e.g., components of the relation evaluation module 143) gen . Finally, the generated relation can be compared with the ground truth 408 expressing the relationship R between the objects. The VISOR score can be computed as described above. Experimental results

[0055] The following experiment used OWL-ViT (Minderer et al., 2022). "Simple open-vocabulary object detection with vision transformers". arXiv preprint arXiv:2205.06230. This experiment employed an open-vocabulary object detector based on the CLIP backbone and the ViT-B / 32 transformer architecture, and set a confidence threshold of 0.1. The open-vocabulary capabilities of OWL-ViT ensure that VISOR is widely applicable to other datasets, categories, and vocabularies. This eliminates the dependence on specific datasets, enabling VISOR to be widely applicable to any free-form text input.

[0056] The following text-to-image synthesis models were investigated as baselines for the following experiments: GLIDE (Nichol et al., 2021). "Glide: Towards photorealistic image generation and editing with text-guided diffusion models" (arXiv preprint arXiv:2112.10741); DALLE-mini (Dayma et al., 2021). "Dall·e mini"; CogView2 (Ding et al., 2022). "Cogview2: Faster and better text-to-image generation via hierarchical transformers" (arXiv preprint arXiv:2204.14217); DALLE-v2 (Ramesh et al., 2022). "Hierarchical text-conditional image generation with clip latents" (arXiv preprint arXiv:2204.06125); and Stable-Diffusion (SD) (Rombach et al., 2022). "High-resolution image synthesis with latent diffusion models" (in the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 10684-10695)), and two versions of the Composable Diffusion Model (GLIDE+CDM and SD+CDM). For each text prompt in the SR2D dataset, each model generated N = 4 images to obtain 126,720 images for each model and compare performance in terms of OA, VISOR, VISOR cond and VISOR 1 / 2 / 3 / 4 aspects.

[0057] Text-to-image synthesis models have mainly been compared in terms of features such as photorealism (purely visual) and human subjective judgments of image quality. The following experiments quantify whether existing automated multimodal metrics are applicable to evaluating the spatial relationships generated by text-to-image synthesis models. The metrics considered include: CLIPScore (Hessel et al., 2021). "Clipscore: A reference-free evaluation metric for image captioning" (arXiv preprint arXiv:2104.08718) (cosine similarity between image and text embeddings), and evaluation based on image captioning (BLEU (Papineni et al., July 2002). "Bleu: A method for automatic evaluation of machine translation" (in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318)), METEOR (Banerjee et al., June 2005). "METEOR: An automatic metric for MT evaluation with improved correlation with human judgments" (in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and / or Summarization (pp. 65-72)), ROUGE (Rouge, L.C., July 2004). "A package for automatic evaluation of summaries" (in Proceedings of the ACL Text Summarization Workshop, Spain), CIDER (Vedantam et al., 2015). "Cider: Consensus-based image description evaluation" (in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4566-4575)), SPICE (Anderson et al., 2016). "Spice: Semantic propositional image caption evaluation" (in Computer Vision - ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V, Volume 14 (pp. 382-398). Springer International Publishing).

[0058] These metrics are used by generating a caption c for the synthetic image x = g(t) and computing a caption score relative to the reference input text t. Note that purely visual metrics (such as FID and Inception Score) ignore the text. Heusel et al., 2017. "Gans trained by a two time-scale update rule converge to a local nash equilibrium" (Advances in Neural Information Processing Systems, 30), Salimans et al., 2016. "Improved techniques for training gans" (Advances in Neural Information Processing Systems, 29). Semantic object accuracy ignores all words except nouns, making it unable to score spatial relationships. Hinz et al., 2020. "Semantic object accuracy for generative text-to-image synthesis" (IEEE Transactions on Pattern Analysis and Machine Intelligence 44(3), pp. 1552-1565).

[0059] Let s t be the score for (x, t), where x is the generated image and t is the input text. Let t flip be the transformed version of t obtained by inverting / flipping the spatial relationships in t (e.g., left → right). Let be the score for (x, t flip ). Then, for each metric, define Δ s as:

[0060] That is, it is the average difference between s t and over the entire SR2D dataset.

[0061] Thus, Δ s captures the metric s' ability to understand spatial relationships. Table 2: shows the values of s and Δ s for each previous metric and each model, while Table 3: Shows the same scores for the VISOR metric. It can be seen that for all the previous metrics, Δ s is almost negligible and close to zero, which means that even if the text is flipped, they return similar scores. In some cases, the difference is negative, indicating that the images have higher scores with the flipped captions. On the other hand, the Δ values for VISOR are higher, meaning that VISOR assigns significantly lower scores to the flipped samples. These results establish the utility of the VISOR evaluation metric as no existing metric can reliably quantify spatial relationships and demonstrate the effectiveness of VISOR for this purpose.

[0062] Table 4: Shows the results of the benchmarks on the SR2D dataset. The first thing to note is that the object accuracies of all models except DALLE-v2 are below 30%. Although DALLE-v2 (63.93%) is significantly better than other models, it still shows a large number of failures in generating the two objects mentioned in the generation prompt. For the unconditional metrics VISOR and VISOR1 / 2 / 3 / 4, DALLE-v2 is the best-performing model. However, in terms of VISOR cond CogView2 performs the best. This means that although CogView2 is better than other models in those examples of generating two objects, the large number of failures of CogView2 in terms of object accuracy (OA) leads to a lower unconditional VISOR score. For all models, including DALLE-v2 (8.54%), VISOR4 is extremely low, revealing a huge gap in performance. Human study

[0063] A human evaluation study was conducted to understand the consistency between the VISOR metric and human judgment and to quantify the gap between object detector performance and human assessment of object presence. In the human study, the following four models were used: CogView2, DALLE-v2, Stable Diffusion (SD), and SD+CDM. Annotators were shown images generated by one of the four models via Amazon Mechanical Turk and were asked to answer seven questions. These questions involved the assessment of image quality, the assessment of scene authenticity (scene likelihood) (using a Likert scale from 1 to 5), the assessment of the number of objects, judging whether an object exists (true / false), selecting valid spatial relationships, and judging whether two objects are merged in the image. A sample size of 1000 images was used for each model, and each sample was evaluated by 3 annotators.

[0064] Figure 7 Shows the response results 700 for each question in the human study. Although DALLE-v2 received the highest image quality rating, SD and SD+CDM received higher scene likelihood ratings. DALLE-v2 also had the largest number of images with object merges (32.46%). The agreement among annotators for all questions was high, whether it was majority agreement (at least 2 out of 3 annotators agreed) or unanimity (all 3 annotators agreed), as shown in Table 5:

[0065] Agreement between VISOR and human responses. Note that in terms of object accuracy (OA) and VISOR, the ranking of the models in the human study and the automated VISOR scores in Table 4 is the same, i.e., DALLE-v2 > SD > SD-CDM > CogView2. Table 6: Shows the percentage of samples where human responses match the automated evaluation using an object detector.

[0066] The experiment also revealed some common types of observed object merges. The common patterns observed included animals being rendered as patterns on inanimate objects, and two objects maintaining their typical shapes but being merged. A significant proportion (over 20%) of the images had object merges - this poses a major challenge for generating different objects and their relationships using text-to-image synthesis models.

[0067] Performance divided by relationship is shown in Table 7:

[0068] Note that five out of the seven models had the best VISOR cond scores for horizontal relationships (left or right). However, five out of the seven models had the best object accuracy for vertical relationships (above or below).

[0069] Performance divided by supercategory. The 80 object categories in SR2D belong to 11 MS-COCO "supercategories". For the best-performing model (DALLE-v2), the VISOR scores for each pair of supercategories were reported. The VISOR scores were highest for co-occurring supercategories such as "animal, outdoor", while the performance was lower for less likely indoor-outdoor object combinations (such as "vehicle, appliance" and "electronics, outdoor").

[0070] Correlation between VISOR and object co-occurrence. The object categories in the dataset involve common objects from MS-COCO, such as wild animals, vehicles, appliances, and humans, which appear in different contexts, including combinations that do not often appear together in real life. For example, an elephant is unlikely to appear indoors near a microwave oven. To understand how object co-occurrence affects VISOR, we first determine P COCO (A, B), that is, the co-occurrence probability of each object pair (A, B), as a proxy for object co-occurrence in the real world. Then, for the pair (A, B) and its P COCO The correlation of the object accuracy of (A, B) is plotted as Figure 8 The object accuracy correlation graph 800 is shown. cond With P COCO The correlation of (A, B) is plotted as VISOR correlation graph 850. cond For both models, the correlations are positive, meaning that the quality of the output is likely to be better for commonly co-occurring objects, clearly demonstrating a bias towards real-world likelihoods. This correlation suggests the difficulty of generating unlikely relationships (such as "an elephant is to the left of a microwave"), although creators may wish to pursue such unlikely combinations for artistic compositions.

[0071] Object generation bias. The object accuracy of generated images is compared for three types of input: (1) single object text, such as "an elephant"; (2) combinations of multiple objects, such as "an elephant and a cat"; and (3) relation text, such as "an elephant is to the right of a cat". Figure 9 The single object vs. multiple object correlation graph 900 shown indicates that, for all models, the OA for a single object is significantly higher than the OA for multiple objects. Therefore, generating a composition using a combination is challenging.

[0072] Text order bias. Figure 10 A first object and a second object correlation diagram 1000 is included, Figure 10 It is shown that for all models, the OA for an object mentioned for the first time in the text (A) is significantly higher than the OA for an object mentioned for the second time in the text (B); generating two objects simultaneously is the most challenging.

[0073] Consistency between equivalent phrases. Ideally, given two equivalent inputs, such as "a cat is above a dog" and "a dog is below a cat", the model should generate images with the same spatial relationship. To evaluate this consistency, consider the case of OA=1 (both objects are detected) and show the consistency for each relationship type in Table 8:

[0074] Note that the best-performing model, DALLE-v2, is the most inconsistent among all models, while CogView2 is the most consistent model. This result suggests that simply rephrasing the input can have a significant effect on the spatial correctness of the output.

[0075] The effect of attributes on spatial understanding. Through a case study using Stable Diffusion (SD), to understand the effect of sentence complexity on the performance of the model VISOR. Via the form [ZA][CA] <r>[ZB][CB] The template randomly assigns two attributes (size Z and color C) to the object categories, increasing the complexity of the text prompts. Eleven object categories, eight colors, and four sizes representing each supercategory in COCO are used. As < / r> Figure 11 As shown in the property effect diagram 1100, among the 15 property combination types, the image generation performance of 13 types decreases compared to the image without properties. Adding the color property (C) will cause a significant decrease in performance. Adding the size descriptor (Z) may improve the performance. This analysis indicates that properties may have a negative effect on spatial combinability. Applications

[0076] The different VISOR metrics described above have various applications. One useful application involves evaluating the text-to-image synthesis model described above. Given a single text-to-image synthesis model, the VISOR metrics can generally convey how well the text-to-image synthesis model performs in generating images that can accurately reflect the spatial relationships between the objects expressed in the text used to generate the images.

[0077] Consider a scenario where the images generated by a given text-to-image synthesis model have relatively low VISOR scores. Developers who hope to improve the text-to-image model can use VISOR cond to gain additional insights into the model's performance. If the text-to-image synthesis model produces relatively low VISOR cond scores while producing high object accuracy scores, this indicates that the development work should focus on revising the text-to-image synthesis model to replicate spatial relationships. On the other hand, if the text-to-image synthesis model produces relatively high VISOR cond scores and relatively low accuracy scores, this indicates that the text-to-image synthesis model can accurately replicate spatial relationships but has difficulty in generating the correct object types. Therefore, the development work should focus on improving the ability of the text-to-image synthesis model to generate the correct object types.

[0078] As another example, consider an overall image quality score based on one or more VISOR metrics and one or more other metrics (such as CLIPScore). Each individual metric can provide different information about the characteristics of the image. Therefore, users who focus on high semantic similarity between the text and the generated image may assign a relatively high weight to CLIPscore and a relatively low weight to the VISOR score when calculating the overall image quality score of a given image. Conversely, users who are more concerned with spatial relationships than semantic similarity may assign a relatively high weight to the VISOR score rather than CLIPscore. This same approach can also be applied to the other image quality metrics described above.

[0079] Separate VISOR scores or overall quality scores with VISOR components can also be employed to rank text-to-image synthesis models for mutual comparison. For example, consider a scenario where multiple text-to-image synthesis models are being considered for deployment in a specific web application. Consider an interior design application where the relative positioning of objects is crucial for generating useful images. For instance, "a sofa next to a window" is an important spatial relationship. In contrast, consider an art application where sometimes unexpected spatial relationships might be considered unexpectedly useful. For example, an artist might find that a sofa in an unusual position in a room has unexpected aesthetic qualities. Thus, by using only the VISOR metric or relatively highly weighting them in a combined metric, a text-to-image synthesis model that can accurately replicate spatial relationships can be ranked higher for interior design applications, while a text-to-image synthesis model that produces unexpected spatial relationships, perhaps with a more realistic or appealing image quality, can be ranked higher for art applications with a relatively lower VISOR metric weight.

[0080] As another example, consider an image search scenario where a user wishes to search for images in a web-based search engine or a local or cloud-based image library (e.g., their own social media images). Recall Figure 3 , assume the user inputs the text string "an orange above a giraffe" as a query and Figure 3 all four images shown for that text string exist in the library. Since the second and third images do not match the spatial relationship expressed in the query, these images can be filtered out so that the user only receives search results that match their query, e.g., the first and fourth images.

[0081] As another example, consider a scenario where a user inputs a caption for an image in their library. If the user inputs the caption "an orange above a giraffe" for the second or third image shown in Figure 3 , an indication that the caption is incorrect can be provided to the user. One way to do this is to automatically suggest an alternative caption, e.g., for the second image shown in Figure 3 , suggest "an orange below a giraffe". One way to generate the suggested caption involves inputting the image into an image-to-text model. Alternatively, a template-based reverse approach can be employed to generate the alternative caption, where the above-described template is used to generate the suggested caption, e.g., by filling the template with the object categories provided by an object detector and the spatial relationship determined using the bounding box centroids determined by the above method.

[0082] In a further embodiment, the caption generated by the image-to-text model can be evaluated by comparing it with the text used to generate the image. If the object categories and spatial relationships generated by the image-to-text model for a given image match the object categories and spatial relationships expressed in the text used to generate the image, this can be considered a positive evaluation of the caption. If not, this can be considered a negative evaluation of the caption. These evaluations can be used to rank or optimize the image-to-text model, for example, using these evaluations as training labels.

[0083] In a further embodiment, the VISOR metric can be employed to clean the corpus of images and matching text. For example, referring back Figure 3 , assume the corpus initially includes all the images shown therein. If this corpus is directly used to train a text-to-image synthesis model, for example, using the text string as input and the image as the training target, then some of the training targets will be images with incorrect spatial relationships. By removing the images with incorrect spatial relationships before training, a clean corpus can be produced, which will prompt the text-to-image synthesis model to learn to generate images with correct spatial relationships.

[0084] A similar approach can be used to generate a clean corpus for training an image-to-text model, where the model receives the images in the corpus as input and the training target is the corresponding text string used to generate the image. By removing the images with incorrect spatial relationships before training, a clean corpus can be produced, which will prompt the image-to-text model to learn to generate text that can correctly describe the spatial relationships in the input image. Technical Effects

[0085] As previously mentioned, traditional metrics for characterizing image quality cannot accurately reflect whether an image correctly reflects the spatial relationships in the text associated with the image. For applications that require accurate characterization of spatial relationships, these metrics are often insufficient. By using the VISOR metric of the present invention, whether used alone or in combination with one or more traditional metrics, the various deficiencies of the previous metrics can be remedied.

[0086] As previously mentioned, using the disclosed VISOR metric enables the manual or automatic selection of a text-to-image synthesis model suitable for a particular application. By selecting a particular text-to-image synthesis model that can accurately generate an image that matches the corresponding spatial relationships expressed in the text, the generation of many irrelevant images can be avoided. The generation of many unrelated images can be avoided. This can save computational resources used to generate the images, such as processor time, memory or storage for storing the images, and bandwidth for transmitting the images.

[0087] For example, as also described above, the image search function can be improved by filtering out images that do not accurately reflect the spatial relationships provided by the query. Thus, the user can receive more relevant images in response to a text query, whether searching their own personal image library or conducting an image search on a web search engine. This can also save computational resources such as memory, storage, or bandwidth that would otherwise be used to store or transmit irrelevant images that would not otherwise be filtered out.

[0088] In addition, as described above, the disclosed VISOR metric can be employed to detect errors in user-generated or automatically-generated captions for a given image. By pointing out incorrect captions and suggesting alternative captions to replace the incorrect ones, the user can obtain more accurate captions from search results and avoid any unexpected errors when adding captions to their own images.

[0089] In addition, as described above, the disclosed VISOR metric can be used to clean the corpus of training data for training text-to-image synthesis models or image-to-text models. First, this reduces the amount of noise in the corpus and may result in a more accurate model. Second, removing noisy examples before training saves computational resources such as memory, storage, or bandwidth that would be used when training on irrelevant images. Device implementation

[0090] As described above with respect to Figure 1 the system 100 includes several devices, including client device 110, server 120, server 130, and server 140. It should also be noted that not all device implementations can be illustrated, and other device implementations should be apparent to those skilled in the art from the above and following descriptions.

[0091] As used herein, the terms "device", "computer", "computing device", "client device", and / or "server device" can refer to any device having some degree of hardware processing capability and / or hardware storage / memory capability. The processing capability can be provided by one or more hardware processors (e.g., hardware processing units / cores) that are capable of executing data in the form of computer-readable instructions to provide functionality. The computer-readable instructions and / or data can be stored on a storage device such as memory / storage and / or a data repository. In this document, the term "system" can refer to a single device, multiple devices, etc.

[0092] The storage resources can be internal or external resources of the respective devices associated therewith. The storage resources can include volatile or non-volatile memories, hard disk drives, flash storage devices, and / or optical storage devices (e.g., CDs, DVDs, etc.). In this document, the term "computer-readable medium" can include signals. In contrast, the term "computer-readable storage medium" does not include signals. Computer-readable storage media include "computer-readable storage devices". Examples of computer-readable storage devices include volatile storage media such as random access memory (RAM), and non-volatile storage media such as hard disk drives, optical discs, and flash memories, etc.

[0093] In some cases, the device is configured with a general-purpose hardware processor and storage resources. In other cases, the device can include a system-on-chip (SoC) type design. In an SoC design implementation, the functions provided by the device can be integrated on a single SoC or multiple coupled SoCs. One or more associated processors can be configured to coordinate with shared resources (e.g., memory, storage, etc.) and / or one or more dedicated resources (e.g., hardware modules configured to perform certain specific functions). Thus, in this document, the terms "processor", "hardware processor", or "hardware processing unit" can also refer to a central processing unit (CPU), a graphics processing unit (GPU), a controller, a microcontroller, a processor core, or other types of processing devices suitable for implementation in traditional computing architectures as well as SoC designs.

[0094] Alternatively or additionally, the functions described herein can be performed at least in part by one or more hardware logic components. By way of example and not limitation, illustrative types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0095] In some configurations, any module / code discussed herein can be implemented in software, hardware, and / or firmware. In any case, these modules / codes can be provided during the device manufacturing process, or by a middleman preparing the device for sale to the end user. In other cases, the end user can install these modules / codes at a later time, e.g., by downloading executable code and installing the executable code on the corresponding device.

[0096] Note also that a device typically has input and / or output functions. For example, a computing device may have various input mechanisms such as a keyboard, a mouse, a touchpad, speech recognition, gesture recognition (e.g., using a depth camera system such as a stereo camera system or a time-of-flight camera system, an infrared camera system, an RGB camera system, or using an accelerometer / gyroscope, face recognition, etc.). A device may also have various output mechanisms such as a printer, a monitor, etc.

[0097] It should also be noted that the devices described herein may operate in an independent or collaborative manner to implement the described techniques. For example, the methods and functions described herein may be executed on a single computing device or distributed across multiple computing devices that communicate with each other via network 150. By way of non-limitation, network 150 may include one or more local area networks (LANs), wide area networks (WANs), the Internet, etc.

[0098] Various examples are described above. Additional examples are described below. One example includes a method that includes: obtaining an image and text associated with the image, detecting two or more objects in the image, determining the respective positions of the two or more detected objects in the image, determining whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text, at least based on the respective positions of the two or more detected objects, and outputting a value reflecting whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text.

[0099] Another example may include any of the above and / or below examples, wherein the value reflects whether the two or more detected objects match the corresponding object categories expressed by the text.

[0100] Another example may include any of the above and / or below examples, wherein the method further includes generating an image from the text by inputting the text into a text-to-image synthesis model.

[0101] Another example may include any of the above and / or below examples, wherein the method further includes generating another value, the other value reflecting whether at least a threshold number of spatial relationships between the detected multiple objects in multiple images match the corresponding spatial relationships expressed by the text, the multiple images being generated by inputting the text into a text-to-image synthesis model.

[0102] Another example may include any of the above and / or below examples, wherein the method further includes generating another value that reflects a conditional probability that a spatial relationship detected among a plurality of objects in a plurality of images generated by inputting a plurality of text instances into a text-to-image synthesis model matches a corresponding spatial relationship expressed by the plurality of text instances, assuming that the detected plurality of objects match corresponding object categories expressed by the plurality of text instances.

[0103] Another example may include any one of the above and / or below examples, wherein the method further includes ranking the text-to-image synthesis model relative to another text-to-image synthesis model based at least on the value and the another value, the another value reflecting whether another spatial relationship detected among two or more objects in another image matches a corresponding spatial relationship expressed by the text, the another image being generated using the text by the another text-to-image synthesis model.

[0104] Another example may include any one of the above and / or below examples, wherein detecting includes inputting the image into an object detector that outputs corresponding categories of two or more detected objects.

[0105] Another example may include any one of the above and / or below examples, wherein the method further includes obtaining bounding boxes around two or more detected objects from the object detector and determining the spatial relationship among the two or more detected objects based on the spatial relationship between the centroids of the bounding boxes.

[0106] Another example includes a system that includes a processor and a storage medium storing instructions that, when executed by the processor, cause the processor to: determine corresponding positions of two or more objects in an image, determine whether a spatial relationship among the two or more objects matches a corresponding spatial relationship expressed by text associated with the image based at least on the corresponding positions of the two or more objects, and output a reflection of whether the spatial relationship among the two or more objects matches the corresponding relationship expressed by the text associated with the image.

[0107] Another example may include any of the above and / or below examples, wherein the instructions, when executed by the processor, cause the processor to: filter searchable images to obtain image search results based at least on whether a spatial relationship among objects in the searchable images matches a spatial relationship expressed in a search query.

[0108] Another example may include any of the above and / or below examples, wherein the searchable images are provided by a web-based search engine or are provided in a local image library or a cloud-based image library associated with a particular user.

[0109] Another example may include any of the above and / or below examples, where the text is a caption for an image.

[0110] Another example may include any of the above and / or below examples, where the caption is generated by a user, and when the instructions are executed by a processor, cause the processor to: output an indication of whether the caption matches the spatial relationship between two or more objects.

[0111] Another example may include any of the above and / or below examples, where when the instructions are executed by a processor, cause the processor to: generate a suggested alternative caption expressing the spatial relationship between two or more objects and output the suggested alternative caption.

[0112] Another example may include any of the above and / or below examples, where the caption is generated by an image-to-text model.

[0113] Another example may include any of the above and / or below examples, where when the instructions are executed by a processor, cause the processor to: perform a comparison of a caption generated by an image-to-text model with another text, the other text being used to generate an image using a text-to-image synthesis model, and improve or evaluate the image-to-text model at least based on the comparison.

[0114] Another example may include any of the above and / or below examples, where when the instructions are executed by a processor, cause the processor to: generate an overall image quality score based on a value and at least one other value reflecting at least one other characteristic of the image.

[0115] Another example may include any of the above and / or below examples, where when the instructions are executed by a processor, cause the processor to: clean a corpus and associated text at least based on a value indicating whether the spatial relationship between objects in images in the corpus reflecting the image matches the spatial relationship expressed by the associated text.

[0116] Another example may include any of the above and / or below examples, where when the instructions are executed by a processor, cause the processor to: train a machine learning model at least based on the cleaned corpus.

[0117] Another example includes a computer-readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform actions including: obtaining an image and text associated with the image, detecting two or more objects in the image, determining corresponding positions of the two or more detected objects in the image, determining whether a spatial relationship between the two or more detected objects matches a corresponding spatial relationship expressed by the text based at least on the corresponding positions of the two or more detected objects, and outputting a value reflecting whether the spatial relationship between the two or more detected objects matches the corresponding spatial relationship expressed by the text. Conclusion

[0118] Although the subject matter is described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the above specific features and acts are disclosed as example forms of implementing the claims, and other features and acts recognizable by those skilled in the art are intended to be included within the scope of the claims.

Claims

1. A method, comprising: Obtaining an image and text associated with the image; Detecting two or more objects in the image; Determining the respective positions of the detected two or more objects in the image; Determining whether a spatial relationship between the detected two or more objects matches a corresponding spatial relationship expressed by the text, at least based on the respective positions of the detected two or more objects; And Outputting a value reflecting whether the spatial relationship between the detected two or more objects matches the corresponding spatial relationship expressed by the text.

2. The method according to claim 1, wherein the value reflects whether the detected two or more objects match corresponding object categories expressed by the text.

3. The method according to claim 1 or 2 further comprises: Generating the image from the text by inputting the text into a text-to-image synthesis model.

4. The method according to claim 3 further comprises: Generating another value, the another value reflecting whether at least a threshold number of spatial relationships between detected multiple objects in multiple images match the corresponding spatial relationships expressed by the text, the multiple images being generated by inputting the text into the text-to-image synthesis model.

5. The method according to claim 3 further comprises: Generating another value, the another value reflecting the conditional probability that spatial relationships between detected multiple objects in multiple images generated by inputting multiple text instances into the text-to-image synthesis model match the corresponding spatial relationships expressed by the multiple text instances, assuming that the detected multiple objects match the corresponding object categories expressed by the multiple text instances.

6. The method according to claim 3 further comprises: Ranking the text-to-image synthesis model relative to another text-to-image synthesis model, at least based on the value and the another value, the another value reflecting whether another spatial relationship between two or more objects detected in another image matches the corresponding spatial relationship expressed by the text, the another image being generated using the text by the another text-to-image synthesis model.

7. The method according to claim 1, wherein the detection includes inputting the image into an object detector, and the object detector outputs the respective categories of the detected two or more objects.

8. The method according to claim 6, further comprising: Obtaining bounding boxes around the detected two or more objects from the object detector; And Determining the spatial relationship between the detected two or more objects based on the spatial relationship between the centroids of the bounding boxes.

9. A system, comprising: A processor; And A storage medium storing instructions that, when executed by the processor, cause the processor to: Determine the respective positions of two or more objects in an image; Determine whether a spatial relationship between the two or more objects matches a corresponding spatial relationship expressed by text associated with the image, at least based on the respective positions of the two or more objects; And Output a value reflecting whether the spatial relationship between the two or more objects matches the corresponding relationship expressed by the text associated with the image.

10. The system according to claim 9, wherein when the instructions are executed by the processor, the processor is caused to: Filter the searchable images to obtain image search results based at least on whether the spatial relationships between objects in the searchable images match the spatial relationships expressed in the search query.

11. The system according to claim 10, wherein the searchable images are provided by a web-based search engine or are provided in a local image library or a cloud-based image library associated with a particular user.

12. The system according to claim 9, wherein the text is a description of the image.

13. The system according to claim 12, wherein the description is user-generated, and when the instructions are executed by the processor, the processor is caused to: Output an indication of whether the description matches the spatial relationship between the two or more objects.

14. The system according to claim 13, wherein when the instructions are executed by the processor, the processor is caused to: Generate a suggested alternative description expressing the spatial relationship between the two or more objects; and Output the suggested alternative description.

15. The system according to claim 12, wherein the description is generated by an image-to-text model.

16. The system according to claim 15, wherein when the instructions are executed by the processor, the processor is caused to: Perform a comparison of the description generated by the image-to-text model with another text that is used to generate the image using a text-to-image synthesis model; and Improve or evaluate the image-to-text model based at least on the comparison.

17. The system according to claim 9, wherein when the instructions are executed by the processor, the processor is caused to: Generate an overall image quality score based on the value and at least one other value reflecting at least one other characteristic of the image.

18. The system according to claim 9, wherein when the instructions are executed by the processor, the processor is caused to: Clean the corpus and the associated text based at least on a value indicating whether the spatial relationships between objects in images in the corpus reflecting the images match the spatial relationships expressed by the associated text.

19. The system according to claim 18, wherein when the instructions are executed by the processor, the processor is caused to: Train a machine learning model based at least on the cleaned corpus.

20. A computer-readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform actions, the actions including: Obtain an image and text associated with the image; Detect two or more objects in the image; Determine the respective positions of the detected two or more objects in the image; Determine whether the spatial relationship between the detected two or more objects matches the corresponding spatial relationship expressed by the text based at least on the respective positions of the detected two or more objects; And Output a value reflecting whether the detected spatial relationship between the two or more objects matches the corresponding spatial relationship expressed by the text.