Text-to-Image Embeddings for Accurate Object Attribute Binding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems using generative neural networks for image generation lack accuracy and flexibility in generating synthetic images from text prompts, particularly when dealing with complex prompts involving multiple objects with different visual attributes, leading to degraded image quality and failure in accurately reflecting the compositionality of the text prompts.
Innovation Solution
A two-stage neural network system that includes an encoder stage to generate a sequence of embeddings with object text embeddings and visual embeddings, replacing the latter with corresponding visual embeddings to accurately generate synthetic images with correct object attributes, utilizing a decoder stage to produce the final image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional text-to-image generative neural networks are used, then image generation capability is provided, but accuracy in reflecting text prompt constraints and object attribute binding deteriorates
Solution Approach 1:
The text prompt is segmented into individual object tokens and attribute tokens separately, rather than processing the entire prompt as a single sequence. This segmentation allows the model to independently encode and track each object and its attributes through the generation process, preventing attribute drift and ensuring accurate binding between objects and their visual characteristics.
Solution Approach 2:
Object embeddings are introduced as intermediary representations that bridge the text prompt and the generated image. These embeddings serve as persistent references throughout the generation process, allowing the model to consistently refer back to specific objects and their attributes, thereby maintaining accurate object-attribute binding and improving reliability in reflecting text prompt constraints.
2Adaptability or versatility
If conventional generative neural networks are used, then synthetic image generation is enabled, but flexibility in handling complex prompts with multiple objects deteriorates
Solution Approach 1:
The prompt processing is segmented into discrete object tokens and attribute tokens, allowing the model to handle complex multi-object prompts systematically. Each object and its attributes are independently encoded and tracked, enabling the model to maintain high image quality even when generating complex scenes with multiple objects, thereby improving both flexibility and manufacturing precision.
Solution Approach 2:
The approach introduces an additional dimensional structure to the prompt encoding by separating objects and attributes into distinct token sequences. This dimensional organization allows the model to navigate complex prompt spaces more effectively, maintaining flexibility in handling diverse prompt configurations while preserving image generation quality through structured attribute binding.
Data Source
AI summary
Methods, systems, and non-transitory computer readable storage media are disclosed for generating digital images via a generative neural network with localized constraints. The disclosed system generates, utilizing one or more encoder neural networks, a sequence of embeddings comprising a prompt embedding representing a text prompt and an object text embedding representing a phrase indicating an object in the text prompt. The disclosed system generates, utilizing the one or more encoder neural networks, a visual embedding representing an object image corresponding to the object. The disclosed system determines a modified sequence of embeddings by replacing the object text embedding with the visual embedding in the sequence of embeddings. The disclosed system also generates, utilizing a generative neural network, a synthetic digital image from the modified sequence of embeddings comprising the visual embedding.


