Shared Attention Text-to-Image Generation for Consistent Subjects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-image diffusion models struggle with maintaining visual consistency of a subject across diverse prompts without requiring subject-specific training or fine-tuning, which constrains creativity and increases training time.
Innovation Solution
A pre-trained text-to-image diffusion model is used with a subject-driven shared attention block and correspondence-based feature injection to promote consistency across images, enabling zero-shot generation of consistent subjects without additional training, and allowing for layout diversity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If subject-specific training or fine-tuning is performed to achieve visual consistency, then subject consistency is improved, but training time and computational resources increase
Solution Approach 1:
The patent performs preliminary action by pre-training the diffusion model on general image data before deployment. The pre-trained model contains general visual knowledge that can be leveraged for consistent subject generation without requiring subject-specific fine-tuning, thus eliminating the need for time-consuming training while maintaining reliability
Solution Approach 2:
The patent uses copying by generating multiple images from the same text prompt and selecting the best matching ones that depict the desired subject consistently. Instead of training the model to learn subject-specific features, the system copies successful generation patterns from the pre-trained model's output space, achieving consistency through selection rather than modification
2Reliability
If subject-specific fine-tuning is applied to learn new words for specific subjects, then subject consistency is improved, but model adaptability to new subjects deteriorates
Solution Approach 1:
The patent applies universality by designing a system where the pre-trained diffusion model serves multiple functions: it can generate images for any subject without retraining, maintain consistent subject representation across different prompts, and adapt to new subjects immediately upon deployment. The subject consistency is achieved through the universal pre-trained knowledge rather than subject-specific adaptation
Solution Approach 2:
The system maintains adaptability by copying successful generation strategies from the pre-trained model rather than modifying the model itself. When new subjects are introduced, the system copies the effective prompt-generation patterns from the pre-trained model's output, allowing immediate adaptation to new subjects without fine-tuning
3Reliability
If conventional personalization techniques are used to generate consistent subject images, then subject consistency is improved, but text alignment and creativity are reduced
Solution Approach 1:
The patent uses copying to generate multiple candidate images from the same text prompt and selects the ones that best align with the desired subject representation. This copying and selection approach maintains text alignment because all candidates are generated from the identical prompt, preserving the original text-image relationship while achieving subject consistency through selection rather than constraint
Solution Approach 2:
The system performs preliminary generation of multiple image candidates before final selection. This preliminary action allows the system to evaluate multiple interpretations of the same prompt and select those that best satisfy both subject consistency and text alignment requirements, rather than being constrained by a single generation pass
Data Source
AI summary
Embodiments of the present disclosure relate to training-free consistent text-to-image generation. A pre-trained text-to-image diffusion model is leveraged to generate images depicting a consistent subject for diverse prompts describing scenes. Inputs to the model are a text description of at least one subject with prompts (scene text descriptions) describing scenes, where each prompt is associated with a different generated image and the text description is used for all images that depict the subject. Internal activations (intermediate data) computed by the model during generation of the different images are shared for generation of the different images. A subject-driven shared attention block and correspondence-based feature injection are incorporated into the model to promote subject consistency within each image and/or between images. Additionally, layout diversity is encouraged while maintaining subject consistency. The model achieves state-of-the-art performance on subject consistency and text alignment, without requiring any optimization and naturally extends to multi-subject scenarios.


