AI Visual Content Generation Overcoming Design Fixation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for visual content generation struggle with design fixation, where users tend to converge on limited solutions, and existing generative visual models fail to facilitate divergent thinking, relying heavily on human effort and limited existing content.

Innovation Solution

A system utilizing a generative language model and a generative visual model to generate semantically diverse texts and images, assisting users in overcoming design fixation by exploring a broader range of possibilities through AI-driven creative processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If users rely on existing generative visual models for content generation, then the generation process is simple, but the output lacks contextual diversity and suffers from design fixation

Engineering Contradiction:
Improvecontextual diversityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

A language model is introduced as an intermediary component between the user prompt and the visual generation model. The language model generates multiple semantically diverse text interpretations of the prompt, which then serve as varied inputs to the visual generation model, thereby increasing contextual diversity without requiring changes to the core visual generation architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The generation process is segmented into distinct stages: prompt interpretation by the language model, generation of multiple diverse text variants, and visual generation from each text variant. This segmentation allows each component to specialize in its function while collectively achieving diverse outputs

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If users manually explore design possibilities, then diverse ideas can be generated, but significant time and labor are required

Engineering Contradiction:
Improvedesign idea diversityVSAvoidgeneration speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs self-service by automatically generating multiple diverse text interpretations and visual variants without requiring manual human intervention at each step. The language model and visual generation model work together autonomously to produce diverse design ideas, significantly reducing the time and labor compared to manual exploration while maintaining high diversity

Inventive Principle:
Principle #25Self-service

3Ease of operation

If existing generative models are used, then the process requires minimal human effort, but the output converges on limited solutions due to design fixation

Engineering Contradiction:
Improveoperational simplicityVSAvoidsolution diversity
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system introduces dynamics by generating multiple different text interpretations of the same prompt and feeding them to the visual generation model. This creates a dynamic exploration of the design space where the same input prompt can lead to multiple diverse visual outcomes, preventing convergence on limited solutions while maintaining ease of operation

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240273308A1System and method for visual content generation and iteration
Publication Date: 2024.08.15 TOYOTA RESEARCH INSTITUTE INC
  • US20240273308A1 patent drawing
  • US20240273308A1 patent drawing
  • US20240273308A1 patent drawing

AI summary

Systems, methods, and other embodiments described herein relate to enhancing and complementing a creative process of a user that includes generating and iterating visual content with an emphasis on diverse design ideas. In one embodiment, a method includes generating a plurality of texts that are related and semantically diverse based on one or more prompts using a generative language model and generating a plurality of images based on at least a portion of the plurality of texts using a generative visual model.