Image Generation Using Language Model Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating image content are inefficient, particularly when multiple pieces of image content need to be created, as they require users to repeatedly select and combine various parts, leading to increased workload and variability in operational efficiency due to differing prior knowledge among users.

Innovation Solution

An image generation apparatus and method that acquire text information and a prompt to derive metadata using a language model, which is then used to generate images by selecting and combining appropriate image parts based on the metadata.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually select and combine multiple parts to generate image content, then image content can be generated with user control, but the workload increases significantly when more pieces of content are generated

Engineering Contradiction:
Improveuser control over image generationVSAvoidtime required for selecting and combining parts
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system enables self-service image generation by automatically selecting and combining image parts based on text prompts. The image generation model autonomously processes the text input and generates the complete image without requiring manual user intervention in the selection and combination steps, thus reducing user workload while maintaining control through the initial text prompt.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary image generation model that acts as a mediator between the user's text prompt and the final generated image. This model translates the text description into appropriate image part selections and combinations, eliminating the need for users to manually perform these complex selection and assembly tasks while still achieving the desired image content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If content generation operations are allocated to multiple users to reduce individual workload, then more content can be generated, but operational efficiency varies depending on user prior knowledge

Engineering Contradiction:
Improvetotal content generation volumeVSAvoidoperational efficiency consistency
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements a universal image generation model that performs multiple functions: text analysis, image part selection, and image combination. This single multi-functional system can be used by any user regardless of their prior knowledge, ensuring consistent operational efficiency across different users while enabling high-volume content generation through automated processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces the mechanical manual process of selecting and combining image parts with an automated computational system. The image generation model uses algorithms to automatically analyze text prompts and generate appropriate images, eliminating the variability in human operational efficiency while maintaining high productivity through scalable automated processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If a large number of image contents are generated by manually selecting parts, then diverse image content can be created, but the workload for selecting parts and generating content increases

Engineering Contradiction:
Improvediversity of generated image contentVSAvoidtime for part selection and content generation
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically selecting diverse image parts and combining them according to the text prompt requirements. The image generation model independently handles the complex task of selecting appropriate parts for each specific image generation request, enabling diverse content creation without increasing user workload.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-training the image generation model on diverse image data and part combinations. This preliminary training enables the model to quickly and accurately select appropriate parts for new image generation tasks without requiring users to manually explore or select from large numbers of parts, thus maintaining diversity while reducing time investment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250166265A1Image generation apparatus, image generation method, and non-transitory computer readable medium
Publication Date: 2025.05.22 RAKUTEN GROUP INC
  • US20250166265A1 patent drawing
  • US20250166265A1 patent drawing
  • US20250166265A1 patent drawing

AI summary

An image generation apparatus acquires text information and a prompt, the prompt being an instruction to output metadata that includes a plurality of attribute values composed of attribute values that respectively correspond to a plurality of attributes and semantically match the text information, derives metadata corresponding to the text information by inputting the text information and the prompt to a language model, generates an image based on the derived metadata.