Scenario Image Generation Using Multimodal Appearance Prompts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation methods using large language models are inefficient and inaccurate, as they require manual input and lack customization, leading to unsatisfactory images and reduced user experience.

Innovation Solution

A method that combines original descriptive text and image information to determine appearance and scenario descriptive texts, using multimodal machine learning models to generate accurate scenario images, incorporating user-specific prompts and pose templates for enhanced customization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual input is used for image generation, then user control is maintained, but generation efficiency and accuracy are reduced

Engineering Contradiction:
Improveimage generation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system automatically generates images by processing uploaded original images and descriptive texts through multimodal machine learning models, eliminating the need for manual prompt input while maintaining customization through automated appearance and scenario text generation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes original images and texts to generate appearance descriptive texts and scenario descriptive texts before final image generation, breaking down the complex generation process into preparatory steps that automate the creation of generation parameters

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If generic image generation is used, then processing speed is maintained, but image accuracy and customization are reduced

Engineering Contradiction:
Improveimage accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The generation process is segmented into distinct stages: obtaining original images and texts, determining appearance descriptive texts, generating scenario descriptive texts, and producing final images. This segmentation allows each stage to be optimized independently, improving overall accuracy without excessive time cost

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts generation parameters by extracting features from original images and texts to create customized appearance and scenario descriptive texts, enabling high accuracy while maintaining reasonable processing time through parameter adaptation rather than generic processing

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If simple text prompts are used, then operation simplicity is maintained, but image customization and precision are reduced

Engineering Contradiction:
Improveoperation simplicityVSAvoidimage customization precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system introduces appearance descriptive texts and scenario descriptive texts as intermediary representations between simple user inputs and final image generation, automatically enriching simple prompts into detailed customization parameters while maintaining ease of use

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transforms simple text prompts into multi-dimensional descriptive texts that include appearance features and scenario contexts, adding dimensional richness to the generation process without requiring users to manually specify these additional dimensions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20260057570A1Method and apparatus, device, medium and program product for generating an image
Publication Date: 2026.02.26 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260057570A1 patent drawing
  • US20260057570A1 patent drawing
  • US20260057570A1 patent drawing

AI summary

Embodiments of the present disclosure relate to a method and apparatus for generating an image, a device, a medium and a program product. The method comprises obtaining first object information for a first object, the first object information including an original descriptive text for the first object and an original image of the first object. The method also comprises determining, based on the original descriptive text and the original image, an appearance descriptive text for the first object. The method further comprises generating, based on the original descriptive text and the appearance descriptive text, a scenario descriptive text of a scenario where the first object is applied. The method also comprises generating, based the scenario descriptive text, a scenario image of the scenario where the first object is applied.