Feature-Guided Image Processing for Direct Creative Photo Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing photo photographing technologies cannot capture images in various styles directly through creative photographing, requiring post-processing for personalized photos.

Innovation Solution

An image processing method and apparatus that utilizes an image engine to determine feature information, combines it with input content to generate a target image, including a first object corresponding to the target feature and a second object based on the content description, using generative models like large language models (LLMs) and image-to-text models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional photo photographing technology is used, then real images of the real world can be captured, but personalized photos of various styles cannot be obtained directly

Engineering Contradiction:
Improvephoto style varietyVSAvoidpost-processing requirement
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system performs preliminary action by pre-training image-to-text models and semantic models with diverse style descriptions and photo templates before actual photo generation. The model library is prepared in advance with multiple style options, enabling the system to directly generate personalized photos of various styles without requiring post-processing operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism - the image-to-text model and semantic model - that bridges between the captured image and the final personalized photo. This intermediary translates image features into text descriptions, matches them with style templates, and generates the target photo, eliminating the need for manual post-processing while providing diverse style options.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If image processing based on image engine is performed to determine feature information, then target features can be extracted, but processing time and computational resources increase

Engineering Contradiction:
Improvefeature extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-training image-to-text models with large amounts of image-text pairs before actual feature extraction. The semantic models and photo templates are prepared in advance, enabling the system to quickly match extracted features with appropriate style templates without requiring extensive real-time computation during the actual photo generation process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic processing by adjusting the level of feature extraction based on the specific photo generation task. The system dynamically selects which features to extract and which pre-trained models to apply, optimizing the balance between extraction accuracy and processing time for different photo styles and requirements.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250252615A1Image processing method and apparatus, and electronic device
Publication Date: 2025.08.07 LENOVO (BEIJING) LTD
  • US20250252615A1 patent drawing
  • US20250252615A1 patent drawing
  • US20250252615A1 patent drawing

AI summary

An image processing method includes obtaining an image, processing the image based on an image engine to determine feature information included in the image, obtaining input content, determining a target feature from the feature information included in the image based on the input content, and generating a target image based at least on the target feature and the input content. The target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input content.