Feature-Guided Image Processing for Direct Creative Photo Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing photo photographing technologies cannot capture images in various styles directly through creative photographing, requiring post-processing for personalized photos.
Innovation Solution
An image processing method and apparatus that utilizes an image engine to determine feature information, combines it with input content to generate a target image, including a first object corresponding to the target feature and a second object based on the content description, using generative models like large language models (LLMs) and image-to-text models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional photo photographing technology is used, then real images of the real world can be captured, but personalized photos of various styles cannot be obtained directly
Solution Approach 1:
The system performs preliminary action by pre-training image-to-text models and semantic models with diverse style descriptions and photo templates before actual photo generation. The model library is prepared in advance with multiple style options, enabling the system to directly generate personalized photos of various styles without requiring post-processing operations.
Solution Approach 2:
The patent introduces an intermediary mechanism - the image-to-text model and semantic model - that bridges between the captured image and the final personalized photo. This intermediary translates image features into text descriptions, matches them with style templates, and generates the target photo, eliminating the need for manual post-processing while providing diverse style options.
2Measurement precision
If image processing based on image engine is performed to determine feature information, then target features can be extracted, but processing time and computational resources increase
Solution Approach 1:
The system performs preliminary action by pre-training image-to-text models with large amounts of image-text pairs before actual feature extraction. The semantic models and photo templates are prepared in advance, enabling the system to quickly match extracted features with appropriate style templates without requiring extensive real-time computation during the actual photo generation process.
Solution Approach 2:
The patent implements dynamic processing by adjusting the level of feature extraction based on the specific photo generation task. The system dynamically selects which features to extract and which pre-trained models to apply, optimizing the balance between extraction accuracy and processing time for different photo styles and requirements.
Data Source
AI summary
An image processing method includes obtaining an image, processing the image based on an image engine to determine feature information included in the image, obtaining input content, determining a target feature from the feature information included in the image based on the input content, and generating a target image based at least on the target feature and the input content. The target image includes a first object corresponding to the target feature and a second object corresponding to a content description of the target image generated based on the input content.


