Controlled Machine Learning Image Generation for Coherent Composites
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models for image generation lack fine-grained control over the integration of new visual elements, requiring careful prompt engineering and struggling to maintain visual coherence and plausibility when blending disparate elements.
Innovation Solution
A controlled image generating ML model is used, guided by content and appearance images, with a mask image defining where visual elements are generated, employing techniques like pre-processing, inference, and post-processing to seamlessly incorporate new elements into existing images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional image editing tools are used, then manual control over image composition is achieved, but the process becomes complex and labour-intensive
Solution Approach 1:
The patent replaces manual mechanical editing operations with an automated machine learning system. The image generation model automatically performs composition tasks that traditionally required manual masking, blending, and adjustment operations, substituting the mechanical editing process with an intelligent automated system that learns from training data.
Solution Approach 2:
The system enables self-service image composition through the machine learning model that autonomously performs editing tasks without requiring manual intervention. The model independently handles object integration, style transfer, and visual coherence maintenance, allowing users to generate composite images through simple prompts rather than complex manual operations.
2Productivity
If machine learning models are used for automated image generation, then productivity is improved, but fine-grained control over visual element integration is lost
Solution Approach 1:
The patent segments the image generation process into distinct controllable components: content images define object structure and layout, appearance images control visual style and texture, and mask images specify spatial regions. This segmentation allows independent control over different aspects of the generated image while maintaining overall productivity through automated processing.
Solution Approach 2:
The system enables fine-grained control by allowing users to modify specific parameters such as the weight and influence of content versus appearance images, adjustment of mask regions, and control over the strength of style transfer. These parameter changes provide precise control over visual element integration without sacrificing the automated productivity benefits.
3Productivity
If current image generation tools are used, then automated composition is achieved, but visual coherence and plausibility are difficult to maintain
Solution Approach 1:
The patent applies preliminary action by pre-processing input images to extract and separate content and appearance features before the main generation process. Content images are prepared to define structural elements, while appearance images are pre-processed to capture style and texture characteristics. This preliminary preparation ensures that the main generation process can reliably maintain visual coherence by working with already-optimized input representations.
Data Source
AI summary
Computer implemented methods and associated systems are described, which have particular application to image generation by machine learning models. A method of generating a composite image is described that is based on two images using a controlled machine learning model. A method of processing a composite image is also described which includes determining that a transition region of the composite image is similar to one of the images on which the composite image was based and using in the transition region visual elements from the basic image. A method for providing a user interface is also described. The method includes displaying representations of images generated using common input and different hyperparameters.


