Tokenized Image Generation via Sketch Edge Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently generating novel images that preserve the style and content semantics of background images, particularly in creating flexible and high-quality image manipulations such as adding copy space or editing backgrounds nondestructively.
Innovation Solution
The method involves using a system of machine learning models to generate novel images. A first model encodes an input image into tokenized representations, a second model encodes a sketch image with edge features into tokenized representations, and a third model predicts subsequent tokenized representations based on both sets of encoded data to reconstruct the image, allowing for the creation of novel images that maintain the original image's style and attributes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional image editing methods are used to manipulate background images, then basic editing operations can be performed, but the ability to preserve original image style and content semantics while generating novel images is limited
Solution Approach 1:
The patent segments the image manipulation task into multiple specialized machine learning models: a first model for encoding input images into tokenized representations, a second model for encoding sketch images with edge features, and a third model for predicting subsequent tokenized representations. This segmentation allows each model to specialize in specific aspects of image manipulation while working together to preserve style and content semantics.
Solution Approach 2:
The patent introduces tokenized representations as an intermediary between the input image/sketch and the output novel image. The first model converts the input image into tokenized representations, the second model processes sketch edge features into tokens, and the third model predicts subsequent tokens based on both inputs. This intermediary representation system enables flexible manipulation while maintaining fidelity to the original image's style and content.
2Manufacturing precision
If high-resolution novel images are generated using machine learning models, then image quality and fidelity are improved, but computational complexity and processing time increase
Solution Approach 1:
The patent divides the complex image generation task into three separate machine learning models, each responsible for a specific function: encoding input images, encoding sketch features, and predicting output representations. This segmentation reduces the complexity of individual models while maintaining high overall fidelity through their coordinated operation.
Solution Approach 2:
The patent transforms images into tokenized representations, changing the parameter space from continuous pixel values to discrete tokens. This parameter transformation enables more efficient processing and prediction while maintaining the ability to generate high-resolution images with faithful reproduction of style and content.
3Adaptability or versatility
If multiple machine learning models are used to encode and predict image representations, then image generation capability is enhanced, but system complexity increases
Solution Approach 1:
The patent creates a universal tokenized representation system that can handle multiple types of image inputs (original images and sketch images) and produce various novel image outputs. The shared tokenization approach and coordinated model system provide multi-functional capability for different image manipulation tasks while maintaining a relatively streamlined architecture.
Data Source
AI summary
Techniques for generating a novel image using tokenized image representations are disclosed. In some embodiments, a method of generating the novel image includes generating, via a first machine learning model, a first sequence of coded representations of a first image having one or more features; generating, via a second machine learning model, a second sequence of coded representations of a sketch image having one or more edge features associated with the one or more features; predicting, via a third machine learning model, one or more subsequent coded representations based on the first sequence of coded representations and the second sequence of coded representations; and based on the subsequent coded representations, generating, via the third machine learning model, a first portion of a reconstructed image having one or more image attributes of the first image, and a second portion of the reconstructed image associated with the one or more edge features.


