Transformer Sketch Generation Model for Geometric Representation Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for converting images into sketches struggle to preserve geometric information and require a sketch dataset, leading to inefficient and unstable processes that are impractical for representation learning.
Innovation Solution
A method and apparatus that convert images into stroke-based sketches using a stroke generation model, which receives initial stroke information and image data to generate final strokes with geometric information, without relying on a separate sketch dataset, utilizing a transformer-based architecture and perceptual loss functions for stable representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional sketch generation methods using optimization-based techniques are employed, then sketches can be generated without a sketch dataset, but the process takes excessively long time and produces numerically unstable results
Solution Approach 1:
The patent pre-trains a transformer model on natural image datasets before fine-tuning it for sketch generation. This preliminary training establishes a robust foundation that enables rapid inference without requiring sketch datasets during the actual generation process, thereby resolving the contradiction between dataset-free generation and generation speed
Solution Approach 2:
The patent replaces the iterative optimization-based mechanical process with a learned transformer model that performs sketch generation through neural network inference. This substitution eliminates the need for thousands of optimization steps, achieving both dataset-free operation and computational efficiency
2Ease of manufacture
If conventional sketch generation methods are used, then sketches can be generated, but they fail to preserve geometric information and do not achieve accurate abstraction
Solution Approach 1:
The patent incorporates a perceptual loss function that provides feedback during training to ensure the generated sketches preserve geometric information. This loss function compares the generated sketches with reference images and adjusts the model parameters to maintain geometric fidelity, thereby resolving the contradiction between generation capability and geometric precision
Solution Approach 2:
The patent changes the parameter space by using a transformer architecture with learnable parameters that are optimized to preserve geometric information. By adjusting the model's internal parameters through perceptual loss guidance, the system achieves accurate geometric abstraction while maintaining sketch generation capability
3Ease of manufacture
If optimization-based sketch generation is employed, then sketches can be generated without a sketch dataset, but the results are variant and numerically unstable
Solution Approach 1:
The patent replaces the unstable iterative optimization process with a deterministic transformer model inference process. The learned model produces consistent and stable results without the numerical instability inherent in optimization-based methods, while still maintaining dataset-free operation
Solution Approach 2:
The extensive pre-training phase serves as a preliminary action that stabilizes the model parameters before deployment. This pre-training establishes a robust parameter configuration that ensures numerical stability during inference, resolving the contradiction between dataset-free generation and reliability
Data Source
AI summary
Proposed herein are an apparatus and method for conversion into sketches. The apparatus for conversion into sketches includes: an input/output interface configured to receive an image and output the results of processing of the image; storage configured to store a program for performing a method for conversion into sketches; and a controller including at least one processor, and configured to convert the image into a sketch by executing the program. The controller receives (i) initial stroke information having attribute parameters and (ii) image information about a target image, converts the target image into a plurality of final strokes having geometric information about the target image by using a stroke generation model, and outputs the results of data processing of the plurality of final strokes.


