Sketch-to-3D Generation Using Semantic Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sketch-to-3D shape generation techniques lack large-scale paired training data, leading to limited generalizability across varying levels of abstraction in sketches and strong inductive biases, constraining their ability to generate 3D shapes from casual doodles to professional drawings.
Innovation Solution
The technique leverages pre-trained image-text models to extract semantic features from input sketches, guiding a generative machine learning model to produce 3D shapes without requiring (sketch, 3D shape) paired training data, enabling 'zero-shot' generation across different levels of complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If existing sketch-to-3D techniques are trained on synthetic datasets or limited paired data, then training feasibility is improved, but generalizability across varying levels of abstraction deteriorates
Solution Approach 1:
The patent introduces an intermediary representation called shape embeddings that mediates between the sketch input and 3D shape output. The model learns to map sketches to shape embeddings, which are then decoded into 3D shapes. This intermediary representation allows the model to generalize across different sketch styles and levels of abstraction without requiring large amounts of paired training data, as the shape embeddings capture essential geometric properties in a standardized form.
Solution Approach 2:
The patent employs pre-trained image encoders and shape embedding decoders that have been trained on large-scale unsupervised datasets before being fine-tuned for the sketch-to-3D task. This preliminary training on abundant unlabeled data provides strong feature representations and geometric priors, enabling the model to achieve good generalization performance with limited paired sketch-shape training data.
2Quantity of substance
If existing techniques are trained on data from only a few categories, then training data requirements are reduced, but ability to handle varying levels of abstraction deteriorates
Solution Approach 1:
The model performs preliminary training on large-scale unsupervised shape datasets to learn general geometric representations and shape priors. This pre-training phase enables the model to understand fundamental shape properties and relationships without being constrained to specific categories or requiring paired sketch-shape data, thereby achieving both data efficiency and broad adaptability to different abstraction levels.
Solution Approach 2:
The shape embedding representation and decoder are designed to be universal and category-agnostic, capable of representing and generating 3D shapes across diverse object categories and abstraction levels. The model learns general shape manifolds that can be applied to any object type, making the system universally applicable rather than specialized for specific categories.
3Speed
If existing techniques use strong inductive biases from training data, then convergence speed is improved, but generalizability across 3D representations deteriorates
Solution Approach 1:
The patent changes the parameterization of the 3D shape representation by using shape embeddings in a latent space rather than direct mesh or voxel representations. This parameter transformation allows the model to learn more flexible and generalizable shape manifolds. Additionally, the use of different loss functions and optimization strategies during training adjusts the learning dynamics to achieve fast convergence without overfitting to specific inductive biases in the training data.
Data Source
AI summary
One embodiment of the present invention sets forth a technique for performing 3D shape generation. This technique includes generating semantic features associated with an input sketch. The technique also includes generating, using a generative machine learning model, a plurality of predicted shape embeddings from a set of fully masked shape embeddings based on the semantic features associated with the input sketch. The technique further includes converting the predicted shape embeddings into one or more 3D shapes. The input sketch may be a casual doodle, a professional illustration, or a 2D CAD software rendering. Each of the one or more 3D shapes may be a voxel representation, an implicit representation, or a 3D CAD software representation.


