Cross-modal 3D shape and color editing via shared latent space

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for editing and coloring 3D objects directly are cumbersome and do not effectively allow users to edit 3D shapes using 2D sketches and 2D RGB views.

Innovation Solution

The use of multi-modal variational auto-decoders (VADs) with a shared latent space enables editing 3D objects by editing 2D sketches and 2D RGB views, where separate VADs for 3D shapes, 2D sketches, and 2D RGB views are trained with paired variational auto-encoders and ground truth triplets, allowing for cross-modal shape and color manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users directly edit 3D objects using 3D editing programs, then the editing capability is provided, but the operation becomes cumbersome and difficult

Engineering Contradiction:
Improveease of editing 3D objectsVSAvoidcomplexity of 3D editing programs
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent uses 2D sketches and 2D RGB views as simplified copies/representations of 3D objects. Users edit these 2D representations instead of directly manipulating complex 3D models. The system maintains a mapping between 2D edits and 3D object changes, allowing users to work with simpler 2D interfaces while still achieving 3D editing goals.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transitions the editing interface from 3D space to 2D space. By allowing users to edit 2D sketches and 2D RGB views, the system reduces the dimensional complexity of the interaction. The underlying 3D object is updated based on 2D edits, effectively solving the contradiction by changing the dimension of the editing interface.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If separate VADs are trained for different modalities, then cross-modal manipulation capability is improved, but the system complexity increases

Engineering Contradiction:
Improvecross-modal manipulation capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a shared latent space that serves multiple functions: it represents 3D shapes, 2D sketches, and 2D RGB views simultaneously. This universal representation enables cross-modal manipulation (editing 3D objects via 2D sketches or RGB views) while avoiding the need for separate, isolated systems for each modality. The shared latent space acts as a common language across different representation types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges multiple VADs (3D shape VAD, 2D sketch VAD, 2D RGB view VAD) into a unified system with a shared latent space. Instead of maintaining completely separate systems for each modality, the VADs are combined through the shared latent space, allowing information to flow between modalities while maintaining the benefits of specialized processing for each type.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12094073B2Cross-modal shape and color manipulation
Publication Date: 2024.09.17 SNAP INC
  • US12094073B2 patent drawing
  • US12094073B2 patent drawing
  • US12094073B2 patent drawing

AI summary

Systems, computer readable media, and methods herein describe an editing system where a three-dimensional (3D) object can be edited by editing a 2D sketch or 2D RGB views of the 3D object. The editing system uses multi-modal (MM) variational auto-decoders (VADs)(MM-VADs) that are trained with a shared latent space that enables editing 3D objects by editing 2D sketches of the 3D objects. The system determines a latent code that corresponds to an edited or sketched 2D sketch. The latent code is then used to generate a 3D object using the MM-VADs with the latent code as input. The latent space is divided into a latent space for shapes and a latent space for colors. The MM-VADs are trained with variational auto-encoders (VAE) and a ground truth.