Multi-Dimensional Video Generation with Tri-Plane Motion Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative adversarial networks (GANs) struggle to generate three-dimensional (3D) videos, particularly for applications like 3D portrait video generation and video manipulation, as they often produce two-dimensional (2D) videos that do not consider underlying 3D geometry, requiring multi-camera systems and large amounts of training data.

Innovation Solution

A multi-dimensionally aware generative model that inverts input data into a latent space, synthesizes content using appearance and motion components, and introduces temporal dynamics to generate multi-dimensional videos, leveraging tri-plane representations and conditioned training on 2D monocular videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If classical graphics techniques with multi-camera systems are used to generate 3D portrait video, then 3D geometry accuracy is improved, but device complexity and training data requirements increase significantly

Engineering Contradiction:
Improve3D geometry accuracyVSAvoidmulti-camera system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses a single-camera system to capture 2D videos and creates synthetic 3D representations by learning from multiple viewpoints implicitly. Instead of requiring actual multi-camera systems to capture 3D geometry, the model learns to generate 3D-consistent video sequences from single-camera input by copying and transforming appearance and motion patterns across virtual viewpoints.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical multi-camera system with a neural network-based generative model. Instead of using physical cameras arranged in specific geometries to capture 3D information, the system uses a trained GAN that learns 3D geometry and appearance relationships from 2D video data and generates multi-view consistent 3D portrait videos through learned transformations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If classical graphics techniques with well-controlled studios are used to generate 3D portrait video, then 3D geometry quality is improved, but ease of operation deteriorates

Engineering Contradiction:
Improve3D geometry qualityVSAvoidstudio setup requirement
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent enables the system to automatically learn 3D geometry and appearance relationships from uncontrolled 2D video environments without requiring manual studio setup. The model performs self-calibration by learning camera poses, lighting conditions, and geometric structures from the input videos themselves, eliminating the need for controlled studio environments with specific lighting and camera arrangements.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the operational parameters from requiring controlled studio conditions to accepting uncontrolled real-world environments. The model learns to handle variations in lighting, camera angles, and background conditions by training on diverse 2D video data, transforming the system into one that operates effectively without specialized studio setups.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If prior GAN models are used to generate video, then ease of operation is maintained, but 3D geometry awareness deteriorates

Engineering Contradiction:
Improvemodel simplicityVSAvoid3D geometry awareness
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent extends standard 2D video generation GANs by incorporating 3D geometric awareness through the use of tri-plane representations. The model maintains the simplicity of GAN operation while adding 3D dimensionality by representing scenes in three orthogonal planes that encode spatial relationships, allowing the network to generate videos with consistent 3D geometry without complicating the basic generator-discriminator framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines multiple components into a composite generative model that integrates appearance encoding, motion encoding, and 3D geometric representation (tri-planes) within the GAN framework. This composite structure allows the model to maintain operational simplicity while incorporating 3D geometry awareness through the synergistic combination of appearance models, motion models, and spatial representations.

Inventive Principle:
Principle #40Composite materials

4Manufacturing precision

If multi-camera systems are used to capture training data for 3D video generation, then 3D geometry accuracy is improved, but loss of substance increases

Engineering Contradiction:
Improve3D geometry accuracyVSAvoidtraining data volume
Core Design Contradiction:
Manufacturing precisionVSLoss of substance

Solution Approach 1:

The patent learns to generate synthetic multi-view training data by copying and transforming appearance and motion patterns from single-camera input videos. The model creates virtual multi-camera viewpoints through learned transformations, eliminating the need to physically capture large volumes of multi-camera training data while still achieving 3D geometry accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the need for extensive multi-camera training data collection with a neural network that learns 3D geometry from 2D videos and generates synthetic training examples. The model substitutes physical data collection infrastructure with learned generative capabilities that can produce unlimited 3D-consistent video sequences from single-camera input.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250259057A1Multi-dimensional generative framework for video generation
Publication Date: 2025.08.14 LEMON INC(GB)
  • US20250259057A1 patent drawing
  • US20250259057A1 patent drawing
  • US20250259057A1 patent drawing

AI summary

Generating a multi-dimensional video using a multi-dimensional video generative model for, including, but not limited to, at least one of static portrait animation, video reconstruction, or motion editing. The method including providing data into the multi-dimensionally aware generator of the multi-dimensional video generative model, and generating the multi-dimensional video from the data by the multi-dimensionally aware generator. The generating of the multi-dimensional video includes inverting the data into a latent space of the multi-dimensionally aware generator, synthesizing content of the multi-dimensional video using an appearance component of the multi-dimensionally aware generator and corresponding camera pose and formulating an intermediate appearance code, developing a synthesis layer for encoding a motion component of the multi-dimensionally aware generator at a plurality of timesteps and formulating an intermediate motion code, introducing temporal dynamics into the intermediate appearance code and the intermediate motion code, and generating multi-dimensionally aware spatio-temporal representations of the data.