Video Motion Embeddings for Appearance-Preserving Motion Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation and editing techniques struggle to control both the appearance and motion in a video in a predictable and fine-grained manner, often requiring complex alignment and control inputs like bounding boxes or manual adjustments, and text-to-video models fail to preserve the appearance and spatial layout of target images.

Innovation Solution

A technique for semantic video motion transfer using motion-textual inversion, where an embedding is determined to encode spatial and temporal attributes of motion, allowing the transfer of motion to an output video with a different appearance without requiring spatial alignment or additional control inputs, by optimizing the embedding based on losses computed between the output and reference videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If dense control inputs (motion vectors, depth maps) are used to control motion in generated videos, then motion control precision is improved, but device complexity and ease of operation deteriorate due to requiring alignment between target and reference videos and manual control inputs

Engineering Contradiction:
Improvemotion control precisionVSAvoidcontrol input complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential motion characteristics from reference videos by computing optical flow and deriving motion embeddings that capture temporal dynamics. This extraction process separates motion information from appearance information, allowing motion control without requiring dense control inputs or alignment between target and reference videos. The motion embedding serves as a compact representation that eliminates the need for complex control mechanisms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces motion embeddings as an intermediary representation between reference videos and generated outputs. These embeddings serve as a mediator that transfers motion characteristics without requiring direct alignment or dense control inputs. The embedding space acts as an intermediate layer that decouples motion control from appearance preservation, simplifying the overall control mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual control inputs (bounding boxes, trajectories) are used for motion control, then motion precision is improved, but ease of operation deteriorates due to significant effort required for complex motions

Engineering Contradiction:
Improvemotion precisionVSAvoidoperation effort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent enables the system to automatically learn and capture motion patterns from reference videos without requiring manual annotation or control input. The motion embeddings are computed automatically through optical flow analysis and neural network processing, allowing the system to serve itself by extracting motion characteristics directly from video data rather than requiring human operators to provide bounding boxes or trajectories.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary computation of motion embeddings from reference videos before the actual video generation process. By pre-computing and storing motion embeddings that capture essential motion dynamics, the system prepares motion control information in advance, eliminating the need for real-time manual control inputs during video generation and reducing operational effort.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If text-to-video models are fine-tuned on motion reference videos to capture motion, then motion representation is improved, but appearance generalization deteriorates due to inadvertently learning appearance of reference video

Engineering Contradiction:
Improvemotion representationVSAvoidappearance generalization
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments motion information and appearance information into separate representations. Motion embeddings are computed independently from appearance features through optical flow analysis, while appearance is controlled separately through image-to-video generation. This segmentation allows the model to learn motion patterns from reference videos without inadvertently learning their appearance, maintaining appearance generalization capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces fine-tuning of text-to-video models with a mechanism that computes motion embeddings through optical flow and neural network processing. Instead of modifying the model weights through fine-tuning on reference videos (which causes appearance leakage), the system substitutes this with an embedding computation approach that extracts motion characteristics without transferring appearance information to the generative model.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Manufacturing precision

If spatial alignment between target image and motion reference video is required, then motion transfer accuracy is improved, but ease of operation and adaptability deteriorate

Engineering Contradiction:
Improvemotion transfer accuracyVSAvoidspatial alignment requirement
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts motion characteristics from reference videos in a spatially invariant manner through optical flow computation and embedding derivation. By focusing on temporal dynamics rather than spatial positioning, the system separates motion information from spatial configuration, allowing motion transfer without requiring alignment between target images and reference videos. This extraction approach captures essential motion patterns independent of their spatial context.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250356506A1Semantic video motion transfer using motion-textual inversion
Publication Date: 2025.11.20 DISNEY ENTERPRISES INC
  • US20250356506A1 patent drawing
  • US20250356506A1 patent drawing
  • US20250356506A1 patent drawing

AI summary

One embodiment of the present invention sets forth a technique for performing motion transfer. The technique includes determining an embedding corresponding to a motion depicted in a first video. The technique also includes generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image.