Video Motion Embeddings for Appearance-Preserving Motion Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video generation and editing techniques struggle to control both the appearance and motion in a video in a predictable and fine-grained manner, often requiring complex alignment and control inputs like bounding boxes or manual adjustments, and text-to-video models fail to preserve the appearance and spatial layout of target images.
Innovation Solution
A technique for semantic video motion transfer using motion-textual inversion, where an embedding is determined to encode spatial and temporal attributes of motion, allowing the transfer of motion to an output video with a different appearance without requiring spatial alignment or additional control inputs, by optimizing the embedding based on losses computed between the output and reference videos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dense control inputs (motion vectors, depth maps) are used to control motion in generated videos, then motion control precision is improved, but device complexity and ease of operation deteriorate due to requiring alignment between target and reference videos and manual control inputs
Solution Approach 1:
The patent extracts the essential motion characteristics from reference videos by computing optical flow and deriving motion embeddings that capture temporal dynamics. This extraction process separates motion information from appearance information, allowing motion control without requiring dense control inputs or alignment between target and reference videos. The motion embedding serves as a compact representation that eliminates the need for complex control mechanisms.
Solution Approach 2:
The patent introduces motion embeddings as an intermediary representation between reference videos and generated outputs. These embeddings serve as a mediator that transfers motion characteristics without requiring direct alignment or dense control inputs. The embedding space acts as an intermediate layer that decouples motion control from appearance preservation, simplifying the overall control mechanism.
2Measurement precision
If manual control inputs (bounding boxes, trajectories) are used for motion control, then motion precision is improved, but ease of operation deteriorates due to significant effort required for complex motions
Solution Approach 1:
The patent enables the system to automatically learn and capture motion patterns from reference videos without requiring manual annotation or control input. The motion embeddings are computed automatically through optical flow analysis and neural network processing, allowing the system to serve itself by extracting motion characteristics directly from video data rather than requiring human operators to provide bounding boxes or trajectories.
Solution Approach 2:
The patent performs preliminary computation of motion embeddings from reference videos before the actual video generation process. By pre-computing and storing motion embeddings that capture essential motion dynamics, the system prepares motion control information in advance, eliminating the need for real-time manual control inputs during video generation and reducing operational effort.
3Reliability
If text-to-video models are fine-tuned on motion reference videos to capture motion, then motion representation is improved, but appearance generalization deteriorates due to inadvertently learning appearance of reference video
Solution Approach 1:
The patent segments motion information and appearance information into separate representations. Motion embeddings are computed independently from appearance features through optical flow analysis, while appearance is controlled separately through image-to-video generation. This segmentation allows the model to learn motion patterns from reference videos without inadvertently learning their appearance, maintaining appearance generalization capability.
Solution Approach 2:
The patent replaces fine-tuning of text-to-video models with a mechanism that computes motion embeddings through optical flow and neural network processing. Instead of modifying the model weights through fine-tuning on reference videos (which causes appearance leakage), the system substitutes this with an embedding computation approach that extracts motion characteristics without transferring appearance information to the generative model.
4Manufacturing precision
If spatial alignment between target image and motion reference video is required, then motion transfer accuracy is improved, but ease of operation and adaptability deteriorate
Solution Approach 1:
The patent extracts motion characteristics from reference videos in a spatially invariant manner through optical flow computation and embedding derivation. By focusing on temporal dynamics rather than spatial positioning, the system separates motion information from spatial configuration, allowing motion transfer without requiring alignment between target images and reference videos. This extraction approach captures essential motion patterns independent of their spatial context.
Data Source
AI summary
One embodiment of the present invention sets forth a technique for performing motion transfer. The technique includes determining an embedding corresponding to a motion depicted in a first video. The technique also includes generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image.


