Neural Network Animation Using Identity-Invariant Feature Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for animating static images into realistic videos face challenges with timewise consistency and effective resolution, often requiring additional rendering computations.
Innovation Solution
An artificial neural-network based method that uses static images of an object and a driving video to generate a realistic animation by extracting identity-invariant features such as pose, lighting, and facial expressions, applying a transformation function to produce animated videos with improved consistency and resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If graphical techniques such as object swapping or face swapping are applied to animate a still image according to a video, then animation can be achieved, but the output animation suffers from poor timewise consistency and poor effective resolution
Solution Approach 1:
The patent segments the animation process into multiple specialized neural network models: a first ML model for extracting identity-invariant features, a second ML model for calculating transformation functions, and a third ML model for generating semantic maps. This segmentation allows each model to specialize in a specific aspect of the animation process, improving both temporal consistency and effective resolution by handling complex transformations in a structured, multi-stage manner
Solution Approach 2:
The patent introduces identity-invariant features as an intermediary representation between the source image and the driving video. These features (extracted by the first ML model) serve as a mediator that captures essential characteristics while being invariant to identity changes, enabling the transformation function to focus on motion and expression transfer without identity confusion, thereby improving temporal consistency and resolution
2Manufacturing precision
If graphical techniques are used to achieve desired quality of animation video, then animation can be produced, but additional rendering computations are required
Solution Approach 1:
The patent replaces traditional graphical rendering mechanics with machine learning-based transformation. Instead of using complex rendering pipelines to achieve quality animation, the system uses trained ML models that directly generate high-quality animated frames by learning the mapping between identity-invariant features and visual appearances, significantly reducing the need for additional rendering computations while maintaining or improving video quality
3Reliability
If identity-invariant features are extracted and transformation functions are applied to animate static images, then temporally consistent animated videos with improved resolution can be generated, but the system requires multiple machine-learning models and feature extraction processes
Solution Approach 1:
The patent creates a universal framework where the first ML model for extracting identity-invariant features serves multiple purposes: it processes both the source image and driving video frames, enables temporal consistency analysis, and provides features for both transformation function calculation and semantic map generation. This multi-functionality reduces overall system complexity despite using multiple models, as each model handles diverse tasks within the animation pipeline
Data Source
AI summary
A system and method of animating an image of an object may include: receiving a first image, depicting a “puppet” object; sampling an input video, depicting a second, “driver” object, to obtain at least one second image; obtaining, by a first machine-learning (ML) model, a first identity-invariant feature of the puppet object, from the first image; obtaining at least one second identity-invariant feature of the driver object, from the respective at least one second image; calculating, by a second ML model, a transformation function, based on the first identity-invariant feature and the at least one second identity-invariant feature; applying the calculated transformation function on the first image, to produce one or more third images, depicting a target object, including at least one identity-invariant feature of the driver object; and appending the one or more third images to produce an output video depicting animation of the puppet object.


