Dynamic Image Generation from Single Reference Image

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for generating digital persons require extensive video collection and data cleaning, making the process cumbersome, costly, and time-consuming, as they necessitate large amounts of high-quality video data for accurate lip-sync animation.

Innovation Solution

A method and apparatus that generate a dynamic image based on a single picture of a target person, using a trained generation network model to extract head pose and expression features, allowing for the creation of a digital person that changes facial expressions based on audio input without the need for video collection or data cleaning, by processing the image through an affine subnetwork and a driving subnetwork.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If extensive video collection and data cleaning are performed to generate digital persons, then the accuracy and quality of lip-sync animation is improved, but the production time and cost increase significantly

Engineering Contradiction:
Improvelip-sync animation qualityVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The method performs preliminary extraction of head pose and expression features from a single reference image before generating the dynamic image. The generation network model is pre-trained to learn the mapping between audio features and facial expression changes, eliminating the need for time-consuming video collection and data cleaning during production

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The method extracts only the necessary features (head pose and expression coefficients) from the reference image using the trained generation network model, rather than collecting and processing extensive video data. This selective extraction of critical features reduces production time while maintaining animation quality

Inventive Principle:
Principle #2Taking out (Extraction)

2Manufacturing precision

If extensive video collection and data cleaning are performed to generate digital persons, then the accuracy and quality of lip-sync animation is improved, but the production cost increases significantly

Engineering Contradiction:
Improvelip-sync animation qualityVSAvoidproduction cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The method extracts only the necessary features (head pose and expression coefficients) from a single reference image using the trained generation network model, rather than collecting and processing extensive video data. This selective extraction of critical features reduces production cost while maintaining animation quality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The method uses a single reference image as a template to generate multiple dynamic images with different facial expressions driven by audio input. The generation network model learns from the reference image and applies the learned patterns to create consistent digital person animations without requiring multiple video samples

Inventive Principle:
Principle #26Copying

3Ease of manufacture

If a single picture is used to generate a dynamic image, then the production process is simplified and cost is reduced, but the complexity of the generation network model increases

Engineering Contradiction:
Improveproduction simplicityVSAvoidgeneration network model complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The generation network model is divided into distinct functional components: an affine subnetwork that processes head pose features through affine transformations, and a driving subnetwork that generates expression changes based on audio features. This modular segmentation manages model complexity while enabling efficient processing from a single reference image

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12260481B1Method for generating a dynamic image based on audio, device, and storage medium
Publication Date: 2025.03.25 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US12260481B1 patent drawing
  • US12260481B1 patent drawing
  • US12260481B1 patent drawing

AI summary

Disclosed are a method for generating a dynamic image based on audio, a device, and a storage medium, relating to the field of natural human-computer interactions. The method includes: obtaining a reference image and reference audio input by a user; determining a target head pose feature and a target expression coefficient feature based on the reference image and a trained generation network model, and adjusting the trained generation network model based on the target head pose feature and the target expression coefficient feature, to obtain a target generation network model; and processing a to-be-processed image based on the reference audio, the reference image, and the target generation network model, to obtain a target dynamic image. An image object in the to-be-processed image is same as that in the reference image. In this case, a corresponding digital person can be obtained based on a single picture of a target person.