Dynamic Facial Patches for Audio-Driven Mouth Animation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio content presentation in content creation platforms is limited to a static image mode, resulting in a poor presentation effect.

Innovation Solution

Generate a dynamic facial patch during audio playback to simulate a singer's singing process, using a target facial patch to represent the mouth shape corresponding to the audio content, and display it in a first facial area to present changing mouth shapes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a static picture is displayed during audio playback, then the presentation mode is simple and easy to implement, but the presentation effect is poor and lacks engagement

Engineering Contradiction:
Improveimplementation simplicityVSAvoidpresentation effect
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent applies the dynamics principle by transforming the static presentation mode into a dynamic one. Specifically, it generates and displays dynamic facial patches that change according to the audio content, making the mouth shapes move and change during playback. This dynamic visualization significantly improves the presentation effect and user engagement while maintaining reasonable implementation complexity through automated generation techniques.

Inventive Principle:
Principle #15Dynamics

2Productivity

If a dynamic facial patch is generated and displayed, then the presentation effect is enhanced and content richness is improved, but the system complexity increases

Engineering Contradiction:
Improvepresentation effectVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs the copying principle by generating facial patches that replicate and simulate real mouth shapes corresponding to audio content. Instead of requiring complex real-time capture and processing of actual facial movements, the system creates synthetic copies of facial expressions that match the audio, thereby enhancing presentation效果 while controlling system complexity through automated generation algorithms.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes by dynamically adjusting the facial patch parameters (such as mouth shape, position, and appearance) based on the audio content being played. This allows the system to transform static image parameters into dynamic ones that respond to audio variations, improving presentation effectiveness without requiring overly complex hardware modifications.

Inventive Principle:
Principle #35Parameter changes

3Power

If static images are used for audio presentation, then the processing requirements are low, but the content diversity and user engagement are limited

Engineering Contradiction:
Improveprocessing powerVSAvoidcontent diversity
Core Design Contradiction:
PowerVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static image processing into dynamic processing by generating facial patches that change over time according to the audio content. This dynamic approach enables the system to provide diverse and engaging content while maintaining efficient processing through optimized generation algorithms, thereby improving both content diversity and processing efficiency simultaneously.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4604064A1Image processing method and apparatus, electronic device, and storage medium
Publication Date: 2025.08.20 BEIJING ZITIAO NETWORK TECH CO LTD
  • EP4604064A1 patent drawingFigure 1~2
  • EP4604064A1 patent drawingFigure 3~4
  • EP4604064A1 patent drawingFigure 5

AI summary

The embodiments of the disclosure provides a method for image processing and an apparatus, an electronic device, a computer readable storage medium, a computer program product and a computer program, and the method includes: generating, during playing audio data, a target facial patch corresponding to a target audio frame, wherein the target facial patch is used to represent a target mouth shape, and the target mouth shape corresponds to audio content of the target audio frame; and displaying the target facial patch in a first facial area of a target image, wherein the first facial area is used to present a change of a mouth shape as the audio data is played. The target facial patch is used to simulate and display the target mouth shape corresponding to the currently played target audio content, so that the mouth shape displayed in the facial area of the target image may change with the change of the audio content, which may imitate the audio process corresponding to the audio data of a real people singing, so that the audio works may present the display effect of video works, and improve the richness and diversity of the display content of audio works.