Dynamic Facial Patches for Audio-Driven Mouth Animation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio content presentation in content creation platforms is limited to a static image mode, resulting in a poor presentation effect.
Innovation Solution
Generate a dynamic facial patch during audio playback to simulate a singer's singing process, using a target facial patch to represent the mouth shape corresponding to the audio content, and display it in a first facial area to present changing mouth shapes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a static picture is displayed during audio playback, then the presentation mode is simple and easy to implement, but the presentation effect is poor and lacks engagement
Solution Approach 1:
The patent applies the dynamics principle by transforming the static presentation mode into a dynamic one. Specifically, it generates and displays dynamic facial patches that change according to the audio content, making the mouth shapes move and change during playback. This dynamic visualization significantly improves the presentation effect and user engagement while maintaining reasonable implementation complexity through automated generation techniques.
2Productivity
If a dynamic facial patch is generated and displayed, then the presentation effect is enhanced and content richness is improved, but the system complexity increases
Solution Approach 1:
The patent employs the copying principle by generating facial patches that replicate and simulate real mouth shapes corresponding to audio content. Instead of requiring complex real-time capture and processing of actual facial movements, the system creates synthetic copies of facial expressions that match the audio, thereby enhancing presentation效果 while controlling system complexity through automated generation algorithms.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting the facial patch parameters (such as mouth shape, position, and appearance) based on the audio content being played. This allows the system to transform static image parameters into dynamic ones that respond to audio variations, improving presentation effectiveness without requiring overly complex hardware modifications.
3Power
If static images are used for audio presentation, then the processing requirements are low, but the content diversity and user engagement are limited
Solution Approach 1:
The patent transforms the static image processing into dynamic processing by generating facial patches that change over time according to the audio content. This dynamic approach enables the system to provide diverse and engaging content while maintaining efficient processing through optimized generation algorithms, thereby improving both content diversity and processing efficiency simultaneously.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
The embodiments of the disclosure provides a method for image processing and an apparatus, an electronic device, a computer readable storage medium, a computer program product and a computer program, and the method includes: generating, during playing audio data, a target facial patch corresponding to a target audio frame, wherein the target facial patch is used to represent a target mouth shape, and the target mouth shape corresponds to audio content of the target audio frame; and displaying the target facial patch in a first facial area of a target image, wherein the first facial area is used to present a change of a mouth shape as the audio data is played. The target facial patch is used to simulate and display the target mouth shape corresponding to the currently played target audio content, so that the mouth shape displayed in the facial area of the target image may change with the change of the audio content, which may imitate the audio process corresponding to the audio data of a real people singing, so that the audio works may present the display effect of video works, and improve the richness and diversity of the display content of audio works.