Facial Feature Video Sync for Audio-Triggered Effects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video applications with artificial intelligence-based human-computer interaction suffer from limited effect props, poor video quality, and synchronization issues between video content and selected effects, leading to a poor user experience.
Innovation Solution
A method and apparatus that determine a target facial feature in a facial image to align with the facial feature presented when a target audio is played, by acquiring a target facial image and audio, determining a key video frame sequence, and generating a target effect audio and video based on the facial feature and audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If effect props are added to video applications, then video content richness is improved, but device computing power requirements increase
Solution Approach 1:
The patent segments the video processing into key frame extraction and non-key frame skipping. By identifying and processing only key frames that contain important facial features, the system reduces the computational load while maintaining video content richness. This segmentation allows effect props to be applied selectively rather than to every frame, resolving the contradiction between content richness and computing power requirements.
Solution Approach 2:
The patent applies local quality by focusing computational resources on specific regions (facial features in key frames) rather than processing the entire video uniformly. The system identifies facial regions in key frames and applies effect props locally to these areas, improving video content richness while minimizing the overall computing power consumption.
2Ease of operation
If facial feature alignment with audio is implemented, then user experience is improved, but processing time increases
Solution Approach 1:
The patent applies preliminary action by extracting and storing key frames with facial features before the actual audio synchronization process. By pre-identifying and preparing the key frames that contain relevant facial information, the system reduces the processing time required during audio playback. This preliminary preparation allows for faster facial feature alignment with audio, improving user experience without excessive processing time.
Data Source
AI summary
Embodiments of the present disclosure provide a video determination method and apparatus, an electronic device, and a storage medium. The method includes: acquiring, in response to an effect trigger operation, a target facial image including a target object; determining a target audio, and determining a key video frame sequence corresponding to the target audio; determining, based on the key video frame sequence and the target facial image, a target facial feature in the target facial image that is presented when the target audio is played; and determining a target effect audio and video based on the target facial feature and the target audio.


