Facial Feature Video Sync for Audio-Triggered Effects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video applications with artificial intelligence-based human-computer interaction suffer from limited effect props, poor video quality, and synchronization issues between video content and selected effects, leading to a poor user experience.

Innovation Solution

A method and apparatus that determine a target facial feature in a facial image to align with the facial feature presented when a target audio is played, by acquiring a target facial image and audio, determining a key video frame sequence, and generating a target effect audio and video based on the facial feature and audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If effect props are added to video applications, then video content richness is improved, but device computing power requirements increase

Engineering Contradiction:
Improvevideo content richnessVSAvoidcomputing power requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the video processing into key frame extraction and non-key frame skipping. By identifying and processing only key frames that contain important facial features, the system reduces the computational load while maintaining video content richness. This segmentation allows effect props to be applied selectively rather than to every frame, resolving the contradiction between content richness and computing power requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by focusing computational resources on specific regions (facial features in key frames) rather than processing the entire video uniformly. The system identifies facial regions in key frames and applies effect props locally to these areas, improving video content richness while minimizing the overall computing power consumption.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If facial feature alignment with audio is implemented, then user experience is improved, but processing time increases

Engineering Contradiction:
Improveuser experienceVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent applies preliminary action by extracting and storing key frames with facial features before the actual audio synchronization process. By pre-identifying and preparing the key frames that contain relevant facial information, the system reduces the processing time required during audio playback. This preliminary preparation allows for faster facial feature alignment with audio, improving user experience without excessive processing time.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260046469A1Video determination method and apparatus, electronic device and storage medium
Publication Date: 2026.02.12 LEMON INC(GB)
  • US20260046469A1 patent drawing
  • US20260046469A1 patent drawing
  • US20260046469A1 patent drawing

AI summary

Embodiments of the present disclosure provide a video determination method and apparatus, an electronic device, and a storage medium. The method includes: acquiring, in response to an effect trigger operation, a target facial image including a target object; determining a target audio, and determining a key video frame sequence corresponding to the target audio; determining, based on the key video frame sequence and the target facial image, a target facial feature in the target facial image that is presented when the target audio is played; and determining a target effect audio and video based on the target facial feature and the target audio.