Facial Recognition Video Conference Bandwidth Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conference systems face issues with transmission delays and poor image quality in low bandwidth networks, especially during busy periods, due to inadequate synchronization of video and audio frames and insufficient quality settings.
Innovation Solution
A facial recognition method using UV mapping to generate 3D body models, which calculates envelope curves for lip movements and transmits dynamic calibration packets to ensure synchronized and high-quality video and audio transmission, even in low bandwidth environments, by filtering audio frequencies and detecting exceptional lip events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If higher quality is set for the video conference system, then image quality is improved, but transmission delay increases and frame per second sequencing deteriorates
Solution Approach 1:
The patent segments the video transmission by extracting only facial regions and lip movements from the full video stream. This segmentation allows selective transmission of critical facial information at higher quality without transmitting the entire video frame, thereby reducing overall data volume and transmission delay while maintaining perceived image quality for the most important elements.
Solution Approach 2:
The patent extracts and transmits only the essential facial recognition elements (face geometry, lip movements) separately from the main video stream. By taking out only the critical components needed for facial recognition and synchronization, the system achieves high-quality transmission of important features while minimizing bandwidth consumption and transmission delay for the overall video conference.
2Loss of time
If lower quality is set for the video conference system, then transmission delay is reduced and frame per second sequencing is improved, but image quality deteriorates
Solution Approach 1:
The patent applies local quality by transmitting high-quality data only for critical regions (facial features, lip movements) while using lower quality or compressed data for the rest of the video stream. This allows the system to maintain acceptable overall transmission performance and frame sequencing while ensuring that the most important visual elements remain high quality for facial recognition and lip synchronization purposes.
3Manufacturing precision
If higher quality is set for video streaming, then image quality is improved, but bandwidth consumption increases
Solution Approach 1:
The patent extracts only the essential facial recognition data (face geometry, lip movement coordinates) from the full video stream and transmits this extracted information separately. This extraction approach significantly reduces bandwidth consumption compared to transmitting the complete high-quality video stream, while still providing sufficient data for accurate facial recognition and lip synchronization.
Solution Approach 2:
The patent segments the video data transmission into two parts: a compressed main video stream for general viewing and a separate high-precision facial feature stream for recognition and synchronization. This segmentation allows the system to optimize bandwidth usage by transmitting only the necessary high-quality data for facial analysis while keeping overall bandwidth consumption manageable.
4Adaptability or versatility
If video and audio frames are transmitted separately, then transmission flexibility is improved, but synchronization accuracy deteriorates
Solution Approach 1:
The patent implements feedback mechanisms where the transmitted facial feature data includes timing information and synchronization markers that allow the receiving end to align video and audio frames accurately. The system uses the extracted facial movement data as a reference to synchronize audio lip movements with video facial expressions, providing feedback-based synchronization that maintains accuracy despite separate transmission of video and audio streams.
Data Source
AI summary
A facial recognition method for video conferencing requiring a reduced bandwidth and transmitting video and audio frames synchronously first determines whether a 3D body model of a first user at a local end has been currently retrieved or is otherwise retrievable from a historical database. Multiple audio frames of first user are collected and audio frequency at a specific range are filtered out. An envelope curve of the first audio frames and multiple attacking time periods and multiple releasing time periods of the envelope curve is calculated and correlated with lip movements of first user. Information packets of same and head-rotating and limb-swinging images of the first user are transmitted to a remote second user so that the 3D body model can simulate and show lip shapes and other movement of the first user.


