Facial Recognition Video Conference Bandwidth Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conference systems face issues with transmission delays and poor image quality in low bandwidth networks, especially during busy periods, due to inadequate synchronization of video and audio frames and insufficient quality settings.

Innovation Solution

A facial recognition method using UV mapping to generate 3D body models, which calculates envelope curves for lip movements and transmits dynamic calibration packets to ensure synchronized and high-quality video and audio transmission, even in low bandwidth environments, by filtering audio frequencies and detecting exceptional lip events.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If higher quality is set for the video conference system, then image quality is improved, but transmission delay increases and frame per second sequencing deteriorates

Engineering Contradiction:
Improveimage qualityVSAvoidtransmission delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the video transmission by extracting only facial regions and lip movements from the full video stream. This segmentation allows selective transmission of critical facial information at higher quality without transmitting the entire video frame, thereby reducing overall data volume and transmission delay while maintaining perceived image quality for the most important elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and transmits only the essential facial recognition elements (face geometry, lip movements) separately from the main video stream. By taking out only the critical components needed for facial recognition and synchronization, the system achieves high-quality transmission of important features while minimizing bandwidth consumption and transmission delay for the overall video conference.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of time

If lower quality is set for the video conference system, then transmission delay is reduced and frame per second sequencing is improved, but image quality deteriorates

Engineering Contradiction:
Improvetransmission delayVSAvoidimage quality
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent applies local quality by transmitting high-quality data only for critical regions (facial features, lip movements) while using lower quality or compressed data for the rest of the video stream. This allows the system to maintain acceptable overall transmission performance and frame sequencing while ensuring that the most important visual elements remain high quality for facial recognition and lip synchronization purposes.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If higher quality is set for video streaming, then image quality is improved, but bandwidth consumption increases

Engineering Contradiction:
Improveimage qualityVSAvoidbandwidth usage
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential facial recognition data (face geometry, lip movement coordinates) from the full video stream and transmits this extracted information separately. This extraction approach significantly reduces bandwidth consumption compared to transmitting the complete high-quality video stream, while still providing sufficient data for accurate facial recognition and lip synchronization.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the video data transmission into two parts: a compressed main video stream for general viewing and a separate high-precision facial feature stream for recognition and synchronization. This segmentation allows the system to optimize bandwidth usage by transmitting only the necessary high-quality data for facial analysis while keeping overall bandwidth consumption manageable.

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If video and audio frames are transmitted separately, then transmission flexibility is improved, but synchronization accuracy deteriorates

Engineering Contradiction:
Improvetransmission flexibilityVSAvoidsynchronization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the transmitted facial feature data includes timing information and synchronization markers that allow the receiving end to align video and audio frames accurately. The system uses the extracted facial movement data as a reference to synchronize audio lip movements with video facial expressions, providing feedback-based synchronization that maintains accuracy despite separate transmission of video and audio streams.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20200364918A1Facial recognition method for video conference and server using the method
Publication Date: 2020.11.19 NANNING FUGUI PRECISION IND CO LTD
  • US20200364918A1 patent drawing
  • US20200364918A1 patent drawing
  • US20200364918A1 patent drawing

AI summary

A facial recognition method for video conferencing requiring a reduced bandwidth and transmitting video and audio frames synchronously first determines whether a 3D body model of a first user at a local end has been currently retrieved or is otherwise retrievable from a historical database. Multiple audio frames of first user are collected and audio frequency at a specific range are filtered out. An envelope curve of the first audio frames and multiple attacking time periods and multiple releasing time periods of the envelope curve is calculated and correlated with lip movements of first user. Information packets of same and head-rotating and limb-swinging images of the first user are transmitted to a remote second user so that the 3D body model can simulate and show lip shapes and other movement of the first user.