Video Conference Layouts via Face Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional videoconference user interface layouts are inconsistent due to varying aspect ratios and frame sizes from different video streams, leading to issues like black areas and inconsistent participant sizing, regardless of the number of people in a frame.
Innovation Solution
A method and system that utilize face detection information to crop and scale video frames based on crop and scale parameters, ensuring consistent frame sizes and layouts by identifying face sizes and positions, and adjusting frames to match a target frame and face region defined by a template, thereby arranging frames according to layout rules such as 'person-to-person' or 'active speaker' layouts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional user interface layouts display video frames directly from different video streams, then the implementation is simple, but the user interface layout becomes inconsistent with black areas and varying frame sizes due to different aspect ratios
Solution Approach 1:
The patent applies preliminary action by performing face detection and calculating crop and scale parameters in advance before displaying video frames. The system analyzes each video frame to identify face positions and sizes, then pre-computes the appropriate cropping and scaling parameters needed to achieve consistent display proportions. This preprocessing step ensures that when frames are displayed, they automatically maintain consistent proportions without requiring complex real-time adjustments during the display process.
2Ease of operation
If video frames are displayed with the same size regardless of content, then the display is uniform, but frames with single participants appear too large while frames with multiple participants appear too small
Solution Approach 1:
The patent applies local quality by making each video frame's display characteristics dependent on its specific content. The system performs face detection on each frame individually to count the number of participants and determine their positions. Based on this local analysis, it calculates customized crop and scale parameters for each frame, allowing frames with single participants to be displayed with appropriate size and frames with multiple participants to be scaled accordingly. This ensures each frame is optimized for its specific content while maintaining overall layout consistency through the standardized parameter application process.
3Ease of operation
If attendees are displayed without adjustment, then the display process is simple, but participants appear with inconsistent sizes due to varying distances from the camera
Solution Approach 1:
The patent applies mechanics substitution by replacing manual or simple mechanical display approaches with an automated computer vision-based face detection and measurement system. Instead of relying on fixed display parameters or manual adjustment, the system uses face detection algorithms to automatically identify faces in each video frame, measure their positions and sizes, and compute the appropriate crop and scale parameters. This automated measurement system accurately determines participant sizes regardless of their distance from the camera, enabling consistent display proportions without manual intervention.
Data Source
AI summary
A method may include obtaining a frame of a video stream of multiple video streams of a video conference, obtaining face detection information identifying a face size and a face position of at least one face detected in the frame, and cropping and scaling the frame according to at least one crop and scale parameter using the face detection information to obtain a modified first frame. The at least one crop and scale parameter is based on frames of the multiple video streams. The frames include the frame. The method may further include presenting the modified frame.


