Depth-Based Video Framing for Conference Room Equity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video systems capture entire conference rooms, leading to uneven head sizes and focus on participants close to the camera, causing meeting inequity and distraction from participants farther away.
Innovation Solution
Implementing depth-based framing using a machine learning model to detect heads, combine bounding boxes by depth distance, and generate head frame definitions for optimal viewing, allowing for individual or combined frames based on proximity and adjacency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If the video system captures the entire conference room, then all participants are included in the view, but participants close to the camera appear larger and receive more focus, causing meeting inequity
Solution Approach 1:
The patent segments the conference room view into multiple depth-based frames, creating separate visual groups for participants at different distances from the camera. This segmentation allows each depth group to be framed uniformly while maintaining overall coverage of all participants.
Solution Approach 2:
The patent introduces depth distance as an additional dimension for organizing the video feed. By using depth information to create separate frames for participants at different distances, the system achieves uniform framing across all participants while maintaining comprehensive coverage.
2Measurement precision
If the video system focuses on participants close to the camera, then the view is more focused and clear, but participants farther away become distracted and less engaged
Solution Approach 1:
The patent segments participants into different depth-based groups, allowing each group to receive focused attention in their respective frames. This eliminates the harm of focusing only on close participants while maintaining focus clarity for each segment.
Solution Approach 2:
The patent applies different framing qualities to different spatial regions. Participants at each depth distance receive optimized framing appropriate to their location, ensuring focus clarity for all groups while eliminating meeting inequity.
3Area of stationary object
If the video system includes all participants in a single frame, then comprehensive coverage is achieved, but uneven head sizes create visual distraction
Solution Approach 1:
The patent segments the single comprehensive frame into multiple depth-based frames. This segmentation eliminates visual distraction from uneven head sizes while preserving comprehensive coverage of all participants across the different frames.
Solution Approach 2:
The patent uses depth distance as an organizing dimension to create separate frames for participants at different distances. This approach maintains comprehensive coverage while eliminating the visual distraction of uneven head sizes by grouping participants of similar size together.
Data Source
AI summary
A method may include obtaining, using a head detection model and for an image of a video stream, head detection information for heads detected in the image. The head detection information may include depth distances of the heads. Method may also include obtaining bounding boxes, where obtaining the bounding boxes may include obtaining head bounding boxes for the heads detected in the image, and combining at least two of the head bounding boxes into a combined bounding box according to the depth distances. Method may furthermore include creating, individually, head frame definitions for the bounding boxes, and processing the video stream using the head frame definitions.


