Video Conferencing Framing Preview with Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing systems require repetitive user inputs or complex image and audio analysis for adjusting camera views, lacking a simplified method to choose and display a customized view of a video conferencing site.
Innovation Solution
A method that involves receiving an image stream, detecting objects, displaying framing previews to include detected objects, and allowing users to select a relevant preview to adjust the camera settings, enabling easy customization of the video conferencing view.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional camera adjustment methods are used (manual remote control or complex image/audio analysis), then the camera view can be adjusted, but the system requires repetitive user inputs or complex analysis, increasing operational complexity and time consumption
Solution Approach 1:
The system automatically detects objects in the video stream and generates framing previews without requiring user intervention. The camera controller autonomously analyzes the video content, identifies objects of interest, and creates multiple framed previews showing different compositions, allowing the system to serve itself rather than requiring repetitive manual adjustments or complex user analysis
Solution Approach 2:
The system performs preliminary object detection and framing preview generation before the user makes a final selection. By pre-processing the video stream to identify objects and create multiple framed options in advance, the system reduces the complexity of real-time decision-making and simplifies the user interaction to merely selecting from pre-prepared options
2Extent of automation
If complex image and audio analysis is performed to automatically adjust the camera, then automation is improved, but the computational complexity and processing requirements increase significantly
Solution Approach 1:
The system extracts only the essential information needed for framing - object detection and positioning - from the video stream, rather than performing comprehensive complex image and audio analysis. By taking out only the necessary object detection function and using it to generate framing previews, the system achieves practical automation without the full computational burden of complete scene understanding
Solution Approach 2:
The system performs partial automation by detecting objects and generating framing previews, but stops short of fully automatic camera control. This partial action approach provides sufficient automation to eliminate repetitive manual adjustments while keeping computational complexity manageable by not implementing complete autonomous decision-making
3Adaptability or versatility
If multiple framing previews are displayed for user selection, then the ability to choose customized views is improved, but the information processing and display complexity increases
Solution Approach 1:
The system segments the video stream into multiple framed previews, each highlighting different objects or compositions. By dividing the single video feed into several discrete framed options, the system makes it easier for users to compare and select their preferred view, transforming a complex continuous adjustment task into a simpler discrete selection task
Data Source
AI summary
In one embodiment, a method includes: receiving an image stream captured by an image capture device associated with one of a plurality of video conferencing endpoints of a video conferencing system; receiving a request to detect objects in the received image stream; upon detecting one or more objects, displaying a first framing preview of the received image stream, wherein the first framing preview is framed to include the detected one or more objects; upon detecting a change in the detected one or more objects, displaying at least one second framing preview of the received image stream, wherein the at least one second framing preview is framed to include the detected change in the detected one or more objects; and receiving an input, the input selecting a relevant framing preview to use, wherein the relevant framing preview is one of the displayed first and at least one second framing previews.


