Gesture Capture in Video Conferencing via Digital Overlay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems are unable to automatically capture and visualize physical gestures made by presenters during presentations or screen sharing, leading to confusion for both presenters and viewers, as these gestures are not digitally recorded or displayed.
Innovation Solution
A system and method that uses a camera to detect gestures through a trained machine learning model, generates a digital drawing of the gestures, and combines this with the video stream for visualization, allowing real-time display during video conferencing sessions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video conferencing systems use traditional video streaming without gesture capture, then the system complexity remains low, but the information completeness deteriorates as gestures are not digitally recorded
Solution Approach 1:
The patent introduces an intermediary gesture recognition system that sits between the video camera and the final video stream. This intermediary layer processes video frames to detect gestures, generates digital drawings of detected gestures, and overlays them onto the original video content, thereby capturing gesture information without fundamentally redesigning the entire video conferencing system
Solution Approach 2:
The patent creates digital copies of physical gestures by generating digital drawings that replicate the gesture patterns detected in video frames. These digital drawings are then overlaid on the video stream, preserving gesture information in a digital format that can be stored, transmitted, and reviewed alongside the video content
2Measurement precision
If the system processes every video frame to detect gestures, then the gesture detection accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent implements periodic gesture detection by processing video frames at specific intervals rather than continuously analyzing every single frame. This periodic processing approach maintains adequate gesture detection accuracy while significantly reducing computational overhead and processing time compared to frame-by-frame analysis
Solution Approach 2:
The system applies partial processing by focusing computational resources only on frames or regions where gestures are likely to occur, rather than uniformly processing all video content. This selective approach reduces overall processing time while maintaining detection accuracy for relevant gestures
3Adaptability or versatility
If the gesture visualization is displayed to all participants, then the viewer engagement improves, but the network bandwidth consumption increases
Solution Approach 1:
The patent implements local quality enhancement by selectively transmitting gesture visualization data only to participants who need or request it, rather than universally broadcasting to all participants. This allows the system to adapt to different viewer needs while optimizing network bandwidth utilization by avoiding redundant transmissions
Data Source
AI summary
A computer-implemented method and system for, using a camera, detecting a gesture during a video stream; using a computing device, generating a digital drawing that corresponds to the gesture and storing the digital drawing in a database as a gesture layer; using the computing device, combining the gesture layer with the video stream to generate a gesture visualization; and using the computing device, causing the gesture visualization to be displayed in one or more displays of one or more other computing devices.


