Automatic Video Stream Selection via Speech Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Handheld wireless communication devices face challenges in transmitting multiple video streams during live video calls due to bandwidth limitations and the need for video editing capabilities, especially when users lack access to computers with video editing software.
Innovation Solution
A handheld device with at least two cameras facing opposite directions that automatically switches between video streams based on speech detection, either through sound direction or lip movement analysis, to generate a multiplexed video stream, which can be transmitted or stored for later viewing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple video streams are transmitted simultaneously during a live video call, then the user can capture multiple different video streams (e.g., self-face and people in front), but the bandwidth required exceeds the available bandwidth
Solution Approach 1:
The patent segments the multiple video streams by capturing them simultaneously with multiple cameras but then processes them separately through speech activity detection. The system divides the video content into speech-related segments and non-speech segments, selecting only relevant streams for transmission based on detected speech activity, thereby reducing overall bandwidth consumption while maintaining versatility.
Solution Approach 2:
The patent implements dynamic video stream selection where the system automatically switches between different video streams based on real-time speech activity detection. The video stream being transmitted is dynamically adjusted according to who is speaking, allowing the system to adapt to changing conditions and transmit only the most relevant video content, thus optimizing bandwidth usage.
2Ease of manufacture
If multiple video streams are uploaded to a computer for editing after the teleconference, then the user can generate a single video stream, but the user may not have access to a computer with video editing capabilities
Solution Approach 1:
The patent enables the handheld device to perform video stream processing and selection autonomously without requiring external computer resources. The device itself conducts speech activity detection, automatically switches between video streams, and generates the final multiplexed video output, making the system self-sufficient and eliminating the need for users to access computers with video editing capabilities.
Solution Approach 2:
The patent integrates multiple functions including video capture, speech activity detection, automatic video stream switching, and video multiplexing into a single handheld device. This multi-functional approach allows the device to perform what would traditionally require separate specialized equipment, making the system more accessible and easier to operate for users without access to professional video editing tools.
3Extent of automation
If the device automatically switches between video streams based on speech detection, then the video is synchronized to the speaker, but the device complexity increases
Solution Approach 1:
The patent replaces manual mechanical switching between video streams with an automated electronic system based on speech activity detection. Instead of physical switches or manual selection, the system uses sensors and processing algorithms to detect speech and automatically control the video stream selection, substituting mechanical operations with electronic automation while managing complexity through integrated design.
Data Source
AI summary
A handheld communication device is used to capture video streams and generate a multiplexed video stream. The handheld communication device has at least two cameras facing in two opposite directions. The handheld communication device receives a first video stream and a second video stream simultaneously from the two cameras. The handheld communication device detects a speech activity of a person captured in the video streams. The speech activity may be detected from direction of sound or lip movement of the person. Based on the detection, the handheld communication device automatically switches between the first video stream and the second video stream to generate a multiplexed video stream. The multiplexed video stream interleaves segments of the first video stream and segments of the second video stream. Other embodiments are also described and claimed.


