Audio-Based Video Feed Selection for Production Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital video production systems require manual selection of video feeds, which can be labor-intensive and may not accurately choose the best feed, especially when multiple devices capture the same event.
Innovation Solution
An automated method for selecting video feeds based on predetermined switching sequences, audio or visual activity detection, or hybrid techniques, to automatically determine the most relevant video stream for production and streaming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual selection of video feeds is used, then the producer can monitor and select video feeds, but the workload of the producer increases and the selection accuracy may decrease
Solution Approach 1:
The system performs automatic video feed selection without requiring manual intervention from the producer. The video production device autonomously monitors multiple video feeds, detects audio activity, and switches between feeds based on detected speakers, thereby serving itself and eliminating the need for manual producer intervention.
Solution Approach 2:
The manual mechanical process of producer monitoring and selection is replaced with an automated electronic system that uses audio processing, signal detection, and algorithmic decision-making to select video feeds based on detected speakers and audio activity levels.
2Adaptability or versatility
If multiple video capture devices are used to capture the same event, then more video feeds are available for selection, but the complexity of selecting the best feed increases
Solution Approach 1:
The system continuously monitors audio feeds from multiple video capture devices and uses this feedback to dynamically determine which device contains the detected speaker. The audio activity detection and speaker identification processes provide real-time feedback that drives automatic switching decisions, simplifying the selection process despite multiple available feeds.
Solution Approach 2:
The system changes the selection parameter from manual producer judgment to objective audio-based metrics such as audio activity detection, speaker identification, and signal strength measurement. This parameter transformation automates the selection process and reduces complexity by using quantifiable audio characteristics rather than subjective manual evaluation.
3Productivity
If automatic video feed selection is implemented, then the producer workload is reduced, but the system complexity increases
Solution Approach 1:
The video production device is designed with multi-functionality, combining video capture, audio processing, speaker detection, and automatic switching capabilities in a single integrated system. This universal design allows the device to perform multiple functions autonomously, reducing the need for separate manual operations while managing system complexity through integration.
Data Source
AI summary
A video production device is deployed to produce a video production stream of an event occurring within an environment that includes a plurality of different video capture devices capturing respective video input streams of the event. The video production device is programmed and operated to: receive video input streams from the video capture devices; determine, for each video capture device, an average root mean square (RMS) audio energy value over a period of time, to obtain device-specific average RMS values for the video capture devices; compare each device-specific average RMS value against a respective device-specific energy threshold value; identify which input stream is associated with an active speaker, based on the comparing step; select one of the identified streams as a current video output stream; and provide the selected stream as the current video output stream.


