Video Audio Tagging for Active Speaker Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Active Speaker Detection (ASD) systems in videoconferencing often erroneously select microphones or cameras picking up audio or video from remote signals, leading to incorrect speaker focus and display issues, especially with high-definition TVs and sound reflections, and there is a lack of automated recording labeling.
Innovation Solution
A videoconferencing system that adds tags to outgoing audio and video signals, allowing the control system to differentiate between local and remote signals, preventing erroneous ASD and enabling accurate speaker selection and camera control, while also embedding metadata for recording purposes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Active Speaker Detection is used to select camera and microphone, then speaker focus is improved, but erroneous selection of remote signals occurs
Solution Approach 1:
The patent introduces audio tagging as an intermediary mechanism between the microphone array and ASD system. Each microphone signal is tagged with metadata indicating its source (local person vs. remote display). This intermediary tagging system allows the ASD to distinguish between genuine local speech and reflected remote audio, resolving the contradiction by providing additional discrimination information without changing the core ASD functionality.
2Measurement precision
If image scan line tracing is used to detect TV sound, then remote signal identification is improved, but it fails with high-definition progressive scan TVs
Solution Approach 1:
The patent changes the detection parameter from visual scan line analysis to audio signal tagging. Instead of relying on the display refresh rate and scan line patterns (which vary by display type), the system embeds audible tags in the remote audio signal and detects these tags in the microphone input. This parameter change makes the detection method independent of display technology, achieving both high precision and broad adaptability.
3Measurement precision
If manual recording labeling is required, then recording accuracy is improved, but human error and time loss occur
Solution Approach 1:
The patent implements self-service automation where the audio tagging system automatically generates and attaches accurate labels to recordings without human intervention. The tags embedded in the audio signal during the conference contain all necessary metadata (participants, topics, timestamps), and this same tagging infrastructure automatically populates recording metadata. This eliminates both human error and manual labeling time while maintaining high accuracy.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A videoconferencing system is described that is configured to select an active speaker while avoiding erroneously selecting a microphone or camera that is picking up audio or video from a connected remote signal. A determination is made whether an audio signal is above a threshold level. If so, then a determination is made as to whether a tag is present in that audio signal. If so, that signal is ignored. If not, a camera is directed toward the sound source identified by the audio signal. A determination is made whether a tag is present in the video signal from that camera. If so, the camera is redirected. If not, local tag(s) are inserted into the audio signal and/or the video signal. The tagged signal(s) are transmitted. Thus, system will ignore sound or video that has an embedded tag from another videoconferencing system.