Audio Augmentation System for Source Identification in Video Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During video conference calls, participants face challenges in identifying who is speaking due to the bundling of video and audio streams, lacking auditory spatial localization, and existing visual-based solutions are inadequate, especially in large or unfamiliar groups.
Innovation Solution
An audio augmentation system that differentiates audio streams from different sources by assigning them distinct output settings, such as volume, speaker channels, or background noises, allowing the audio output device to acoustically distinguish between them, independent of the content, and optionally incorporating user commands or AI for intelligent assignment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple audio streams are bundled together for transmission, then network efficiency is improved, but the ability to identify speaking sources is worsened
Solution Approach 1:
The system segments the bundled audio streams by assigning each stream to different spatial zones or audio channels. The audio output device divides the mixed audio signal into separate directional outputs, allowing users to distinguish between different speaking sources while maintaining the efficiency of bundled transmission.
Solution Approach 2:
The system introduces an intermediary audio processing layer that sits between the bundled audio streams and the output device. This intermediary applies spatial audio processing and directional signal separation to preserve source identification information without requiring changes to the underlying bundled transmission protocol.
2Loss of information
If visual graphic indicators are used to identify speaking sources, then source identification is improved, but user attention is diverted from other visual content
Solution Approach 1:
The system replaces the visual-based identification mechanism with an acoustic-based mechanism. Instead of requiring users to visually monitor graphic indicators, the system uses spatial audio cues, directional sound output, and acoustic source separation to enable intuitive identification of speaking sources through hearing alone, freeing visual attention for other content.
3Device complexity
If audio streams are emitted without differentiation, then device complexity is reduced, but user comprehension of audio sources is worsened
Solution Approach 1:
The system applies local quality differentiation to audio streams by assigning distinct spatial characteristics, directional properties, or channel-specific features to each audio source. This allows the audio output device to emit streams with differentiated local qualities (such as directionality or spatial position) while maintaining relatively simple overall device architecture.
Data Source
AI summary
An audio augmentation system includes a memory and one or more processors that obtain a first audio stream generated by a first remote audio input device and a second audio stream generated by a second remote audio input device. The first and second audio streams are tagged with respective first and second source information. The processors assign the first audio stream to a first output setting based on the first source information, and assign the second audio stream to a different, second output setting based on the second source information. The processors control an audio output device to audibly emit the first audio stream according to the first output setting and the second audio stream according to the second output setting to acoustically differentiate the first audio stream from the second audio stream, independent of content of the first and second audio streams.


