Synchronized Audience Audio Mixing for Mass Media Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer applications face challenges in effectively displaying and navigating large amounts of media information, as they often rely on summaries or small images that may not provide sufficient detail, leading to a non-intuitive user experience, especially when trying to convey information like product details or reactions to media.
Innovation Solution
A computer-implemented method that synchronizes audio reactions with mass media presentations by mixing audio from client devices, using a mixer server to create a crowd effect, and organizing clients into virtual rooms for volume control and spatialization, allowing for intuitive and natural interaction with media content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If larger images or videos are used to provide sufficient information, then information completeness is improved, but device complexity and navigation difficulty worsen
Solution Approach 1:
The patent segments the audience reaction into multiple independent audio channels (e.g., applause, laughter, cheers) that can be independently controlled and mixed. This allows the system to provide comprehensive reaction information without requiring a single large complex media file, as each reaction type can be independently managed and combined.
Solution Approach 2:
The patent adds the temporal dimension to audio reactions by synchronizing them with specific timestamps in the video content. This allows reactions to be associated with particular moments in the media without requiring larger spatial representation, effectively adding information density through time-based organization rather than increasing media size.
2Adaptability or versatility
If more audio reactions are mixed from more client devices, then audience engagement is improved, but loss of useful information worsens due to echo and noise
Solution Approach 1:
The patent extracts and removes echo components from the captured audio reactions using echo cancellation technology. By identifying and removing the harmful echo portion while preserving the original audience reaction, the system maintains high engagement levels without the degradation of useful information that would otherwise occur from mixing multiple client device inputs.
Solution Approach 2:
The patent converts the potentially harmful effect of echo and noise from multiple client devices into a benefit by using these signals to identify and isolate the true audience reaction. The system uses the mixed audio input to detect reaction patterns, then enhances and cleanses the signal, turning the noise problem into an opportunity for more accurate reaction detection.
3Adaptability or versatility
If synthesized sounds are added to audio reactions, then audience effect is improved, but device complexity worsens
Solution Approach 1:
The patent merges real captured audio reactions with synthesized sound effects into a unified audio output stream. By combining these different audio sources through mixing and spatialization techniques, the system creates an enhanced audience effect without requiring separate complex processing systems, as both real and synthesized sounds are integrated into the same audio pipeline.
Data Source
AI summary
Systems and methods of the present disclosure provide a plurality of audio reactions from a plurality of client devices. The audio reactions are captured by microphones on the client devices and are time-stamped. The method also includes mixing the audio reactions by a mixer server to form a mixed audio reaction, and sending the mixed audio reaction to at least one of the client devices. The client device is adapted to play the mixed audio reaction and a mass media presentation. The mixed audio reaction and the mass media presentation are synchronized to create an audience effect for the mass media presentation. The present technology also provides echo removal, volume balancing, compression, and time stamping of an audio stream by the client device. Reactions from at least one of buttons and gestures to activate synthesized sounds, for example clapping, booing, and cheering, which are mixed into the mixed audio reaction.


