Personalized Audio Rendering for Clear Dialogue in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio streaming solutions limit personalization of the audio experience by relying on a single audio bitstream per streaming application, requiring users to manually switch streams and incur delays, and fail to adapt to environmental noise, affecting dialogue intelligibility.
Innovation Solution
A system and method for personalized audio streaming and rendering that creates an accessible audio mix using enhanced dialogue clarity and dynamic range compression, allowing simultaneous playback of multiple audio streams across devices, and enables real-time manual adjustment of the dialogue and non-dialogue ratio without switching streams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple pre-defined dialogue enhanced versions are provided, then users can select from available options, but users must manually switch streams which causes buffering delays and cannot adapt to changing environmental noise
Solution Approach 1:
The system dynamically adjusts the dialogue enhancement level in real-time based on environmental noise conditions without requiring stream switching. The accessible audio mix continuously adapts its dialogue-to-non-dialogue ratio according to current listening conditions, eliminating the static nature of pre-defined streams.
Solution Approach 2:
The system changes the mixing parameter (dialogue-to-non-dialogue ratio) dynamically rather than switching between discrete pre-encoded streams. By adjusting the blend ratio in real-time, the system provides continuous adaptation to environmental noise without the time penalty of stream switching and buffering.
2Adaptability or versatility
If a single audio bitstream is used per streaming application, then system complexity is reduced, but personalization of audio experience is limited
Solution Approach 1:
The system segments the audio processing into distinct functional components: the original cinematic mix, the accessible audio mix with enhanced dialogue, and the cross-fade rendering engine. This segmentation allows personalization through selective blending while maintaining a relatively simple streaming architecture.
Solution Approach 2:
The system adds a new dimension to audio delivery by providing both the original mix and an accessible mix as separate layers that can be blended. This dimensional approach enables personalization without fundamentally changing the core streaming system, as users can adjust the mix ratio rather than requiring multiple complete stream versions.
3Measurement precision
If TV devices apply post-processing to decoded audio bitstream, then dialogue clarity is improved, but integrity of original sound mix and creative intent are not preserved
Solution Approach 1:
The accessible audio mix with enhanced dialogue clarity is prepared in advance during the audio encoding stage, not as post-processing on the decoded stream. This preliminary enhancement preserves the integrity of the original creative mix while providing improved dialogue clarity in the accessible version.
Solution Approach 2:
The system introduces an intermediary accessible audio mix that serves as a bridge between the original cinematic mix and the user's listening needs. This intermediary track contains pre-enhanced dialogue while maintaining the ability to blend with the original mix, thus preserving creative intent while improving accessibility.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Systems and methods to provide personalized audio streaming and rendering include receiving a cinematic audio track for selected content and a maximum accessible audio track for the content, and adjustably combining the cinematic audio track and accessible audio track to provide an improved dialogue audio track which allows a user/listener to hear the dialogue over an environmental noise floor, the combining being provided by a cross-fade renderer adjusted manually by the user/listener or automatically adjusted based on a measured noise floor, and optionally providing a personalized equalizer for hearing impairments. It also allows a content provider to create the maximum accessible audio track for a given content separate from the user/listener, which may also use a X-Fade renderer to verify quality of the accessible audio track.