Scene-Based Audio Mixing for Accessible Narration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processes for creating audio descriptions (AD) are time-consuming, costly, and labor-intensive, and they struggle to maintain a non-disruptive user experience due to varying sound loudness levels across different scenes in multimedia content.
Innovation Solution
The implementation of scene-based audio mixing algorithms using machine learning and audio processing techniques, which automatically generate AD content by adjusting loudness levels of AD narration and OV audio based on scene-by-scene compression, ensuring the AD narration is audible without disrupting the original audio experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual audio mixing processes are used to create audio descriptions, then quality control is maintained, but production time and cost increase significantly
Solution Approach 1:
The patent replaces manual mechanical mixing processes with an automated digital system that uses audio analysis algorithms and machine learning models to perform scene detection, loudness measurement, and adaptive mixing operations, thereby reducing production time while maintaining quality standards
Solution Approach 2:
The system enables audio descriptions to mix themselves by automatically analyzing the original audio content, detecting scene boundaries, measuring loudness levels, and adjusting mixing parameters without requiring manual intervention at each step, significantly reducing production time
2Device complexity
If uniform loudness levels are applied across all scenes, then mixing complexity is reduced, but audio descriptions become disruptive during high-loudness scenes
Solution Approach 1:
The patent applies different loudness levels and mixing parameters to different scenes based on their specific characteristics. The system analyzes each scene's loudness profile and adjusts the audio description levels locally to match the scene requirements, preventing disruption during high-loudness scenes while maintaining visibility during low-loudness scenes
Solution Approach 2:
The system dynamically adjusts mixing parameters based on real-time analysis of scene characteristics. Rather than using static uniform levels, the audio description loudness is continuously adapted to match the original audio's loudness profile, ensuring optimal blending across varying scene intensities
3Loss of information
If audio descriptions are added to all scenes, then information completeness is improved, but the original audio experience is disrupted
Solution Approach 1:
The system changes the loudness parameter of audio descriptions based on the original audio's characteristics. By measuring the loudness of each scene and adjusting the description levels accordingly, the system ensures descriptions are audible during quiet scenes but do not overpower loud scenes, maintaining both accessibility and experience quality
4Manufacturing precision
If scene-by-scene audio analysis is performed, then mixing precision is improved, but processing complexity increases
Solution Approach 1:
The patent divides the audio content into discrete scenes and processes each scene independently. By segmenting the continuous audio stream into manageable units with distinct characteristics, the system can apply specific mixing parameters to each scene while using automated algorithms to reduce the overall processing complexity
Data Source
AI summary
The present disclosure generally relates to systems and methods for generating an AD content. In some implementation examples, an AD content system obtains and input audio and an AD narration, and normalizes a loudness of a section of the AD narration using a loudness of the input audio during a scene that the section corresponds to for generating a normalized section. Based on a loudness of the normalized section, the AD content system compresses a first audio channel of the input audio during the scene to generate a first compressed audio channel, and mix the normalized section to the first compressed audio channel during the scene to generate a first sound channel of the AD content.


