Scene-Based Audio Mixing for Accessible Narration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current processes for creating audio descriptions (AD) are time-consuming, costly, and labor-intensive, and they struggle to maintain a non-disruptive user experience due to varying sound loudness levels across different scenes in multimedia content.

Innovation Solution

The implementation of scene-based audio mixing algorithms using machine learning and audio processing techniques, which automatically generate AD content by adjusting loudness levels of AD narration and OV audio based on scene-by-scene compression, ensuring the AD narration is audible without disrupting the original audio experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual audio mixing processes are used to create audio descriptions, then quality control is maintained, but production time and cost increase significantly

Engineering Contradiction:
Improveaudio description qualityVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical mixing processes with an automated digital system that uses audio analysis algorithms and machine learning models to perform scene detection, loudness measurement, and adaptive mixing operations, thereby reducing production time while maintaining quality standards

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables audio descriptions to mix themselves by automatically analyzing the original audio content, detecting scene boundaries, measuring loudness levels, and adjusting mixing parameters without requiring manual intervention at each step, significantly reducing production time

Inventive Principle:
Principle #25Self-service

2Device complexity

If uniform loudness levels are applied across all scenes, then mixing complexity is reduced, but audio descriptions become disruptive during high-loudness scenes

Engineering Contradiction:
Improvemixing complexityVSAvoidaudio disruption
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The patent applies different loudness levels and mixing parameters to different scenes based on their specific characteristics. The system analyzes each scene's loudness profile and adjusts the audio description levels locally to match the scene requirements, preventing disruption during high-loudness scenes while maintaining visibility during low-loudness scenes

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts mixing parameters based on real-time analysis of scene characteristics. Rather than using static uniform levels, the audio description loudness is continuously adapted to match the original audio's loudness profile, ensuring optimal blending across varying scene intensities

Inventive Principle:
Principle #15Dynamics

3Loss of information

If audio descriptions are added to all scenes, then information completeness is improved, but the original audio experience is disrupted

Engineering Contradiction:
Improvevisual information accessibilityVSAvoidaudio experience disruption
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The system changes the loudness parameter of audio descriptions based on the original audio's characteristics. By measuring the loudness of each scene and adjusting the description levels accordingly, the system ensures descriptions are audible during quiet scenes but do not overpower loud scenes, maintaining both accessibility and experience quality

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If scene-by-scene audio analysis is performed, then mixing precision is improved, but processing complexity increases

Engineering Contradiction:
Improvemixing precisionVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the audio content into discrete scenes and processes each scene independently. By segmenting the continuous audio stream into manageable units with distinct characteristics, the system can apply specific mixing parameters to each scene while using automated algorithms to reduce the overall processing complexity

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250142139A1Scene based audio mixing for generating audio description content
Publication Date: 2025.05.01 AMAZON TECH INC
  • US20250142139A1 patent drawing
  • US20250142139A1 patent drawing
  • US20250142139A1 patent drawing

AI summary

The present disclosure generally relates to systems and methods for generating an AD content. In some implementation examples, an AD content system obtains and input audio and an AD narration, and normalizes a loudness of a section of the AD narration using a loudness of the input audio during a scene that the section corresponds to for generating a normalized section. Based on a loudness of the normalized section, the AD content system compresses a first audio channel of the input audio during the scene to generate a first compressed audio channel, and mix the normalized section to the first compressed audio channel during the scene to generate a first sound channel of the AD content.