Multi-Band Audio Normalization for Sudden Volume Spikes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video content distribution technologies, particularly those using broadcast networks or IP-based networks, often result in sudden volume changes, which can be unpleasant for individuals with PTSD, hearing aids users, and those with autism, as dynamic ad insertion techniques frequently cause advertisements to be significantly louder than the content, leading to unpleasant viewing experiences that conventional manual adjustments and audio processing techniques fail to address effectively.
Innovation Solution
A system that dynamically normalizes and compresses audio in video streams by splitting audio into multiple bands based on decibel levels, determining acceptable audio ranges using national hearing data and other standards, and generating a second audio track that maintains the integrity of the audio while reducing loud noises and enhancing quiet portions, which can be selectively played back by users without requiring end-user processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If dynamic ad insertion techniques are used to insert advertisements into video content, then advertising revenue and content distribution efficiency are improved, but sudden volume changes occur that cause unpleasant viewing experiences for users with PTSD, hearing aids, autism, and other disabilities
Solution Approach 1:
The system performs preliminary audio analysis and normalization during the video encoding process, preparing normalized audio tracks in advance before distribution. This allows the content to be delivered with pre-processed audio that mitigates volume issues, rather than requiring real-time processing at the user end.
Solution Approach 2:
The patent introduces an intermediary audio processing system that creates multiple audio tracks including normalized versions. This intermediary layer between content creation and user consumption enables volume normalization without affecting the original content or requiring user-side processing.
2Stability of the object's composition
If conventional audio processing techniques are used to normalize audio levels, then audio consistency is improved, but all sounds are enhanced which may cause clipping and reduce quiet audio portions to inaudible levels
Solution Approach 1:
The system segments the audio signal into multiple frequency bands and processes each band separately with different normalization parameters. This allows selective adjustment of audio levels across different frequency ranges, preventing uniform enhancement that causes clipping while maintaining audio quality and intelligibility.
Solution Approach 2:
Different normalization strategies are applied to different portions of the audio signal based on their characteristics. Quiet portions receive greater gain while loud portions receive less gain, with dynamic range compression applied selectively to prevent clipping. This local quality approach preserves audio quality while achieving consistency.
3Ease of operation
If manual volume adjustment is used to compensate for volume changes, then user control is improved, but the complexity of operation increases and does not effectively mitigate volume changes during dynamic ad insertion
Solution Approach 1:
The system provides self-service audio normalization by automatically processing audio tracks during encoding and delivery. Multiple normalized audio tracks are included in the distributed content, allowing users to select pre-processed tracks that meet their needs without requiring manual adjustment or complex device operations.
Data Source
AI summary
Methods, systems, and apparatuses are described herein for improved processing audio in a video stream. A system may split audio in a frame of video content into multiple bands based on their audio levels. The system may then dynamically compress and dynamically normalize the audio level in each band. When dynamically compressing the bands, the system may determine, based on stored information, what audio level range is acceptable for an end user and may smooth and maintain the ranges of the audio to be within the acceptable range. The system may include the dynamically normalized and dynamically compressed frames as a second audio track in the video content. A computing device receiving the video content may select the second audio track during playback. If an end user selects the second audio track, the video is delivered with the modified sound of the second audio track.


