Dynamic Audio Normalization Using Multi-Band Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video content distribution methods, particularly those using broadcast networks or IP-based networks, often result in sudden volume changes, which can be unpleasant for individuals with PTSD, hearing aids users, and those with autism, as dynamic ad insertion techniques frequently cause advertisements to be significantly louder than the content, leading to unpleasant viewing experiences that conventional manual adjustments and audio processing techniques fail to address effectively.
Innovation Solution
A system that dynamically normalizes and compresses audio in video streams by splitting audio into multiple bands based on decibel levels, determining acceptable audio ranges using national hearing data and other standards, and generating a second audio track that maintains the integrity of the audio while reducing loud noises and enhancing quiet portions, which can be selected by users during playback without requiring end-user processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If dynamic ad insertion techniques are used to insert advertisements into video content, then advertising revenue and content variety are improved, but audio volume becomes significantly louder than the rest of the content, causing unpleasant viewing experiences
Solution Approach 1:
The system applies dynamic audio processing that adjusts compression and normalization parameters in real-time based on the detected audio characteristics of each segment. The processor continuously monitors audio levels and adapts the processing intensity to match the content type, allowing advertisements and program content to be delivered at consistent perceived volumes despite their inherently different audio profiles.
Solution Approach 2:
The system changes audio processing parameters dynamically based on content analysis. It detects whether a segment is an advertisement or program content and adjusts normalization targets, compression ratios, and limiting thresholds accordingly. This allows the system to compensate for the typically higher volume characteristics of advertisements while maintaining the integrity of program content audio levels.
2Reliability
If conventional audio processing techniques are used to normalize audio levels, then overall audio consistency is improved, but all sounds are enhanced uniformly, causing quiet audio portions to be reduced to inaudible levels or causing clipping
Solution Approach 1:
The system segments the audio spectrum into multiple frequency bands and processes each band independently with tailored compression and normalization parameters. This prevents uniform processing from either clipping loud sounds or making quiet sounds inaudible, as each frequency range receives appropriate attention based on its specific characteristics and the overall audio balance requirements.
Solution Approach 2:
The system applies different processing qualities to different portions of the audio signal. Rather than uniform normalization, it identifies specific audio segments that require attention and applies targeted processing only where needed, preserving the natural dynamic range in areas that don't require intervention while correcting problematic sections.
3Ease of operation
If manual volume adjustment is used to compensate for loud advertisements, then individual user control is improved, but the solution is impractical for continuous viewing and does not address the root cause
Solution Approach 1:
The system performs automatic audio level normalization and compression without requiring user intervention. The processor continuously monitors and adjusts audio levels in real-time, compensating for loud advertisements and sudden volume changes automatically, freeing users from the need to manually adjust volume controls during viewing.
4Object-affected harmful factors
If conventional compression techniques are used to reduce audio dynamic range, then loud sounds are reduced, but clipping occurs and quiet audio portions are reduced to inaudible levels
Solution Approach 1:
The system employs dynamic compression with continuously variable parameters that adapt to the instantaneous audio signal characteristics. The compression ratio, attack time, and release time are adjusted in real-time based on the detected audio levels and content type, preventing both clipping of loud sounds and excessive reduction of quiet portions that would occur with fixed-ratio compression.
Solution Approach 2:
The system uses feedback mechanisms where the output of the compression process is continuously monitored and fed back to adjust the compression parameters. This closed-loop control prevents clipping by detecting when compression is approaching problematic levels and automatically reducing the compression ratio, while simultaneously ensuring quiet portions remain audible by adjusting the overall gain structure based on the compressed signal characteristics.
Data Source
AI summary
Methods, systems, and apparatuses are described herein for improved processing audio in a video stream. A system may split audio in a frame of video content into multiple bands based on their audio levels. The system may then dynamically compress and dynamically normalize the audio level in each band. When dynamically compressing the bands, the system may determine, based on stored information, what audio level range is acceptable for an end user and may smooth and maintain the ranges of the audio to be within the acceptable range. The system may include the dynamically normalized and dynamically compressed frames as a second audio track in the video content. A computing device receiving the video content may select the second audio track during playback. If an end user selects the second audio track, the video is delivered with the modified sound of the second audio track.


