Adaptive Audio Stream Chunking for Multi-Track Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current adaptive media streaming systems require multiple encodings for different languages, leading to high processing resources and large data storage needs, as they typically deliver audio and video in a single track, which is inefficient and resource-intensive.
Innovation Solution
Implementing a system that encodes multiple audio tracks into a single adaptive video stream, allowing clients to request specific audio tracks using byte range requests, and separating audio content from video segments to reduce data transfer and storage requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple encodings are used for different languages, then multiple language programs can be delivered, but processing resources and data storage requirements increase significantly
Solution Approach 1:
The patent merges multiple audio track encodings into a single adaptive stream by chunking audio data and interleaving it with video segments. Instead of creating separate encoded files for each language, the system combines all audio tracks into one stream where clients can select and download only the audio chunks they need, significantly reducing server storage requirements while maintaining multi-language support.
Solution Approach 2:
The patent segments audio content into small chunks that can be independently requested and combined with video segments. By dividing the audio stream into manageable pieces that correspond to video segment boundaries, the system enables clients to selectively download only the audio chunks they need for their chosen language, rather than downloading complete separate audio files for each language option.
2Adaptability or versatility
If multiple encodings are used for different languages, then multiple language programs can be delivered, but processing resources increase
Solution Approach 1:
The system merges the encoding process into a single operation that produces an adaptive stream containing multiple audio tracks. Instead of running separate encoding processes for each language, the encoder processes all audio content once and creates an adaptive manifest that allows clients to select their preferred language, dramatically reducing CPU usage and processing time on both encoder and client sides.
Solution Approach 2:
The patent implements dynamic audio track selection where the client can choose which audio chunks to download based on their language preference. The adaptive stream structure allows the client to dynamically select only the necessary audio components rather than processing or downloading complete separate encodings, improving processing efficiency while maintaining flexibility.
3Ease of operation
If audio content is included in video segments, then single stream delivery is simplified, but data transfer volume increases
Solution Approach 1:
The patent segments audio and video content into synchronized chunks that can be independently requested. By organizing audio chunks to align with video segment boundaries and using byte range requests, the system enables clients to download only the specific audio portions they need rather than receiving complete audio tracks, reducing overall data transfer while maintaining simple stream delivery through a unified adaptive manifest.
4Adaptability or versatility
If separate audio segments are used, then audio can be removed or moved, but system complexity increases
Solution Approach 1:
The patent merges the audio selection mechanism into the existing video segment structure by using byte range requests on video segments to retrieve audio content. This approach maintains a relatively simple unified stream structure while enabling flexible audio selection, as clients can request specific byte ranges from video segments that contain their desired audio tracks, avoiding the need for completely separate audio segment management systems.
Data Source
AI summary
Systems, devices and methods are provided to support multiple audio tracks in an adaptive media stream. Segments of the adaptive stream are encoded so that the player is able to locate and request a specific one of the available audio tracks using byte range requests or the like. Audio content can be removed from video segments, or at least moved to the end of the segments so that a byte range request obtains just the video content when the default audio is not desired. The audio content can be obtained from a separate audio segment. Indeed, multiple audio tracks can be packaged into a common audio segment so that byte range requests can obtain just the particular audio track desired.


