Audio Fragment Alignment for Video Streaming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content delivery systems face challenges in maintaining audio-visual synchronization, particularly when secondary content is inserted, due to non-integer frame rates and variable video fragment durations, leading to misalignment between audio and video streams.

Innovation Solution

The implementation of a video fragment aware audio packaging service that determines the number of audio frames for corresponding video fragments and generates audio fragments to ensure periodic and best alignment, even for non-integer frame rates and variable durations, by using a manifest to identify multimedia representations and adjusting audio fragment durations to match video fragment boundaries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If audio and video streams are packaged independently with standard frame rates, then processing simplicity is maintained, but audio-visual synchronization deteriorates when secondary content is inserted

Engineering Contradiction:
Improveprocessing simplicityVSAvoidaudio-visual synchronization
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments both audio and video streams into fragments that are packaged independently. Each audio fragment is associated with specific video fragments through timing information, allowing flexible combination while maintaining synchronization. This segmentation enables independent processing of audio and video while resolving the synchronization issue through precise timing metadata.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic timing adjustment mechanisms where audio fragment durations are adjusted based on the corresponding video fragment durations. The system uses可变 (variable) audio fragment lengths that adapt to non-integer video frame rates, allowing the audio stream to dynamically resynchronize with video even when secondary content is inserted, thereby maintaining audio-visual synchronization without requiring complex coordinated processing.

Inventive Principle:
Principle #15Dynamics

2Productivity

If fixed audio fragment durations are used, then packaging efficiency is improved, but synchronization precision deteriorates for non-integer frame rates

Engineering Contradiction:
Improvepackaging efficiencyVSAvoidsynchronization precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic audio fragment duration adjustment where the audio fragment duration is calculated based on the video fragment duration and video frame rate. For non-integer frame rates (e.g., 23.976 fps, 29.97 fps), the system computes precise audio fragment lengths that maintain synchronization. This dynamic adjustment preserves packaging efficiency while achieving high synchronization precision for various frame rates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the audio fragment duration parameter to match the video fragment characteristics. By calculating audio fragment duration as a function of video frame rate and video fragment length, the system adapts audio parameters to the specific video stream properties. This parameter adjustment enables precise synchronization for non-standard frame rates while maintaining efficient packaging through automated calculation.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If audio alignment is adjusted for each video fragment, then synchronization accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary calculation of audio fragment durations and timing information during the packaging phase, before playback occurs. The system pre-computes the relationship between audio and video fragments and embeds this timing metadata in the packaged output. This preliminary action eliminates the need for complex real-time adjustment during playback, reducing system complexity while maintaining high synchronization accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates feedback mechanisms where timing information from video fragments is used to adjust audio fragment packaging. The system uses timing metadata and duration information as feedback to automatically calculate appropriate audio fragment lengths and synchronization parameters. This feedback-driven approach achieves high synchronization accuracy through automated adjustment rather than complex manual coordination.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11317172B1Video fragment aware audio packaging service
Publication Date: 2022.04.26 AMAZON TECH INC
  • US11317172B1 patent drawing
  • US11317172B1 patent drawing
  • US11317172B1 patent drawing

AI summary

Techniques for video fragment aware audio packaging that ensure a periodic and best alignment of audio and video fragments at any corresponding audio and video fragments are described. As one example, a video fragment aware audio packaging service determines a number of audio frames for a corresponding video fragment of video frames and generates an audio fragment that includes those audio frames, with flexible choices of video fragment duration, which may be considered and decided for device compatibility or content encoding optimization purposes.