Pitch-Preserving Audio Tracks for Lip-Sync During Video Rate Changes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio-video stream synchronization technologies experience delays during stream transitions, such as channel changes, leading to desynchronization of audio and video, which affects the user experience.
Innovation Solution
The implementation of audio-video pacing systems that generate and deliver pitch-preserving audio tracks synchronized with video streams, allowing for buffer expansion or contraction while maintaining synchronized playback, by slowing down the video decoding rate to match the pitch-preserving audio track, thereby eliminating initial delays and ensuring uninterrupted, lip-synchronized playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional A/V stream synchronization is used during channel changes, then the system can maintain simple architecture, but audio and video become desynchronized causing lip-sync issues and user experience degradation
Solution Approach 1:
The system pre-generates and stores pitch-preserving audio tracks at multiple playback rates before they are needed. When a channel change occurs, the appropriate pre-generated audio track is immediately selected and played back at the required rate, eliminating the need for real-time pitch transformation and ensuring immediate audio-video synchronization without adding complex real-time processing infrastructure
Solution Approach 2:
The patent introduces an intermediary component (audio processing system) that sits between the video stream and the audio output. This intermediary receives video streams, generates corresponding pitch-preserving audio tracks at multiple rates, and selectively plays them back to match the video playback rate. This intermediary handles all the complexity of pitch preservation and rate matching, leaving the rest of the system simple while ensuring reliable synchronization
2Reliability
If the video decoding rate is slowed down to match pitch-preserving audio tracks, then audio-video synchronization is achieved without delays, but the video playback duration is extended
Solution Approach 1:
The system changes the playback rate parameter of pre-generated audio tracks to match the video decoding rate. By having audio tracks generated at multiple different rates in advance, the system can select and play back the audio track at the exact rate needed to synchronize with the video, whether that rate is normal speed or slowed down, without extending the overall playback duration beyond what is necessary for synchronization
Solution Approach 2:
The system dynamically adjusts the audio playback rate to match the video decoding rate in real-time. By having multiple pre-generated audio tracks at different rates, the system can dynamically select and switch between them based on the current video playback speed, allowing the video to be played back at any speed while maintaining synchronization without permanently extending the playback duration
3Reliability
If buffer expansion is used to accommodate pitch-preserving audio tracks, then synchronized playback is maintained during transitions, but initial buffer size requirements increase
Solution Approach 1:
The system pre-generates and pre-loads pitch-preserving audio tracks into buffers before they are needed for playback. By preparing these audio tracks in advance and storing them in buffers, the system ensures that when a channel change or video rate adjustment occurs, the synchronized audio is already available in the buffer and can be played back immediately without requiring large real-time buffer expansions
Solution Approach 2:
The system uses pre-generated audio tracks as a cushion or backup that is prepared in advance and stored in buffers. This beforehand preparation creates a cushion of pre-processed audio data that can be drawn from during transitions or rate changes, eliminating the need for large real-time buffer expansions and ensuring synchronized playback without increasing initial buffer size requirements
Data Source
AI summary
In one method embodiment, providing a multiplex of compressed versions of a first video stream and a first audio stream, each corresponding to an audiovisual (A/V) program, the first video stream and the first audio stream each corresponding to a first playout rate and un-synchronized with each other for an initial playout portion; and providing a compressed version of a second audio stream, the second audio stream corresponding to a pitch-preserving, second playout rate different than the first playout rate, the second audio stream synchronized to the initial playout portion of the first video stream when the first video stream is played out at the second playout rate, the first audio stream replaceable by the second audio stream for the initial playout portion.


