Audio Bitstream Synchronization for High Frame Rate Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio and video frame rates in commercial applications are not harmonized, leading to synchronization issues and metadata loss during audiovisual data stream processing, especially when dealing with high frame rates and legacy infrastructure.

Innovation Solution

An audio bitstream format is proposed that allows for synchronization with video frames by encoding audio signals into bitstream frames with a higher frame rate, maintaining audio-visual synchronicity without increasing bitrate consumption, using signal analysis with a basic stride and forming N bitstream frames from a decodable set of audio data, which can be reconstructed efficiently on the decoder side.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio frames are decoded into baseband format for synchronization, then synchronization flexibility is improved, but metadata integrity is lost

Engineering Contradiction:
Improvesynchronization flexibilityVSAvoidmetadata integrity
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The audio signal is divided into multiple segments that can be independently processed and synchronized. Each segment maintains its metadata intact while allowing flexible synchronization operations. This segmentation enables synchronization without requiring full decode-encode cycles that would destroy metadata.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A synchronization mechanism is introduced as an intermediary between the audio frames and the synchronization process. This intermediary allows audio frames to be synchronized with video frames without decoding them into baseband format, thus preserving metadata while achieving synchronization flexibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If audio frame rate is increased to match high video frame rates, then audio-visual synchronicity is improved, but bitrate consumption increases

Engineering Contradiction:
Improveaudio-visual synchronicityVSAvoidbitrate consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The audio frame rate parameter is changed to match high video frame rates (e.g., 60fps or 120fps) while maintaining efficient encoding. The system adjusts the frame rate parameter without proportionally increasing bitrate, achieving high synchronicity with controlled bandwidth usage through optimized encoding parameters.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If audio frames are duplicated or dropped for synchronization, then video-to-video synchronicity is improved, but audio-to-video lag increases

Engineering Contradiction:
Improvevideo-to-video synchronicityVSAvoidaudio-to-video lag
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary alignment of audio and video frames before synchronization operations. By pre-aligning the frames and using frame-accurate positioning, the system avoids the need for post-synchronization lag compensation, achieving synchronicity without introducing audio-to-video lag.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10304471B2Encoding and decoding of audio signals
Publication Date: 2019.05.28 DOLBY INTERNATIONAL AB
  • US10304471B2 patent drawing
  • US10304471B2 patent drawing
  • US10304471B2 patent drawing

AI summary

An audio signal (X) is represented by a bitstream (B) segmented into frames. An audio processing system (500) comprises a buffer (510) and a decoding section (520). The buffer joins sets of audio data (D1; D2, . . . , DN) carried by N respective frames (F1, F2, . . . , FN) into one decodable set of audio data (D) corresponding to a first frame rate and to a first number of samples of the audio signal per frame. The frames have a second frame rate corresponding to a second number of samples of the audio signal per frame. The first number of samples is N times the second number of samples. The decoding section decodes the decodable set of audio data into a segment of the audio signal by at least employing signal synthesis, based on the decodable set of audio data, with a stride corresponding to the first number of samples of the audio signal.