Audio Video Sync Detection via ML Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Maintaining synchronization between audio and video streams in media content delivery is challenging, as consumers experience dissatisfaction due to excessive lag or lead, necessitating techniques to identify and correct audio-visual desynchronization.

Innovation Solution

The use of machine learning models to determine correlations between video frames and audio bins by generating confidence scores, allowing for the adjustment of either the audio or video components to improve synchronization, with methods including shifting frames and bins relative to their common media timeline.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional audio-video synchronization methods are used, then the system complexity is low, but the synchronization precision deteriorates leading to consumer dissatisfaction

Engineering Contradiction:
Improvesynchronization precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces machine learning models as intermediary components that analyze correlations between audio and video streams. These models act as mediators that automatically detect desynchronization and generate correction instructions, resolving the contradiction by providing high-precision synchronization through intelligent intermediaries rather than simple time-stamp matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical synchronization methods (based on fixed time codes and manual adjustment) with machine learning-based correlation analysis. This substitution enables adaptive, high-precision synchronization that automatically adjusts to varying content characteristics, overcoming the limitations of rigid mechanical approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If machine learning models are used to detect desynchronization, then the synchronization accuracy is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improvedesynchronization detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies machine learning models to generate confidence scores that predict potential desynchronization issues before they significantly impact playback. By performing preliminary analysis on audio-video correlations, the system can proactively adjust synchronization parameters, reducing the need for real-time computational overhead during actual playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The machine learning models are designed to autonomously analyze audio-video correlations and self-correct synchronization issues without requiring continuous external intervention. The models learn from patterns in the data and automatically adjust timing parameters, reducing the need for complex external control systems and minimizing processing delays.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11659217B1Event based audio-video sync detection
Publication Date: 2023.05.23 AMAZON TECH INC
  • US11659217B1 patent drawing
  • US11659217B1 patent drawing
  • US11659217B1 patent drawing

AI summary

Techniques are described for detecting desynchronization between an audio component and a video component of a media presentation. Feature sets may be determined for portions of the audio component and portions of the video component, which may then be used to generate correlations between portions of the audio component and portions of the video component. Synchronization may then be assessed based on the correlations.