Audio Video Synchronization Perceptual Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional audio and video synchronization techniques fail to consider the emotional impact on listeners and are time-consuming and computationally intensive, relying heavily on user input.

Innovation Solution

An audio and video synchronizing perceptual model that identifies perceptual characteristics of audio data indicative of emotional impact, allowing for automatic synchronization of audio and video to achieve a specific emotional effect by determining transition points based on relative emotional impact assessment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional audio and video synchronization techniques are used, then synchronization can be achieved, but the process is time-consuming and computationally intensive

Engineering Contradiction:
Improvesynchronization efficiencyVSAvoidsynchronization time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical/audio-based synchronization methods with a perceptual model that simulates human brain processing. This model uses perceptual characteristics (temporal envelope, spectral flux, loudness) to automatically identify emotionally significant moments, eliminating the need for time-consuming manual analysis and complex computational algorithms while achieving faster, more intuitive synchronization.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If traditional synchronization techniques are used, then technical analysis can be performed, but user input is heavily required

Engineering Contradiction:
Improveautomatic synchronizationVSAvoiduser input requirement
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The perceptual model operates autonomously by automatically extracting perceptual characteristics from audio and video signals, identifying emotionally significant moments, and determining synchronization points without user intervention. The system serves itself by using built-in algorithms to analyze temporal envelope, spectral flux, and loudness characteristics, eliminating the need for users to manually mark or select synchronization points.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a perceptual model as an intermediary between raw audio/video signals and synchronization output. This model acts as a mediator that processes signals through simulated human perception mechanisms, translating technical audio characteristics into emotionally meaningful synchronization decisions without requiring direct user input or complex manual operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If traditional audio analysis is used, then sound characteristics can be analyzed, but emotional impact on listeners is not considered

Engineering Contradiction:
Improveemotional impact informationVSAvoidemotional impact measurement
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent transforms traditional audio analysis parameters into perceptual characteristics that reflect human emotional response. By computing temporal envelope (amplitude modulation), spectral flux (frequency changes), and loudness (perceived intensity) with specific weighting and time-windowing, the system converts raw audio data into emotionally meaningful metrics that capture nostalgic, dramatic, or intense moments, enabling precise measurement of emotional impact.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10559323B2Audio and video synchronizing perceptual model
Publication Date: 2020.02.11 ADOBE INC
  • US10559323B2 patent drawing
  • US10559323B2 patent drawing
  • US10559323B2 patent drawing

AI summary

An audio and video synchronizing perceptual model is described that is based on how a person perceives audio and/or video (e.g., how the brain processes sound and/or visual content). The relative emotional impact associated with different audio portions may be employed to determine transition points to facilitate automatic synchronization of audio data to video data to create a production that achieves a particular overall emotional effect on the listener/viewer. Various processing techniques of the perceptual model may utilize perceptual characteristics within the audio portions to determine a transition point for automatic synchronization with video data.