Stereo Audio Time Offset Compensation via Relative Delay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Stereo audio files with temporal disparities between channels result in incorrect sound localization, as delays in one channel relative to the other can intentionally or unintentionally alter the perceived sound source location, leading to inconsistencies in audio playback.

Innovation Solution

A method to detect and correct time offsets between successive temporal windows of stereo audio signals by evaluating candidate offsets, applying relative delays, and generating output signals with reduced temporal disparities, using both envelope and sample modes of operation to optimize correlation and power analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If time offsets between stereo audio channels are not corrected, then the audio processing is simple, but sound localization accuracy deteriorates

Engineering Contradiction:
Improvesound localization accuracyVSAvoidaudio processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple temporal windows for analysis. The method processes each window separately to detect time offsets, allowing localized correction without requiring complex global processing. This segmentation enables accurate sound localization by analyzing specific time segments independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method performs preliminary detection of time offsets between stereo channels before final audio playback. By evaluating candidate offsets and determining the most likely offset value in advance, the system prepares correction data that is then applied to synchronize the channels, ensuring accurate sound localization without adding complexity during real-time playback.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If time offset detection is performed for every temporal window, then sound localization accuracy improves, but processing time increases

Engineering Contradiction:
Improvetime offset detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The method evaluates a limited set of candidate offset values (e.g., -10ms to +10ms in 1ms steps) rather than searching through all possible time shifts. This partial action approach focuses computational effort on the most likely offset range, achieving sufficient detection accuracy without exhaustive processing of every possible time window variation.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The time offset detection is performed periodically across temporal windows rather than continuously for every single sample. The method analyzes representative segments of the audio signal at regular intervals, maintaining accurate time offset detection while reducing overall processing time through this periodic sampling approach.

Inventive Principle:
Principle #19Periodic action

3Measurement precision

If candidate offset evaluation is performed exhaustively, then offset detection precision improves, but computational complexity increases

Engineering Contradiction:
Improveoffset detection precisionVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method applies different processing strategies to different parts of the offset evaluation process. For each temporal window, it locally evaluates candidate offsets based on the specific characteristics of that window's audio content, rather than applying a uniform complex algorithm throughout. This local quality approach maintains precision while reducing overall computational complexity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The method changes the parameter space by evaluating offset values in discrete steps (e.g., 1ms increments) within a bounded range, rather than attempting continuous precision measurement. This parameter discretization achieves sufficient detection precision while dramatically reducing computational complexity compared to exhaustive continuous analysis.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If temporal windows are made larger to improve offset detection reliability, then detection reliability improves, but temporal resolution deteriorates

Engineering Contradiction:
Improveoffset detection reliabilityVSAvoidtemporal resolution
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The audio signal is segmented into multiple overlapping temporal windows of moderate size. Each window provides sufficient data for reliable offset detection while maintaining temporal resolution through the overlapping structure. This segmentation allows the system to achieve both reliability and temporal resolution that would be difficult to obtain with a single large window.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method dynamically adjusts the analysis by processing multiple temporal windows with different time positions and potentially different characteristics. This dynamic approach allows the system to maintain reliable detection across varying audio conditions while preserving temporal resolution through the distributed window structure, adapting to local signal properties in each window.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11212632B2Audio processing to compensate for time offsets
Publication Date: 2021.12.28 SONY EUROPE BV
  • US11212632B2 patent drawing
  • US11212632B2 patent drawing
  • US11212632B2 patent drawing

AI summary

A method of processing each of a first plurality of temporal windows of first and second input audio signals to generate first and second output audio signals comprises (a) detecting a time offset between respective portions of the first and second input audio signals corresponding to a given temporal window by: (i) detecting a correlation between one or more properties of the respective portions according to each of a group of candidate time offsets under test; and (ii) selecting, as a detected time offset for the given temporal window, an offset for which the detecting step (i) detects a correlation which meets a predetermined criterion such as greatest correlation; and (b) for each of a second plurality of temporal windows, generating a portion of the first and second output signals by applying a relative delay between portions of the first and second input audio signals in order to correct one or both of the input audio signals to generate a pair of output audio signals (such as a stereo pair) having a reduced temporal disparity between the audio content of the two signals.