Audio Delay Determination via Sub-Fingerprint Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio delay compensation methods in karaoke systems, such as those based on time domain prediction, suffer from poor antinoise performance, leading to inaccurate delay predictions and unsatisfactory synchronization between singing sounds and accompaniments, resulting in a dual sound phenomenon.

Innovation Solution

A method and device that determine audio delay by extracting sub-fingerprint sequences from both the accompaniment and singing sounds, calculating similarities through relative shifting operations, and determining a matching degree to accurately compensate for delays, thereby improving synchronization and reducing the dual sound effect.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If time domain prediction methods (energy method, autocorrelation method, or contour method) are used for delay compensation, then the singing sound can be synchronized with the accompaniment to some extent, but the antinoise performance is poor leading to inaccurate delay prediction

Engineering Contradiction:
Improvedelay prediction accuracyVSAvoidantinoise performance
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces time domain prediction methods with frequency domain analysis using audio fingerprints. Instead of analyzing temporal patterns directly, the system transforms audio signals into frequency domain representations and compares spectral characteristics, thereby avoiding the poor antinoise performance of time domain methods while achieving accurate delay prediction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the analysis parameters from time domain features (energy, autocorrelation, contour) to frequency domain features (spectral fingerprints, frequency bin patterns). This parameter transformation enables the system to achieve both high measurement precision and reliability by comparing frequency spectral patterns that are more robust to noise.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If delay compensation is performed using conventional methods, then the singing sound can be aligned with the accompaniment, but the dual sound phenomenon occurs due to inaccurate delay prediction

Engineering Contradiction:
Improvesynchronization accuracyVSAvoiddual sound phenomenon
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The patent substitutes conventional time domain delay compensation with frequency domain fingerprint matching. By comparing spectral patterns and identifying the time shift that maximizes fingerprint similarity, the system achieves precise synchronization without the dual sound phenomenon that plagues conventional methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements a feedback mechanism where the delay compensation result is continuously refined through iterative fingerprint comparison. The system calculates similarities at different delay offsets and uses this feedback to identify the optimal delay value that eliminates dual sound while maintaining synchronization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3493198B1Method and device for determining delay of audio
Publication Date: 2022.11.16 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3493198B1 patent drawingFigure 1~2
  • EP3493198B1 patent drawingFigure 3(a)~5
  • EP3493198B1 patent drawingFigure 6

AI summary

A method and device for determining the delay of an audio. The method comprises: obtaining inputted audios to be adjusted, which are a first audio and a second audio (101); extracting a first sub-fingerprint sequence of the first audio and a second sub-fingerprint sequence of the second audio (102), the first sub-fingerprint sequence comprising at least one first sub-fingerprint, and the second sub-fingerprint sequence comprising at least one second sub-fingerprint; determining multiple similarities between the first sub-fingerprint sequence and the second sub-fingerprint sequence (103); determining the matching degree between the first sub-fingerprint sequence and the second sub-fingerprint sequence according to the multiple similarities (104); and determining the delay of the second audio relative to the first audio according to the matching degree (105). By means of the method and the device, the computing precision can be improved, and accordingly the effect of delay compensation can be improved and the phenomena of sound overlapping can be reduced.