Audio Content Recognition Using Variable Frame Shifts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio fingerprinting methods face high computational complexity and long processing times, leading to increased network resource loads and reduced matching accuracy.

Innovation Solution

An apparatus and method that vary the frame shift size between frames in a section with significant information, allowing for overlapping frames and adaptive signal-to-noise ratio-based section determination, coupled with a two-stage matching process for efficient and accurate content recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional fingerprint generation and matching methods are used, then matching accuracy can be maintained, but computational complexity is high and processing time is long

Engineering Contradiction:
Improvematching accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio signal processing into distinct stages: audio signal input, feature extraction, fingerprint generation, and matching. By dividing the complex processing task into manageable segments, the system reduces overall computational complexity while maintaining matching accuracy through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies parameter changes by adjusting fingerprint generation parameters such as frame size, hop size, and frequency resolution based on the characteristics of the audio signal. This allows the system to optimize computational complexity for different signal types while preserving matching accuracy through adaptive parameter selection.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If conventional fingerprint generation and matching methods are used, then matching accuracy can be maintained, but processing time is excessively long

Engineering Contradiction:
Improvematching accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-processing the audio signal to extract features and generate fingerprints before the actual matching operation. This preliminary fingerprint generation stores essential characteristics in advance, allowing rapid matching operations without re-processing the entire audio signal, thus reducing processing time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential features from the audio signal to create compact fingerprints, taking out only the most discriminative characteristics needed for matching. This extraction process eliminates redundant information, reducing processing time for both fingerprint generation and matching while preserving matching accuracy through selective feature extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If the number of fingerprints is decreased to reduce computational load, then processing time is reduced, but matching accuracy deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidmatching accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by generating variable numbers of fingerprints at different stages of the matching process. Initial broad matching uses fewer fingerprints for fast screening, while subsequent refined matching uses more fingerprints for accurate verification. This localized adjustment of fingerprint quantity optimizes both processing speed and matching accuracy at different stages.

Inventive Principle:
Principle #3Local quality

4Device complexity

If the matching process is simplified to reduce computational complexity, then processing time is reduced, but matching accuracy deteriorates

Engineering Contradiction:
Improvematching complexityVSAvoidmatching accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements dynamics in the matching process by adaptively adjusting the matching algorithm's complexity based on the characteristics of the audio signals being compared. For highly similar signals, more complex matching is applied to ensure accuracy, while for dissimilar signals, simpler matching suffices, reducing overall processing time while maintaining accuracy when needed.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP2685450B1Device and method for recognizing content using audio signals
Publication Date: 2020.04.22 ENSWERS CO LTD
  • EP2685450B1 patent drawingFigure 1
  • EP2685450B1 patent drawingFigure 2
  • EP2685450B1 patent drawingFigure 3

AI summary

The present invention relates to an apparatus and method for recognizing content using an audio signal. The content recognition apparatus includes a query fingerprint extraction unit for forming frames having a preset frame length for an audio signal, and generating frame-based feature vectors for respective frames, thus extracting a query fingerprint. A reference fingerprint DB stores reference fingerprints to be compared with the query fingerprint and pieces of content information corresponding to the reference fingerprints. A fingerprint matching unit determines a reference fingerprint matching the query fingerprint. In this case, the query fingerprint extraction unit forms the frames while varying a frame shift size that is an interval between start points of neighboring frames in a partial section. According to the present invention, there can be provided a content recognition apparatus and method which can maintain the accuracy and reliability of matching while promptly providing results.