Harmonogram Audio Matching for Noisy Live Performances

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio identification systems, such as those using audio fingerprinting, are ineffective when given short and noisy audio excerpts, particularly in live performances, as they fail to accurately identify cover versions or variations in tempo, instrumentation, and other musical aspects.

Innovation Solution

The development of a harmonogram-based fingerprinting technique that generates a two-dimensional representation of audio data, focusing on dominant frequencies and their harmonics, allowing for robust comparison of live audio performances against reference versions, even with variations in tempo, vocal timbre, and instrumentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional audio fingerprinting is used to identify audio content, then the system can identify exact matches of recorded songs, but it fails to accurately identify cover versions or live performances with variations in tempo, instrumentation, and vocal timbre

Engineering Contradiction:
Improveaudio identification accuracyVSAvoidability to handle cover versions and live performances
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the audio representation from traditional fingerprinting parameters to harmonogram parameters (dominant frequencies and their harmonics across time). This parameter transformation enables the system to capture essential musical characteristics while being tolerant of variations in tempo, instrumentation, and vocal timbre, thus resolving the contradiction between identification accuracy and adaptability to different performance types

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a new dimensional representation of audio data through the harmonogram, which organizes audio information by dominant frequencies and their harmonics across time slices. This dimensional transformation allows the system to compare live performances against reference recordings by matching harmonic structures rather than exact waveforms, enabling accurate identification of cover versions and live performances that traditional fingerprinting cannot handle

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If the system is designed to be robust to audio degradations like noise and encoding, then it can handle degraded audio quality, but it still cannot accurately identify songs from short and noisy excerpts in live performance settings

Engineering Contradiction:
Improverobustness to audio degradationVSAvoididentification accuracy from short noisy excerpts
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent extracts only the essential harmonic information (dominant frequencies and their harmonics) from the audio signal, discarding non-essential details such as specific instrumentation timbres and exact temporal positioning. This extraction of core harmonic features allows the system to maintain high identification accuracy even from short and noisy excerpts, as the essential musical identity is preserved while noise and degradation are filtered out

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

By changing from traditional fingerprinting parameters to harmonogram parameters that focus on dominant frequencies and harmonics, the system achieves both robustness to degradation and high precision in short excerpt identification. The harmonogram representation inherently filters noise while preserving the essential harmonic structure that defines the song identity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If the system compares full audio recordings for identification, then it can achieve high accuracy with clean recordings, but it becomes inoperative or inaccurate when given short and noisy excerpts

Engineering Contradiction:
Improveidentification accuracyVSAvoidaudio data length and quality
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts essential harmonic features from audio data, allowing identification to proceed with only the most critical information (dominant frequencies and harmonics). This extraction enables the system to achieve high identification accuracy even with short excerpts, as it relies on the essential harmonic identity rather than requiring complete or high-quality audio data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial fingerprinting by focusing only on dominant frequencies and their harmonics rather than analyzing the complete audio spectrum. This partial analysis approach allows accurate identification from short and noisy excerpts by concentrating computational resources on the most discriminative features while ignoring redundant or noisy information

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11366850B2Audio matching based on harmonogram
Publication Date: 2022.06.21 GRACENOTE INC
  • US11366850B2 patent drawing
  • US11366850B2 patent drawing
  • US11366850B2 patent drawing

AI summary

Apparatus, articles of manufacture, and systems for audio matching based on a harmonogram are disclosed. An example apparatus includes memory, and hardware to execute instructions to determine a first dominant frequency in a time slice of audio data based on a segment of a first spectrogram associated with the audio data, the first dominant frequency indicative of a first harmonic component of the time slice, determine a second dominant frequency indicative of a second harmonic component of the time slice, the second harmonic component less dominant than the first, generate a query harmonogram of the audio data, different segments of the query harmonogram representative of aggregate energy values of dominant frequencies in different time slices of the audio data, the dominant frequencies including at least one of the first or second dominant frequencies, and identify query sound based on a comparison of the query harmonogram to a reference harmonogram.