Intervalgram Audio Representation for Key-Invariant Melody Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated song recognition systems struggle to identify songs across different recordings, key transpositions, and instrumentation changes, making them inefficient for large media libraries.
Innovation Solution
The system generates an intervalgram representation of audio clips, which represents pitch intervals over time, allowing for robust recognition of melodies across variations, and creates a database of reference fingerprints for matching against input clips.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional exact match recognition is used, then precision of song identification is improved, but adaptability to different recordings and key transpositions deteriorates
Solution Approach 1:
The patent transforms the audio representation from raw waveform or spectrum to pitch intervals extracted from chroma features. This parameter transformation enables the system to represent melodies in a key-invariant manner, allowing the same melody to be recognized across different key transpositions while maintaining identification precision
Solution Approach 2:
The patent introduces a new dimensional representation by computing pitch intervals between consecutive chroma vectors over time. This temporal-difference dimension captures melodic contours independently of absolute pitch, enabling recognition across different recordings and instrumentation while preserving identification accuracy
2Productivity
If automated song recognition is implemented, then productivity of media library management is improved, but reliability of recognition across variations deteriorates
Solution Approach 1:
The patent extracts only the essential melodic information by computing pitch intervals from chroma features, discarding irrelevant information such as absolute pitch, timbre, and tempo. This extraction approach enables automated processing of large media libraries while maintaining reliable recognition across variations in recording, instrumentation, and performance
Solution Approach 2:
The patent segments the audio signal into sequential chroma frames and computes pitch intervals between consecutive frames. This temporal segmentation captures the melodic contour as a sequence of interval transitions, enabling reliable automated recognition across different recordings while maintaining high processing productivity
Data Source
AI summary
A system, method, and computer readable storage medium generates an audio fingerprint for an input audio clip that is robust to differences in key, instrumentation, and other performance variations. The audio fingerprint comprises a sequence of intervalgrams that represent a melody in an audio clip according pitch intervals between different time points in the audio clip. The fingerprint for an input audio clip can be compared to a set of reference fingerprints in a reference database to determine a matching reference audio clip.


