Audio Key Estimation Using Learned Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques face challenges in accurately estimating the key of musical pieces, especially when the frequency of the key note is low or when the power of each pitch name is not correlated with the tonality type, leading to difficulties in identifying the key for various types of musical pieces.

Innovation Solution

An audio analysis method that uses a learned model to generate key information by inputting a time series of feature amounts from audio signals, which has learned the relationship between keys and time series of feature amounts, allowing for accurate key estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If rule-based methods are used to estimate key from audio signals, then the system complexity is low, but the key estimation accuracy deteriorates for various types of musical pieces

Engineering Contradiction:
Improvekey estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the key estimation problem from a rule-based approach to a parameter-based statistical model approach. By changing the fundamental parameter from simple frequency counting to statistical distribution analysis of pitch names, the system achieves higher accuracy across diverse musical pieces while managing complexity through probabilistic modeling.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical rule-based system with a statistical estimation model. Instead of applying fixed rules for key determination, the system uses statistical analysis of pitch name distributions and their relationships with tonality types, substituting deterministic mechanics with probabilistic reasoning to improve accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If the frequency of the key note is used to determine the key, then the method is simple, but the key estimation accuracy deteriorates when the key note frequency is low

Engineering Contradiction:
Improvekey estimation accuracyVSAvoidmethod simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent creates a universal statistical model that handles multiple scenarios simultaneously - it works whether the key note frequency is high or low, and adapts to different tonality types. The model's multi-functionality allows it to process various musical piece types without requiring separate methods for different frequency conditions, maintaining both accuracy and reasonable simplicity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If the power of each pitch name is used to identify tonality type, then the method is straightforward, but the key estimation accuracy deteriorates when power and tonality type are not correlated

Engineering Contradiction:
Improvekey estimation accuracyVSAvoidanalysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary statistical relationship between pitch names and tonality types. Instead of directly using pitch power to determine tonality, the system uses statistical models that capture the probabilistic relationships between pitch name distributions and tonality types, serving as an intermediary layer that handles cases where direct correlation is weak or absent.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12014705B2Audio analysis method and audio analysis device
Publication Date: 2024.06.18 YAMAHA CORP
  • US12014705B2 patent drawing
  • US12014705B2 patent drawing
  • US12014705B2 patent drawing

AI summary

An audio analysis method is realized by a computer and includes generating key information which represents a key, by inputting a time series of a feature amount of an audio signal into a learned model that has learned a relationship between keys and time series of feature amounts of audio signals.