Audio Source Separation via Logarithmic Frequency Image Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional electronic keyboard instruments lack effective methods for teaching melody playing, and existing audio processing techniques using machine learning struggle with separating specific audio components from mixed audio data, particularly in musical contexts.

Innovation Solution

A machine learning method that transforms audio data into image data using a logarithmic frequency axis, allowing for the separation of audio components by training a model to generate image data from mixed audio inputs, which is then used to separate and light corresponding keys on a luminescent keyboard.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio data is transformed into image data using traditional linear frequency axis methods, then the audio components can be visualized, but the separation accuracy of specific audio components from mixed audio data remains insufficient

Engineering Contradiction:
Improveaudio component separation accuracyVSAvoidmachine learning model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the frequency axis from linear to logarithmic scale, fundamentally changing the parameter representation. This logarithmic transformation better represents human auditory perception and improves the separability of audio components in the frequency domain, allowing the machine learning model to more effectively distinguish between different audio sources without requiring excessive model complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces image data as an intermediary representation between raw audio data and the separation task. By converting mixed audio data into image representations with logarithmic frequency axes, the system creates a more suitable intermediate form that enhances the effectiveness of subsequent machine learning processing for audio component separation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If conventional audio processing methods are used in electronic keyboard instruments, then the instrument can play music, but it lacks effective melody teaching functionality

Engineering Contradiction:
Improveteaching functionalityVSAvoiduser learning difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces traditional mechanical audio processing methods with machine learning-based image processing. By substituting conventional signal processing with neural network-based image analysis, the system achieves superior audio component separation, enabling the keyboard to effectively isolate and teach melody lines from mixed music recordings

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent enables the electronic keyboard to perform multiple functions: it can both play music conventionally and provide intelligent melody teaching through audio separation. The same machine learning model that separates audio components is used to identify and highlight melody lines, allowing the instrument to adapt to different teaching scenarios and user needs

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If audio components are separated from mixed audio data, then melody extraction is possible, but the existing methods struggle with accuracy in musical contexts

Engineering Contradiction:
Improvemelody extraction accuracyVSAvoidaudio component information loss
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The logarithmic frequency transformation preserves more information about different frequency components by distributing them more evenly across the image representation. This parameter change ensures that both low and high frequency components maintain their distinguishability, reducing information loss during the transformation and subsequent separation process while improving melody extraction accuracy

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11568857B2Machine learning method, audio source separation apparatus, and electronic instrument
Publication Date: 2023.01.31 CASIO COMPUTER CO LTD
  • US11568857B2 patent drawing
  • US11568857B2 patent drawing
  • US11568857B2 patent drawing

AI summary

A machine learning method for training a learning model includes: transforming a first audio type of audio data into a first image type of image data, wherein a first audio component and a second audio component are mixed in the first audio type of audio data, and the first image type of image data corresponds to the first audio type of audio data; transforming a second audio type of audio data into a second image type of image data, wherein the second audio type of audio data includes the first audio component without mixture of the second audio component, and the second image type of image data corresponds to the second audio type of audio data; and performing machine learning on the learning model with training data including sets of the first image type of image data and the second image type of image data.