Audio Source Separation via Logarithmic Frequency Image Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional electronic keyboard instruments lack effective methods for teaching melody playing, and existing audio processing techniques using machine learning struggle with separating specific audio components from mixed audio data, particularly in musical contexts.
Innovation Solution
A machine learning method that transforms audio data into image data using a logarithmic frequency axis, allowing for the separation of audio components by training a model to generate image data from mixed audio inputs, which is then used to separate and light corresponding keys on a luminescent keyboard.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio data is transformed into image data using traditional linear frequency axis methods, then the audio components can be visualized, but the separation accuracy of specific audio components from mixed audio data remains insufficient
Solution Approach 1:
The patent transforms the frequency axis from linear to logarithmic scale, fundamentally changing the parameter representation. This logarithmic transformation better represents human auditory perception and improves the separability of audio components in the frequency domain, allowing the machine learning model to more effectively distinguish between different audio sources without requiring excessive model complexity
Solution Approach 2:
The patent introduces image data as an intermediary representation between raw audio data and the separation task. By converting mixed audio data into image representations with logarithmic frequency axes, the system creates a more suitable intermediate form that enhances the effectiveness of subsequent machine learning processing for audio component separation
2Adaptability or versatility
If conventional audio processing methods are used in electronic keyboard instruments, then the instrument can play music, but it lacks effective melody teaching functionality
Solution Approach 1:
The patent replaces traditional mechanical audio processing methods with machine learning-based image processing. By substituting conventional signal processing with neural network-based image analysis, the system achieves superior audio component separation, enabling the keyboard to effectively isolate and teach melody lines from mixed music recordings
Solution Approach 2:
The patent enables the electronic keyboard to perform multiple functions: it can both play music conventionally and provide intelligent melody teaching through audio separation. The same machine learning model that separates audio components is used to identify and highlight melody lines, allowing the instrument to adapt to different teaching scenarios and user needs
3Manufacturing precision
If audio components are separated from mixed audio data, then melody extraction is possible, but the existing methods struggle with accuracy in musical contexts
Solution Approach 1:
The logarithmic frequency transformation preserves more information about different frequency components by distributing them more evenly across the image representation. This parameter change ensures that both low and high frequency components maintain their distinguishability, reducing information loss during the transformation and subsequent separation process while improving melody extraction accuracy
Data Source
AI summary
A machine learning method for training a learning model includes: transforming a first audio type of audio data into a first image type of image data, wherein a first audio component and a second audio component are mixed in the first audio type of audio data, and the first image type of image data corresponds to the first audio type of audio data; transforming a second audio type of audio data into a second image type of image data, wherein the second audio type of audio data includes the first audio component without mixture of the second audio component, and the second image type of image data corresponds to the second audio type of audio data; and performing machine learning on the learning model with training data including sets of the first image type of image data and the second image type of image data.


