ML Audio Signal Separation and Enhancement for Easy Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional audio processing tools require specialized personnel and significant hardware resources, leading to high costs and complexity for non-specialized users.

Innovation Solution

An electronic device equipped with machine learning models that can perform various audio processing operations autonomously, including signal separation, transcription, and audio enhancement, with user-friendly interfaces, allowing non-specialized users to perform these tasks efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional audio processing tools are used, then audio processing functionality is provided, but specialized personnel intervention and high hardware requirements are needed

Engineering Contradiction:
Improveease of audio processing operationVSAvoidhardware requirements and personnel expertise
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical/audio engineering systems with machine learning-based automated processing. The system uses neural networks to automatically perform audio separation, transcription, and enhancement tasks that previously required specialized audio engineers and complex hardware setups, thereby reducing device complexity and personnel requirements while maintaining processing capabilities

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service audio processing by allowing users to directly interact with machine learning models through user-friendly interfaces. Users can upload audio files and receive automated processing results without needing specialized knowledge or intervention from audio engineers, making the system self-sufficient and easy to operate

Inventive Principle:
Principle #25Self-service

2Productivity

If traditional audio processing tools are used, then audio processing functionality is provided, but high licensing costs are incurred

Engineering Contradiction:
Improveaudio processing capabilityVSAvoidlicensing cost
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The patent employs open-source machine learning models and frameworks that can be freely downloaded and used without expensive licensing fees. Instead of requiring costly commercial software licenses, the system uses freely available AI models that provide comparable audio processing capabilities at minimal cost, thereby reducing the financial barrier for users

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Extent of automation

If machine learning models are applied for audio processing, then automated processing is achieved, but computational resources are required

Engineering Contradiction:
Improveautomated audio processingVSAvoidcomputational resource consumption
Core Design Contradiction:
Extent of automationVSUse of energy by moving object

Solution Approach 1:

The system applies machine learning models selectively and efficiently by processing audio data in optimized batches and using computational resources only when and where needed. The architecture processes audio signals through multiple stages, applying heavy computational models only to critical processing steps while using lighter processing for less demanding tasks, thereby achieving automation while managing computational resource consumption

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12499902B2Intelligent audio processing
Publication Date: 2025.12.16 SONY GROUP CORP
  • US12499902B2 patent drawing
  • US12499902B2 patent drawing
  • US12499902B2 patent drawing

AI summary

An electronic device and method for intelligent audio processing is provided. The electronic device receives a first user input for selection of a source audio signal. The electronic device receives a second user input for selection of a first operation to be performed on the source audio signal. The electronic device receives a third user input for selection of at least a first audio signal. The electronic device applies a first machine learning (ML) model on the source audio signal. The electronic device further extracts at least the first audio signal from the source audio signal based on the application of the first ML model on the source audio signal. The electronic device further applies a second ML model on the extracted at least first audio signal from the plurality of audio signals based on the first operation. The electronic device further generates output information as per the first operation.