Singing Guidance Audio Separation for Real-Time Vocal Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio playback systems do not effectively support users who want to sing along with music, as their voice interferes with the original recording, and less experienced singers lack guidance for improving their performance.

Innovation Solution

An electronic device performs audio source separation to isolate vocals and accompaniment signals, and conducts confidence analysis on the user's voice to provide real-time guidance through visual and audio cues, adjusting the original vocals based on the user's performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio source separation is performed to separate vocals and accompaniment signals, then the ability to provide targeted guidance to users is improved, but the device complexity increases

Engineering Contradiction:
Improvevoice signal analysis accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is segmented into distinct components (vocals and accompaniment) through source separation, allowing independent analysis and processing of each component. This enables precise voice signal analysis while managing complexity by handling separated signals rather than the full mixed audio stream.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary processing stage is introduced that takes the separated vocals signal as input and generates guidance information as output. This intermediary layer bridges the gap between complex source separation and simple guidance provision, allowing the system to maintain high measurement precision while managing overall device complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If real-time confidence analysis is performed on user's voice to provide guidance, then the usefulness of the system for improving singing performance is improved, but the processing time and computational load increase

Engineering Contradiction:
Improvesinging guidance capabilityVSAvoidreal-time processing delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing the reference vocals signal through source separation to extract the vocals component before the user sings. This preparation allows the confidence analysis to operate more efficiently during real-time user performance, reducing processing delays while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A feedback loop is established where the user's voice is continuously analyzed against the reference vocals, and guidance is provided based on the confidence analysis results. This feedback mechanism enables real-time adaptation to user performance while managing computational load by focusing analysis on relevant parameters.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If the vocals signal is adjusted based on confidence analysis to provide guidance, then the user's singing experience is improved, but the audio signal processing complexity increases

Engineering Contradiction:
Improvesinging along experienceVSAvoidaudio signal processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The vocals signal is dynamically adjusted based on real-time confidence analysis results rather than using a static processing approach. The system adapts the guidance signal characteristics according to the user's performance level, improving ease of operation while managing complexity through adaptive rather than exhaustive processing.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Specific parameters of the vocals signal (such as gain, pitch, or timing) are changed based on confidence analysis outcomes to provide targeted guidance. By modifying only relevant parameters rather than the entire signal, the system improves user experience while keeping audio signal processing complexity manageable.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12633230B2Electronic device, method and computer program
Publication Date: 2026.05.19 SONY GROUP CORP
  • US12633230B2 patent drawing
  • US12633230B2 patent drawing
  • US12633230B2 patent drawing

AI summary

An electronic device having a circuitry configured to perform audio source separation on an audio input signal to obtain a vocals signal and an accompaniment signal and to perform a confidence analysis on a user's voice signal based on the vocals signal to provide guidance to the user.