Audio Signal Processing for Vocal Quality Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing terminal technologies struggle to produce high-quality audio signals when users with poor singing skills record songs, as they directly use the user's audio signal without enhancing its quality.

Innovation Solution

An audio signal processing method that acquires a user's audio signal, extracts the user's timbre information, acquires intonation information from a standard audio signal, and generates a new audio signal by synthesizing the timbre and intonation information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the terminal directly uses the user's audio signal for recording, then the operation is simple, but the audio quality is poor when the user has poor singing skills

Engineering Contradiction:
Improveoperation simplicityVSAvoidaudio signal quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent introduces an audio signal processing module as an intermediary between the user's raw audio signal and the final recorded output. This module extracts timbre information from the user's voice and combines it with reference audio signals to generate enhanced output, thereby improving audio quality without significantly complicating the user operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses reference audio signals (copies of professional singing performances) as a basis for enhancing the user's audio recording. By copying the timbre characteristics from professional singers and combining them with the user's vocal performance, the system generates high-quality audio output even when the user has poor singing skills

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If the terminal extracts timbre information and synthesizes audio signals, then the audio quality is improved, but the processing complexity increases

Engineering Contradiction:
Improveaudio signal qualityVSAvoidsignal processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio processing task into distinct functional modules: an audio signal acquisition module that captures user input, a timbre information extraction module that analyzes vocal characteristics, a reference signal selection module that chooses appropriate reference audio, and an audio signal generation module that synthesizes the final output. This segmentation allows each module to perform its specific function efficiently while keeping the overall system manageable

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs automatic timbre extraction and audio synthesis without requiring manual intervention from the user. The audio signal processing module autonomously analyzes the user's voice characteristics, selects appropriate reference signals, and generates the enhanced audio output, thereby reducing the operational burden on the user despite the complex processing involved

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3614383B1Audio data processing method and apparatus, and storage medium
Publication Date: 2025.05.07 GUANGZHOU KUGOU COMP TECH CO LTD
  • EP3614383B1 patent drawingFigure 1
  • EP3614383B1 patent drawingFigure 2
  • EP3614383B1 patent drawingFigure 3~4

AI summary

The present disclosure discloses an audio signal processing method and apparatus, and a storage medium and belongs to the field of terminal technologies. The audio signal processing method includes: acquiring a first audio signal of a target song sung by a user; extracting timbre information of the user from the first audio signal; acquiring intonation information of a standard audio signal of the target song; and generating a second audio signal of the target song based on the timbre information and the intonation information. Since the second audio signal of the target song is generated based on the timbre information of the standard audio signal and the intonation information of the user, even if the user's singing skills are poor, a high-quality audio signal may still be generated. Thus, the quality of the generated audio signal is improved.