Singing Synthesis Parameter Estimation for Human-Like Voice

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current singing synthesis systems struggle to iteratively estimate and modify pitch and dynamics parameters of an input singing voice, leading to difficulties in producing a high-quality, human-like singing voice that accurately matches the user's desired expression, especially when conditions such as sound source data or synthesis systems change.

Innovation Solution

A singing synthesis parameter data estimation system that analyzes input singing voice audio signals to iteratively update pitch and dynamics parameters, using a database of singing sound source data and lyric information to synthesize a voice that closely matches the input, with features like pitch parameter estimation, dynamics parameter conversion, and lyric alignment to ensure accuracy and adaptability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual parameter adjustment is used for fine-tuning singing expression, then user control over singing voice quality is improved, but operation complexity and time consumption increase

Engineering Contradiction:
Improvesinging voice qualityVSAvoidparameter adjustment complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system automatically estimates and adjusts singing synthesis parameters by analyzing the input audio signal characteristics. The parameter estimation section extracts features such as pitch, dynamics, and vibrato from the user's singing voice and automatically generates appropriate synthesis parameters, eliminating the need for manual parameter tuning while maintaining high singing voice quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system compares the characteristics of the input singing voice with the synthesized output and iteratively adjusts parameters to minimize the difference. By using the actual singing performance as feedback, the system automatically optimizes parameters such as pitch contour, dynamics, and vibrato to match the user's intended expression.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If singing synthesis parameters are adjusted for different sound source data or synthesis systems, then adaptability to various systems is improved, but parameter adjustment time and complexity increase

Engineering Contradiction:
Improvesystem compatibilityVSAvoidparameter re-adjustment time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The parameter estimation section is designed to extract universal singing characteristics from the input audio signal that are independent of the specific sound source data or synthesis system being used. By focusing on fundamental acoustic features such as pitch, dynamics, and vibrato rather than system-specific parameters, the estimated parameters can be adapted to different synthesis systems without requiring re-adjustment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system automatically adapts parameters by changing them based on the characteristics of the input singing voice and the target synthesis system. The parameter estimation section generates optimized parameter values that are suited to the specific sound source data or synthesis system being used, eliminating the need for manual re-adjustment when switching systems.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If automated parameter estimation is implemented, then operation ease is improved, but measurement precision of singing parameters may deteriorate

Engineering Contradiction:
Improveautomatic parameter estimationVSAvoidparameter estimation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system replaces manual parameter tuning with automated acoustic analysis. The parameter estimation section uses signal processing techniques to extract singing parameters such as pitch, dynamics, and vibrato directly from the audio waveform and spectral characteristics, providing accurate measurements without manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The parameter estimation section acts as an intermediary between the raw audio input and the synthesis system. It extracts meaningful singing parameters from the complex audio signal and transforms them into format-appropriate parameters for the synthesis system, maintaining precision through specialized analysis algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8244546B2Singing synthesis parameter data estimation system
Publication Date: 2012.08.14 NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY
  • US8244546B2 patent drawing
  • US8244546B2 patent drawing
  • US8244546B2 patent drawing

AI summary

There is provided a singing synthesis parameter data estimation system that automatically estimates singing synthesis parameter data for automatically synthesizing a human-like singing voice from an audio signal of input singing voice. A pitch parameter estimating section 9 estimates a pitch parameter, by which the pitch feature of an audio signal of synthesized singing voice is got closer to the pitch feature of the audio signal of input singing voice based on at least both of the pitch feature and lyric data with specified syllable boundaries of the audio signal of input singing voice. A dynamics parameter estimating section 11 converts the dynamics feature of the audio signal of input singing voice to a relative value with respect to the dynamics feature of the audio signal of synthesized singing voice, and estimates a dynamics parameter, by which the dynamics feature of the audio signal of synthesized singing voice is got close to the dynamics feature of the audio signal of input singing voice that has been converted to the relative value.