Voice Synthesis Device Morphing Intermediate Voice Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice synthesis devices have limited freedom in voice quality conversion, suffer from poor sound quality due to excessive dynamic range processing, and may inaccurately specify reference points for morphing, resulting in poor sound quality and increased calculation complexity.

Innovation Solution

A voice synthesis device that stores multiple voice element information for different qualities, uses a designating unit to set fixed points for morphing, and generates intermediate voice information by calculating intermediate values of characteristic parameters, allowing for greater voice quality freedom without excessive dynamic range processing and reducing calculation complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional voice synthesis devices use preset voice element databases with fixed voice qualities, then the device complexity is reduced, but the freedom of voice quality conversion is limited

Engineering Contradiction:
Improvefreedom of voice quality conversionVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The voice element information is segmented into multiple characteristic parameters (pitch, intensity, timbre, etc.) that can be independently adjusted. This allows the system to create intermediate voice qualities by combining parameters from different preset voice elements, thereby increasing voice quality conversion freedom without requiring a complete redesign of the voice synthesis system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention changes the approach from selecting entire preset voice elements to adjusting individual characteristic parameters of voice elements. By modifying parameters such as pitch, intensity, and timbre independently, the system can generate intermediate voice qualities that were not previously available in the preset databases, resolving the contradiction between versatility and complexity.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the dynamic range in the spectrum is increased to improve voice quality adjustment, then the freedom of voice quality conversion is improved, but the sound quality deteriorates

Engineering Contradiction:
Improvefreedom of voice quality adjustmentVSAvoidsound quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

Instead of uniformly increasing the dynamic range across the entire spectrum, the invention applies local adjustments to specific characteristic parameters (pitch, intensity, timbre) at different time points. This localized approach allows for flexible voice quality adjustment while maintaining the natural spectral characteristics and avoiding sound quality deterioration.

Inventive Principle:
Principle #3Local quality

3Productivity

If a morphing process is carried out on waveform data with automatically specified reference points, then the processing speed is improved, but the sound quality becomes poor due to incorrect specification

Engineering Contradiction:
Improveprocessing speedVSAvoidsound quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The invention performs preliminary action by pre-specifying reference points in the waveform data before the morphing process begins. These reference points are carefully selected to correspond to important acoustic features such as peaks and transitions. By preparing these reference points in advance, the system ensures accurate morphing without sacrificing processing speed, as the reference specification does not need to be performed during the actual morphing operation.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If intermediate voice qualities are generated by calculating intermediate values of characteristic parameters, then the sound quality is maintained, but the calculation complexity increases

Engineering Contradiction:
Improvesound qualityVSAvoidcalculation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The invention simplifies the calculation by focusing changes on a limited set of key characteristic parameters (pitch, intensity, timbre) rather than processing the entire waveform or spectrum. By calculating intermediate values only for these essential parameters, the system maintains sound quality while significantly reducing calculation complexity compared to full waveform or spectral processing.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7571099B2Voice synthesis device
Publication Date: 2009.08.04 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US7571099B2 patent drawing
  • US7571099B2 patent drawing
  • US7571099B2 patent drawing

AI summary

A voice synthesis device for generating synthetic voice having great freedom in voice quality and good sound quality from text data is provided.The voice synthesis device is provided with: voice synthesis DBs (110a and 101z); a voice synthesis unit (103) for acquiring a text (10) and generating a voice synthesis parameter value sequence (11) of voice quality A which corresponds to a character that is included in the text (10) from the voice synthesis DB (101a); a voice synthesis unit (103) generating a voice synthesis parameter value sequence (11) of voice quality Z which corresponds to a character that is included in the text (10) from the voice synthesis DB (101z); a voice morphing unit (105) for generating a intermediate voice synthesis parameter value sequence (13) for indicating synthetic voice of intermediate voice quality between voice quality A and voice quality Z which corresponds to a character that is included in the text (10) from the voice synthesis parameter value sequence (11) of voice quality A and voice quality Z; and a speaker (107) for converting the generated intermediate voice synthesis parameter value sequence (13) to the synthetic voice and outputting the resulting synthetic voice.