Speech Synthesizer Multiple-Acoustic Feature Parameter Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesizers face issues with naturalness due to over-smoothing in HMM-based systems and discontinuity distortion in waveform-concatenation methods, resulting in unnatural sound quality.

Innovation Solution

A speech synthesizer that generates a multiple-acoustic feature parameter sequence based on context information and HMM sequences, using a combination of statistical-model and waveform generators to produce a distribution sequence for natural and realistic voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If HMM-based speech synthesis uses averaged feature parameters, then the synthesis process is simplified and computation is reduced, but over-smoothing occurs resulting in unnatural sound quality

Engineering Contradiction:
Improvesynthesis process simplicityVSAvoidsound quality naturalness
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the speech synthesis process into multiple stages: generating multiple candidate acoustic feature parameter sequences from the HMM distribution sequence, then selecting the most appropriate sequence based on evaluation criteria. This segmentation allows the system to maintain computational efficiency while avoiding over-smoothing by preserving multiple candidate sequences rather than immediately averaging them.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter selection approach by introducing multiple acoustic feature parameter sequences with different characteristics (e.g., different levels of smoothing, different source segments) and dynamically selecting or combining them based on the specific synthesis context. This parameter change enables the system to adapt to different sound quality requirements while maintaining synthesis efficiency.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If waveform-concatenation method selects speech elements from database, then high quality synthesized speech similar to recorded speech can be obtained, but discontinuity distortion occurs between adjacent speech segments

Engineering Contradiction:
Improvesynthesized speech qualityVSAvoidcontinuity between segments
Core Design Contradiction:
Manufacturing precisionVSStability of the object's composition

Solution Approach 1:

The patent introduces an intermediary evaluation and selection mechanism that bridges the HMM-based parameter generation and waveform synthesis stages. By evaluating multiple candidate acoustic feature parameter sequences and selecting those that ensure smooth transitions, the system mediates between the statistical model output and the waveform concatenation process, reducing discontinuity distortion.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary action by generating and evaluating multiple candidate acoustic feature parameter sequences before final waveform synthesis. This pre-evaluation allows the system to identify and select parameter sequences that will result in smooth transitions between speech segments, preventing discontinuity distortion before it occurs.

Inventive Principle:
Principle #10Preliminary action

3Stability of the object's composition

If multiple speech segments are selected for each section, then sense of stability is enhanced, but complexity of the synthesis process increases

Engineering Contradiction:
Improvesense of stabilityVSAvoidsynthesis process complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent applies partial action by generating multiple candidate acoustic feature parameter sequences (excessive) but then selecting only the most appropriate one or a limited subset for final synthesis. This approach captures the stability benefits of multiple segments while avoiding the full complexity of processing and combining all candidates, achieving a balance between stability and computational efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10529314B2Speech synthesizer, and speech synthesis method and computer program product utilizing multiple-acoustic feature parameters selection
Publication Date: 2020.01.07 KK TOSHIBA
  • US10529314B2 patent drawing
  • US10529314B2 patent drawing
  • US10529314B2 patent drawing

AI summary

A speech synthesizer includes a statistical-model sequence generator, a multiple-acoustic feature parameter sequence generator, and a waveform generator. The statistical-model sequence generator generates, based on context information corresponding to an input text, a statistical model sequence that comprises a first sequence of a statistical model comprising a plurality of states. The multiple-acoustic feature parameter sequence generator, for each speech section corresponding to each state of the statistical model sequence, selects a first plurality of acoustic feature parameters from a first set of acoustic feature parameters extracted from a first speech waveform stored in a speech database and generates a multiple-acoustic feature parameter sequence that comprises a sequence of the first plurality of acoustic feature parameters. The waveform generator generates a distribution sequence based on the multiple-acoustic feature parameter sequence and generates a second speech waveform based on a second set of acoustic feature parameters generated based on the distribution sequence.