Speech Synthesis Model Training Using Pitch Synchronous Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesis technologies face issues with acoustic quality deterioration due to fixed frame rate analysis and unnatural phoneme duration caused by pitch mismatch during training and synthesis.

Innovation Solution

A speech synthesis device and method that employs pitch synchronous analysis to generate acoustic feature parameters, using Hidden Semi-Markov Models (HSMM) to model duration distribution based on timing parameters, ensuring accurate representation of speech waveforms and preventing phoneme duration mismatch.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If fixed frame rate speech analysis is used, then the analysis process is simple, but acoustic quality deteriorates

Engineering Contradiction:
Improveanalysis process simplicityVSAvoidacoustic quality
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes the fundamental parameter of speech analysis from fixed frame rate to pitch-synchronous analysis. By aligning analysis frames with pitch periods rather than using uniform time intervals, the system achieves both improved acoustic quality and maintained computational efficiency. The analysis frame rate dynamically adapts to the speech signal's pitch characteristics.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If pitch synchronous analysis is used, then acoustic quality improves, but phoneme duration becomes unnatural due to pitch mismatch

Engineering Contradiction:
Improveacoustic qualityVSAvoidphoneme duration accuracy
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent introduces a pitch-cycle waveform count as an intermediary parameter that bridges pitch-synchronous analysis and natural phoneme duration. This count represents the number of pitch cycles within each phoneme, allowing the system to maintain pitch-synchronous analysis benefits while ensuring that synthesized phonemes have natural durations by controlling the number of pitch cycles rather than fixed time duration.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If pitch-cycle waveform count is decided using duration distribution and pitch information, then phoneme duration accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvephoneme duration accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training duration distributions from speech corpus data before synthesis. The system learns typical duration patterns for different phonemes and contexts during training, then uses this pre-learned knowledge during synthesis to quickly determine appropriate pitch-cycle waveform counts without complex real-time calculations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback mechanisms where the decided pitch-cycle waveform count is evaluated against the output distribution of pitch feature parameters. This feedback loop ensures that the chosen waveform count is consistent with the synthesized pitch information, allowing the system to adjust and refine duration decisions based on the actual pitch characteristics of the generated speech.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11423874B2Speech synthesis statistical model training device, speech synthesis statistical model training method, and computer program product
Publication Date: 2022.08.23 KK TOSHIBA
  • US11423874B2 patent drawing
  • US11423874B2 patent drawing
  • US11423874B2 patent drawing

AI summary

A speech synthesis model training device includes one or more hardware processors configured to perform the following. Storing, in a speech corpus storing unit, speech data, and pitch mark information and context information of the speech data. From the speech data, analyzing acoustic feature parameters at each pitch mark timing in pitch mark information. From the acoustic feature parameters analyzed, training a statistical model which has a plurality of states and which includes an output distribution of acoustic feature parameters including pitch feature parameters and a duration distribution based on timing parameters.