Statistical Parameter Model for Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional statistical parameter models for speech synthesis suffer from information loss during feature extraction, resulting in synthesized timbre that is not full enough and has an obvious machine-like voice quality.

Innovation Solution

A method that trains a statistical parameter model using a non-linear mapping calculation on a vector matrix formed by matching text feature sample points with speech sample points, avoiding explicit feature extraction and optimizing model parameters to minimize differences between predicted and original speech sample points, thereby improving the naturalness and saturation of synthesized speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit feature extraction is performed on speech to obtain pitch frequency, sounding duration, and spectrum characteristic, then the statistical parameter model can process the extracted features for synthesis, but information loss occurs during conversion which makes synthesized timbre not full enough and produces obvious machine voice

Engineering Contradiction:
Improvefeature extraction accuracyVSAvoidspeech information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts only the necessary text feature sample points from the complete speech signal, rather than extracting all speech features. This selective extraction approach obtains the essential information needed for synthesis while avoiding the information loss that occurs when converting the entire speech signal into traditional acoustic features like pitch frequency, sounding duration, and spectrum characteristic.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces text feature sample points as an intermediary representation between the original speech and the synthesized speech. Instead of directly extracting and processing traditional speech features, the system uses text feature sample points that retain more original speech information while still serving as effective inputs for the statistical parameter model.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional speech feature extraction and conversion is used, then the processing pipeline is well-established, but the synthesized speech has poor naturalness and insufficient timbre saturation

Engineering Contradiction:
Improvesynthesis stabilityVSAvoidmachine voice quality
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent changes the parameter representation from traditional acoustic features (pitch frequency, sounding duration, spectrum characteristic) to text feature sample points. This parameter transformation allows the system to maintain processing stability while significantly improving the naturalness and timbre saturation of synthesized speech by preserving more original speech information.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If the statistical parameter model processes extracted speech features, then the synthesis process can proceed, but the timbre is not full enough and machine voice characteristics are obvious

Engineering Contradiction:
Improvesynthesis efficiencyVSAvoidmachine voice quality
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent creates a closer copy of the original speech by using text feature sample points that retain more of the original speech's characteristics. Instead of processing heavily processed and information-loss-prone traditional features, the system works with a representation that better copies the original speech, resulting in more natural synthesized output with improved timbre saturation and reduced machine voice qualities.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3614376B1Speech synthesis method, server and storage medium
Publication Date: 2022.10.19 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3614376B1 patent drawingFigure 1~2
  • EP3614376B1 patent drawingFigure 3
  • EP3614376B1 patent drawingFigure 4

AI summary

A statistical parameter modeling method is disclosed, including: obtaining, by a server, model training data, the model training data including a text feature sequence and a corresponding original speech sample sequence; inputting, by the server, an original vector matrix formed by matching a text feature sample point in the text feature sample sequence with a speech sample point in the original speech sample sequence into a statistical parameter model for training; performing, by the server, non-linear mapping calculation on the original vector matrix in a hidden layer, to output a corresponding prediction speech sample point; and determining, by the server, a model parameter of the statistical parameter model according to the prediction speech sample point and a corresponding original speech sample point by using a smallest difference principle, to obtain a corresponding target statistical parameter model.