Glottal Pulse Vector Database for Speech Synthesis Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis systems, particularly Hidden Markov Model based statistical parametric speech synthesis, fail to effectively capture low frequency source content in excitation signals, leading to suboptimal quality of synthetic speech.

Innovation Solution

The method involves computing a glottal pulse distance metric, clustering a glottal pulse database, forming a vector database by associating vectors with each glottal pulse based on centroid pulses and distance metrics, and using Eigenvectors to form parametric models that generate excitation signals, thereby capturing low frequency information and improving speech synthesis quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional excitation signal methods are used in HMM-based speech synthesis, then the system complexity remains low, but the quality of synthetic speech deteriorates due to inability to capture low frequency source content

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The glottal pulse database is segmented into multiple clusters based on similarity metrics, allowing the system to capture diverse low frequency characteristics. Each cluster represents a distinct group of glottal pulses with similar properties, enabling selective retrieval of appropriate excitation signals for different speech contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary clustering and vector database formation during the training phase, organizing glottal pulses into structured groups before synthesis. This pre-processing creates ready-to-use excitation signal repositories that can be efficiently queried during speech generation without adding real-time computational complexity.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a simple excitation model is used, then the system is easier to implement, but the low frequency source content is not captured accurately

Engineering Contradiction:
Improvelow frequency content accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates vector representations (copies) of glottal pulses that capture their essential low frequency characteristics. These vectors serve as compact proxies that preserve the critical information needed for accurate excitation signal selection without requiring storage or processing of the complete high-resolution waveforms.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms glottal pulses from time-domain waveforms into frequency-domain parameter representations through clustering analysis. This parameter transformation enables accurate capture of low frequency content by focusing on spectral characteristics rather than raw signal values.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If no clustering is performed on glottal pulses, then the database structure remains simple, but the ability to select appropriate excitation signals deteriorates

Engineering Contradiction:
Improveexcitation signal selection capabilityVSAvoiddatabase structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The glottal pulse database is segmented into multiple clusters based on similarity metrics, allowing the system to capture diverse low frequency characteristics. Each cluster represents a distinct group of glottal pulses with similar properties, enabling selective retrieval of appropriate excitation signals for different speech contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The vector database serves as an intermediary structure between the raw glottal pulse database and the speech synthesis process. It provides a structured interface for querying and selecting excitation signals based on spectral similarity, mediating between the unorganized pulse data and the synthesis requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3149727B1Method for forming the excitation signal for a glottal pulse model based parametric speech synthesis system
Publication Date: 2021.01.27 INTERACTIVE INTELLIGENCE GROUP INC
  • EP3149727B1 patent drawingFigure 1
  • EP3149727B1 patent drawingFigure 2
  • EP3149727B1 patent drawingFigure 3

AI summary

A method is presented for forming the excitation signal for a glottal pulse model based parametric speech synthesis system. In one embodiment, fundamental frequency values are used to form the excitation signal. The excitation is modeled using a voice source pulse selected from a database of a given speaker. The voice source signal is segmented into glottal segments, which are used in vector representation to identify the glottal pulse used for formation of the excitation signal. Use of a novel distance metric and preserving the original signals extracted from the speakers voice samples helps capture low frequency information of the excitation signal. In addition, segment edge artifacts are removed by applying a unique segment joining method to improve the quality of synthetic speech while creating a true representation of the voice quality of a speaker.