Glottal Pulse Vector Database for Speech Synthesis Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis systems, particularly Hidden Markov Model based statistical parametric speech synthesis, fail to effectively capture low frequency source content in excitation signals, leading to suboptimal quality of synthetic speech.
Innovation Solution
The method involves computing a glottal pulse distance metric, clustering a glottal pulse database, forming a vector database by associating vectors with each glottal pulse based on centroid pulses and distance metrics, and using Eigenvectors to form parametric models that generate excitation signals, thereby capturing low frequency information and improving speech synthesis quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional excitation signal methods are used in HMM-based speech synthesis, then the system complexity remains low, but the quality of synthetic speech deteriorates due to inability to capture low frequency source content
Solution Approach 1:
The glottal pulse database is segmented into multiple clusters based on similarity metrics, allowing the system to capture diverse low frequency characteristics. Each cluster represents a distinct group of glottal pulses with similar properties, enabling selective retrieval of appropriate excitation signals for different speech contexts.
Solution Approach 2:
The system performs preliminary clustering and vector database formation during the training phase, organizing glottal pulses into structured groups before synthesis. This pre-processing creates ready-to-use excitation signal repositories that can be efficiently queried during speech generation without adding real-time computational complexity.
2Measurement precision
If a simple excitation model is used, then the system is easier to implement, but the low frequency source content is not captured accurately
Solution Approach 1:
The system creates vector representations (copies) of glottal pulses that capture their essential low frequency characteristics. These vectors serve as compact proxies that preserve the critical information needed for accurate excitation signal selection without requiring storage or processing of the complete high-resolution waveforms.
Solution Approach 2:
The system transforms glottal pulses from time-domain waveforms into frequency-domain parameter representations through clustering analysis. This parameter transformation enables accurate capture of low frequency content by focusing on spectral characteristics rather than raw signal values.
3Adaptability or versatility
If no clustering is performed on glottal pulses, then the database structure remains simple, but the ability to select appropriate excitation signals deteriorates
Solution Approach 1:
The glottal pulse database is segmented into multiple clusters based on similarity metrics, allowing the system to capture diverse low frequency characteristics. Each cluster represents a distinct group of glottal pulses with similar properties, enabling selective retrieval of appropriate excitation signals for different speech contexts.
Solution Approach 2:
The vector database serves as an intermediary structure between the raw glottal pulse database and the speech synthesis process. It provides a structured interface for querying and selecting excitation signals based on spectral similarity, mediating between the unorganized pulse data and the synthesis requirements.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method is presented for forming the excitation signal for a glottal pulse model based parametric speech synthesis system. In one embodiment, fundamental frequency values are used to form the excitation signal. The excitation is modeled using a voice source pulse selected from a database of a given speaker. The voice source signal is segmented into glottal segments, which are used in vector representation to identify the glottal pulse used for formation of the excitation signal. Use of a novel distance metric and preserving the original signals extracted from the speakers voice samples helps capture low frequency information of the excitation signal. In addition, segment edge artifacts are removed by applying a unique segment joining method to improve the quality of synthetic speech while creating a true representation of the voice quality of a speaker.