Speech Spectral Envelope Parameterization Using Local Domain Bases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spectral envelope parameters for speech synthesis fail to achieve high quality, effectiveness, and easy band processing, as they are often influenced by spectral fine structures and do not efficiently represent spectral information with a reduced dimensionality.
Innovation Solution
A speech processing apparatus models the logarithmic spectral envelope as a linear combination of local domain bases, where each basis represents a frequency band with a peak frequency and zero values outside the band, allowing for minimal distortion and efficient parameter calculation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional spectral envelope parameters are used for speech synthesis, then spectral information can be represented, but the parameters are influenced by spectral fine structures and do not achieve high quality with reduced dimensionality
Solution Approach 1:
The spectrum is divided into multiple frequency bands, with each band represented by a localized basis function centered at a specific peak frequency. This segmentation allows the spectral envelope to be captured without being influenced by fine structures in other frequency regions, achieving both high quality representation and dimensionality reduction.
Solution Approach 2:
Each basis function is designed to be localized in the frequency domain, representing only a specific frequency band with its peak frequency. This local quality approach ensures that each parameter captures spectral information from a specific region, improving measurement precision while reducing the total number of parameters needed to represent the entire spectrum.
2Reliability
If spectral parameters are designed to represent the entire frequency range, then complete spectral information is captured, but processing becomes difficult and less effective
Solution Approach 1:
The frequency spectrum is segmented into multiple bands, each handled by a separate basis function. This segmentation makes band processing straightforward, as operations can be performed independently on each frequency band and then combined, improving ease of operation while maintaining complete spectral representation through the combination of all basis functions.
Data Source
AI summary
An information extraction unit extracts spectral envelope information of L-dimension from each frame of speech data by discrete Fourier transform. The spectral envelope information is represented by L points. A basis storage unit stores N bases (L>N>1). Each basis is differently a frequency band having a maximum as a peak frequency in a spectral domain having L-dimension. A value corresponding to a frequency outside the frequency band along a frequency axis of the spectral domain is zero. Two frequency bands of which two peak frequencies are adjacent along the frequency axis partially overlap. A parameter calculation unit minimizes a distortion between the spectral envelope information and a linear combination of each basis with a coefficient for each of L points of the spectral envelope information by changing the coefficient, and sets the coefficient of each basis from which the distortion is minimized to a spectral envelope parameter of the spectral envelope information.


