Speech Synthesis Neural Network Phoneme Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis methods rely on vocoders to convert acoustic characteristics into speech, requiring manual alignment and segmentation of phonemes and speech waveforms, which is inefficient and labor-intensive.
Innovation Solution
A method and apparatus for speech synthesis using a pre-trained end-to-end neural network that determines phoneme sequences, extracts acoustic characteristics, and synthesizes speech waveform units based on a preset index and cost function, eliminating the need for vocoders and manual alignment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a vocoder is used to convert acoustic characteristics into speech, then speech can be generated, but the process becomes inefficient and labor-intensive due to manual alignment and segmentation requirements
Solution Approach 1:
The patent replaces the mechanical vocoder-based conversion process with a neural network model that directly generates speech waveforms from acoustic characteristics. This substitution eliminates the need for manual alignment and segmentation operations, automating the speech synthesis process and significantly improving efficiency while reducing operational complexity
Solution Approach 2:
The neural network model performs automatic alignment and segmentation of phonemes and speech waveforms internally during the synthesis process. The system serves itself by handling these previously manual tasks through learned patterns in the training data, eliminating the need for external manual intervention and improving overall productivity
2Manufacturing precision
If manual alignment and segmentation of phonemes and speech waveforms is performed, then accurate speech synthesis can be achieved, but the process becomes labor-intensive and inefficient
Solution Approach 1:
The neural network model is pre-trained on large datasets containing aligned phoneme-speech waveform pairs. This preliminary training action embeds alignment knowledge into the model parameters, enabling the system to perform accurate alignment automatically during inference without requiring manual preprocessing for each synthesis task
Solution Approach 2:
The manual alignment and segmentation process is replaced by the neural network's learned mapping between acoustic characteristics and speech waveforms. The model automatically performs the alignment function that previously required manual labor, maintaining accuracy while eliminating time loss associated with manual processing
Data Source
AI summary
A method of speech synthesis is provided, which comprises: determining a phoneme sequence of a to-be-processed text; inputting the phoneme sequence into a pre-trained speech model to obtain an acoustic characteristic corresponding to each phoneme in the phoneme sequence, where the speech model is used for characterizing a corresponding relationship between each phoneme in the phoneme sequence and the acoustic characteristic; determining, for each phoneme in the phoneme sequence, at least one speech waveform unit corresponding to each phoneme based on a preset index of phonemes and speech waveform units, and determining a target speech waveform unit of the at least one speech waveform unit based on the acoustic characteristic corresponding to the phoneme and a preset cost function; and synthesizing the target speech waveform unit corresponding to each phoneme in the phoneme sequence to generate a speech.


