Speech Synthesis Neural Network Phoneme Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis methods rely on vocoders to convert acoustic characteristics into speech, requiring manual alignment and segmentation of phonemes and speech waveforms, which is inefficient and labor-intensive.

Innovation Solution

A method and apparatus for speech synthesis using a pre-trained end-to-end neural network that determines phoneme sequences, extracts acoustic characteristics, and synthesizes speech waveform units based on a preset index and cost function, eliminating the need for vocoders and manual alignment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a vocoder is used to convert acoustic characteristics into speech, then speech can be generated, but the process becomes inefficient and labor-intensive due to manual alignment and segmentation requirements

Engineering Contradiction:
Improvespeech synthesis efficiencyVSAvoidmanual alignment and segmentation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical vocoder-based conversion process with a neural network model that directly generates speech waveforms from acoustic characteristics. This substitution eliminates the need for manual alignment and segmentation operations, automating the speech synthesis process and significantly improving efficiency while reducing operational complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network model performs automatic alignment and segmentation of phonemes and speech waveforms internally during the synthesis process. The system serves itself by handling these previously manual tasks through learned patterns in the training data, eliminating the need for external manual intervention and improving overall productivity

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual alignment and segmentation of phonemes and speech waveforms is performed, then accurate speech synthesis can be achieved, but the process becomes labor-intensive and inefficient

Engineering Contradiction:
Improvephoneme-speech waveform alignment accuracyVSAvoidmanual processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The neural network model is pre-trained on large datasets containing aligned phoneme-speech waveform pairs. This preliminary training action embeds alignment knowledge into the model parameters, enabling the system to perform accurate alignment automatically during inference without requiring manual preprocessing for each synthesis task

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The manual alignment and segmentation process is replaced by the neural network's learned mapping between acoustic characteristics and speech waveforms. The model automatically performs the alignment function that previously required manual labor, maintaining accuracy while eliminating time loss associated with manual processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10553201B2Method and apparatus for speech synthesis
Publication Date: 2020.02.04 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US10553201B2 patent drawing
  • US10553201B2 patent drawing
  • US10553201B2 patent drawing

AI summary

A method of speech synthesis is provided, which comprises: determining a phoneme sequence of a to-be-processed text; inputting the phoneme sequence into a pre-trained speech model to obtain an acoustic characteristic corresponding to each phoneme in the phoneme sequence, where the speech model is used for characterizing a corresponding relationship between each phoneme in the phoneme sequence and the acoustic characteristic; determining, for each phoneme in the phoneme sequence, at least one speech waveform unit corresponding to each phoneme based on a preset index of phonemes and speech waveform units, and determining a target speech waveform unit of the at least one speech waveform unit based on the acoustic characteristic corresponding to the phoneme and a preset cost function; and synthesizing the target speech waveform unit corresponding to each phoneme in the phoneme sequence to generate a speech.