Concatenation-Sensitive Neural Networks for Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current unit-selection text-to-speech synthesis methods require complex algorithms to optimize weights for concatenation and target costs, making the process cumbersome and prone to inaccuracies in selecting natural-sounding speech segments.
Innovation Solution
The use of concatenation-sensitive neural networks determines predicted acoustic model parameters for speech segments based on both linguistic and acoustic features, eliminating the need for separate concatenation cost optimization and simplifying the unit-selection process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a sophisticated unit-selection algorithm is implemented to select the most suitable speech units, then the acoustic quality and naturalness of synthesized speech is improved, but the algorithm complexity and computational overhead increases
Solution Approach 1:
The patent replaces the traditional mechanical unit-selection algorithm with a neural network-based acoustic model. The neural network learns optimal speech unit selection through training on acoustic data, substituting the manual rule-based mechanical selection process with an intelligent system that automatically identifies and selects the most suitable speech units based on acoustic similarity, thereby maintaining high acoustic quality while reducing algorithmic complexity
Solution Approach 2:
The patent transforms the unit-selection problem from a complex multi-parameter optimization task into a simpler acoustic similarity matching task. By changing the selection criterion from multiple weighted parameters (phonetic accuracy, prosody, timing) to a single dominant acoustic similarity metric computed by the neural network, the system achieves high speech quality with reduced computational complexity
2Manufacturing precision
If separate concatenation cost optimization is performed to ensure smooth transitions between speech segments, then the naturalness of synthesized speech is improved, but the overall system complexity and number of optimization parameters increases
Solution Approach 1:
The patent merges the previously separate concatenation cost optimization into the unified acoustic similarity metric computed by the neural network. The acoustic model is trained to simultaneously optimize both the acoustic quality of individual speech units and the smoothness of transitions between them, eliminating the need for separate concatenation cost calculations and reducing the total number of optimization parameters
Solution Approach 2:
The neural network-based acoustic model serves multiple functions simultaneously: it performs speech unit selection, acoustic quality assessment, and concatenation smoothness evaluation. This multi-functional approach replaces the traditional specialized algorithms for each task, reducing system complexity while maintaining or improving speech segment continuity and naturalness
Data Source
AI summary
Systems and processes for performing unit-selection text-to-speech synthesis are provided. In one example process, a sequence of target units can represent a spoken pronunciation of text. A set of predicted acoustic model parameters of a second target unit can be determined using a set of acoustic features of a first candidate speech segment of a first target unit and a set of linguistic features of the second target unit. A likelihood score of the second candidate speech segment with respect to the first candidate speech segment can be determined using the set of predicted acoustic model parameters of the second target unit and a set of acoustic features of the second candidate speech segment of the second target unit. The second candidate speech segment can be selected for speech synthesis based on the determined likelihood score. Speech corresponding to the received text can be generated using the selected second candidate speech segment.


