Concatenation-Sensitive Neural Networks for Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current unit-selection text-to-speech synthesis methods require complex algorithms to optimize weights for concatenation and target costs, making the process cumbersome and prone to inaccuracies in selecting natural-sounding speech segments.

Innovation Solution

The use of concatenation-sensitive neural networks determines predicted acoustic model parameters for speech segments based on both linguistic and acoustic features, eliminating the need for separate concatenation cost optimization and simplifying the unit-selection process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a sophisticated unit-selection algorithm is implemented to select the most suitable speech units, then the acoustic quality and naturalness of synthesized speech is improved, but the algorithm complexity and computational overhead increases

Engineering Contradiction:
Improveacoustic qualityVSAvoidalgorithm complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the traditional mechanical unit-selection algorithm with a neural network-based acoustic model. The neural network learns optimal speech unit selection through training on acoustic data, substituting the manual rule-based mechanical selection process with an intelligent system that automatically identifies and selects the most suitable speech units based on acoustic similarity, thereby maintaining high acoustic quality while reducing algorithmic complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the unit-selection problem from a complex multi-parameter optimization task into a simpler acoustic similarity matching task. By changing the selection criterion from multiple weighted parameters (phonetic accuracy, prosody, timing) to a single dominant acoustic similarity metric computed by the neural network, the system achieves high speech quality with reduced computational complexity

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If separate concatenation cost optimization is performed to ensure smooth transitions between speech segments, then the naturalness of synthesized speech is improved, but the overall system complexity and number of optimization parameters increases

Engineering Contradiction:
Improvespeech segment continuityVSAvoidoptimization parameters
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the previously separate concatenation cost optimization into the unified acoustic similarity metric computed by the neural network. The acoustic model is trained to simultaneously optimize both the acoustic quality of individual speech units and the smoothness of transitions between them, eliminating the need for separate concatenation cost calculations and reducing the total number of optimization parameters

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network-based acoustic model serves multiple functions simultaneously: it performs speech unit selection, acoustic quality assessment, and concatenation smoothness evaluation. This multi-functional approach replaces the traditional specialized algorithms for each task, reducing system complexity while maintaining or improving speech segment continuity and naturalness

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9697820B2Unit-selection text-to-speech synthesis using concatenation-sensitive neural networks
Publication Date: 2017.07.04 APPLE INC
  • US9697820B2 patent drawing
  • US9697820B2 patent drawing
  • US9697820B2 patent drawing

AI summary

Systems and processes for performing unit-selection text-to-speech synthesis are provided. In one example process, a sequence of target units can represent a spoken pronunciation of text. A set of predicted acoustic model parameters of a second target unit can be determined using a set of acoustic features of a first candidate speech segment of a first target unit and a set of linguistic features of the second target unit. A likelihood score of the second candidate speech segment with respect to the first candidate speech segment can be determined using the set of predicted acoustic model parameters of the second target unit and a set of acoustic features of the second candidate speech segment of the second target unit. The second candidate speech segment can be selected for speech synthesis based on the determined likelihood score. Speech corresponding to the received text can be generated using the selected second candidate speech segment.