Sequential Prosody Feature Text-to-Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech synthesis methods struggle to finely control prosody, leading to loss of information and inaccurate representation of emotions or intentions, especially when applying fixed-length prosody features to varying text lengths and across different speakers.
Innovation Solution
A method and system that input sequential prosody features to an artificial neural network text-to-speech synthesis model, using an attention module to match variable-length prosody features to the input text and normalize embedding vectors, allowing for precise control of prosody and adaptation across speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If fixed-length prosody features are applied to text-to-speech synthesis, then the synthesis process is simplified, but information loss occurs and prosody control precision deteriorates
Solution Approach 1:
The patent transforms the static fixed-length prosody feature into a dynamic variable-length feature that adapts to different text lengths. The prosody feature is extracted and processed to maintain its sequential structure, allowing it to dynamically adjust to the input text length while preserving all original prosodic information including pitch, energy, and timing characteristics.
Solution Approach 2:
The patent segments the prosody feature into multiple dimensional components including pitch, energy, and timing information. By dividing the prosody feature into these segments and processing each dimension separately through the neural network, the system preserves detailed prosodic information while maintaining computational efficiency.
2Ease of operation
If fixed-length prosody features are used, then processing is easier, but prosody control precision at specific time points deteriorates
Solution Approach 1:
The system employs dynamic variable-length prosody features that maintain temporal correspondence with the input text. The prosody feature length automatically adapts to match the text length, enabling precise control of prosody at specific time points while preserving ease of processing through the neural network's inherent ability to handle variable-length sequences.
Solution Approach 2:
The patent introduces a temporal dimension to the prosody feature by maintaining its sequential structure. This allows the system to control prosody not only in terms of overall characteristics but also at specific time points within the speech sequence, achieving fine-grained prosodic control without complicating the processing pipeline.
3Device complexity
If prosody features are applied without preprocessing, then the process is simpler, but synthesis quality deteriorates when pitch ranges differ between speakers
Solution Approach 1:
The patent applies parameter transformation to normalize prosody features from different speakers. By changing the parameter space of the prosody features through statistical normalization and speaker-adaptive processing, the system maintains high synthesis quality when transferring prosody between speakers with different pitch ranges, while avoiding overly complex preprocessing pipelines.
4Manufacturing precision
If variable-length sequential prosody features are used, then prosody control precision is improved, but device complexity increases
Solution Approach 1:
The patent replaces complex mechanical preprocessing systems with a neural network-based approach. The neural network automatically handles variable-length sequential prosody features through its inherent ability to process sequences, eliminating the need for manual feature alignment and length normalization while maintaining high prosody control precision.
Data Source
AI summary
The present disclosure relates to a text-to-speech synthesis method using machine learning based on a sequential prosody feature. The text-to-speech synthesis method includes receiving input text, receiving a sequential prosody feature, and generating output speech data for the input text reflecting the received sequential prosody feature by inputting the input text and the received sequential prosody feature to an artificial neural network text-to-speech synthesis model.


