Speech Synthesis Acoustic Feature Adjustment for Pronunciation and Prosody
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deep neural network (DNN) speech synthesis technologies face challenges in efficiently adjusting synthetic speech for incorrect pronunciation and unnatural prosody, requiring repetitive user input for similar adjustments.
Innovation Solution
A speech synthesis device that utilizes encoder-decoder neural networks to convert speech unit attributes into intermediate representations, applying adjustment dictionaries to adjust duration and acoustic features based on attribute information and cluster numbers, thereby reducing the need for repetitive user input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional DNN speech synthesis is used to achieve high sound quality, then synthetic speech quality is improved, but adjustment efficiency deteriorates due to difficult and repetitive user input requirements
Solution Approach 1:
The system performs preliminary analysis of the input text to automatically identify sections requiring adjustment (such as potential misreadings or prosody issues) before synthesis. By pre-defining adjustment targets based on text analysis, the system eliminates the need for users to repeatedly manually specify adjustment regions, thereby improving adjustment efficiency while maintaining speech quality
Solution Approach 2:
The speech synthesis device automatically performs adjustment processing without requiring continuous user input. The system uses its own text analysis capabilities to identify problems and applies adjustments autonomously based on predefined rules or models, enabling the system to serve itself in the adjustment process and significantly reducing user burden
2Manufacturing precision
If manual adjustment processing is performed for each synthesis case, then speech quality can be improved, but user workload increases due to repetitive input
Solution Approach 1:
The system implements a feedback mechanism where the text analysis results automatically inform the adjustment processing. The analysis output (identifying problematic sections) feeds into the adjustment module, which applies corrections without requiring user confirmation for each case. This automated feedback loop maintains speech quality while dramatically reducing user workload
Solution Approach 2:
The patent introduces an intermediate text analysis step that acts as a mediator between the input text and the synthesis process. This intermediary analysis automatically identifies sections needing adjustment and prepares adjustment instructions, serving as a bridge that eliminates the need for direct user intervention in each synthesis case while maintaining precision
Data Source
AI summary
A speech synthesis device according to an embodiment includes a memory and a hardware processor connected to the memory. The processor executes encoder processing with a first neural network to convert attribute information of a speech unit into an intermediate representation. The processor executes decoder processing with a second neural network to generate an acoustic feature from the intermediate representation. The processor executes adjustment processing by using an adjustment dictionary in which at least the attribute information of the speech unit is set as a key and an adjustment instruction to the acoustic feature is set as a value. The processor executes the adjustment processing by defining, by the key, a section to which the adjustment instruction is applied, and adjusting the acoustic feature in the defined section based on the adjustment instruction.


