Speech Synthesis Weight Optimization via Mean Dissimilarity Matrix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies based on waveform concatenation require manual adjustment of weights for optimal continuity, which is time-consuming and often results in inaccurate weights, leading to unsatisfactory speech continuity effects.
Innovation Solution
A method that acquires training speech data by concatenating speech segments with the lowest target cost, extracts training speech segments with superior continuity, calculates a mean dissimilarity matrix, and generates a concatenation cost model with target weights to improve the accuracy of speech synthesis, reducing the need for manual adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual adjustment of weights in concatenation cost model is performed, then speech continuity effect can be improved, but time consumption and complexity increase significantly
Solution Approach 1:
The system performs self-learning by automatically adjusting weights through machine learning models based on training data, eliminating the need for manual weight adjustment. The concatenation cost model learns optimal weights independently by analyzing training speech data and continuously improving speech continuity effects through automated feedback mechanisms.
Solution Approach 2:
The system performs preliminary training by pre-processing training speech data, extracting acoustic features, and pre-adjusting weights before actual speech synthesis operations. This preliminary action prepares the concatenation cost model with optimized parameters in advance, reducing time consumption during runtime operations.
2Reliability
If manual adjustment of weights is performed multiple times, then speech continuity may be improved, but manufacturing precision and reliability of weight accuracy remain insufficient
Solution Approach 1:
The system implements feedback mechanisms where the machine learning model continuously evaluates speech synthesis results and automatically adjusts weights based on performance metrics. Training speech data provides feedback signals that guide the optimization process, enabling the concatenation cost model to learn optimal weights through iterative improvement rather than random manual adjustments.
Solution Approach 2:
The system dynamically changes weight parameters in the concatenation cost model based on learned patterns from training data. Instead of fixed manual weights, the system adjusts acoustic feature weights, phoneme weights, and other parameters automatically according to the specific characteristics of the input speech data, achieving higher precision through adaptive parameter optimization.
3Loss of time
If automated machine learning methods are used for weight adjustment, then time consumption is reduced, but the complexity of the system increases
Solution Approach 1:
The system performs complex training operations in advance by pre-processing training speech data, extracting acoustic features, and training the machine learning model before deployment. This preliminary action shifts computational complexity from runtime to setup phase, reducing time consumption during actual speech synthesis operations while managing system complexity through structured preprocessing pipelines.
Data Source
AI summary
A method is performed by at least one processor, and includes acquiring training speech data by concatenating speech segments having a lowest target cost among candidate concatenation solutions, and extracting training speech segments of a first annotation type, from the training speech data, the first annotation type being used for annotating that a speech continuity of a respective one of the training speech segments is superior to a preset condition. The method further includes calculating a mean dissimilarity matrix, based on neighboring candidate speech segments corresponding to the training speech segments before concatenation, the mean dissimilarity matrix representing a mean dissimilarity in acoustic features of groups of the neighboring candidate speech segments belonging to a same type of concatenation combination relationship, and generating a concatenation cost model having a target concatenation weight, based on the mean dissimilarity matrix, the concatenation cost model corresponding to the same type of concatenation combination relationship.


