Acoustic Model Learning Device for Natural Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice synthesis devices using DNN acoustic models generate synthetic voices that lack natural intonation, requiring additional processing like post-filtering, and adversarial learning can fail to accurately determine speaker information when dealing with diverse feature distributions, leading to deteriorated synthetic voices.
Innovation Solution
An acoustic model learning device that incorporates a first learning unit for estimating synthetic acoustic feature values using a voice determination model and a speaker determination model, along with a second learning unit to determine the naturalness of synthetic voices and a third learning unit to identify the speaker, all based on acoustic and language feature values, to generate high-quality synthetic voices with intonation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a DNN acoustic model is constructed to minimize mean squared error between acoustic feature values, then the synthetic acoustic feature values become smoother, but the synthetic voice loses natural voice feeling and intonation
Solution Approach 1:
The patent introduces a determination model as an intermediary component that works in conjunction with the acoustic model. The determination model receives both the synthetic acoustic feature values from the acoustic model and the target acoustic feature values, then determines whether they match. This intermediary structure allows the system to maintain the smoothing function of the acoustic model while adding a verification layer that ensures natural voice characteristics are preserved, thereby resolving the contradiction between smoothness and natural voice feeling
Solution Approach 2:
The patent implements a feedback mechanism where the determination model's output is used to guide the acoustic model's learning process. The loss function incorporates both the mean squared error term and a determination accuracy term, creating a feedback loop that continuously adjusts the acoustic model to produce synthetic feature values that are both smooth and naturally sounding. This feedback mechanism resolves the contradiction by making the system self-correcting rather than relying on one-sided optimization
2Reliability
If adversarial learning is applied to learn acoustic model and determination model alternately, then the synthetic voice quality improves, but speaker information cannot be accurately determined when feature distributions are diverse
Solution Approach 1:
The patent segments the determination task into two distinct models: a voice determination model that assesses whether synthetic acoustic feature values match target values, and a speaker determination model that identifies the speaker. This segmentation allows each model to specialize in its specific function, preventing the voice quality optimization from interfering with speaker identification accuracy, thereby resolving the contradiction between improved synthetic voice quality and accurate speaker determination
Solution Approach 2:
The patent extends the determination process from a single dimension (voice quality assessment) to multiple dimensions by introducing separate determination models for voice quality and speaker identification. This dimensional expansion allows the system to optimize voice quality through the voice determination model while simultaneously maintaining speaker identification accuracy through the speaker determination model, effectively resolving the contradiction by operating in multiple independent determination dimensions
3Reliability
If separate processing such as post-filtering is applied to generate natural-sounding voice, then the voice quality improves, but additional computational processing and time are required
Solution Approach 1:
The patent performs the determination of voice quality and speaker identity during the acoustic model learning phase itself, rather than as separate post-processing steps. The loss function incorporates determination accuracy terms that guide the acoustic model to produce naturally sounding output with accurate speaker characteristics directly during synthesis. This preliminary action eliminates the need for subsequent post-filtering and separate determination processing, resolving the contradiction by embedding the quality assurance functions within the core synthesis process
Solution Approach 2:
The patent merges the voice quality assessment and speaker identification functions into the acoustic model learning process itself. The determination models are integrated into the training loop, and their loss functions are combined with the acoustic model's loss function to create a unified optimization objective. This merging eliminates the need for separate post-processing stages, reducing device complexity while maintaining high natural voice quality through the unified multi-objective learning framework
Data Source
AI summary
An acoustic model learning device is provided for obtaining an acoustic model used to synthesize voice signals with intonation. The device includes a first learning unit that learns the acoustic model to estimate synthetic acoustic feature values using voice and speaker determination models based on acoustic feature values of speakers, language feature values corresponding to the acoustic feature values and speaker data items, a second learning unit that learns the voice determination model to determine whether the synthetic acoustic feature value is a predetermined acoustic feature value or not based on the acoustic feature values and the synthetic acoustic feature values, and a third learning unit that learns the speaker determination model to determine whether the speaker of the synthetic acoustic feature value is a predetermined speaker or not based on the acoustic feature values and the synthetic acoustic feature values.


