Multi-Voice Font Interpolation for TTS Emotion Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech (TTS) systems rely on single voice fonts trained with specific recording corpora, limiting flexibility and making it costly and impractical to achieve a variety of emotions and styles, especially when attempting to transplant emotions or speaking styles between voice fonts.
Innovation Solution
A multi-voice font interpolation engine that uses a text parser, characteristic predictors, and interpolators to combine and weight parameters from multiple voice fonts, allowing for the generation of speech with desired speaker characteristics and prosody, enabling the adaptation of voice styles and emotions while retaining base sound qualities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single voice font trained with a specific recording corpus is used, then the voice font strongly corresponds to the prosody and characteristics of the voice talent, but the flexibility to express rich emotion types and speaking styles is limited
Solution Approach 1:
The patent combines multiple voice fonts representing different emotions and speaking styles into a unified TTS system. By merging several pre-trained voice fonts (e.g., happy, sad, neutral styles) into a single system, the patent achieves both high voice quality accuracy (from professionally trained voice talents) and flexible emotion/style adaptation (by selecting or combining different voice fonts). This resolving the contradiction between maintaining authentic voice characteristics and enabling diverse emotional expression.
Solution Approach 2:
The TTS system is designed to perform multiple functions by supporting multiple voice fonts within a single system. The system can adaptively select or combine different voice fonts based on conversational context, user preferences, or emotional requirements, making it universally applicable across various interaction scenarios. This multi-functionality allows the system to maintain high voice quality while adapting to different emotional and stylistic needs.
2Adaptability or versatility
If recordings covering a variety of emotions and styles are obtained for multiple voices, then the flexibility and variety of TTS voices is improved, but the cost and practicality deteriorate
Solution Approach 1:
The patent creates synthetic voice representations by copying and combining characteristics from a limited set of professionally recorded voice fonts. Instead of requiring extensive original recordings for every emotion and style, the system synthesizes new voice variations by interpolating between existing voice font characteristics. This copying approach maintains high voice quality while dramatically reducing the need for costly new recording sessions.
Solution Approach 2:
The system achieves voice variety by changing parameters of existing voice fonts through interpolation techniques. By adjusting prosody parameters, pitch contours, and temporal characteristics of base voice fonts, the system generates diverse emotional and stylistic variations without requiring new recordings. This parameter-based approach allows cost-effective creation of multiple voice personalities from a limited recording corpus.
3Adaptability or versatility
If conventional voice adaptation techniques are used to transplant emotion or speaking style between voice fonts, then some adaptability is achieved, but the quality of the resulting voice fonts deteriorates
Solution Approach 1:
The patent creates composite voice representations by combining multiple voice font characteristics in a unified acoustic model. Instead of简单地 transposing emotions between separate voice fonts, the system integrates prosodic, spectral, and temporal characteristics from multiple source voice fonts into a composite representation. This composite approach maintains the natural quality of original voice talents while successfully transplanting emotional and stylistic characteristics, resolving the quality degradation issue of conventional adaptation methods.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Multi-voice font interpolation is provided. A multi-voice font interpolation engine allows the production of computer generated speech with a wide variety of speaker characteristics and/or prosody by interpolating speaker characteristics and prosody from existing fonts. Using prediction models from multiple voice fonts, the multi-voice font interpolation engine predicts values for the parameters that influence speaker characteristics and/or prosody for the phoneme sequence obtained from the text to spoken. For each parameter, additional parameter values are generated by a weighted interpolation from the predicted values. Modifying an existing voice font with the interpolated parameters changes the style and/or emotion of the speech while retaining the base sound qualities of the original voice. The multi-voice font interpolation engine allows the speaker characteristics and/or prosody to be transplanted from one voice font to another or entirely new speaker characteristics and/or prosody to be generated for an existing voice font.