Text-to-Speech Context Vector Determination via Phoneme and Semantic Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for converting text information into voice information through machine learning suffer from inaccuracies in determining context vectors, leading to poor sound quality and rhythm in synthesized voice.
Innovation Solution
A text information processing method that acquires phoneme and semantic vectors, determines a context vector by processing semantic and phoneme information, and synthesizes voice information using a Mel spectrum network, thereby improving the accuracy of voice synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning methods are used to convert text to voice, then the process can be completed, but the context vector determination is inaccurate leading to poor sound quality and rhythm
Solution Approach 1:
The patent segments the context vector determination process into multiple components: phoneme vector encoding, semantic vector encoding, and attention mechanism-based context vector generation. By dividing the determination process, each component can be optimized independently to improve overall accuracy
Solution Approach 2:
The patent introduces phoneme vectors and semantic vectors as intermediary elements between text input and voice output. These vectors serve as mediators that capture linguistic features and semantic information, enabling more accurate context vector determination through the attention mechanism
2Ease of operation
If phoneme information is encoded to obtain semantic information, then voice can be synthesized, but the sound quality and rhythm are poor
Solution Approach 1:
The patent implements an attention mechanism that provides feedback by dynamically weighting phoneme and semantic vectors based on their relevance to the current synthesis context. This feedback loop allows the system to adjust context vector determination in real-time, improving sound quality and rhythm
Solution Approach 2:
The patent changes the parameters used in context vector determination by incorporating multiple vector types (phoneme vectors, semantic vectors) and using attention scores as dynamic parameters. This multi-parameter approach enables more precise control over voice synthesis quality
Data Source
AI summary
Embodiments of the present application provide a text information processing method and apparatus, the method includes: acquiring a phoneme vector corresponding to an individual phoneme and a semantic vector corresponding to the individual phoneme in text information; acquiring first semantic information output at a last moment, wherein the first semantic information is semantic information corresponding to part of the text information in the text information, and the part of the text information is text information that has been converted into voice information; determining a context vector corresponding to a current moment according to the first semantic information, the phoneme vector corresponding to the individual phoneme and the semantic vector corresponding to the individual phoneme; and determining voice information at the current moment according to the context vector and the first semantic information.


