Text-to-Speech Context Vector Determination via Phoneme and Semantic Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for converting text information into voice information through machine learning suffer from inaccuracies in determining context vectors, leading to poor sound quality and rhythm in synthesized voice.

Innovation Solution

A text information processing method that acquires phoneme and semantic vectors, determines a context vector by processing semantic and phoneme information, and synthesizes voice information using a Mel spectrum network, thereby improving the accuracy of voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning methods are used to convert text to voice, then the process can be completed, but the context vector determination is inaccurate leading to poor sound quality and rhythm

Engineering Contradiction:
Improvecontext vector determination accuracyVSAvoidvoice synthesis quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the context vector determination process into multiple components: phoneme vector encoding, semantic vector encoding, and attention mechanism-based context vector generation. By dividing the determination process, each component can be optimized independently to improve overall accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phoneme vectors and semantic vectors as intermediary elements between text input and voice output. These vectors serve as mediators that capture linguistic features and semantic information, enabling more accurate context vector determination through the attention mechanism

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If phoneme information is encoded to obtain semantic information, then voice can be synthesized, but the sound quality and rhythm are poor

Engineering Contradiction:
Improvevoice synthesis processVSAvoidsound quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements an attention mechanism that provides feedback by dynamically weighting phoneme and semantic vectors based on their relevance to the current synthesis context. This feedback loop allows the system to adjust context vector determination in real-time, improving sound quality and rhythm

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameters used in context vector determination by incorporating multiple vector types (phoneme vectors, semantic vectors) and using attention scores as dynamic parameters. This multi-parameter approach enables more precise control over voice synthesis quality

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12266344B2Text information processing method and apparatus
Publication Date: 2025.04.01 BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
  • US12266344B2 patent drawing
  • US12266344B2 patent drawing
  • US12266344B2 patent drawing

AI summary

Embodiments of the present application provide a text information processing method and apparatus, the method includes: acquiring a phoneme vector corresponding to an individual phoneme and a semantic vector corresponding to the individual phoneme in text information; acquiring first semantic information output at a last moment, wherein the first semantic information is semantic information corresponding to part of the text information in the text information, and the part of the text information is text information that has been converted into voice information; determining a context vector corresponding to a current moment according to the first semantic information, the phoneme vector corresponding to the individual phoneme and the semantic vector corresponding to the individual phoneme; and determining voice information at the current moment according to the context vector and the first semantic information.