Speech Synthesis Using Unified Decoding and Attention Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies face issues with rhythm inconsistency and poor speech quality due to independent processes for generating recordings and synthesized speech, leading to discrepancies in speech speed and tone.

Innovation Solution

A method and apparatus that utilize a preset decoding model, attention model, and feature vector set to predict acoustic features for synthesized speech, incorporating both recorded and query result statements, ensuring consistent rhythm and improved quality by performing feature conversion and synthesis on predicted acoustic features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If independent processes are used for generating recording and synthesized speech, then the processes can be processed separately and efficiently, but the speech speed and tone become inconsistent leading to rhythm inconsistency

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidrhythm consistency
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent merges the previously independent recording generation process and synthesized speech generation process into a unified speech synthesis system. The decoding model integrates both processes, using the same acoustic feature predictions for both the recorded statement portion and the query result statement portion, ensuring consistent speech speed and tone throughout the complete speech.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If recording and synthesized speech are generated independently, then processing can be simplified, but transition time uncertainty increases and speech quality deteriorates

Engineering Contradiction:
Improveprocess complexityVSAvoidspeech quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the decoding model uses attention models to dynamically adjust acoustic feature predictions based on the input symbol sequence. The model continuously refines its predictions by comparing expected outputs with actual outputs, ensuring consistent speech characteristics across the transition from recorded to synthesized portions, thereby improving speech quality and reducing transition time uncertainty.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12198674B2Speech synthesis method and apparatus, and storage medium
Publication Date: 2025.01.14 BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
  • US12198674B2 patent drawing
  • US12198674B2 patent drawing
  • US12198674B2 patent drawing

AI summary

Disclosed are a speech synthesis method and apparatus, and a storage medium. The method comprises: acquiring a symbol sequency of a statement to be synthesized, wherein the statement to be synthesized comprises a recorded statement characterizing a target object and a query result statement for the target object; encoding the symbol sequence by using a pre-set encoding model, in order to obtain a feature vector set; acquiring recording acoustic features corresponding to the recorded statement; predicting, according to a pre-set decoding model, the feature vector set, a pre-set attention model and the recording acoustic features, acoustic features corresponding to the statement to be synthesized, in order to obtain predicted acoustic features corresponding to the statement to be synthesized, wherein the pre-set attention model is a model that uses the feature vector set to generate a context vector used for decoding, and the predicted acoustic features are composed of at least one associated acoustic feature; and performing feature conversion and synthesis on the predicted acoustic features to obtain a speech corresponding to the sentence to be synthesized.