Speech Synthesis Splicing Acoustic Features for Authenticity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis methods lack authenticity and naturalness in the synthesized speech, failing to effectively replicate the speech features of a target user.

Innovation Solution

A speech synthesis method that involves obtaining text and speech features of a target user, predicting acoustic features based on the text and speech features, extracting acoustic features from a template audio, splicing these features to generate target acoustic features, and performing speech synthesis to produce authentic and natural speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing speech synthesis methods are used to convert text into audio with target user speech features, then the synthesis process can be completed, but the authenticity and naturalness of the synthesized speech are poor

Engineering Contradiction:
Improveauthenticity and naturalness of synthesized speechVSAvoidcomplexity of speech synthesis process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the acoustic feature generation process into multiple components: extracting acoustic features from template audio, predicting acoustic features from text and speech features, and splicing these features together. This segmentation allows each component to be optimized independently while achieving superior overall speech synthesis quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates composite acoustic features by splicing together different acoustic feature segments from template audio and predicted acoustic features. This composite approach combines the strengths of both sources to generate authentic and natural speech that accurately reflects the target user's speech characteristics.

Inventive Principle:
Principle #40Composite materials

2Manufacturing precision

If template audio is used to extract acoustic features, then the authenticity of speech can be improved, but the process complexity increases

Engineering Contradiction:
Improveauthenticity of synthesized speechVSAvoidcomplexity of feature extraction and splicing process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces acoustic feature extraction as an intermediary process between template audio and the final synthesized speech. By extracting acoustic features from template audio and then splicing them with predicted acoustic features, the system achieves authentic speech while managing process complexity through structured intermediate steps.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary extraction of acoustic features from template audio before the final synthesis process. This preliminary action prepares the acoustic feature data in advance, allowing for more efficient and accurate splicing with predicted acoustic features to generate authentic speech.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12211485B2Speech synthesis method, and electronic device
Publication Date: 2025.01.28 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12211485B2 patent drawing
  • US12211485B2 patent drawing
  • US12211485B2 patent drawing

AI summary

The disclosure provides a speech synthesis method, and an electronic device. The technical solution is described as follows. A text to be synthesized and speech features of a target user are obtained. Predicted first acoustic features based on the text to be synthesized and the speech features are obtained. A target template audio is obtained from a template audio library based on the text to be synthesized. Second acoustic features of the target template audio are extracted. Target acoustic features are generated by splicing the first acoustic features and the second acoustic features. Speech synthesis is performed on the text to be synthesized based on the target acoustic features and the speech features, to generate a target speech of the text to be synthesized.