Speech Library Expansion for Coherent Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies, particularly in star and personalized scenarios, face challenges with coherence and naturalness due to limited speech data, making it impractical to record large-scale corpora, which results in incoherent synthesis effects.

Innovation Solution

Expanding a speech library using a pre-trained speech synthesis model and synthesized text, incorporating manually-collected original materials and updating the library with synthesized speech, allowing for more extensive language materials and improved synthesis quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large-scale speech corpus is recorded to improve synthesis quality, then the coherence and naturalness of speech synthesis is improved, but the cost and time required for data collection increases significantly

Engineering Contradiction:
Improvecoherence and naturalness of speech synthesisVSAvoidtime required for data collection
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training a speech synthesis model using available speech data before performing the actual speech splicing and synthesis. This pre-trained model is then used to generate synthesized speech that expands the limited speech library, enabling high-quality synthesis without requiring extensive manual data collection at the time of use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by generating synthesized speech copies from the pre-trained model and adding them to the speech library as synthesized language materials. These synthesized speech segments serve as virtual copies that expand the effective size of the speech library without requiring additional physical recording sessions.

Inventive Principle:
Principle #26Copying

2Reliability

If a large-scale speech corpus is recorded to improve synthesis quality, then the coherence and naturalness of speech synthesis is improved, but the cost of recording increases

Engineering Contradiction:
Improvecoherence and naturalness of speech synthesisVSAvoidcost of data collection
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent uses copying by generating synthesized speech copies from the pre-trained model and adding them to the speech library as synthesized language materials. These synthesized speech segments serve as virtual copies that expand the effective size of the speech library without requiring additional physical recording sessions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system applies self-service by using the pre-trained speech synthesis model to automatically generate the necessary speech data. The model serves itself by producing synthesized speech that expands the library, eliminating the need for external human recorders and manual data collection processes.

Inventive Principle:
Principle #25Self-service

3Productivity

If only a small amount of speech data is collected to reduce cost and time, then the cost and time are reduced, but the speech synthesis effect becomes incoherent and unnatural

Engineering Contradiction:
Improveefficiency of data collectionVSAvoidcoherence and naturalness of speech synthesis
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training a speech synthesis model using available speech data before performing the actual speech splicing and synthesis. This pre-trained model is then used to generate synthesized speech that expands the limited speech library, enabling high-quality synthesis without requiring extensive manual data collection at the time of use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by transforming the speech library from a static collection of manually recorded segments to a dynamic expanded library that includes synthesized speech generated by the pre-trained model. This changes the composition and size parameters of the speech library, enabling sufficient speech segments for coherent synthesis even when original recorded data is limited.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If manually-collected original language materials are used to train the speech synthesis model, then the model can be trained with available data, but the speech library remains too small for effective splicing and synthesis

Engineering Contradiction:
Improvefeasibility of model trainingVSAvoidsize of speech library
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training a speech synthesis model using available speech data before performing the actual speech splicing and synthesis. This pre-trained model is then used to generate synthesized speech that expands the limited speech library, enabling high-quality synthesis without requiring extensive manual data collection at the time of use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by generating synthesized speech copies from the pre-trained model and adding them to the speech library as synthesized language materials. These synthesized speech segments serve as virtual copies that expand the effective size of the speech library without requiring additional physical recording sessions.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10803851B2Method and apparatus for processing speech splicing and synthesis, computer device and readable medium
Publication Date: 2020.10.13 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10803851B2 patent drawing
  • US10803851B2 patent drawing
  • US10803851B2 patent drawing

AI summary

The present disclosure provides a method for processing speech splicing and synthesis and apparatus, a computer device and a readable medium. The method comprises: expanding a speech library according to a pre-trained speech synthesis model and an obtained synthesized text; the speech library before the expansion comprises manually-collected original language materials; using the expanded speech library to perform speech splicing and synthesis processing. According to the technical solution of the present embodiment, the speech library is expanded so that the speech library includes sufficient language materials. As such, when speech splicing processing is performed according to the expanded speech library, it is possible to select more speech segments, and thereby improve coherence and naturalness of the effect of speech synthesis so that the speech synthesis effect is very coherent with very good naturalness and can sufficiently satisfy the user's normal use.