Speech Feature Library for Sustained Voice Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for dubbing and singing in films and TV dramas are limited by the aging, sickness, or death of voice actors, as they require real people to perform personalized speeches, which are not sustainable over time.

Innovation Solution

A method and apparatus for building a speech feature library that converts speech recordings into personalized textual and audio features, allowing for the analysis and storage of context and semantically identical textual information, enabling the synthesis of personalized speech even after the original voice actor is no longer available.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real voice actors are used for personalized speech in films and TV dramas, then the speech quality and emotional expression are improved, but the reliability of personalized speech is worsened due to aging, sickness, or death of the voice actors

Engineering Contradiction:
Improvespeech qualityVSAvoidpersonalized speech availability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent creates a speech feature library that stores and reproduces voice actors' speech characteristics through extracted features (pitch, energy, spectral characteristics). This digital copy allows personalized speech to be generated without the original voice actor, resolving the reliability issue while maintaining speech quality through sophisticated feature synthesis and reconstruction algorithms.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary extraction and storage of speech features from voice actors during a preparation phase, building a speech feature library before the actors are needed for production. This advance preparation ensures that personalized speech can be reliably generated later even if the original actors are no longer available, as their speech characteristics have already been captured and stored.

Inventive Principle:
Principle #10Preliminary action

2Duration of action of stationary object

If speech features are extracted and stored in a speech feature library, then the sustainability of personalized speech is improved, but the device complexity is worsened due to the need for feature extraction, storage, and synthesis systems

Engineering Contradiction:
Improvepersonalized speech sustainabilityVSAvoidspeech feature library system
Core Design Contradiction:
Duration of action of stationary objectVSDevice complexity

Solution Approach 1:

The speech signal is segmented into multiple independent features including pitch contour, energy contour, and spectral characteristics. Each feature is extracted, stored, and processed separately in the speech feature library, allowing for more efficient storage and manipulation while reducing the overall system complexity compared to storing raw speech signals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the speech signal from its original complex waveform form into a set of simplified parametric representations (pitch, energy, spectral features). This parameter transformation reduces storage requirements and enables more efficient synthesis operations, balancing sustainability benefits with reduced system complexity.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If audio sampling is performed to obtain audio sample values, then the accuracy of speech synthesis is improved, but the loss of time is worsened due to the sampling and processing required

Engineering Contradiction:
Improvespeech synthesis accuracyVSAvoidsampling and processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Audio sampling and feature extraction are performed in advance during the speech feature library construction phase, not during real-time synthesis. The sampled audio data is processed into feature representations and stored for later retrieval, which eliminates the time-consuming sampling and processing steps from the actual synthesis operation and maintains high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs comprehensive audio sampling at high resolution during the preparation phase to ensure maximum accuracy, even though this takes significant time. The trade-off is acceptable because this one-time excessive action during library construction enables fast, accurate synthesis operations thereafter, improving overall system efficiency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9697819B2Method for building a speech feature library, and method, apparatus, device, and computer readable storage media for speech synthesis
Publication Date: 2017.07.04 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US9697819B2 patent drawing
  • US9697819B2 patent drawing
  • US9697819B2 patent drawing

AI summary

The present invention provides a method for building a speech feature library, as well as a method, an apparatus, a device and corresponding non-volatile, non-transitory computer readable storage media for speech synthesis. Because the speech feature library used in the present invention saves at least one context corresponding to each piece of personalized textual information and at least one piece of textual information semantically identical to the personalized textual information, when performing speech synthesis, even if the provided textual information is not personalized textual information corresponding to the desired personalized speech, personalized textual information semantically identical to the textual information to be subject to speech synthesis may be first found in the speech feature library to thereby achieve personalized speech synthesis, such that use of the personalized speech will not be restricted by aging, sickness, and death of a person.