Speech Feature Library for Sustained Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for dubbing and singing in films and TV dramas are limited by the aging, sickness, or death of voice actors, as they require real people to perform personalized speeches, which are not sustainable over time.
Innovation Solution
A method and apparatus for building a speech feature library that converts speech recordings into personalized textual and audio features, allowing for the analysis and storage of context and semantically identical textual information, enabling the synthesis of personalized speech even after the original voice actor is no longer available.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real voice actors are used for personalized speech in films and TV dramas, then the speech quality and emotional expression are improved, but the reliability of personalized speech is worsened due to aging, sickness, or death of the voice actors
Solution Approach 1:
The patent creates a speech feature library that stores and reproduces voice actors' speech characteristics through extracted features (pitch, energy, spectral characteristics). This digital copy allows personalized speech to be generated without the original voice actor, resolving the reliability issue while maintaining speech quality through sophisticated feature synthesis and reconstruction algorithms.
Solution Approach 2:
The system performs preliminary extraction and storage of speech features from voice actors during a preparation phase, building a speech feature library before the actors are needed for production. This advance preparation ensures that personalized speech can be reliably generated later even if the original actors are no longer available, as their speech characteristics have already been captured and stored.
2Duration of action of stationary object
If speech features are extracted and stored in a speech feature library, then the sustainability of personalized speech is improved, but the device complexity is worsened due to the need for feature extraction, storage, and synthesis systems
Solution Approach 1:
The speech signal is segmented into multiple independent features including pitch contour, energy contour, and spectral characteristics. Each feature is extracted, stored, and processed separately in the speech feature library, allowing for more efficient storage and manipulation while reducing the overall system complexity compared to storing raw speech signals.
Solution Approach 2:
The patent transforms the speech signal from its original complex waveform form into a set of simplified parametric representations (pitch, energy, spectral features). This parameter transformation reduces storage requirements and enables more efficient synthesis operations, balancing sustainability benefits with reduced system complexity.
3Measurement precision
If audio sampling is performed to obtain audio sample values, then the accuracy of speech synthesis is improved, but the loss of time is worsened due to the sampling and processing required
Solution Approach 1:
Audio sampling and feature extraction are performed in advance during the speech feature library construction phase, not during real-time synthesis. The sampled audio data is processed into feature representations and stored for later retrieval, which eliminates the time-consuming sampling and processing steps from the actual synthesis operation and maintains high accuracy.
Solution Approach 2:
The system performs comprehensive audio sampling at high resolution during the preparation phase to ensure maximum accuracy, even though this takes significant time. The trade-off is acceptable because this one-time excessive action during library construction enables fast, accurate synthesis operations thereafter, improving overall system efficiency.
Data Source
AI summary
The present invention provides a method for building a speech feature library, as well as a method, an apparatus, a device and corresponding non-volatile, non-transitory computer readable storage media for speech synthesis. Because the speech feature library used in the present invention saves at least one context corresponding to each piece of personalized textual information and at least one piece of textual information semantically identical to the personalized textual information, when performing speech synthesis, even if the provided textual information is not personalized textual information corresponding to the desired personalized speech, personalized textual information semantically identical to the textual information to be subject to speech synthesis may be first found in the speech feature library to thereby achieve personalized speech synthesis, such that use of the personalized speech will not be restricted by aging, sickness, and death of a person.


