Representative Speech Unit Waveform Storage for Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis methods require manual division of speech into units, leading to inefficient usage and increased data storage needs, as the same speech units are often reused across sentences.
Innovation Solution
A method and apparatus that automatically divide speech waveforms into units based on phonologic and prosody information, identifying and storing representative units with similar waveforms to reduce data storage and enhance efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual division of speech into units is used, then speech units can be created, but usage efficiency is low and data storage requirements increase
Solution Approach 1:
The patent creates representative speech units that serve as templates or copies for multiple similar speech instances. Instead of storing every individual speech unit, the system identifies similar speech patterns and stores only the representative version, which can then be reused multiple times through concatenation to synthesize different sentences.
Solution Approach 2:
The patent makes speech units universally applicable by creating representative units that can serve multiple functions. A single representative speech unit can be used in various contexts and combined with different other units to generate multiple different sentences, thereby increasing usage efficiency and reducing the total number of units needed.
2Reliability
If all sentences are stored as speech, then complete speech output is achieved, but data storage quantity becomes excessively large
Solution Approach 1:
The patent divides speech into smaller modular units that can be independently stored and recombined. Instead of storing complete sentences as single speech files, the system segments speech into reusable units that can be concatenated in various sequences to form different sentences, reducing storage requirements while maintaining output completeness.
Solution Approach 2:
The patent implements a hierarchical structure where representative speech units contain or represent multiple similar speech instances. This nested approach allows the system to store a compact representative version that embodies multiple variations, enabling complete speech output through systematic combination rather than storing every possible sentence individually.
Data Source
AI summary
According to one embodiment, a method for editing speech is disclosed. The method can generate speech information from a text. The speech information includes phonologic information and prosody information. The method can divide the speech information into a plurality of speech units, based on at least one of the phonologic information and the prosody information. The method can search at least two speech units from the plurality of speech units. At least one of the phonologic information and the prosody information in the at least two speech units are identical or similar. In addition, the method can store a speech unit waveform corresponding to one of the at least two speech units as a representative speech unit into a memory.


