Representative Speech Unit Waveform Storage for Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesis methods require manual division of speech into units, leading to inefficient usage and increased data storage needs, as the same speech units are often reused across sentences.

Innovation Solution

A method and apparatus that automatically divide speech waveforms into units based on phonologic and prosody information, identifying and storing representative units with similar waveforms to reduce data storage and enhance efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual division of speech into units is used, then speech units can be created, but usage efficiency is low and data storage requirements increase

Engineering Contradiction:
Improveusage efficiency of speech unitsVSAvoiddata storage requirements
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent creates representative speech units that serve as templates or copies for multiple similar speech instances. Instead of storing every individual speech unit, the system identifies similar speech patterns and stores only the representative version, which can then be reused multiple times through concatenation to synthesize different sentences.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent makes speech units universally applicable by creating representative units that can serve multiple functions. A single representative speech unit can be used in various contexts and combined with different other units to generate multiple different sentences, thereby increasing usage efficiency and reducing the total number of units needed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If all sentences are stored as speech, then complete speech output is achieved, but data storage quantity becomes excessively large

Engineering Contradiction:
Improvespeech output completenessVSAvoiddata storage quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent divides speech into smaller modular units that can be independently stored and recombined. Instead of storing complete sentences as single speech files, the system segments speech into reusable units that can be concatenated in various sequences to form different sentences, reducing storage requirements while maintaining output completeness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a hierarchical structure where representative speech units contain or represent multiple similar speech instances. This nested approach allows the system to store a compact representative version that embodies multiple variations, enabling complete speech output through systematic combination rather than storing every possible sentence individually.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS8868422B2Storing a representative speech unit waveform for speech synthesis based on searching for similar speech units
Publication Date: 2014.10.21 TOSHIBA DIGITAL SOLUTIONS CORP
  • US8868422B2 patent drawing
  • US8868422B2 patent drawing
  • US8868422B2 patent drawing

AI summary

According to one embodiment, a method for editing speech is disclosed. The method can generate speech information from a text. The speech information includes phonologic information and prosody information. The method can divide the speech information into a plurality of speech units, based on at least one of the phonologic information and the prosody information. The method can search at least two speech units from the plurality of speech units. At least one of the phonologic information and the prosody information in the at least two speech units are identical or similar. In addition, the method can store a speech unit waveform corresponding to one of the at least two speech units as a representative speech unit into a memory.