Prosody Generator Using Density-Based Method Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating prosody information in speech synthesis, whether using rules or statistical techniques, face challenges in producing natural-sounding speech, particularly with sparse data, leading to unnatural sound and displaced accent positions, and require large datasets that are difficult to collect.

Innovation Solution

A prosody generator that divides the data space into subspaces and extracts density information to selectively choose between statistical and rule-based methods for prosody information generation, using statistical techniques for dense areas and rule-based methods for sparse areas, thereby avoiding the need for extensive data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If statistical techniques are used for prosody generation, then human-like prosody information is generated, but sparse data portions cause incorrect F0 patterns and disturbed accent positions

Engineering Contradiction:
Improveaccuracy of F0 patternsVSAvoidquantity of learning data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The learning data space is divided into multiple subspaces based on feature quantity characteristics. Each subspace is processed independently with appropriate method selection, allowing dense subspaces to use statistical techniques while sparse subspaces use rule-based methods, thereby resolving the contradiction between data quantity requirements and accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different prosody generation methods are applied to different subspaces based on their local data density characteristics. Dense subspaces receive statistical processing while sparse subspaces receive rule-based processing, ensuring each region gets the most appropriate treatment for its specific conditions rather than applying a uniform approach.

Inventive Principle:
Principle #3Local quality

2Ease of manufacture

If rule-based methods are used for prosody generation, then simple F0 pattern generation is achieved, but the synthesized speech sounds mechanical and unnatural

Engineering Contradiction:
Improvesimplicity of F0 pattern generationVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The prosody generation process is segmented into different method applications: rule-based methods handle simple, sparse cases while statistical methods handle complex, dense cases. This segmentation allows the system to leverage the simplicity of rules where appropriate while using statistics where naturalness is critical.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects between rule-based and statistical methods based on the characteristics of each subspace. This dynamic adaptation allows the system to optimize between simplicity and naturalness in real-time, choosing the most appropriate method for each specific situation rather than being fixed to one approach.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If data space is divided into clusters based on information quantity, then statistical processing becomes feasible, but sparse and dense portions are created causing inconsistent prosody quality

Engineering Contradiction:
Improveflexibility of statistical processingVSAvoidconsistency of prosody quality
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The system applies different processing qualities to different subspaces: statistical processing for dense subspaces and rule-based processing for sparse subspaces. This local differentiation maintains consistency within each subspace while accommodating the varying data densities across the entire data space.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

By segmenting the data space and applying appropriate methods to each segment, the system maintains stable, consistent prosody quality within each subspace. The segmentation prevents the inconsistency that would arise from applying a single method uniformly across subspaces with vastly different data densities.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9324316B2Prosody generator, speech synthesizer, prosody generating method and prosody generating program
Publication Date: 2016.04.26 NEC CORP
  • US9324316B2 patent drawing
  • US9324316B2 patent drawing
  • US9324316B2 patent drawing

AI summary

There is provided a prosody generator that generates prosody information for implementing highly natural speech synthesis without unnecessarily collecting large quantities of learning data. A data dividing means 81 divides into subspaces the data space of a learning database as an assembly of learning data indicative of the feature quantities of speech waveforms. A density information extracting means 82 extracts density information indicative of the density state in terms of information quantity of the learning data in each of the subspaces divided by the data dividing means 81. A prosody information generating method selecting means 83 selects either a first method or a second method as a prosody information generating method based on the density information, the first method involving generating the prosody information using a statistical technique, the second method involving generating the prosody information using rules based on heuristics.