Voice Morphing Database Reduction via Phoneme Unit Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice morphing systems face challenges in creating natural-sounding synthetic speech due to incomplete phoneme and diphone sets for target speakers, requiring large databases and resulting in unnatural outputs when insufficient recording data is available.

Innovation Solution

A method and apparatus that reduces the size of recorded data required by using a computer system to select and concatenate the best matching speech units from a database based on phonetic transcription, pitch, speaking rate, and formant differences, allowing for more efficient voice conversion without significant digital signal processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If unit selection synthesis uses large databases of recorded speech to achieve naturalness, then the output quality is improved, but the database size and recording time requirements increase significantly

Engineering Contradiction:
Improveoutput qualityVSAvoiddatabase size
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments speech into phoneme-level units rather than storing complete utterances or larger segments. By breaking down speech into fundamental phonemic building blocks and using concatenative synthesis with phoneme-based units, the system achieves natural speech output with a significantly reduced database size compared to traditional unit selection methods that require storing entire utterances or large speech segments.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If unit selection algorithms select segments from a place that results in less than ideal synthesis, then the database coverage is improved, but the output quality deteriorates

Engineering Contradiction:
Improvedatabase coverageVSAvoidoutput quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent transforms the selection criteria from simple unit matching to a multi-parameter optimization process. The system evaluates candidate phoneme sequences based on multiple parameters including phonetic transcription accuracy, pitch contour matching, speaking rate consistency, and formant differences. By changing the selection parameters from basic to multi-dimensional, the system achieves both broad database coverage and high output quality simultaneously.

Inventive Principle:
Principle #35Parameter changes

3Loss of time

If the target speaker records less than the requisite amount of data, then the recording time is reduced, but missing units result in incomplete or unnatural output

Engineering Contradiction:
Improverecording timeVSAvoidoutput completeness
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent makes the phoneme database universally applicable across different target speakers. Instead of requiring each speaker to record complete speech databases, the system uses a shared phoneme inventory that can be adapted to any speaker through acoustic parameter transformation. This multi-functional approach allows the same core database to serve multiple speakers, dramatically reducing individual recording requirements while maintaining output completeness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If digital signal processing is applied to recorded speech to achieve voice conversion, then the voice transformation capability is improved, but the naturalness of the output deteriorates

Engineering Contradiction:
Improvevoice transformation capabilityVSAvoidnaturalness
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical digital signal processing methods with a phoneme-based concatenative synthesis approach. Instead of using DSP filters and transformations that can introduce artifacts and reduce naturalness, the system concatenates recorded phoneme segments that naturally occur in speech. This substitution of the synthesis mechanism preserves the natural acoustic characteristics of human speech while still achieving voice transformation through selective unit concatenation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10008216B2Method and apparatus for exemplary morphing computer system background
Publication Date: 2018.06.26 SPEECH MORPHING SYST
  • US10008216B2 patent drawing
  • US10008216B2 patent drawing
  • US10008216B2 patent drawing

AI summary

Method and apparatus for reducing a size of databases required for recorded speech data.