Voice Quality Preference Learning via Eigenvoice Dimensionality Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis technologies require significant trial and error to find a user-preferred voice quality due to the large number of adjustable parameters, making it inefficient to create personalized voices.

Innovation Solution

A voice quality preference learning device that dimensionally reduces acoustic models into a lower-dimensional voice quality space using eigenvoices, allowing users to input preferences and learn preference models to efficiently synthesize desired voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the number of voice quality parameters is increased to provide more adjustable control, then the flexibility and adaptability of voice synthesis is improved, but the complexity of the system increases and enormous trial and error is required

Engineering Contradiction:
Improvevoice quality control flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dimensionality reduction by transforming the high-dimensional acoustic model parameter space into a lower-dimensional voice quality space. This allows the system to maintain flexibility in voice quality control while reducing system complexity. The dimensionality reduction technique projects multiple voice quality parameters onto a smaller set of representative vectors, enabling efficient manipulation of voice characteristics without requiring adjustment of numerous individual parameters.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the number of voice quality parameters is increased to create diverse voice combinations, then the versatility of voice synthesis is improved, but the time required to find a preferred voice increases significantly

Engineering Contradiction:
Improvevoice quality varietyVSAvoidtime for preference learning
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-computing and storing representative voice quality vectors in advance. Instead of requiring users to perform trial and error adjustments in real-time, the system prepares a set of representative voice quality representations beforehand. During actual use, the system quickly matches user preferences against these pre-computed vectors, dramatically reducing the time required to find a preferred voice while maintaining diverse voice quality options.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If high-dimensional acoustic models are used to achieve precise voice quality control, then the manufacturing precision of synthetic voices is improved, but the ease of operation deteriorates due to difficulty in managing numerous parameters

Engineering Contradiction:
Improvevoice quality precisionVSAvoidparameter adjustment ease
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent extracts the essential voice quality characteristics from high-dimensional acoustic models by identifying and retaining only the most significant representative vectors. This extraction process separates the crucial voice quality information from the redundant parameters, allowing precise voice quality control to be achieved through manipulation of a small number of extracted features rather than numerous detailed parameters, thereby improving ease of operation.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10930264B2Voice quality preference learning device, voice quality preference learning method, and computer program product
Publication Date: 2021.02.23 TOSHIBA DIGITAL SOLUTIONS CORP
  • US10930264B2 patent drawing
  • US10930264B2 patent drawing
  • US10930264B2 patent drawing

AI summary

A voice quality preference learning device according to an embodiment includes a storage, a user interface system, and a learning processor. The storage stores a plurality of acoustic models. The user interface system receives an operation input indicating a voice quality preference of a user for voice quality. The learning processor learns a preference model corresponding to the voice quality preference of the user based at least in part on the operation input, the operation input associated with a voice quality space, wherein the voice quality space is obtained by dimensionally reducing the plurality of acoustic models.