Voice Quality Preference Learning via Eigenvoice Dimensionality Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis technologies require significant trial and error to find a user-preferred voice quality due to the large number of adjustable parameters, making it inefficient to create personalized voices.
Innovation Solution
A voice quality preference learning device that dimensionally reduces acoustic models into a lower-dimensional voice quality space using eigenvoices, allowing users to input preferences and learn preference models to efficiently synthesize desired voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the number of voice quality parameters is increased to provide more adjustable control, then the flexibility and adaptability of voice synthesis is improved, but the complexity of the system increases and enormous trial and error is required
Solution Approach 1:
The patent applies dimensionality reduction by transforming the high-dimensional acoustic model parameter space into a lower-dimensional voice quality space. This allows the system to maintain flexibility in voice quality control while reducing system complexity. The dimensionality reduction technique projects multiple voice quality parameters onto a smaller set of representative vectors, enabling efficient manipulation of voice characteristics without requiring adjustment of numerous individual parameters.
2Adaptability or versatility
If the number of voice quality parameters is increased to create diverse voice combinations, then the versatility of voice synthesis is improved, but the time required to find a preferred voice increases significantly
Solution Approach 1:
The patent implements preliminary action by pre-computing and storing representative voice quality vectors in advance. Instead of requiring users to perform trial and error adjustments in real-time, the system prepares a set of representative voice quality representations beforehand. During actual use, the system quickly matches user preferences against these pre-computed vectors, dramatically reducing the time required to find a preferred voice while maintaining diverse voice quality options.
3Manufacturing precision
If high-dimensional acoustic models are used to achieve precise voice quality control, then the manufacturing precision of synthetic voices is improved, but the ease of operation deteriorates due to difficulty in managing numerous parameters
Solution Approach 1:
The patent extracts the essential voice quality characteristics from high-dimensional acoustic models by identifying and retaining only the most significant representative vectors. This extraction process separates the crucial voice quality information from the redundant parameters, allowing precise voice quality control to be achieved through manipulation of a small number of extracted features rather than numerous detailed parameters, thereby improving ease of operation.
Data Source
AI summary
A voice quality preference learning device according to an embodiment includes a storage, a user interface system, and a learning processor. The storage stores a plurality of acoustic models. The user interface system receives an operation input indicating a voice quality preference of a user for voice quality. The learning processor learns a preference model corresponding to the voice quality preference of the user based at least in part on the operation input, the operation input associated with a voice quality space, wherein the voice quality space is obtained by dimensionally reducing the plurality of acoustic models.


