Synthetic Voice Parameter Search for Minimal Recording Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for creating synthetic voices require direct interaction or access to voice recordings of the person, which can be problematic when recordings are unavailable.
Innovation Solution
A method for determining synthetic voice parameters by iteratively searching, mixing, and adjusting parameterized voices within a 2D search space to mimic a target voice, using techniques like UMAP projection and LSTM networks to identify and adjust underlying voice parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional voice parameterization methods are used requiring direct access to voice recordings, then voice authenticity is improved, but accessibility and ease of operation deteriorate when recordings are unavailable
Solution Approach 1:
The patent introduces an intermediary approach by using a small set of reference voice samples (as little as 5 seconds) combined with extensive parameterization techniques to create a comprehensive voice model. This intermediary method bridges the gap between having no recordings and having extensive recordings, allowing authentic voice synthesis with minimal input data through iterative parameter adjustment and mixing.
2Measurement precision
If extensive voice recordings are collected for accurate parameterization, then voice resemblance quality is improved, but time consumption and complexity increase
Solution Approach 1:
The patent applies preliminary action by pre-processing and pre-analyzing the minimal voice samples to extract comprehensive parameter sets in advance. The system performs upfront parameter identification, feature extraction, and model training on the small reference set, creating a ready-to-use parameterized voice model that can be quickly adjusted and mixed without requiring additional time-consuming data collection during the synthesis process.
3Ease of operation
If minimal voice samples are used for parameterization, then ease of operation is improved, but voice parameterization accuracy deteriorates
Solution Approach 1:
The patent employs parameter changes by extensively manipulating and adjusting numerous voice parameters through iterative processes. The system varies parameters such as pitch, timbre, prosody, and spectral characteristics through mixing different reference samples and adjusting parameter values to achieve accurate voice resemblance. This comprehensive parameter exploration compensates for the minimal input data, transforming a small sample set into a richly parameterized voice model through systematic parameter variation and optimization.
Data Source
AI summary
Methods, device and system for determining a target voice parameters. A location within a 2D search space is assigned to parameterized voices, perceptually similar voices being proximate. Candidate-voices are inserted into a candidate list when a resemblance threshold is reached; A choice between two unmixed voices is received. The plurality of underlying parameters of the unmixed voices are mixed into a mixed voice towards the target-voice. The plurality of underlying parameters from the candidate list are identified. The unadjusted voice is adjusted into an adjusted voice by altering values of the plurality of underlying parameters towards the target-voice. A user interface module receives a choice of a candidate-voice from the 2D search space. An audio playback device plays back at least a portion of the candidate-voice.


