Synthetic Voice Parameter Search for Minimal Recording Input

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating synthetic voices require direct interaction or access to voice recordings of the person, which can be problematic when recordings are unavailable.

Innovation Solution

A method for determining synthetic voice parameters by iteratively searching, mixing, and adjusting parameterized voices within a 2D search space to mimic a target voice, using techniques like UMAP projection and LSTM networks to identify and adjust underlying voice parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional voice parameterization methods are used requiring direct access to voice recordings, then voice authenticity is improved, but accessibility and ease of operation deteriorate when recordings are unavailable

Engineering Contradiction:
Improvevoice authenticityVSAvoidaccessibility when recordings unavailable
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces an intermediary approach by using a small set of reference voice samples (as little as 5 seconds) combined with extensive parameterization techniques to create a comprehensive voice model. This intermediary method bridges the gap between having no recordings and having extensive recordings, allowing authentic voice synthesis with minimal input data through iterative parameter adjustment and mixing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive voice recordings are collected for accurate parameterization, then voice resemblance quality is improved, but time consumption and complexity increase

Engineering Contradiction:
Improvevoice resemblance qualityVSAvoidtime for data collection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing and pre-analyzing the minimal voice samples to extract comprehensive parameter sets in advance. The system performs upfront parameter identification, feature extraction, and model training on the small reference set, creating a ready-to-use parameterized voice model that can be quickly adjusted and mixed without requiring additional time-consuming data collection during the synthesis process.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If minimal voice samples are used for parameterization, then ease of operation is improved, but voice parameterization accuracy deteriorates

Engineering Contradiction:
Improveease of voice sample collectionVSAvoidvoice parameterization accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent employs parameter changes by extensively manipulating and adjusting numerous voice parameters through iterative processes. The system varies parameters such as pitch, timbre, prosody, and spectral characteristics through mixing different reference samples and adjusting parameter values to achieve accurate voice resemblance. This comprehensive parameter exploration compensates for the minimal input data, transforming a small sample set into a richly parameterized voice model through systematic parameter variation and optimization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260045266A1Voice parameter determination methods, system and device
Publication Date: 2026.02.12 INCUBATEUR TECHNOLOGIQUE INOVUM
  • US20260045266A1 patent drawing
  • US20260045266A1 patent drawing
  • US20260045266A1 patent drawing

AI summary

Methods, device and system for determining a target voice parameters. A location within a 2D search space is assigned to parameterized voices, perceptually similar voices being proximate. Candidate-voices are inserted into a candidate list when a resemblance threshold is reached; A choice between two unmixed voices is received. The plurality of underlying parameters of the unmixed voices are mixed into a mixed voice towards the target-voice. The plurality of underlying parameters from the candidate list are identified. The unadjusted voice is adjusted into an adjusted voice by altering values of the plurality of underlying parameters towards the target-voice. A user interface module receives a choice of a candidate-voice from the 2D search space. An audio playback device plays back at least a portion of the candidate-voice.