Voice Timbre Vector Spaces for Cadence-Preserving Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice conversion technologies struggle to isolate and accurately transform the timbre of a voice while maintaining the original cadence, rhythm, and accent, as they often incorporate accent and cadence characteristics into the conversion process, limiting control over timbre.
Innovation Solution
A voice-to-voice conversion system using a timbre vector space and machine learning to extract and transform speech segments into a target voice by filtering timbre data with a temporal receptive field, maintaining source cadence and accent, and employing a generative neural network to produce a new target voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing voice conversion technologies are used to transform timbre, then timbre transformation is achieved, but cadence and accent are inadvertently altered along with timbre
Solution Approach 1:
The patent segments voice characteristics into distinct components: timbre, cadence, and accent. By separating these elements, the system can selectively transform timbre while preserving cadence and accent characteristics from the source voice. This is achieved through separate neural network pathways that process different voice attributes independently.
Solution Approach 2:
The patent extracts only the timbre component from the target voice for transformation purposes, while explicitly excluding cadence and accent characteristics. The source voice's cadence and accent are extracted and preserved separately, then recombined with the transformed timbre to produce the final output.
2Adaptability or versatility
If traditional voice conversion methods are applied, then voice transformation is achieved, but control over specific timbre characteristics is limited
Solution Approach 1:
The patent introduces a vector space representation that adds a new dimension to voice transformation. By mapping timbre characteristics into a multi-dimensional vector space, the system enables precise control over specific timbre attributes through vector operations, allowing granular adjustment of voice characteristics beyond traditional binary transformation approaches.
Solution Approach 2:
The patent transforms voice timbre by manipulating parameters in the vector space representation. By changing specific vector parameters that correspond to different timbre characteristics, the system can selectively adjust aspects such as pitch, tone quality, and spectral features while maintaining control over the transformation process.
Data Source
AI summary
A method of building a new voice having a new timbre using a timbre vector space includes receiving timbre data filtered using a temporal receptive field. The timbre data is mapped in the timbre vector space. The timbre data is related to a plurality of different voices. Each of the plurality of different voices has respective timbre data in the timbre vector space. The method builds the new timbre using the timbre data of the plurality of different voices using a machine learning system.


