Voice Timbre Vector Mapping for Accent-Preserving Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice conversion technologies struggle to isolate and accurately transform the timbre of a voice while maintaining the original cadence and accent, often incorporating unintended accent and cadence characteristics due to long receptive fields.

Innovation Solution

A voice-to-voice conversion system using a timbre vector space and a generative neural network to filter and map voice data, allowing for real-time transformation of speech segments into a target voice while preserving the source's cadence and accent, utilizing a temporal receptive field and machine learning to refine frequency components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a long receptive field is used in voice conversion, then the system can capture more context information, but the system incorporates unintended accent and cadence characteristics into the transformed voice

Engineering Contradiction:
Improvecontext informationVSAvoidtimbre transformation accuracy
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

The patent divides the voice conversion process into separate functional modules: a timbre extraction module that isolates only the timbre characteristics from the source voice, and a timbre generation module that synthesizes the target voice using only extracted timbre features. This segmentation prevents contamination from unwanted accent and cadence information while maintaining accurate timbre transformation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the essential timbre characteristics from the source voice signal, separating them from accompanying accent and cadence information. By taking out only the necessary timbre components and discarding the rest, the system achieves precise timbre conversion without introducing unintended characteristics from the source speaker.

Inventive Principle:
Principle #2Taking out (Extraction)

2Manufacturing precision

If existing voice conversion technologies are used, then the system can transform voice timbre, but it cannot maintain the original cadence and accent while converting timbre

Engineering Contradiction:
Improvetimbre conversion accuracyVSAvoidcadence and accent preservation
Core Design Contradiction:
Manufacturing precisionVSStability of the object's composition

Solution Approach 1:

The patent extracts and isolates the timbre component from the source voice, separating it from cadence and accent information. The extracted timbre is then applied to the target voice without the original cadence and accent constraints, allowing independent control of timbre transformation while preserving the desired cadence and accent characteristics.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different aspects of the voice signal: high-precision timbre extraction and transformation for the timbre component, while maintaining the original cadence and accent patterns unchanged. This local quality approach allows selective optimization of timbre conversion without compromising cadence and accent stability.

Inventive Principle:
Principle #3Local quality

3Productivity

If a neural network with temporal receptive field is used, then the system can process voice data effectively, but the receptive field length affects the balance between capturing context and avoiding unwanted characteristics

Engineering Contradiction:
Improvevoice data processing efficiencyVSAvoidtimbre isolation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the voice conversion task into distinct stages: timbre extraction, timbre generation, and output synthesis. By dividing the processing functionally rather than relying solely on receptive field length, the system achieves effective processing while maintaining precise timbre isolation through architectural design rather than temporal window size alone.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250378841A1System and method for creating timbres
Publication Date: 2025.12.11 MODULATE INC
  • US20250378841A1 patent drawing
  • US20250378841A1 patent drawing
  • US20250378841A1 patent drawing

AI summary

A method of building a new voice having a new timbre using a timbre vector space includes receiving timbre data filtered using a temporal receptive field. The timbre data is mapped in the timbre vector space. The timbre data is related to a plurality of different voices. Each of the plurality of different voices has respective timbre data in the timbre vector space. The method builds the new timbre using the timbre data of the plurality of different voices using a machine learning system.