Voice Cloning via Timbre Extraction and Inversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice cloning technologies face challenges in ensuring consistency between a user's recorded voice and the text being synthesized, requiring significant user effort and affecting user experience.

Innovation Solution

A voice processing method that involves performing voice conversion processing based on a user's voice and specified timbre information, training a voice conversion model, and using a voice synthesis model to generate a target synthesized voice that matches the user's timbre, thereby simplifying the voice cloning process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional voice cloning is used where a user provides recorded voice and corresponding text, then voice customization is achieved, but the reading consistency between recorded voice and text cannot be guaranteed requiring cleaning and correction operations

Engineering Contradiction:
Improvevoice-text consistencyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Instead of requiring the user to provide both voice and text (traditional approach), the patent inverts the process by extracting text from the user's voice recording through speech-to-text conversion. This ensures the text automatically matches the voice content, eliminating consistency issues without requiring manual cleaning and correction operations.

Inventive Principle:
Principle #13The other way round (Inversion)

2Adaptability or versatility

If traditional voice cloning requires user to provide recorded voice and corresponding text, then voice customization is achieved, but user effort and recording requirements increase

Engineering Contradiction:
Improvevoice customizationVSAvoiduser operation simplicity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent extracts text information directly from the user's voice recording using speech-to-text conversion technology. By taking out the text from the voice rather than requiring separate text input, the system reduces user effort and simplifies the operation process while maintaining voice customization capability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If voice conversion processing is performed based on user voice and specified timbre information, then voice cloning efficiency is improved, but model training complexity increases

Engineering Contradiction:
Improvevoice cloning efficiencyVSAvoidmodel training complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary voice conversion processing to generate a converted voice sample that matches the target timbre before final voice cloning. This preliminary action creates a reference that guides the model training process, improving efficiency by providing a clear target for the conversion model to learn from, thereby reducing the overall training complexity despite the additional processing step.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250149051A1Voice processing methods, apparatuses, computer devices, and computer-readable storage media
Publication Date: 2025.05.08 NETEASE (HANGZHOU) NETWORK CO LTD
  • US20250149051A1 patent drawing
  • US20250149051A1 patent drawing
  • US20250149051A1 patent drawing

AI summary

A voice processing method includes: performing voice conversion processing based on a user voice of a target user and specified timbre information to obtain a specified converted voice having a specified timbre; training a voice conversion model based on the user voice of the target user and the specified converted voice to obtain a target voice conversion model; inputting a target text for voice synthesis and the specified timbre information into a voice synthesis model to generate an intermediate voice having the specified timbre; and performing voice conversion processing on the intermediate voice by the target voice conversion model to generate a target synthesized voice that matches a timbre of the target user.