Voice Cloning via Timbre Extraction and Inversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice cloning technologies face challenges in ensuring consistency between a user's recorded voice and the text being synthesized, requiring significant user effort and affecting user experience.
Innovation Solution
A voice processing method that involves performing voice conversion processing based on a user's voice and specified timbre information, training a voice conversion model, and using a voice synthesis model to generate a target synthesized voice that matches the user's timbre, thereby simplifying the voice cloning process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice cloning is used where a user provides recorded voice and corresponding text, then voice customization is achieved, but the reading consistency between recorded voice and text cannot be guaranteed requiring cleaning and correction operations
Solution Approach 1:
Instead of requiring the user to provide both voice and text (traditional approach), the patent inverts the process by extracting text from the user's voice recording through speech-to-text conversion. This ensures the text automatically matches the voice content, eliminating consistency issues without requiring manual cleaning and correction operations.
2Adaptability or versatility
If traditional voice cloning requires user to provide recorded voice and corresponding text, then voice customization is achieved, but user effort and recording requirements increase
Solution Approach 1:
The patent extracts text information directly from the user's voice recording using speech-to-text conversion technology. By taking out the text from the voice rather than requiring separate text input, the system reduces user effort and simplifies the operation process while maintaining voice customization capability.
3Productivity
If voice conversion processing is performed based on user voice and specified timbre information, then voice cloning efficiency is improved, but model training complexity increases
Solution Approach 1:
The patent performs preliminary voice conversion processing to generate a converted voice sample that matches the target timbre before final voice cloning. This preliminary action creates a reference that guides the model training process, improving efficiency by providing a clear target for the conversion model to learn from, thereby reducing the overall training complexity despite the additional processing step.
Data Source
AI summary
A voice processing method includes: performing voice conversion processing based on a user voice of a target user and specified timbre information to obtain a specified converted voice having a specified timbre; training a voice conversion model based on the user voice of the target user and the specified converted voice to obtain a target voice conversion model; inputting a target text for voice synthesis and the specified timbre information into a voice synthesis model to generate an intermediate voice having the specified timbre; and performing voice conversion processing on the intermediate voice by the target voice conversion model to generate a target synthesized voice that matches a timbre of the target user.


