Personalized Voice Synthesis via Acoustic Model Fine-Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic devices face challenges in providing personalized voice synthesis using a small amount of user input data, which affects user experience and efficiency.
Innovation Solution
An electronic device with a processor and memory that receives a voice input, extracts features, selects an acoustic model, and performs fine-tuning to learn the user's voice, allowing for personalized voice synthesis with minimal user inconvenience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional TTS models are used with large amounts of training data, then voice synthesis quality is improved, but user convenience and data collection burden worsen
Solution Approach 1:
The patent applies partial action by using only a small subset of training data (e.g., a few seconds of user voice recording) instead of requiring large amounts of data. The fine-tuning process selectively adapts pre-trained acoustic models with minimal user-specific data, achieving personalized voice synthesis without the burden of extensive data collection.
Solution Approach 2:
The patent implements preliminary action by pre-training acoustic models on large datasets before deployment. The heavy lifting of learning general speech patterns is done in advance, allowing the model to be quickly adapted to individual users with minimal data through fine-tuning, thus resolving the contradiction between quality and convenience.
2Measurement precision
If acoustic models are fine-tuned with more user data, then personalized voice accuracy is improved, but processing time and computational resources worsen
Solution Approach 1:
The patent uses partial action by fine-tuning only specific parameters of pre-trained acoustic models rather than retraining entire models from scratch. This selective adaptation focuses computational resources on user-specific voice characteristics while maintaining the efficiency of pre-trained components, achieving accuracy without excessive processing time.
Solution Approach 2:
The patent applies parameter changes by modifying specific parameters of acoustic models during fine-tuning to adapt to user voice characteristics. This approach changes only the necessary parameters to achieve personalization while keeping the overall model structure and most parameters intact, thus maintaining processing efficiency.
3Stability of the object's composition
If voice synthesis models are highly specialized for individual users, then user experience consistency is improved, but model complexity and adaptability worsen
Solution Approach 1:
The patent implements universality by using a single acoustic model architecture that can serve multiple users through fine-tuning. The pre-trained model provides universal speech synthesis capabilities, while selective fine-tuning adapts it to individual users, maintaining both consistency for each user and adaptability across different users without requiring separate models.
Solution Approach 2:
The patent applies dynamics by making the model adaptable through fine-tuning based on user feedback and usage patterns. The system dynamically adjusts to individual users while maintaining a consistent underlying framework, allowing the same base model to provide consistent experiences across different users through adaptive parameter adjustment.
Data Source
AI summary
An electronic device is provided. The electronic device includes a processor and a memory operatively connected to the processor. The memory may store instructions that, when executed, cause the processor to receive a voice input of a user, to extract a feature from the voice input of the user, to select an acoustic model through comparison with the extracted feature, and to learn the feature of the voice input by performing fine-tuning on the selected acoustic model.


