AI Voice Modulation Using Spectrogram Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices using artificial intelligence for voice feedback provide a uniform, mechanical sound, lacking user familiarity, which is a limitation in providing a friendly user experience, especially in AI assistant services.
Innovation Solution
An electronic apparatus and method that convert a user's voice into another user's voice by transforming the spectrogram of the first user's voice into the spectrogram of the target user's voice using a trained AI model, specifically employing techniques like discrete wavelet transform, Variational Auto-encoder, or Generative Adversarial Networks, and converting it back into audio using algorithms like Griffi-Lim.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a uniform machine voice is used for AI feedback, then the system complexity is reduced and implementation is easier, but the user familiarity and friendliness deteriorate
Solution Approach 1:
The patent creates virtual voice models by copying and learning from actual user voice recordings. The system captures voice characteristics, pitch, tone, and speech patterns to generate synthetic voice models that replicate individual user voices, enabling personalized feedback without requiring complex hardware modifications
Solution Approach 2:
The system modulates voice parameters such as pitch, frequency, timbre, and speech rate to transform the AI assistant's uniform voice into personalized voices matching specific users. By dynamically adjusting these acoustic parameters based on the selected user profile, the system achieves voice customization while maintaining a manageable technical architecture
2Adaptability or versatility
If voice modulation using AI models is implemented, then user experience and familiarity are improved, but the processing time and computational resources increase
Solution Approach 1:
The system performs preliminary voice model training during an initial setup phase, where user voice recordings are captured and processed to create personalized voice models. These pre-trained models are stored for rapid retrieval and application during subsequent interactions, eliminating the need for real-time training and significantly reducing processing delays during actual voice modulation operations
Solution Approach 2:
The patent replaces traditional mechanical voice synthesis methods with deep learning-based neural network models. These AI models, once trained, can generate personalized voice outputs extremely quickly by processing input text through learned patterns, achieving both high-quality voice modulation and fast processing speeds that mechanical systems cannot match
Data Source
AI summary
The present disclosure relates to an artificial intelligence (AI) system utilizing a machine learning algorithm such as deep learning, etc. and an application thereof. In particular, a controlling method of an electronic apparatus includes obtaining a user voice of a first user, converting the voice of the first user into a first spectrogram, obtaining a second spectrogram by inputting the first spectrogram to a trained model through an artificial intelligence algorithm, converting the second spectrogram into a voice of a second user, and outputting the converted second user voice. Here, the trained model is a model trained to obtain a spectrogram of a style of the second user voice by inputting a spectrogram of a style of the first user voice. In particular, at least part of the controlling method of the electronic apparatus uses an artificial intelligence model trained according to at least one of machine learning, a neural network or a deep learning algorithm.


