AI Voice Modulation Using Spectrogram Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices using artificial intelligence for voice feedback provide a uniform, mechanical sound, lacking user familiarity, which is a limitation in providing a friendly user experience, especially in AI assistant services.

Innovation Solution

An electronic apparatus and method that convert a user's voice into another user's voice by transforming the spectrogram of the first user's voice into the spectrogram of the target user's voice using a trained AI model, specifically employing techniques like discrete wavelet transform, Variational Auto-encoder, or Generative Adversarial Networks, and converting it back into audio using algorithms like Griffi-Lim.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a uniform machine voice is used for AI feedback, then the system complexity is reduced and implementation is easier, but the user familiarity and friendliness deteriorate

Engineering Contradiction:
Improveuser familiarityVSAvoidvoice modulation system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates virtual voice models by copying and learning from actual user voice recordings. The system captures voice characteristics, pitch, tone, and speech patterns to generate synthetic voice models that replicate individual user voices, enabling personalized feedback without requiring complex hardware modifications

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system modulates voice parameters such as pitch, frequency, timbre, and speech rate to transform the AI assistant's uniform voice into personalized voices matching specific users. By dynamically adjusting these acoustic parameters based on the selected user profile, the system achieves voice customization while maintaining a manageable technical architecture

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If voice modulation using AI models is implemented, then user experience and familiarity are improved, but the processing time and computational resources increase

Engineering Contradiction:
Improvevoice style variationVSAvoidvoice conversion processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary voice model training during an initial setup phase, where user voice recordings are captured and processed to create personalized voice models. These pre-trained models are stored for rapid retrieval and application during subsequent interactions, eliminating the need for real-time training and significantly reducing processing delays during actual voice modulation operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical voice synthesis methods with deep learning-based neural network models. These AI models, once trained, can generate personalized voice outputs extremely quickly by processing input text through learned patterns, achieving both high-quality voice modulation and fast processing speeds that mechanical systems cannot match

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11398223B2Electronic device for modulating user voice using artificial intelligence model and control method thereof
Publication Date: 2022.07.26 SAMSUNG ELECTRONICS CO LTD
  • US11398223B2 patent drawing
  • US11398223B2 patent drawing
  • US11398223B2 patent drawing

AI summary

The present disclosure relates to an artificial intelligence (AI) system utilizing a machine learning algorithm such as deep learning, etc. and an application thereof. In particular, a controlling method of an electronic apparatus includes obtaining a user voice of a first user, converting the voice of the first user into a first spectrogram, obtaining a second spectrogram by inputting the first spectrogram to a trained model through an artificial intelligence algorithm, converting the second spectrogram into a voice of a second user, and outputting the converted second user voice. Here, the trained model is a model trained to obtain a spectrogram of a style of the second user voice by inputting a spectrogram of a style of the first user voice. In particular, at least part of the controlling method of the electronic apparatus uses an artificial intelligence model trained according to at least one of machine learning, a neural network or a deep learning algorithm.