Imagined Speech Voice Synthesis From Brain Wave Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing brain-computer interface technologies for imagined speech have limitations in the degree of freedom for communication, as they are confined to the number of classified classes, hindering intuitive and diverse speech synthesis.
Innovation Solution
A method and apparatus that utilize deep learning-based speech synthesis from brain waves, involving generator and discriminator training on actual speech, followed by transfer learning on imagined speech, to generate mel-spectrograms and convert them into voice using a vocoder, enhancing the degree of freedom in speech synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If classification-based user intention recognition is used, then the system can recognize user intentions, but the degree of freedom for communication is confined to the number of classified classes
Solution Approach 1:
The patent replaces the traditional classification-based mechanical system with a deep learning-based neural network system that uses generator and discriminator models to synthesize speech from brain waves, enabling continuous and high-degree freedom communication rather than discrete class selection
Solution Approach 2:
The patent changes the fundamental parameter from discrete class labels to continuous embedding vectors generated by deep learning models, allowing the system to represent and synthesize speech with high degree of freedom based on brain wave patterns
2Quantity of substance
If imagined speech brain wave patterns are used, then the system can increase the number of classes and improve recognition, but the degree of freedom remains limited by classification constraints
Solution Approach 1:
The patent transitions from low-dimensional class labels to high-dimensional embedding vectors generated by deep learning models, adding dimensional complexity that enables continuous variation and high degree of freedom in speech synthesis while maintaining the ability to recognize numerous speech patterns
Data Source
AI summary
The present invention relates to a method and an apparatus for synthesizing the voice based on brain waves during imagined speech. The method for synthesizing the voice based on brain waves during imagined speech according to an embodiment of the present invention may include the following steps: a step to obtain the user's brain waves during imagined speech; a step to convert the above-mentioned brain waves of imagined speech into embedding vectors; a step to generate the mel-spectrograms based on the above-mentioned embedding vectors; a step to generate the voice using the above-mentioned mel-spectrograms; a step to output the above-mentioned voice.


