Imagined Speech Voice Synthesis From Brain Wave Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing brain-computer interface technologies for imagined speech have limitations in the degree of freedom for communication, as they are confined to the number of classified classes, hindering intuitive and diverse speech synthesis.

Innovation Solution

A method and apparatus that utilize deep learning-based speech synthesis from brain waves, involving generator and discriminator training on actual speech, followed by transfer learning on imagined speech, to generate mel-spectrograms and convert them into voice using a vocoder, enhancing the degree of freedom in speech synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If classification-based user intention recognition is used, then the system can recognize user intentions, but the degree of freedom for communication is confined to the number of classified classes

Engineering Contradiction:
Improvedegree of freedom for communicationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces the traditional classification-based mechanical system with a deep learning-based neural network system that uses generator and discriminator models to synthesize speech from brain waves, enabling continuous and high-degree freedom communication rather than discrete class selection

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter from discrete class labels to continuous embedding vectors generated by deep learning models, allowing the system to represent and synthesize speech with high degree of freedom based on brain wave patterns

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If imagined speech brain wave patterns are used, then the system can increase the number of classes and improve recognition, but the degree of freedom remains limited by classification constraints

Engineering Contradiction:
Improvenumber of classesVSAvoiddegree of freedom
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent transitions from low-dimensional class labels to high-dimensional embedding vectors generated by deep learning models, adding dimensional complexity that enables continuous variation and high degree of freedom in speech synthesis while maintaining the ability to recognize numerous speech patterns

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12488779B2Method and apparatus for voice synthesis based on brain waves during imagined speech
Publication Date: 2025.12.02 KOREA UNIV RES & BUSINESS FOUND
  • US12488779B2 patent drawing
  • US12488779B2 patent drawing
  • US12488779B2 patent drawing

AI summary

The present invention relates to a method and an apparatus for synthesizing the voice based on brain waves during imagined speech. The method for synthesizing the voice based on brain waves during imagined speech according to an embodiment of the present invention may include the following steps: a step to obtain the user's brain waves during imagined speech; a step to convert the above-mentioned brain waves of imagined speech into embedding vectors; a step to generate the mel-spectrograms based on the above-mentioned embedding vectors; a step to generate the voice using the above-mentioned mel-spectrograms; a step to output the above-mentioned voice.