Multi-mode assisted non-invasive speech brain-computer interface decoding method

Through a multimodal-assisted non-invasive speech brain-computer interface decoder, combining text and audio information, and using contrast learning to train the decoder, the problem of excessive dependence on prior knowledge and poor decoding performance in the existing technology is solved, and efficient speech BCI decoding is achieved, especially in Chinese language BCI.

CN120354166APending Publication Date: 2025-07-22HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510448064.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing verbal BCI decoding methods rely on too much prior knowledge, complex manual feature extraction, high risk of invasive brain-computer interfaces, traditional methods are susceptible to noise interference and poor decoding performance, making it difficult to effectively decode speech activities.

Method used

A multimodal-assisted non-invasive speech brain-computer interface decoder is adopted. Through the brain signal feature extraction module and classification module, combined with text and audio information, the decoder is trained using contrast learning to reduce dependence on prior knowledge and improve decoding performance.

Benefits of technology

Good decoding performance is achieved on limited data sets, which significantly improves the classification accuracy of speech BCI, and is suitable for non-invasive decoding in multiple languages, especially the decoding accuracy of Chinese language BCI has been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354166A_ABST
    Figure CN120354166A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode assisted non-invasive speech brain-computer interface decoding method, which belongs to the field of speech brain-computer interface signal decoding, and comprises the following steps of: assisting learning of a brain signal feature extractor by using a task related mode, comprehensively considering a speech sounding mechanism, converting label information into text and voice information, and decoding the text and voice information by using a multi-mode assisted non-invasive speech brain-computer interface. The training of a brain signal feature extractor is guided by utilizing comparative learning; and inputting the extracted brain features into a classification head, and further improving the gradient descent of classification loss to complete the training of a decoder. The method does not excessively depend on priori knowledge, and can simply and efficiently decode brain speech activities. In actual use, speech decoding can be realized only by collecting brain signals of a user. An effective decoding mode is provided for the speech brain-computer interface, help is provided for further exploration, and the speech brain-computer interface system is expected to be widely applied in the fields of medical rehabilitation, information communication, entertainment and the like in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech brain-computer interface signal decoding, and more specifically, relates to a non-invasive speech brain-computer interface decoding method assisted by multi-modalities. Background Art

[0002] A brain-computer interface (BCI) establishes a direct connection path between the human brain and external devices, aiming to control electronic devices or computers through brain signals to reflect the user's intention. It is widely used in neuroscience research and can help paralyzed patients recover motor ability and improve the ability to communicate with the outside world. Common BCI paradigms include steady-state visually evoked potential (SSVEP), event-related potential (ERP), motor imagery (MI), etc. However, since ERP and SSVEP are likely to cause visual fatigue in users and cannot be used by patients with visual impairments; the MI paradigm only includes several limited instructions and requires training of users, and some users are BCI-blind and cannot use it. Therefore, a new BCI paradigm has emerged. As an emerging BCI paradigm, a speech brain-computer interface can convert brain signals generated during the speech activities of a subject into text, sound, or facial expressions for output, directly reflecting what the subject hears and thinks. Compared with traditional BCI paradigms, this paradigm has a sufficient number of instructions and has the characteristics of spontaneous generation, no need for stimulation, and being user-friendly, etc., and can better communicate with the outside world or achieve device control. Its appearance is equivalent to installing a "language prosthesis" for patients who have lost their language ability due to nerve damage, bringing good news to patients with severe language disorders.

[0003] Traditional research generally believes that the auditory language center (a part of Wernicke's area) is located in the posterior part of the superior temporal gyrus, the motor language center (Broca's area) is located in the posterior part of the inferior frontal gyrus, the reading center (a part of Wernicke's area and the angular gyrus above it) is located in the angular gyrus of the inferior parietal lobe, and the writing language center is located in the middle frontal gyrus. However, the latest research points out that Broca's area contains little or no information about orofacial movements, phonemes, or words, and lacks neural activities related to speech production. Thus, it can be seen that the speech area is widely distributed and no unified conclusion has been formed yet, and the prior knowledge about speech decoding is not sufficient enough.

[0004] Traditional machine learning methods usually require manual feature extraction, which mainly relies on expert knowledge and experience. Speech BCI started relatively late, and the related research is not yet mature. It may overlook potential useful information, and the manual feature extraction process is complex and time-consuming, and is easily affected by noise and interference. End-to-end deep learning networks can directly learn and extract features from raw brain signals, reducing manual intervention and the complexity of feature extraction. They can process high-dimensional and non-linear data and capture the complex patterns and hidden information of brain signals. Most of the existing speech BCI decoding methods refer to the language generation mechanism and are constructed starting from the smallest unit of language generation. Specifically, a model is used to capture the phoneme classification probability, a sentence is formed through a given matching rule, and then the result is corrected by a language large model; this method has high requirements for the phoneme decoding model and is easily affected by the "hallucination" of the large model, thus ignoring the real brain speech information.

[0005] In addition, most current research is based on invasive brain-computer interfaces, which require surgical implantation, have relatively high risks, and the postoperative care is cumbersome. Summary of the Invention

[0006] In view of the above defects or improvement requirements of the prior art, the present invention provides a multi-modal assisted non-invasive speech brain-computer interface decoding method, which can simply and efficiently decode brain speech activities without relying too much on prior knowledge.

[0007] To achieve the above object, according to the first aspect of the present invention, there is provided a method for constructing a multi-modal assisted non-invasive speech brain-computer interface decoder, including:

[0008] Construct a decoder and train the decoder using a data set;

[0009] Wherein, the decoder includes:

[0010] A brain signal feature extraction module, including a first and a second feature extraction network and a splicing layer; the first and second feature networks are respectively used to extract brain signal features including text information and audio information from brain signals The splicing layer is used to Splice to obtain Z b ; the brain signal is magnetoencephalogram signal;

[0011] A classification module, used to classify Z b And output the corresponding text category;

[0012] The data set includes brain signals of subjects during speech activities collected under text stimuli, text signals and audio signals corresponding to the brain signals, and text signals and audio signals not corresponding to the brain signals; the audio signal is obtained by performing TTS processing on the text signal;

[0013] The training process includes:

[0014] Maximizing The similarity between The corresponding text features, And the similarity between The corresponding audio features, while minimizing The similarity between The non - corresponding text features, And the similarity between The non - corresponding audio features, and minimizing the classification loss as the objective, training the decoder to obtain a trained decoder; the same The corresponding text features, the same The corresponding audio features are obtained by respectively inputting the text signal and the audio signal corresponding to the brain signal into the text feature extraction module and the audio feature extraction module for extraction; the same The non - corresponding text features, the same The non - corresponding audio features are obtained by respectively inputting the text signal and the audio signal not corresponding to the brain signal into the text feature extraction module and the audio feature extraction module for extraction.

[0015] According to the second aspect of the present invention, a multi - modal assisted non - invasive speech brain - computer interface decoding method is provided, including:

[0016] Sequentially inputting the brain signal to be recognized into the decoder constructed by using the method described in the first aspect to obtain the corresponding text category.

[0017] According to the third aspect of the present invention, a multi - modal assisted non - invasive Chinese speech brain - computer interface decoding method is provided, including:

[0018] Sequentially inputting the brain signal to be recognized into the decoder constructed by using the method described in the first aspect to obtain the corresponding text category; wherein, when constructing the decoder, the text information is Chinese text information and the text signal is a Chinese text signal.

[0019] According to the fourth aspect of the present invention, an electronic device is provided, including: a computer - readable storage medium and a processor;

[0020] The computer - readable storage medium is used for storing executable instructions;

[0021] The processor is used for reading the executable instructions stored in the computer - readable storage medium and executing the method described in the first aspect, the second aspect or the third aspect.

[0022] According to the fifth aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to execute the method as described in the first aspect, the second aspect or the third aspect.

[0023] According to the sixth aspect of the present invention, there is provided a computer program product including a computer program or instructions, which when executed by a processor implement the method as described in the first aspect, the second aspect or the third aspect.

[0024] Generally speaking, compared with the prior art, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0025] The method provided by the present invention introduces task-related text and speech modalities, and uses contrastive learning to constrain the consistency between brain signals and task-related modalities, thereby assisting the training of the brain signal feature extraction module, and can effectively improve the decoding performance; compared with traditional machine learning methods and methods for decoding starting from acoustic units, the method provided by the present invention can make full use of the training label information, greatly reduce the dependence on prior knowledge, have low requirements for the decoding model, and can achieve good decoding performance on a limited data set. Verified by experimental tests, the method provided by the present invention has significant and stable performance in speech BCI classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic structural diagram of a multi-modal assisted non-invasive Chinese speech brain-computer interface decoder provided by an embodiment of the present invention;

[0027] Figure 2 It is a schematic diagram of the classification accuracy rate (%) of different Chinese speech brain-computer interface methods on different users; among them, the highest result on each user is marked in bold. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0029] Magnetoencephalography (MEG) records the magnetic field activities around the scalp of the brain and is a non-invasive brain-computer interface technology. Compared with the commonly used electroencephalography (EEG), the signals collected by MEG have higher spatio-temporal resolution and more stable signal quality.

[0030] Based on this, an embodiment of the present invention provides a method for constructing a non-invasive speech brain-computer interface decoder assisted by multiple modalities, including:

[0031] Construct a decoder and train the decoder using a dataset;

[0032] Among them, the decoder includes:

[0033] A brain signal feature extraction module, including a first and a second feature extraction network and a splicing layer; the first feature network is used to extract brain signal features including text information from brain signals The second feature network is used to extract brain signal features including audio information from brain signals The splicing layer is used to splice the brain signal features including text information and audio information to obtain Z b ; the brain signal is a magnetoencephalogram signal;

[0034] A classification module, used to classify Z b and output the corresponding text category;

[0035] The dataset includes brain signals of subjects during speech activities collected under text stimuli, text signals and audio signals corresponding to the brain signals, and text signals and audio signals not corresponding to the brain signals; the audio signal is obtained by performing TTS processing on the text signal;

[0036] The process of the training includes:

[0037] Taking maximizing the similarity between the corresponding text features, the similarity between the corresponding audio features, minimizing the similarity between the non-corresponding text features, the similarity between the non-corresponding audio features, and minimizing the classification loss as the objective, training the decoder to obtain a trained decoder; the corresponding text features, the corresponding audio features are obtained by respectively inputting the text signal and the audio signal corresponding to the brain signal into the text feature extraction module and the audio feature extraction module; the non-corresponding text features, the non-corresponding audio features are obtained by respectively inputting the text signal and the audio signal not corresponding to the brain signal into the text feature extraction module and the audio feature extraction module.

[0038] The method provided by the present invention, considering the relationship between language and brain activity, adopts Text-to-Speech (TTS) technology to convert the text library corresponding to the labels into an audio library for subsequent use. Using X b ∈R B×C×T to represent the brain signal, where B represents the batch size, C represents the magnetoencephalogram channels, T represents the sampling time of each trial, and the text signal for stimulation is denoted as X t . The corresponding audio signal X s is obtained through the one-to-one mapping relationship between the text library and the audio library.

[0039] That is, in order to fully extract the acoustic features contained in the brain signals of speech activities, the TTS technology is used to introduce the speech modality X s . TTS parses the input text through front-end processing, constructs a statistical model based on the text information to predict speech parameters (such as fundamental frequency, formant frequency, etc.), and then converts these parameters into a model through a vocoder module to generate speech. Voice conversion is performed with the help of speech synthesis technologies such as Baidu Smart Cloud, etc., to convert the text library into an audio library for subsequent extraction of audio features.

[0040] The structures of each module are introduced separately below.

[0041] 1. Brain Signal Feature Extraction Module

[0042] The first and second feature networks can adopt common neural networks for processing brain signals. For example, EEGNet, DeepConvNet, ShallowConvNet, Conformer, etc. Considering the recognition accuracy, both the first and second feature networks are preferably EEGNet.

[0043] Use EEGNet as the brain signal feature extraction module f b to extract the brain feature Z b = f b (X b ). EEGNet extracts spatio-temporal features simultaneously by using depthwise separable convolution and has been proven effective in decoding in various BCI paradigms. Use two EEGNets to extract and respectively, and then and are concatenated to obtain

[0044] 2. Text Feature Extraction Module

[0045] Use the text pre-training model as the text feature extraction module f t to extract the text feature For example, fastText and BERT are respectively used to extract text features. fastText splits the word sequence into n-gram sequences, generates word vectors for each n-gram sequence, and combines them to represent features, which is very beneficial for Chinese characters composed of multiple single characters. BERT is stacked by multiple Transformer encoder layers, and comprehensively captures the semantic and context information of the text by combining word embeddings, segment embeddings, and position embeddings, etc.

[0046] 3. Audio Feature Extraction Module

[0047] Use a traditional acoustic model or a speech pre-training model as the audio feature extraction module f s to extract audio features Mel Frequency Cepstrum Coefficient (MFCC), wav2vec2.0, or HuBERT can be used to extract audio features. Since the human ear is more sensitive to low-frequency sounds and less sensitive to high-frequency sounds, by converting linear frequency to mel scale, the behavior of the human auditory system can be better simulated. Through cepstrum analysis, useful features can be extracted from the speech signal. Both wav2vec2.0 and HuBERT are trained through self-supervised learning. wav2vec2.0 predicts the audio features in the unlabeled speech through masking, while HuBERT first performs discrete clustering on the audio and then learns the audio features by reconstructing these discrete representations.

[0048] 4. Training Objectives

[0049] In the training stage, the brain signals and the corresponding text signals and audio signals are fed into the decoder, and feature extraction is respectively performed through the brain signal feature extraction module, the text feature extraction module, and the audio feature extraction module. Contrastive learning is used to guide the training of the brain signal feature extraction module; the extracted brain features are input into the classification head (i.e., the classification module), and the decoder training is completed by further improving through gradient descent of the classification loss. In the application stage, the brain signals to be classified are input into the trained decoder to obtain the corresponding text categories.

[0050] The training of the entire decoder mainly relies on contrastive loss and classification loss. N-pair Loss, SupConLoss, or InfoNCE can be used as the contrastive loss function. First, the outputs of different modalities are normalized, and then the cosine similarities between the text-brain signal features and the audio-brain signal features are calculated respectively, and the smoothness of the distribution is controlled by temperature scaling. Through training iterations, the distance between the representations of positive sample pairs can be maximized, and the distance between the representations of negative sample pairs can be minimized, so as to learn effective feature representations.

[0051] Taking the use of InfoNCE as the contrast loss function as an example, during the training process, the loss function adopted is

[0052] where is the contrast loss function between and the corresponding text features, is the i-th brain signal feature including text information in B is the batch size of the training set, is the corresponding text feature, is all the corresponding and non-corresponding text features, and τ is the temperature coefficient; is the contrast loss function between and the corresponding audio features, is the i-th brain signal feature including audio information in is the corresponding audio feature, is all the corresponding and non-corresponding text features; is the classification loss function; taking the use of cross-entropy loss function as the classification loss function as an example, is b the i-th sample in X y i is the true class of the i-th sample, Other common classification loss functions can also be adopted, which will not be elaborated here.

[0053] Considering the case of less data volume, in order to further improve the decoding accuracy, data augmentation is performed by adding noise in the time-frequency domain respectively. Salt-and-pepper noise is used to randomly select a certain proportion of signal points, and the values of these points are set to the maximum or minimum value to simulate discrete interference. A small number of data points are selected in each trial and replaced with the maximum or minimum value of the whole trial as the noise signal, which is superimposed on the original signal to increase the training data volume.

[0054] The embodiment of the present invention provides a multi-modal assisted non-invasive speech brain-computer interface decoding method, including:

[0055] Sequentially inputting the brain signal to be recognized into the decoder constructed by using the method described in any one of the above embodiments to obtain the corresponding text category.

[0056] It can be understood that the decoding method provided by the present invention can be used for speech BCI decoding of any type of language, and the text information and text signals used in training the decoder can adopt the corresponding language type.

[0057] Considering that Chinese is one of the most widely used languages in the world, although speech BCI has been developed for some time, the research on Chinese speech BCI is still scarce. Most of the research remains at the phoneme level or the decoding of a small number of Chinese characters, making it difficult to meet the basic communication needs. Based on this, an embodiment of the present invention provides a non-invasive Chinese speech brain-computer interface decoding method assisted by multi-modalities, including:

[0058] Sequentially inputting the brain signals to be recognized into the decoder constructed by using the construction method described in any of the above embodiments to obtain the corresponding text categories; wherein, when constructing the decoder, the text information is Chinese text information, and the text signal is Chinese text signal, as Figure 1 shown.

[0059] Verified by experiments, on the MEG Chinese speech BCI data based on a Chinese 48-word vocabulary, the method proposed by the present invention can effectively improve the decoding accuracy, as Figure 2 shown. It can be observed that compared with the non-modal assisted decoding, the multi-modal assisted methods under different combinations of feature extraction methods can all effectively improve the decoding accuracy. The average top5 of 48-classification has increased by 7.95%, and the top1 has increased by 3.29% on average, proving that the method of using multi-modal assisted speech imagination decoding can effectively guide the learning of the brain signal feature extractor (i.e., the brain signal feature extraction module). After using data augmentation, the decoding accuracy can be further improved. The average top5 accuracy of 48-classification can reach up to 35.80% at most, and the top1 average accuracy can reach up to 10.69% at most. Among them, the top5 of subject No. 9 can reach up to 46.21% at most, and the top1 can reach up to 15.32% at most.

[0060] In summary, the method provided by the present invention comprehensively considers the vocalization and semantic activities involved during speech activities, introduces text and audio signals, extracts the corresponding features, uses task-related modalities to constrain the brain feature representation, and assists in the training of the brain signal feature extraction module. This method has relatively low requirements for prior knowledge, can achieve good decoding performance on a relatively simple decoding network, provides an effective decoding method for speech brain-computer interfaces, helps with further exploration, and is expected to promote the wider application of speech brain-computer interface systems in the fields of medical rehabilitation, information communication, entertainment, etc. in the future.

[0061] An embodiment of the present invention provides an electronic device, including: a computer-readable storage medium and a processor;

[0062] The computer-readable storage medium is used to store executable instructions;

[0063] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the construction method or the decoding method as described in any of the above embodiments.

[0064] An embodiment of the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to execute the construction method or the decoding method as described in any of the above embodiments.

[0065] An embodiment of the present invention provides a computer program product including a computer program or instructions that, when executed by a processor, implement the construction method or the decoding method as described in any of the above embodiments.

[0066] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A construction method of a multi-modal assisted non-invasive speech brain-computer interface decoder, characterized in that, Including: Construct a decoder and train the decoder using a dataset; Wherein, the decoder includes: The brain signal feature extraction module includes a first and a second feature extraction network and a splicing layer; the first and second feature networks are respectively used to extract brain signal features including text information and audio information from brain signals The splicing layer is used to perform splicing to obtain Z b ; The brain signal is a magnetoencephalogram signal; Classification module for classifying Z b and outputting the corresponding text category; The dataset includes brain signals of a subject during speech activities collected under text stimuli, text signals and audio signals corresponding to the brain signals, and text signals and audio signals not corresponding to the brain signals; the audio signal is obtained by performing TTS processing on the text signal; The process of the training includes: To maximize the similarity between the corresponding text features, and the similarity between the corresponding audio features, while minimizing the similarity between the non-corresponding text features, and the similarity between the non-corresponding audio features, and minimizing the classification loss as the objective, train the decoder to obtain a trained decoder; the same corresponding text features, the same corresponding audio features are obtained by respectively inputting the text signal and the audio signal corresponding to the brain signal into the text feature extraction module and the audio feature extraction module; the same non-corresponding text features, the same non-corresponding audio features are obtained by respectively inputting the text signal and the audio signal not corresponding to the brain signal into the text feature extraction module and the audio feature extraction module.

2. The method according to claim 1, wherein Both the first and second feature extraction networks are EEGNet.

3. The method according to claim 1 or 2, characterized in that, The text feature extraction module uses a text pre-training model; The audio feature extraction module uses a traditional acoustic model or a pre-trained speech model.

4. The method according to claim 1, characterized in that The loss function used during the training process Among them, is the contrast loss function between and the corresponding text features, is the i-th brain signal feature including text information in B is the batch size of the training set, is the corresponding text feature, is all corresponding and non-corresponding text features, τ is the temperature coefficient; is the contrast loss function between and the corresponding audio features, is the i-th brain signal feature including audio information in is the corresponding audio feature, is all corresponding and non-corresponding text features, B is the batch size; is the classification loss function.

5. A non-invasive speech brain-computer interface decoding method assisted by multi-modalities, characterized in that, Including: Sequentially input the brain signal to be recognized into a decoder constructed by using the method according to any one of claims 1-4 to obtain the corresponding text category.

6. A non-invasive Chinese speech brain-computer interface decoding method assisted by multi-modalities, characterized in that, Including: Sequentially input the brain signal to be recognized into a decoder constructed by using the method according to any one of claims 1-4 to obtain the corresponding text category; wherein, when constructing the decoder, the text information is Chinese text information and the text signal is a Chinese text signal.

7. An electronic device, characterized in that, Including: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the construction method according to any one of claims 1-4 or the decoding method according to claim 5 or 6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to execute the construction method according to any one of claims 1-4 or the decoding method according to claim 5 or 6.

9. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, it implements the construction method according to any one of claims 1-4 or the decoding method according to claim 5 or 6.