A speech brain-computer interface decoding system and method based on knowledge data fusion

CN122526413APending Publication Date: 2026-08-07HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-05-13
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0009]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于知识数据融合的言语脑机接口解码系统及方法,由此解决现有言语脑机接口解码技术存在难以显式利用多尺度调制规律,造成时间结构信息丢失,难以充分挖掘与言语活动相关的关键空间特征,难以将局部调制特征提取、尺度相关空间建模以及全局时序依赖建模有效结合,导致解码精度低的技术问题

Benefits of technology

(1)本发明在多尺度时域卷积模块中使用多个不同时间感受野的并行卷积分支显式地将输入信号分解到多个时间尺度上进行处理,显式利用多尺度调制规律,避免了单一尺度造成的时间结构信息丢失。尺度特异性通道注意力模块在各尺度上对通道权重进行自适应学习与特征重标定,以突出不同时间尺度下的关键电极信息并建模尺度相关的通道协同关系,充分挖掘与言语活动相关的关键空间特征。两阶段多尺度时空融合模块为了防止多尺度特征过早混叠,先通过分组卷积进行内部压缩对齐,再通过标准卷积实现跨尺度、跨通道的深度融合,同时保留各尺度特征的独立性和跨尺度交互能力。本发明通过多尺度时域卷积模块实现局部调制特征提取,通过尺度特异性通道注意力模块和两阶段多尺度时空融合模块联合完成尺度相关空间建模,通过Transformer 编码模块实现全局时序依赖建模,由此将局部调制特征提取、尺度相关空间建模以及全局时序依赖建模有效结合,有效提升言语脑机接口解码精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526413A_ABST
    Figure CN122526413A_ABST
Patent Text Reader

Abstract

The application discloses a speech brain-computer interface decoding system and method based on knowledge data fusion, and belongs to the field of speech brain-computer interface decoding. In the multi-scale time domain convolution module, a plurality of different time receptive field parallel convolution branches are used to explicitly decompose the input signal into a plurality of time scales for processing, and the multi-scale modulation law is explicitly used. The scale-specific channel attention module adaptively learns the channel weight and recalibrates the feature at each scale, fully excavating the key spatial features related to speech activity. The application realizes local modulation feature extraction through the multi-scale time domain convolution module, completes scale-related spatial modeling through the scale-specific channel attention module and the two-stage multi-scale space-time fusion module, and realizes global time sequence dependency modeling through the Transformer coding module. Thus, the local modulation feature extraction, scale-related spatial modeling and global time sequence dependency modeling are effectively combined, and the speech brain-computer interface decoding precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech brain-computer interface decoding, and more specifically, relates to a speech brain-computer interface decoding system and method based on knowledge data fusion. Background Technology

[0002] Brain-computer interfaces (BCIs) establish a direct communication pathway between the human brain and external devices, aiming to control electronic devices or computers through brain signals to reflect user intentions. They have significant application value in neuroscience research, motor function assistance, consciousness assessment, and human-computer interaction. Common BCI paradigms include steady-state visual evoked potentials, event-related potentials, and motor imagery. However, these paradigms suffer from significant visual fatigue, limited command sets, high user training costs, and difficulties in stable use by some users, making it difficult to meet the needs of natural and efficient communication.

[0003] Speech-based brain-computer interfaces (BCIs) are an emerging paradigm that directly converts speech-related neural activity signals into language information (such as text, speech, or control commands), thus providing a new communication pathway for patients with severe speech disorders. Compared to traditional BCI paradigms, speech-based BCIs offer advantages such as not requiring prolonged gaze stimulation, providing a richer command space, and exhibiting more natural interaction, making them particularly suitable for scenarios involving silent communication and severe motor disorders.

[0004] Existing EEG decoding methods can be broadly categorized into traditional machine learning methods and deep learning methods. Traditional machine learning methods typically rely on manual feature extraction, bandpass filtering, spatial filtering, and classifier combinations, requiring significant human experience and struggling to capture the complex hierarchical spatiotemporal patterns and multi-scale evolutionary laws of non-stationary features in EEG signals. While deep learning methods can directly learn features from raw signals, most models follow general EEG modeling frameworks, such as EEGNet and EEG Conformer, relying heavily on single-scale convolutions or directly using general attention structures. They fail to adequately utilize the hierarchical modulation rhythms and scale-related spatial relationships inherent in speech-related neural activities and lack dedicated modeling mechanisms specific to the characteristics of speech brain-computer interface decoding tasks.

[0005] During speech production and speech imagery, brain signals often exhibit multi-timescale modulation characteristics. Different modulation frequencies correspond to different temporal levels of language; for example, higher modulation frequencies are closer to phonemes and transient speech changes, medium modulation frequencies are closer to subsyllabic structures and syllable rhythms, and lower modulation frequencies are related to prosody or phrase-level structures. If the decoding model cannot explicitly utilize this multi-scale modulation pattern, temporal structural information is easily lost.

[0006] The contributions between EEG electrode channels are not constant, and the connections between key brain regions and channels may differ at different time scales. However, existing methods often employ a uniform spatial convolution or a uniform attention mechanism, lacking a mechanism to model channel dependencies separately at different time scales, making it difficult to fully explore key spatial features related to speech activities.

[0007] Meanwhile, speech EEG decoding also needs to consider both local pattern extraction and long-range temporal dependency modeling. While convolutional networks are good at extracting local spatiotemporal features, their ability to model longer-range contexts is limited; while Transformers are prone to insufficient feature representation on raw noisy EEG signals. The challenge lies in effectively combining local modulation feature extraction, scale-related spatial modeling, and global temporal dependency modeling, which leads to low decoding accuracy in speech EEG interfaces.

[0008] It is evident that existing speech brain-computer interface decoding technologies suffer from several technical problems, including difficulty in explicitly utilizing multi-scale modulation patterns, resulting in the loss of temporal structure information, difficulty in fully mining key spatial features related to speech activities, and difficulty in effectively combining local modulation feature extraction, scale-related spatial modeling, and global temporal dependency modeling, leading to low decoding accuracy. Summary of the Invention

[0009] To address the above-mentioned deficiencies or improvement needs of existing technologies, this invention provides a speech brain-computer interface decoding system and method based on knowledge data fusion. This solves the technical problems of existing speech brain-computer interface decoding technologies, such as difficulty in explicitly utilizing multi-scale modulation rules, resulting in the loss of temporal structure information, difficulty in fully mining key spatial features related to speech activities, and difficulty in effectively combining local modulation feature extraction, scale-related spatial modeling, and global temporal dependency modeling, leading to low decoding accuracy.

[0010] To achieve the above objectives, according to a first aspect of the present invention, a speech brain-computer interface decoding system based on knowledge data fusion is provided, comprising: a multi-scale temporal convolution module, a scale-specific channel attention module, a two-stage multi-scale spatiotemporal fusion module, a Transformer encoding module, and a classification output module. The multi-scale temporal convolution module uses parallel convolution branches with multiple different temporal receptive fields to extract temporal features at multiple different time scales from speech-related neural activity signals. The scale-specific channel attention module is used to adaptively learn and recalibrate the channel weights of temporal features at multiple different time scales to obtain recalibrated temporal features at multiple different time scales. The two-stage multi-scale spatiotemporal fusion module is used to perform cross-channel fusion of recalibrated temporal features at the same time scale in the first stage using grouped one-dimensional convolution to obtain cross-channel fused features. In the second stage, standard one-dimensional convolution is used to perform cross-scale fusion of cross-channel fused features at multiple different time scales and then map them to a unified dimension embedding space to obtain embedded features. The Transformer encoding module is used to segment and tokenize the embedded features along the time axis to capture long-range temporal dependencies and global context information, thereby obtaining global features; The classification output module uses global features to classify and predict speech categories.

[0011] Furthermore, the multi-scale temporal convolution module includes an original signal branch and multiple parallel convolution branches with different temporal receptive fields. The original signal branch directly outputs the speech-related neural activity signal to the scale-specific channel attention module. The multiple parallel convolution branches with different temporal receptive fields perform temporal modeling on the speech-related neural activity signal through one-dimensional depth convolution with multiple different convolution kernels to obtain temporal features at multiple different time scales, which are then output to the scale-specific channel attention module.

[0012] Furthermore, the scale-specific channel attention module provides a corresponding channel attention module for the original signal branch and the parallel convolutional branch of each different temporal receptive field. Each channel attention module performs global average pooling along the temporal dimension of the input feature to obtain the channel description vector. Then, a one-dimensional convolution is applied to the channel description vector in the channel dimension and the channel weights are obtained through the Sigmoid function. Finally, the channel weights are used to recalibrate the input feature channel by channel.

[0013] Furthermore, the Transformer encoding module performs segmented average pooling of the embedded features along the time dimension according to a fixed window size to obtain multiple tokens. A learnable classification label is added before each token, and then a learnable positional encoding is superimposed to construct the Transformer encoder input. The Transformer encoder includes multiple stacked encoding modules. Each encoding module includes a multi-head self-attention sub-layer and a feedforward network sub-layer, and is combined with residual connections and layer normalization. After the Transformer encoder input is encoded through multiple layers, global features are obtained.

[0014] According to a second aspect of the present invention, a training method for a speech brain-computer interface decoding system based on knowledge data fusion is provided, comprising: A speech brain-computer interface decoding system based on knowledge data fusion was trained using a speech brain-computer interface dataset; The error between the predicted speech category and the actual speech category of the neural activity signal in the speech brain-computer interface dataset is used as the loss function. The parameters are updated by backpropagation and trained until convergence to obtain a well-trained speech brain-computer interface decoding system.

[0015] Furthermore, prior to the training, the neural activity signals in the speech brain-computer interface dataset are processed as follows: The neural activity signal is sequentially subjected to detrending, rereference, high-pass filtering, notch filtering, Hilbert transform, and instantaneous amplitude envelope extraction to obtain the envelope signal, which serves as the input to the speech brain-computer interface decoding system.

[0016] Furthermore, the training also includes: The target subject dataset is formed by randomly selecting 1%-10% of the subject data from the speech brain-computer interface dataset, and the remaining subject data from the speech brain-computer interface dataset is formed by the source domain subject dataset. The speech brain-computer interface decoding system was first pre-trained using the source domain subject dataset, and then fine-tuned using the target subject dataset to obtain the trained speech brain-computer interface decoding system.

[0017] According to a third aspect of the present invention, a speech brain-computer interface decoding method based on knowledge data fusion is provided, comprising: The neural activity signal to be decoded is input into a training method for a speech brain-computer interface decoding system based on knowledge data fusion. The resulting speech brain-computer interface decoding system predicts speech categories.

[0018] According to a fourth aspect of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program, and when executed by the processor, the computer program is used to implement a training method for a speech brain-computer interface decoding system based on knowledge data fusion or a speech brain-computer interface decoding method based on knowledge data fusion.

[0019] According to a fifth aspect of the present invention, a computer product is provided, which, when running, enables the computer to execute a training method for a speech brain-computer interface decoding system based on knowledge data fusion or a speech brain-computer interface decoding method based on knowledge data fusion.

[0020] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) In this invention, the input signal is explicitly decomposed into multiple time scales for processing by using parallel convolutional branches with multiple different temporal receptive fields in the multi-scale temporal convolution module. This explicitly utilizes the multi-scale modulation rules and avoids the loss of temporal structure information caused by a single scale. The scale-specific channel attention module performs adaptive learning and feature recalibration on the channel weights at each scale to highlight the key electrode information at different time scales and model the scale-related channel collaboration relationship, fully mining the key spatial features related to speech activities. In order to prevent premature aliasing of multi-scale features, the two-stage multi-scale spatiotemporal fusion module first performs internal compression and alignment through group convolution, and then achieves deep fusion across scales and channels through standard convolution, while preserving the independence of features at each scale and the ability to interact across scales. This invention achieves local modulation feature extraction through the multi-scale temporal convolution module, completes scale-related spatial modeling through the joint use of the scale-specific channel attention module and the two-stage multi-scale spatiotemporal fusion module, and achieves global temporal dependency modeling through the Transformer encoding module. This effectively combines local modulation feature extraction, scale-related spatial modeling, and global temporal dependency modeling, thereby effectively improving the decoding accuracy of speech brain-computer interface.

[0021] (2) The multi-scale temporal convolution module of the present invention includes an original signal branch and multiple parallel convolution branches with different temporal receptive fields. The multiple parallel convolution branches with different temporal receptive fields explicitly incorporate the modulation frequency of speech-related neural activities into the network structure prior. The original branch retains the original signal to capture slower prosody and phrase information. Each branch is independent to avoid irrelevant aliasing of different electrode channels in the early stage, while reducing the number of parameters and computational complexity.

[0022] (3) To enhance the modeling capability of key brain regions and electrode channels at different time scales, this invention does not use a single set of channel weights for all scales, but instead learns scale-specific channel weights for different scale branches. Therefore, it can accurately reflect the differences in the contributions of key electrodes at different time scales. Global average pooling is used to compress the input features into channel description vectors, and then one-dimensional convolution is used to mine the local correlations between adjacent channels on this vector, thereby generating more accurate channel weights. Channel weights are used to recalibrate the input features channel by channel. Before the features enter the two-stage multi-scale spatiotemporal fusion module, the importance of each channel has been dynamically adjusted, ensuring that the subsequent two-stage multi-scale spatiotemporal fusion module processes high-value features that have been selected.

[0023] (4) This invention performs segmented average pooling of embedded features along the time dimension according to a fixed window size to obtain multiple tokens, which is the basis for realizing long-range dependency capture. By segmented average pooling, the continuous temporal representation is discretized into a token sequence, enabling the self-attention mechanism to directly capture the long-range dependency relationship between any two positions in the sequence. Each encoding module includes a multi-head self-attention sub-layer and a feedforward network sub-layer. The essence of the self-attention mechanism is to calculate the correlation between any two positions in the sequence. This mechanism overcomes the limitations of traditional convolutional neural networks (CNNs) that easily ignore global information and recurrent neural networks (RNNs) that easily forget early information, and breaks through the limitation of fixed receptive fields on long-range modeling. The multi-head mechanism captures dependencies at different levels in parallel from multiple subspaces, and can capture information at different levels of phonemes, words, and sentences. The stacking of multiple encoders realizes the layer-by-layer transmission and recombination of information, and finally generates a sequence-level representation containing global context.

[0024] (5) In training the speech brain-computer interface decoding system, this invention performs specific preprocessing on the data based on the characteristics of the speech signal. First, the original signal undergoes detrending and rereference processing to reduce slow drift and interference introduced by the reference electrode. High-pass filtering is applied to the signal to highlight high-frequency cortical neural activity related to speech production and motor processes. Simultaneously, notch filtering is performed to remove 50 Hz power frequency noise and its odd harmonic components. After filtering, Hilbert transform is performed on the signal and the instantaneous amplitude envelope is extracted to capture the modulation information of the signal amplitude changing over time.

[0025] (6) This invention makes some improvements to the training strategy. To address the issue of individual differences in cross-subject scenarios, a "pre-training + fine-tuning" training mode is designed for inter-subject tasks. Pre-training is first completed on source domain subject data, and then fine-tuning is performed using a small amount of target subject data. This invention employs a "pre-training + fine-tuning" transfer learning strategy. By learning common task features on source domain data and implementing personalized parameter adaptation on target subjects, it significantly reduces the calibration cost and deployment time of the speech brain-computer interface system, effectively improving the model's generalization ability and practicality in cross-subject application scenarios. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the overall network architecture of the speech brain-computer interface decoding system provided in an embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0028] A speech brain-computer interface decoding system based on knowledge data fusion includes: a multi-scale temporal convolution module, a scale-specific channel attention module, a two-stage multi-scale spatiotemporal fusion module, a Transformer encoding module, and a classification output module. The multi-scale temporal convolution module uses parallel convolution branches with multiple different temporal receptive fields to extract temporal features at multiple different time scales from speech-related neural activity signals. The scale-specific channel attention module is used to adaptively learn and recalibrate the channel weights of temporal features at multiple different time scales to obtain recalibrated temporal features at multiple different time scales. The two-stage multi-scale spatiotemporal fusion module is used to perform cross-channel fusion of recalibrated temporal features at the same time scale in the first stage using grouped one-dimensional convolution to obtain cross-channel fused features. In the second stage, standard one-dimensional convolution is used to perform cross-scale fusion of cross-channel fused features at multiple different time scales and then map them to a unified dimension embedding space to obtain embedded features. The Transformer encoding module is used to segment and tokenize the embedded features along the time axis to capture long-range temporal dependencies and global context information, thereby obtaining global features; The classification output module uses global features to classify and predict speech categories.

[0029] This invention addresses the diversity of temporal structures through multi-scale convolution, the selectivity of spatial features through channel attention, the collaborative modeling of spatiotemporal features through two-stage fusion, and finally, the understanding of global semantics through Transformer. These modules work together to form a highly efficient and accurate speech brain-computer interface decoding system, fundamentally overcoming the limitations of existing technologies.

[0030] Example 1 This invention proposes a knowledge-data fusion-based speech brain-computer interface decoding system (NeuralModulation-aware Spatial-Temporal Transformer, NMST-Former), which integrates prior knowledge in the speech domain into the model design and combines channel attention mechanism and Transformer architecture to further model spatiotemporal features.

[0031] like Figure 1 As shown, NMST-Former combines prior knowledge in the speech domain with data-driven modeling. It achieves targeted modeling of the spatiotemporal features of speech EEG signals by constructing a modulation frequency-driven multi-scale temporal convolution module, a scale-specific channel attention module, a two-stage multi-scale spatiotemporal fusion module, a Transformer encoding module, and a classification output module. After introducing prior knowledge, it further performs end-to-end optimization using training data, thereby improving the accuracy and overall performance of speech EEG decoding.

[0032] Next, we will introduce the structure of each module.

[0033] 1. Multi-scale temporal convolution module This invention explicitly incorporates the prior modulation frequencies of speech-related neural activities into the network structure. This invention preserves the original signal branches. Furthermore, three parallel convolutional branches with different temporal receptive fields are constructed based on this, each using a different kernel size. One-dimensional depthwise convolution is used to model the signal in the temporal domain, corresponding to time lengths of 500 ms, 250 ms, and 125 ms, extracting temporal information related to phonemes, subsyllabic structures, and syllable rhythms. The original signal branch retains information at slower time scales, corresponding to prosody and phrase-level structures at approximately 2 Hz. Each scale branch employs depthwise convolution, performing independent convolutions channel by channel to avoid cross-channel aliasing and reduce the number of parameters, resulting in scale-specific temporal features. The output of this module is... A multi-granularity time-domain representation consisting of four branches.

[0034] Each temporal convolution branch preferably employs a channel-independent depthwise convolution to avoid irrelevant aliasing between different electrode channels in the early stages, while also reducing the number of parameters and computational complexity. After this module, a multi-granularity temporal feature representation consisting of the original signal branch and three parallel convolution branches with different temporal receptive fields can be obtained.

[0035] 2. Scale-Specific Channel Attention Module To enhance the modeling capability of key brain regions and electrode channels at different time scales, this invention introduces scale-specific channel attention (Efficient Channel Attention, ECA) into the original signal branch and the parallel convolutional branches of three different temporal receptive fields. Given the first... Input features of each scale branch ECA first performs global average pooling along the time dimension to obtain the channel description vector. Subsequently, a one-dimensional convolution is applied to the channel dimension, and the channel weights are obtained using the Sigmoid function. Finally, the input features are recalibrated channel by channel. ECA employs a strategy of adaptively calculating the convolutional kernel size based on the number of channels C, enabling the attention mechanism to automatically adjust the local receptive field according to different channel sizes, balancing efficiency and expressive power.

[0036] This invention does not use a single set of channel weights for all scales, but instead learns scale-specific channel weights for different scale branches. Therefore, it can more accurately reflect the differences in the contribution of key electrodes under different modulation time scales.

[0037] 3. Two-stage multi-scale spatiotemporal fusion module After obtaining the features of each scale branch recalibrated by channel attention, this invention splices the four branches along the channel dimension to obtain... To fully utilize features across different scales and obtain a unified dimensional embedding representation, this invention employs a two-stage multi-scale feature fusion module. The first stage uses grouped one-dimensional convolutions for intra-scale fusion, ensuring that the convolution only combines channels within each scale block. The second stage uses standard one-dimensional convolutions (kernel size 3, output channels E) for cross-scale fusion along the channel dimension, mapping features to a unified embedding dimension. ,get This completes the integration of cross-scale and cross-channel information.

[0038] This two-stage fusion approach is beneficial for preserving the independence of features at each scale and the ability to interact across scales, making it suitable for processing hierarchical multi-timescale information in speech EEG.

[0039] 4. Transformer Encoding Module Features after cross-scale and cross-channel information fusion Along the time dimension according to a fixed window size Perform segmented average pooling to divide the continuous time series into segments. A time period, and the tensor rearranged to obtain Each token represents a local embedding feature within a corresponding time window. A learnable classification label is then added before the token sequence. And superimposed with learnable positional encoding Construct the Transformer encoder input .

[0040] The Transformer encoder consists of L stacked encoding modules, each layer containing a multi-head self-attention sublayer and a feedforward network sublayer, along with residual connections and layer normalization. After layer encoding, extract the sequence-level representation corresponding to the [CLS] Token. As global features, they are mapped to the category space through the classification head to obtain the final prediction result.

[0041] 5. Classification Output Module After the Transformer encoder outputs, the feature vector corresponding to the [CLS] Token is extracted as the global feature for the entire trial, and the prediction result is obtained by mapping it to the class space through the classification head.

[0042] NMST-Former makes three core improvements to its network structure to address the physiological characteristics of speech neural activity: ① Knowledge-driven multi-scale temporal convolution: Based on the prior knowledge that "low-frequency modulation signals represent the hierarchical structure of language" in speech neural activity, three parallel branches are designed. Convolutional kernels of sizes 32, 64, and 128 are used respectively to capture temporal dynamic features corresponding to the phoneme (approximately 16 Hz), subsyllable (approximately 8 Hz), and syllable rhythm (approximately 4 Hz) levels, while preserving the original signal to capture slower prosody and phrase information. ② Scale-specific channel attention module (ECA): A lightweight attention module is introduced into each scale branch. It adaptively learns the weights of electrode channels for different time scales, thereby accurately identifying key brain region spatial features related to speech tasks. ③ Two-stage feature fusion mechanism: To prevent premature aliasing of multi-scale features, the model first performs internal compression and alignment through grouped convolution, then achieves deep fusion across scales and channels through standard convolution, and finally maps to the Transformer's embedding space.

[0043] Example 2 A training method for a speech brain-computer interface decoding system based on knowledge data fusion includes: Train a speech brain-computer interface decoding system using a speech brain-computer interface dataset; The cross-entropy between the predicted speech category and the actual speech category is used as the loss function. The parameters are updated by backpropagation, and the training continues until convergence to obtain a well-trained speech brain-computer interface decoding system.

[0044] Speech brain-computer interface datasets include: BCI2020 Speech Imagined Dataset, VocalMind Dataset, Full Spectrum Chinese Decoding Dataset, Brain-to-Text Dataset (2024 / 2025), Chinese-MEG 48 Dataset, Chisco (Chinese Imagined Speech Corpus), and ChineseEEG Dataset.

[0045] This invention employs specific preprocessing techniques based on the characteristics of speech signals. First, the original signal undergoes detrending and rereference processing to reduce slow drift and interference introduced by the reference electrodes. A 30 Hz high-pass filter is applied to highlight high-frequency cortical neural activity related to speech production and motor processes. Simultaneously, a notch filter is performed to remove 50 Hz power frequency noise and its odd harmonic components. After filtering, a Hilbert transform is applied to the signal, and the instantaneous amplitude envelope is extracted to capture the modulation information of the signal amplitude over time. The extracted envelope signal serves as the system input for subsequent speech decoding analysis.

[0046] The target subject dataset is formed by randomly selecting 1%-10% of the subject data from the speech brain-computer interface dataset, and the remaining subject data from the speech brain-computer interface dataset is formed by the source domain subject dataset. The speech brain-computer interface decoding system was first pre-trained using the source domain subject dataset, and then fine-tuned using the target subject dataset to obtain the trained speech brain-computer interface decoding system.

[0047] This indicates that the present invention has made some improvements in the training strategy. Addressing the common issue of individual differences in brain-computer interfaces, a "pre-training + fine-tuning" training mode was designed for the inter-subject task. Pre-training is first completed on source domain subject data, and then fine-tuning is performed using a small amount of data from the target subject.

[0048] During the training phase, the cross-entropy loss function is used to optimize the entire network end-to-end. The optimizer is Adam, and an early stopping strategy is introduced during training to prevent overfitting.

[0049] Example 3 In Embodiment 3 of this invention, three evaluation tasks are set up: Task 1 (within-subject classification) is trained and evaluated separately for each subject to evaluate the model's ability to decode brain signal patterns of a specific individual; Task 2 (mixed-subject classification) trains a unified model for all subjects to evaluate the model's learning ability under mixed data distribution of multiple subjects; Task 3 (cross-subject fine-tuning classification) adopts a leave-one-out-of-subject strategy, randomly selecting one subject's data from the speech brain-computer interface dataset as the target data, first training with the remaining 14 subjects' data, and then fine-tuning with 30% of the target subject's training data to evaluate the model's cross-subject transferability and rapid adaptation ability.

[0050] Evaluation metrics include: ACC (accuracy), Macro-F1 (macro-average F1 score), Kappa (Kappa coefficient), and Macro-AUC (macro-average AUC). ACC focuses on the overall correctness and error rate, intuitively reflecting the overall decoding success rate. Macro-F1 focuses on the balance of each category. The Kappa coefficient is used to evaluate the consistency between the model's prediction results and the true labels. Macro-AUC focuses on the discrimination of each category, evaluating whether the model's discrimination boundary for each category is clear.

[0051] As shown in Table 1, the experimental results demonstrate that, on the speech brain-computer interface dataset, the NMST-Former proposed in this invention has the following advantages compared to traditional convolutional neural network models (EEGNet, DeepConvNet, ShallowConvNet), multi-scale and multi-view convolutional decoding models (ADFCNN, FBCNet, IFNet, EEGWaveNet), and decoding models based on the Transformer structure (Conformer, Deformer, CTNet, MSVTNet, TMSA-Net, MSCFormer, DBConformer, FAST). These advantages are mainly reflected in the following three aspects: 1. Leading in overall performance NMST-Former achieved optimal results on the vast majority of evaluation metrics, demonstrating the effectiveness of its architecture.

[0052] In Task 1, NMST-Former ranked first in all three key metrics: ACC, Kappa, and Macro-AUC, with scores of 58.44%, 0.4806, and 0.8551, respectively, significantly outperforming the suboptimal model.

[0053] In Task 2: NMST-Former also achieved the best results in ACC, Kappa and Macro-AUC, at 54.84%, 0.4356 and 0.8258 respectively.

[0054] In Task 3: Among the four metrics (ACC 44.04%, Macro-F1 0.4305, Kappa 0.3006, Macro-AUC 0.7025), Macro-AUC was basically on par with the optimal value (Conformer, 0.7030), while the other three metrics achieved the best results.

[0055] This demonstrates that NMST-Former delivers the most stable and robust decoding performance regardless of the task settings.

[0056] 2. Effectively solves the problem of category imbalance When dealing with multi-class classification problems, accuracy and macro-average F1 score are equally important. NMST-Former performs well on both metrics, demonstrating that the model can stably capture the temporal and spatial features specific to each class in multi-class speech decoding tasks.

[0057] Compared to the unbalanced baseline models, NMST-Former effectively overcomes the class bias that models are prone to exhibit in small-sample multi-class tasks. For example, in the within-subjects classification task (Task 1), the existing TMSA-Net model achieves an accuracy of 42.04%, but its Macro-F1 score is only 0.3995, reflecting its limited discriminative performance on minority class samples.

[0058] Experimental results show that NMST-Former achieved an accuracy of 58.44% and a macro-average F1 score of 0.5724 in Task 1, both of which outperform the current state-of-the-art models. This indicates that NMST-Former, through multi-scale feature extraction and global modeling, effectively suppresses the model's tendency to overfit high-frequency class samples, achieving balanced recognition of various speech intentions and significantly improving the system's classification reliability and practical value in real-world interactive scenarios.

[0059] 3. Significantly superior to various baseline models By comparing it with models using different technical approaches, the superiority of NMST-Former can be clearly seen.

[0060] Compared to traditional CNN models: NMST-Former improved the accuracy of task 1 by about 20 percentage points compared to traditional convolutional models such as EEGNet and DeepConvNet, which shows that the introduction of multi-scale and Transformer structure brought about a huge performance leap.

[0061] Compared to multi-scale CNN models: Compared to models such as FBCNet and IFNet that also employ the multi-scale approach, NMST-Former still maintains an accuracy advantage of about 8-10 percentage points, indicating that it is more efficient in multi-scale feature extraction and fusion.

[0062] Compared to other Transformer models: In comparisons with state-of-the-art Transformer-based models such as Conformer and Deformer, NMST-Former demonstrates an accuracy advantage of approximately 5-8 percentage points in Tasks 1 and 2, with a more significant advantage in Task 3. This validates that the specific module combination proposed by NMST-Former (multi-scale temporal convolution, scale-specific channel attention, and two-stage fusion) is more effective than existing Transformer variants in capturing the spatiotemporal features of EEG signals.

[0063] Table 1

[0064] In the three tasks, the method proposed in this invention achieved the best results on most evaluation metrics, and only achieved the second-best result on the Macro-AUC metric in Task 3, but it was basically on par with the best result, indicating that the method has good stability and generalization ability.

[0065] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A speech brain-computer interface decoding system based on knowledge data fusion, characterized in that, include: The system includes a multi-scale temporal convolution module, a scale-specific channel attention module, a two-stage multi-scale spatiotemporal fusion module, a Transformer encoding module, and a classification output module. The multi-scale temporal convolution module uses parallel convolution branches with multiple different temporal receptive fields to extract temporal features at multiple different time scales from speech-related neural activity signals. The scale-specific channel attention module is used to adaptively learn and recalibrate the channel weights of temporal features at multiple different time scales to obtain recalibrated temporal features at multiple different time scales. The two-stage multi-scale spatiotemporal fusion module is used to perform cross-channel fusion of recalibrated temporal features at the same time scale in the first stage using grouped one-dimensional convolution to obtain cross-channel fused features. In the second stage, standard one-dimensional convolution is used to perform cross-scale fusion of cross-channel fused features at multiple different time scales and then map them to a unified dimension embedding space to obtain embedded features. The Transformer encoding module is used to segment and tokenize the embedded features along the time axis to capture long-range temporal dependencies and global context information, thereby obtaining global features; The classification output module uses global features to classify and predict speech categories.

2. The speech brain-computer interface decoding system based on knowledge data fusion as described in claim 1, characterized in that, The multi-scale temporal convolution module includes an original signal branch and multiple parallel convolution branches with different temporal receptive fields. The original signal branch directly outputs speech-related neural activity signals to the scale-specific channel attention module. The multiple parallel convolution branches with different temporal receptive fields perform temporal modeling of speech-related neural activity signals through one-dimensional depth convolution with multiple different convolution kernels to obtain temporal features at multiple different time scales, which are then output to the scale-specific channel attention module.

3. The speech brain-computer interface decoding system based on knowledge data fusion as described in claim 2, characterized in that, The scale-specific channel attention module provides a corresponding channel attention module for the original signal branch and the parallel convolutional branch of each different temporal receptive field. Each channel attention module performs global average pooling along the temporal dimension of the input feature to obtain the channel description vector. Then, a one-dimensional convolution is applied to the channel description vector in the channel dimension and the channel weights are obtained through the Sigmoid function. Finally, the channel weights are used to recalibrate the input feature channel by channel.

4. A speech brain-computer interface decoding system based on knowledge data fusion as described in any one of claims 1-3, characterized in that, The Transformer encoding module performs segmented average pooling of the embedded features along the time dimension according to a fixed window size to obtain multiple tokens. A learnable classification label is added before each token, and then a learnable positional encoding is superimposed to construct the input of the Transformer encoder. The Transformer encoder includes multiple stacked encoding modules. Each encoding module includes a multi-head self-attention sub-layer and a feedforward network sub-layer, and is combined with residual connections and layer normalization. After the Transformer encoder input is encoded through multiple layers, global features are obtained.

5. A training method for a speech brain-computer interface decoding system based on knowledge data fusion, characterized in that, include: The speech brain-computer interface decoding system based on knowledge data fusion as described in any one of claims 1-4 is trained using a speech brain-computer interface dataset. The error between the predicted speech category and the actual speech category of the neural activity signal in the speech brain-computer interface dataset is used as the loss function. The parameters are updated by backpropagation and trained until convergence to obtain a well-trained speech brain-computer interface decoding system.

6. The training method for a speech brain-computer interface decoding system based on knowledge data fusion as described in claim 5, characterized in that, Prior to training, the neural activity signals in the speech brain-computer interface dataset were processed as follows: The neural activity signal is sequentially subjected to detrending, rereference, high-pass filtering, notch filtering, Hilbert transform, and instantaneous amplitude envelope extraction to obtain the envelope signal, which serves as the input to the speech brain-computer interface decoding system.

7. A training method for a speech brain-computer interface decoding system based on knowledge data fusion as described in claim 5 or 6, characterized in that, The training also includes: The target subject dataset is formed by randomly selecting 1%-10% of the subject data from the speech brain-computer interface dataset, and the remaining subject data from the speech brain-computer interface dataset is formed by the source domain subject dataset. The speech brain-computer interface decoding system was first pre-trained using the source domain subject dataset, and then fine-tuned using the target subject dataset to obtain the trained speech brain-computer interface decoding system.

8. A speech brain-computer interface decoding method based on knowledge data fusion, characterized in that, include: The neural activity signal to be decoded is input into the training method of a speech brain-computer interface decoding system based on knowledge data fusion as described in any one of claims 5-7, and the resulting speech brain-computer interface decoding system is trained to predict the speech category.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, is used to implement the training method of the speech brain-computer interface decoding system based on knowledge data fusion as described in any one of claims 5-7 or the speech brain-computer interface decoding method based on knowledge data fusion as described in claim 8.

10. A computer product, characterized in that, When the product is running, it enables the computer to execute the training method of the speech brain-computer interface decoding system based on knowledge data fusion as described in any one of claims 5-7, or the speech brain-computer interface decoding method based on knowledge data fusion as described in claim 8.