Speech imagination decoding method fusing brain region dynamic routing and frequency-space semantics

By integrating brain region dynamic routing and frequency-space semantics, the problems of data sparsity and cross-domain generalization in EEG speech imagery recognition were solved, achieving high-precision and robust decoding results and improving the model's expressive power and cross-domain generalization performance.

CN120974276APending Publication Date: 2025-11-18HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511109210.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing EEG speech imagery recognition technology faces problems such as small data scale, sparse samples, low signal-to-noise ratio, and poor cross-subject and cross-scene transfer performance, making it difficult to achieve high accuracy and robust cross-domain generalization.

Method used

By employing a method that integrates dynamic brain region routing and frequency-space semantics, and combining a brain region perception and labeling routing module and a frequency-space semantic feature module with a sparse gating mechanism, we can achieve region perception and heterogeneous feature modeling, thereby improving the robustness and cross-domain generalization ability of the model.

Benefits of technology

It improves the accuracy and robustness of EEG-based verbal imagery recognition, enhances the model's expressive power and cross-domain generalization performance, and achieves high-precision decoding results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974276A_ABST
    Figure CN120974276A_ABST
Patent Text Reader

Abstract

A speech imagination decoding method fusing brain region dynamic routing and frequency-space semantics comprises the steps that a brain region perception mark routing module is constructed, and a region class mark generation unit generates class marks including a brain region class mark and a global class mark according to a functional brain region; the region perception attention unit is used for controlling interaction between class marks and outputting class mark feature representation by constructing a multi-head attention mechanism; the gating expert network unit receives the class tag feature representation and outputs fusion features through a sparse gating mechanism; the brain region perception classification unit outputs category prediction according to the fusion features; the frequency-space semantic feature module is constructed, and the frequency-space semantic feature module comprises a heterogeneous feature expert selection unit which divides the fusion frequency spectrum-space features into a plurality of groups along the channel dimension and selects weights, and then sends the groups to a sparsely gated hybrid expert module for modeling and then outputs splicing features; the frequency-space semantic classification unit is used for outputting a classification result according to the splicing features; and carrying out weighted fusion on category prediction and classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electroencephalogram (EEG) signal processing technology, and in particular to a speech imagery decoding method that integrates dynamic brain region routing and frequency-space semantics. Background Technology

[0002] EEG-based speech imagery recognition is a cutting-edge research area in brain-computer interface (BCI) technology. It aims to automatically identify and decode an individual's silent speech by analyzing the brain's electrical activity during speech imagery. Compared to traditional speech recognition systems, speech imagery recognition requires no sound output and holds significant promise for applications in areas such as assisted communication, silent control, and interaction with people with disabilities. However, current research on EEG-based speech imagery faces several key challenges that severely limit its accuracy and generalization capabilities in practical applications.

[0003] First, due to the high cost of EEG acquisition and the complexity of experimental procedures, currently available EEG datasets for verbal imagery are generally small in size and sparsely distributed, making it difficult to support large-scale training of deep learning models. This data limitation easily leads to overfitting on the training set, making it difficult to generalize to new subjects or tasks. Furthermore, EEG signals themselves are highly susceptible to interference, resulting in a low signal-to-noise ratio, further exacerbating the performance degradation of models in cross-subject or cross-scene transfer. Therefore, there is an urgent need to develop models with robust feature extraction capabilities and cross-domain generalization abilities to improve stability and generalization performance under multi-source data.

[0004] Secondly, different brain regions perform different functional roles in speech processing, with some regions responding more significantly to semantic tasks. Simultaneously, due to differences in physiological structure, cognitive strategies, and emotional states, there are significant individual differences in brain region activation patterns among different subjects. This leads to data distribution shifts between the training and test sets, especially in cross-subject or cross-task scenarios, making it difficult for models to generalize to new domains or users, thus becoming a core challenge in current "EEG domain generalization" tasks.

[0005] Furthermore, brain activity during verbal imagination is highly nonlinear and dynamic, and its semantic intent exhibits a multidimensional and complex distribution in the spectral-spatial telecommunication pattern. Traditional single-viewpoint or shallow feature extraction methods are insufficient to effectively characterize its potential semantic subspace, limiting the model's ability to recognize and transfer fine-grained differences, and further exacerbating the performance degradation of the model during cross-domain generalization.

[0006] Therefore, how to integrate the grouping modeling mechanism of brain region structure priors and latent semantics, and introduce dynamic sparse expert routing strategies to achieve region perception and heterogeneous feature modeling, has become a key technical challenge to improve the domain generalization ability, model robustness and interpretability of EEG speech imagery recognition. Summary of the Invention

[0007] This invention proposes a speech imagery decoding method that integrates brain region dynamic routing and frequency-space semantics. It designs a dual-path dynamic sparse gating expert mechanism that combines region priors and latent semantic division. While reducing annotation costs, it achieves high-precision, robust, and interpretable EEG speech imagery decoding results, effectively improving the model's expressive power and cross-domain generalization performance.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics includes the following steps:

[0010] Step S1: Collect and preprocess raw EEG data to generate EEG samples;

[0011] Step S2: Extract features from the EEG sample to generate differential entropy features;

[0012] Step S3: Construct a brain region perception labeling routing module, including:

[0013] Brain region segmentation units are used to segment functional brain regions containing multiple channels;

[0014] The region class label generation unit generates class labels including brain region class labels and global class labels based on the functional brain regions;

[0015] The region-aware attention unit is used to control the interaction between the class labels and output the class label feature representation by constructing a multi-head attention mechanism;

[0016] The gated expert network unit receives the class-labeled feature representation and outputs fused features through a sparse gating mechanism;

[0017] The brain region perception and classification unit outputs a category prediction based on the fused features;

[0018] Step S4: Construct the frequency-space semantic feature module, including:

[0019] The frequency-space attention feature extraction unit takes the differential entropy feature as input and outputs the fused frequency-space feature.

[0020] The heterogeneous feature expert selection unit divides the fused spectral-spatial features into multiple groups along the channel dimension and selects weights, then sends them to the sparsely gated hybrid expert module for modeling and outputs the spliced ​​features.

[0021] The frequency-space semantic classification unit outputs the classification result based on the concatenated features;

[0022] Step S5: Perform a weighted fusion of the category prediction and the classification result, and output the final fusion output.

[0023] Preferably, in the region class label generation unit, the class label is initialized; the global class label, the brain region class label, and the channel label are concatenated to form a Transformer input sequence.

[0024] Preferably, in the region-aware attention unit, the interaction between the class tags is controlled by constructing an attention mask matrix; the global class tag can interact bidirectionally with all the class tags; the brain region class tag can interact bidirectionally with its corresponding channel tag; the channel tags can interact bidirectionally with each other; and all class tags can access themselves.

[0025] Preferably, the gated expert network unit inputs the class label into a shared gated network, calculates its selection weight at each expert, and uses a Top-k strategy to select the top k experts activated for each class label; stacks the feature representations output by each expert to generate a representation vector; and fuses the representation vector with the gated weights to generate a weighted fusion feature.

[0026] Preferably, in the brain region perception classification unit, the weighted fusion features are concatenated along the feature dimension to obtain a fusion feature representation; the fusion feature representation is classified and predicted by a two-layer perceptron classifier, each layer of the perceptron classifier including layer normalization and nonlinear transformation; the category prediction matrix is ​​calculated and output in the output layer.

[0027] Preferably, the category prediction matrix is ​​input into the cross-entropy loss function to obtain the optimized main classification loss; the expert average weight distribution is obtained according to the expert selection weights; the balance of the expert average weight distribution is measured using the entropy function to construct the normalized balance loss; and the total loss function is constructed based on the normalized balance loss and the optimized main classification loss.

[0028] Preferably, in step S5, the final fusion output is:

[0029] Z fused =α·Z route +(1-α)·Z subspace

[0030] Where α is the adaptive weight, which is automatically updated after each training round, and Z... route This represents the category prediction Z after passing through the brain region perception label routing module. subspace This represents the classification result after passing through the frequency-space semantic feature module.

[0031] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0032] 1. Compared with traditional domain generalization methods that use channels as the smallest modeling unit, this invention introduces brain region structure priors, divides channels into functional brain regions, and designs region-aware class labels to guide attention to relevant channels within local brain regions. This effectively enhances the model's adaptability to differences in neural structure and its cross-subject generalization ability, thereby improving the model's robustness and interpretability in speech decoding.

[0033] 2. Compared with the existing general fully connected attention mechanism, this invention adopts a structured attention mask to explicitly constrain the attention flow path between global class labels, brain region class labels and channel features, and realizes a hierarchical modeling strategy of global aggregation, intra-regional focusing and inter-channel interaction, thereby improving the structural clarity of feature representation and the attention controllability of the model.

[0034] 3. Compared with the existing strategy of using fixed-parameter networks for unified modeling, this invention introduces a hybrid expert network and combines it with a sparse gating mechanism to achieve dynamic selection of experts at the sample level. This enables different semantic heterogeneous features or brain region representations to activate the most suitable expert subnetwork, thereby improving the model's expressive diversity and generalization elasticity, and alleviating the problem of significant semantic shifts between different subjects.

[0035] 4. Compared with EEG generalization models based on a single feature path, this invention integrates a brain region dynamic routing module and a frequency-space semantic feature module. By dynamically integrating the classification results of the two modules through an adaptive weighting strategy, it achieves complementary synergy between structural perception and semantic expression, thereby enhancing the model's ability to represent fine-grained semantic differences in EEG signals and its generalization performance across semantic tasks. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description represent only some embodiments of the present invention. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0037] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention. Detailed Implementation

[0038] To make the technical means, inventive features, objectives, and effects of the invention readily understandable, the invention is further described below with reference to specific illustrations. However, the invention is not limited to the embodiments described below.

[0039] It should be noted that the structures, proportions, sizes, etc., illustrated in the accompanying drawings of this specification are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0040] Example 1:

[0041] This method first introduces a brain region perception labeling mechanism structurally, dividing channels into multiple class labels according to brain regions. These labels, along with global labels, are input into a shared Transformer encoder, achieving intra-regional modeling and inter-regional interference suppression through prior attention masks. Each label output is dynamically routed to its corresponding expert via sparse gating, enabling region-guided fine-grained modeling. Then, a spectral-spatial attention module extracts high-dimensional semantic features from the fusion of spectral and spatial channels, dividing these features into multiple heterogeneous feature blocks. These blocks are then input into a hybrid expert module for differentiated modeling, enhancing the expression and generalization ability of fine-grained semantic features. The outputs of the two sub-modules are jointly decided through an adaptive fusion strategy, effectively improving the model's expressive power and cross-domain generalization performance.

[0042] like Figure 1 The method for decoding speech imagery, which integrates dynamic brain region routing and frequency-space semantics, includes the following steps:

[0043] Step S1: Electrodes are placed on the scalp using specialized equipment to collect EEG signals from multiple subjects at different time points using the same paradigm in a non-invasive manner. The collected EEG data is filtered and downsampled to reduce noise and improve signal quality.

[0044] This experiment processed the EEG signals of 10 subjects during verbal imagery tasks (up, down, left, right). In the first and second phases, each subject completed 80 tests; in the third phase, the number of tests completed by each subject varied due to individual circumstances. EEG data was acquired using a 128-channel system at a sampling rate of 1024Hz and processed using a zero-phase bandpass finite impulse response filter, with a low-frequency cutoff frequency set to 0.5Hz and a high-frequency cutoff frequency set to 100Hz. To ensure data quality, the acquired EEG signals underwent two-step preprocessing. First, the sampling rate was reduced from 1024Hz to 256Hz, and a 1-50Hz bandpass filter was applied to remove potential noise and artifacts.

[0045] Step S2: Extract features from the EEG samples to generate differential entropy features; and assign speech imagination category labels to some EEG samples. The collected EEG samples are divided into two parts: (1) Labeled samples L: EEG samples containing known speech imagination category labels, used for the initial training of the model; (2) Unlabeled samples U: EEG samples containing no assigned speech imagination category labels. These samples will be selectively labeled through intelligent sampling strategies in the subsequent active learning process.

[0046] Step S3: Construct a brain region perception labeling routing module;

[0047] Step S31: In the brain region segmentation unit, based on neuroscience knowledge, the 62 channels of the EEG are divided into 16 functional brain regions according to their spatial distribution and functional association with brain regions. Each brain region contains several channels, specifically: area_1 contains channels AF3, FP1, FPZ, FP2, and AF4; area_2 contains channels F7 and F5; area_3 contains channels F3, F1, FZ, F2, and F4; area_4 contains channels F6 and F8; area_5 contains channels FT7, FC5, T7, C5, TP7, and CP5; area_6 contains channels FC3, FC1, FCZ, FC2, and FC4; area_7 contains channels FC6, FT8, C6, T8, CP6, and TP8; area_8 contains channels C3, C1, and C2. Z, C2, C4, area_9 contains channels CP3, CP1, CPZ, CP2, CP4, area_10 contains channels P7, P5, area_11 contains channels P3, P1, PZ, P2, P4, area_12 contains channels P6, P8, area_13 contains channels PO7, PO5, CB1, area_14 contains channels PO3, POZ, PO4, area_15 contains channels PO6, PO8, CB2, area_16 contains channels O1, OZ, O2. The 128-channel system can be seen as an interpolated extension of the traditional 62-channel system.

[0048] Step S32: The region class label generation unit initializes the class labels, network parameters θ, and optimizer, and sets the base learning rate η = [1e...]. -3 5e -4 ,1e -4 Batch size B = 32;

[0049] Initialize one global class label and A brain region class labels, and copy them to the batch dimension:

[0050]

[0051] Where d represents the dimension of the differential entropy feature after being embedded by brain region encoding.

[0052] Then, the global class marker, the region class marker, and the channel marker are concatenated to form the Transformer input sequence:

[0053]

[0054] in, This indicates a channel-level marker.

[0055] Step S33: Construct region-aware attention units;

[0056] By constructing an attention mask matrix M to control different labels, the total number of labels is:

[0057] N = 1 + A + C

[0058] Where 1 represents the global class label, A represents the number of brain regions, and C represents the number of channels;

[0059] Let the attention mask be

[0060] M∈{0,1} B×N×N

[0061] The global class marker (position 0) is interactive with all markers in both directions:

[0062]

[0063] Let the set of channel indices for the i-th brain region be . Bidirectional attention between brain region class markers (i-th, position 1+i) and their corresponding channel markers:

[0064]

[0065] All channel markers can access each other:

[0066]

[0067] All tags are accessible to themselves, meaning the diagonal is set to 1:

[0068]

[0069] Step S34: Multi-head attention encoding under attention mask;

[0070] To improve the feature representation ability and robustness of EEG signals, the differential entropy features were first analyzed. Perform channel-level linear embedding transformation to obtain a uniform-dimensional embedding representation:

[0071]

[0072] Where F represents the dimension of the differential entropy feature, to avoid the model's excessive dependence on specific channel labels, which could lead to overfitting, this invention applies a Dropout operation to the embedded features, randomly zeroing out some channel dimension features. This operation can be mathematically described as follows:

[0073] X in =X'☉P

[0074] Where P∈{0,1} B×C×d Let be the Dropout mask matrix, and let ⊙ denote element-wise multiplication, satisfying:

[0075] Pr[P b,c,d =1]=1-p

[0076] p∈(0,1) is the set probability of dropping.

[0077] The input sequence is linearly mapped to generate the query, key, and value matrix W required for the multi-head attention mechanism. Q W K W V , and input X in Multiply them to generate query, key, and value vectors respectively:

[0078]

[0079] Where h represents the number of attention heads, d h For each head dimension, d = h·d h .

[0080] Then, Q, K, and V were split and rearranged in multiple directions:

[0081]

[0082] The same applies to K and V.

[0083] For each attention head, compute the scaled dot product attention:

[0084]

[0085] To guide attention in accordance with brain region structural information, an attention mask matrix M∈{0,1} is introduced here. B×N×N In the attention score, assign a minimum value (e.g., -10) to non-associated locations. 9 To achieve path suppression:

[0086]

[0087] The obtained attention scores are then Softmax normalized, and the attention weights are calculated:

[0088]

[0089] Multiplying this by the value vector V yields the context representation:

[0090]

[0091] After concatenating all the headers, linearly transform them back to the original dimensions:

[0092]

[0093] Then through a linear layer The projection of this image yields the final attention output:

[0094] X attn =Z cat ·W O

[0095] To alleviate the vanishing gradient problem in deep networks and accelerate the convergence process, a residual connection mechanism is introduced:

[0096] X1 = LayerNorm(X attn +X in )

[0097] LayerNorm represents the layer normalization operation.

[0098] After adding attention-enhanced features while preserving the original information, the data is fed into a position-independent feedforward sub-network FNN to improve the nonlinear modeling capability of the features.

[0099] X ffn =W4·σ(W3·X1+b1)+b2

[0100] in, Let d be the projection matrix, d′ represent the intermediate layer dimension, σ(·) represent the nonlinear activation function GELU, and b1 and b2 are bias terms.

[0101] The FNN output is then summed with the input residual to form the final output X. out This further enhances deep stability and nonlinear modeling capabilities:

[0102] X out =LayerNorm(X ffn +X1)

[0103] Repeat the above process L times to deepen the model's representational capabilities, where L is the number of Transformer layers.

[0104] Extract the first 1+A labels from the final output as global label feature representation and brain region class feature representation:

[0105] T cls =X out [:,:1+A,:]

[0106] Used for subsequent gating expert network units.

[0107] Step S35: Construct a gating expert network unit;

[0108] First, calculate the i-th tag based on the gating network. The selection weight α for each expert i :

[0109]

[0110] Among them, g i For feature projection, These are gating network parameters, b g It is the gating network bias term, and E indicates that all labels share E expert networks.

[0111] To enable different experts to gradually develop expressive abilities specializing in different brain regions during training, this invention employs a Top-k based sparse gating strategy, activating only the few expert networks most relevant to the current input:

[0112]

[0113] according to Perform sparse weighted normalization:

[0114]

[0115] Here, ∈ represents the minimum value, to avoid the denominator being 0.

[0116] Each expert receives the same input h. i And output a feature representation:

[0117] o i,e =f e (h,), for e=1,…E

[0118] Where f e (·) represents the e-th expert network.

[0119] Stack the feature representations output by all experts to generate a representation vector:

[0120]

[0121] Where d″ represents the output dimension of the expert network.

[0122] The representation vector is fused with a gated weighted output to generate a weighted fused feature:

[0123]

[0124] Step S36, Design of Brain Region Perception Classification Units:

[0125] By concatenating all the expert outputs along the feature dimension, we obtain the fused feature representation:

[0126] y concat =[y1||y2|| … ||y 1+A ]

[0127] To enhance feature discriminative power, a two-layer perceptron classifier is designed to classify and predict the above fused features. Each layer includes layer normalization and nonlinear transformation.

[0128]

[0129] Among them W s It is the projection matrix of the classifier, b s It is the bias vector.

[0130] After passing through two layers of perceptron, the category prediction matrix Z is calculated in the output layer. route :

[0131] Z route =concat(h1,h2,…h m )

[0132] Where C1 is the number of categories in the hidden speech classification, and m is the number of labeled samples.

[0133] The above prediction Z route Input the cross-entropy loss function, and define the true sample labels as one-hot encoded. This invention yields an optimized main classification loss:

[0134]

[0135] To mitigate the problem of excessive reliance on certain experts that may be caused by sparse gating mechanisms and to improve model stability, an expert-balanced regularized loss was also constructed.

[0136] The present invention has obtained the selection weight α of each expert in the gating network. i From this invention, the average weight distribution of the selected experts can be statistically determined:

[0137]

[0138] Use the entropy function to measure its equilibrium and construct the normalized equilibrium loss:

[0139]

[0140] Where ∈ represents a local minimum.

[0141] Finally, the model jointly optimizes the main classification loss and the expert balance regularization term to construct the total loss function:

[0142]

[0143] Where λ bal These are the weighting coefficients for the balancing term.

[0144] Step S4: Construct the frequency-space semantic feature module;

[0145] Step S41, Frequency-space attention feature extraction unit;

[0146] First, input differential entropy features. The fused spectral-spatial features are extracted using the frequency-space attention feature extraction unit.

[0147]

[0148] Where DAttention represents the frequency-space attention feature extraction unit, D represents the output layer dimension of the DAttention module, and the attention weights are R = {R f R c} represent frequency band and channel (spatial) attention, respectively.

[0149] The frequency band attention weight R in the DAttention module f A global frequency band description F is extracted using adaptive average pooling and max pooling to extract differential entropy features. avg F avg :

[0150] F avg =AvgPool(X)

[0151] F avg =MaxPool(X)

[0152] Where X represents the differential entropy feature of the original data.

[0153] After fusing the two shared 1×1 convolutional layers, we get:

[0154] R avg =W2(ReLU(W1F) avg ))

[0155] R max =W2(ReLU(W1F) max ))

[0156] W1 and W2 are both compression matrices, and ReLU is a widely used activation function.

[0157] Frequency band attention weights R are generated using the Sigmoid activation function. f :

[0158] R f =σ(R) avg +R max )

[0159] Where σ represents the Sigmoid activation function, and the output is the sum of the residual of the original input X and the frequency band attention weighting result to realize the importance modeling of the input frequency band.

[0160] Similar to the frequency band attention modeling described above, the channel attention weights R are then concatenated. c This enables the DAttention module.

[0161] Step S42: Construct an expert network with sparse gating for heterogeneous features;

[0162] The above Z att Divide into G subgroups along the channel dimension:

[0163]

[0164] For each subgroup Z g Through its corresponding gating network G g (·) Generate expert selection weights:

[0165]

[0166] Select the top-k experts based on their weights and perform sparse weighted routing to construct a sparse mask M. g ∈{0,1} B×E If only the top-k values ​​are retained, then the weight of the e-th expert is:

[0167]

[0168] According to the gating weight W g The weighted summation yields the merged output:

[0169]

[0170] Among them, Expert e Let e ​​represent the e-th expert model.

[0171] Finally, the outputs of all subgroups are concatenated and normalized to obtain the output of this module:

[0172]

[0173] Z moe =LayerNorm(Z) moe )

[0174] Step S43: Classification Module Design and Loss Construction

[0175] The design of the classification module and the main classification loss Same as step S36 above. Characterize Z. moe After the classification module, the prediction matrix Z is obtained. subspace Based on this, and to avoid over-activation by some experts, load balancing loss is defined as follows:

[0176]

[0177] Finally, the model jointly optimizes the main classification loss and the expert balance regularization term to construct the total loss function:

[0178]

[0179] Step S5: Fusion of classification results based on adaptive weights;

[0180] To achieve effective fusion of classification results, a learnable fusion coefficient is introduced, which can dynamically adjust the weight distribution of the two classification results, thereby adaptively outputting a comprehensive discrimination result. This strategy can automatically update the fusion parameters based on data feedback during training, effectively improving the overall classification accuracy and generalization ability of the system.

[0181] Z fused =α·Z route +(1-α)·Z subspace

[0182] The learnable weight parameter α is restricted to the range [0, 1].

[0183] α = clamp(α, 0.0, 1.0)

[0184] During the testing phase, we input unlabeled samples U into the trained model to obtain their corresponding predicted categories or classification probabilities, which are used to evaluate the model's generalization ability in the target domain.

[0185] The proposed speech imagery decoding method, which integrates dynamic brain region routing and frequency-spatial semantics, was tested on an EEG dataset (10 participants) and its performance was compared with recently proposed methods within a cross-participant recognition framework. The proposed method significantly improves the classification accuracy of the speech imagery task. The obtained recognition accuracy is shown in Table 1 below:

[0186] Table 1. Classification accuracy of the verbal imagination dataset

[0187]

Claims

1. A speech imagery decoding method integrating dynamic brain region routing and frequency-spatial semantics, characterized in that, Includes the following steps: Step S1: Collect and preprocess raw EEG data to generate EEG samples; Step S2: Extract features from the EEG sample to generate differential entropy features; Step S3: Construct a brain region perception labeling routing module, including: Brain region segmentation units are used to segment functional brain regions containing multiple channels; The region class label generation unit generates class labels including brain region class labels and global class labels based on the functional brain regions; A region-aware attention unit is used to control the interaction between the class labels and output class label feature representations by constructing a multi-head attention mechanism; a gating expert network unit receives the class label feature representations and outputs fused features through a sparse gating mechanism. The brain region perception and classification unit outputs a category prediction based on the fused features; Step S4: Construct the frequency-space semantic feature module, including: The frequency-space attention feature extraction unit takes the differential entropy feature as input and outputs the fused frequency-space feature. The heterogeneous feature expert selection unit divides the fused spectral-spatial features into multiple groups along the channel dimension and selects weights, then sends them to the sparsely gated hybrid expert module for modeling and outputs the spliced ​​features. The frequency-space semantic classification unit outputs the classification result based on the concatenated features; Step S5: Perform a weighted fusion of the category prediction and the classification result, and output the final fusion output.

2. The speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics according to claim 1, characterized in that, In the region class label generation unit, the class label is initialized; the global class label, the brain region class label, and the channel label are concatenated to form the Transformer input sequence.

3. The speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics according to claim 1, characterized in that, In the region-aware attention unit, the interaction between the class tags is controlled by constructing an attention mask matrix; the global class tag can interact bidirectionally with all the class tags; the brain region class tag can interact bidirectionally with its corresponding channel tag; and the channel tags can interact bidirectionally with each other. All class tags are accessible to themselves.

4. The speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics according to claim 1, characterized in that, The gated expert network unit inputs the class label into the shared gated network, calculates its selection weight at each expert, and selects the top k experts activated for each class label using a Top-k strategy; stacks the feature representations output by each expert to generate a representation vector; and fuses the representation vector with the gated weights to generate a weighted fusion feature.

5. The speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics according to claim 4, characterized in that, In the brain region perception classification unit, the weighted fusion features are concatenated along the feature dimension to obtain a fusion feature representation; the fusion feature representation is classified and predicted by a two-layer perceptron classifier, each layer of the perceptron classifier including layer normalization and nonlinear transformation; the category prediction matrix is ​​calculated and output at the output layer.

6. The speech imagery decoding method integrating dynamic brain region routing and frequency-spatial semantics according to claim 5, characterized in that, The category prediction matrix is ​​input into the cross-entropy loss function to obtain the optimized main classification loss; the expert average weight distribution is obtained according to the expert selection weight; the balance of the expert average weight distribution is measured using the entropy function, and a normalized balance loss is constructed. Construct a total loss function based on the normalized balance loss and the optimized main classification loss.

7. The speech imagery decoding method integrating dynamic brain region routing and frequency-space semantics according to claim 1, characterized in that, In step S5, the final fusion output is: WITH fused =α·Z route +(1-α)·Z subspace Where α is the adaptive weight, which is automatically updated after each training round, and Z... route This represents the category prediction Z after passing through the brain region perception label routing module. subspace This represents the classification result after passing through the frequency-space semantic feature module.