A three-dimensional convolution method for imagined language electroencephalography topography decoding

By constructing a frequency-band specific rectangular BEAM map using a three-dimensional convolution method and designing a dual-branch network, the problem of insufficient decoding accuracy in traditional methods is solved, achieving efficient decoding of the spatial structure of EEG signals and improving classification accuracy.

CN120995220BActive Publication Date: 2025-12-23CHANGCHUN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511476192.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-12-23
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing technologies for EEG decoding of imaginary language suffer from problems related to signal level and feature representation complexity. In particular, the geometric mismatch and single-dimensional feature extraction of the traditional BEAM method lead to a decrease in decoding accuracy, making it difficult to effectively mine the spatial structural information of EEG signals.

Method used

A three-dimensional convolutional method is adopted, and a rectangular BEAM graph is constructed by frequency band-specific kriging interpolation. Combining three-dimensional continuous spatial-frequency band and spatial-temporal representations, a two-branch three-dimensional convolutional network is designed. Attention weights are used for feature fusion decoding to improve decoding accuracy.

Benefits of technology

It improves the ability to recognize language-based imagery EEG patterns and significantly enhances classification accuracy, especially in decoding multi-band and multi-time-step EEG signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995220B_ABST
    Figure CN120995220B_ABST
Patent Text Reader

Abstract

A three-dimensional convolution method for imagined language electroencephalography (EEG) topography decoding. It solves the problem of the deficiency of the existing convolutional neural network method in modeling ability for multi-rhythm EEG topography, which leads to a significant decline in decoding accuracy in imagined language. In the aspect of rectangular BEAM reconstruction, a frequency-specific adaptive variogram kriging interpolation model is built, which automatically selects the optimal variogram function according to the spatial variation characteristics of the EEG power of each frequency band, and constructs a high-resolution BEAM map sequence under multi-frequency and multi-time steps, taking into account the spatial continuity and the adaptability of the CNN structure. A double-branch three-dimensional convolutional neural network is designed to extract high-dimensional features from spatial-time and spatial-frequency tensors, respectively, to explore the multi-scale structure and dynamic expression ability of neural activity. The discriminative ability of the two feature vector branches is dynamically weighted to improve the recognition ability of the classifier for the EEG pattern of language imagination.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of imagined language decoding in electroencephalogram (EEG) topography, and particularly relates to a three-dimensional convolution method for imagined language EEG topography decoding. BACKGROUND

[0002] In the field of brain-computer interface (BCI) research, decoding imagined language as a natural and intuitive interaction method has received extensive attention in recent years. Compared with motor imagery, imagined language is closer to human actual expression intention, has higher information capacity and practical potential. Due to this more natural expression characteristic, imagined language decoding is considered as a key direction to break through the efficiency bottleneck of traditional BCI interaction, but it also faces more complex technical problems.

[0003] The challenges of imagined language decoding based on electroencephalogram (EEG) signals mainly exist in the signal level and feature expression. The EEG signals generated in the imagined language process usually have the characteristics of low signal-to-noise ratio, high nonlinearity and non-stationarity, making it extremely difficult to strip out stable language-related features from the EEG data mixed with physiological noise. In addition, there are significant differences in language strategies and brain activation regions among different individuals. When the neural response mode difference between individuals exceeds the adaptation range of the model, a unified decoding framework is difficult to work, making it more difficult to build a universal model.

[0004] To cope with the above challenges, researchers initially focused on the one-dimensional time dynamic characteristics of EEG signals, using recurrent neural networks, variant long short-term memory networks, gated recurrent units and other models to learn the time-dependent relationship in the EEG signals, and to classify the time series of the EEG signals. However, the design focus of these models is concentrated in the time dimension, and there are limitations in capturing the spatial topological information of the EEG signals, making it difficult to fully exploit the potential spatial structure information in the EEG signals.

[0005] To fully extract the spatial features of the EEG signal, researchers have begun to turn to the modeling of two-dimensional spatial features, among which the brain electrical activity mapping (BEAM) has become a core tool for spatial information analysis because it can intuitively present the spatial distribution of brain activation. Researchers usually generate event-related potential topography by using the nearest neighbor, linear or spline interpolation method with the help of the topoplot function of EEGLAB, or convert EEG signals into two-dimensional image inputs by the EEG2Image technology. Some studies process EEG signals from a multidimensional perspective, such as generating task-specific topography according to energy differences, generating average, standard deviation and peak distribution maps based on the power spectral density of EEG signals, extracting frequency band power using wavelet packet transform and mapping it to the head shape space, and constructing frequency domain BEAM maps.

[0006] Although the existing methods have made up for the lack of spatial information capture to some extent, the inherent defects of traditional BEAM still restrict the decoding performance: (1) The traditional BEAM is mostly circular in structure, which is geometrically mismatched with the rectangular input format of the convolutional neural network, leading to distortion of edge information in the convolution operation and reducing the utilization efficiency of spatial features; (2) Studies on imagined language decoding mainly use one-dimensional or two-dimensional convolutional networks, focusing on the single-dimensional extraction of one-dimensional time series signals or two-dimensional EEG spatial features (circular EEG topography), and lacking in-depth mining of the sub-space representation of multi-band EEG time-space-frequency features. SUMMARY

[0007] To solve the problem of the lack of modeling ability of the existing convolutional network method for multi-rhythm EEG topography, which leads to a significant decrease in decoding accuracy in imagined language, the present application provides a three-dimensional convolution method for imagined language EEG topography decoding, which comprises the following steps:

[0008] S1, collecting a real imagined language EEG signal dataset using electrodes, including an imagined category;

[0009] S2, constructing a frequency-specific rectangular BEAM sequence by using a frequency-specific Kriging interpolation method on the EEG signal, mapping the original electrode spatial layout to a fixed-size two-dimensional rectangular grid, and obtaining the BEAM graph of the EEG signal;

[0010] S3, stacking the BEAM graphs of multiple frequency bands at the same time point of the EEG signal by using a three-dimensional continuous space-frequency representation method;

[0011] S4, stacking the BEAM graphs of multiple time points at the same frequency band of the EEG signal by using a three-dimensional continuous space-time representation method;

[0012] S5. A dual-branch 3D imaginary language encoding method is adopted, integrating the discriminative features from the perspectives of steps S3 and S4. An attention-weighted feature fusion decoding method is used to achieve adaptive weighted fusion of the importance of features from different modalities. Then, a 2-layer MLP network is used to map the fused features to... Given a classification result, obtain the probability corresponding to each classification result.

[0013] Furthermore, in step S2, when constructing a frequency-specific rectangular BEAM sequence using the frequency-specific Kriging interpolation method on the EEG signal, the spatial structure characteristics of different frequency bands are set as linear, Gaussian, exponential, or spherical variation models. A leave-one-out cross-validation strategy is adopted, and the mean square error is selected as the index. The variation model with the smallest mean square error is selected for full-image interpolation.

[0014] Furthermore, in step S2, when mapping the original electrode spatial layout to a fixed-size two-dimensional rectangular mesh, an edge compensation strategy is proposed to extract the radius of the furthest point from the electrode. edge area , Represents a point in a two-dimensional rectangular grid. The width of the edge region, for all Calculate the minimum distance from the point to the edge point. Find the nearest edge value Applying the exponential decay interpolation formula To achieve a natural transition from a circular beam diagram to a rectangular beam diagram, the formula is as follows: The power value is the power value at the nearest boundary point. The normalized distance from the external point to the nearest point on the boundary. This represents the power estimate for the region outside the circular mask in the rectangular BEAM diagram.

[0015] Furthermore, step S3 specifically involves: processing the EEG signals corresponding to each electrode. Power spectral density was extracted from six typical frequency bands using a bandpass filter, yielding the characteristic representation of each channel in different frequency bands. , Number of frequency bands;

[0016] The six typical frequency bands are, in order: delta: 1–4 Hz, theta: 4–8 Hz, alpha: 8–13 Hz, beta: 13–30 Hz, low-gamma: 30–55 Hz, and high-gamma: 55–100 Hz, respectively. , , , , and correspond;

[0017] Then for each typical frequency band Frequency-specific rectangular EEG topographic maps are constructed using the interpolation method in step S2. , It is the size of the spatially interpolated image. After stacking all the frequency band images, a spatial-frequency joint tensor is formed. : .

[0018] Furthermore, step S4 specifically involves: continuously The electrode amplitude at each time point within each time step is interpolated in step S2 to generate a rectangular EEG topography map of consecutive frames. This indicates the electric field distribution at this point in time. Indicates the number of electrodes, including all The frames are stacked in chronological order to form a space-time joint tensor. , .

[0019] Furthermore, in step S5, the dual-branch 3D imaginative language encoding method is performed through a spatial-temporal domain branch network and a spatial-frequency domain branch network, and the attention-weighted feature fusion decoding method is performed through an attention-weighted feature fusion classification network.

[0020] Furthermore, in the space-time domain branching network, the input is a five-dimensional tensor. Five-dimensional tensor Yes The tensor obtained after expanding the batch and channel dimensions. For batch size, For channel dimensions;

[0021] First, a layer-by-layer 3D convolution operation is performed:

[0022] ;

[0023] Get output ,in, Number of output channels Input the number of channels. The kernel size is the convolution kernel size. Indicates input, This represents layer-by-layer 3D convolution. Represents the convolution kernel. Indicates offset, subscript Indicates the first 3D convolution;

[0024] Then batch normalization and activation, obtaining , wherein is an activation function, is a three-dimensional batch normalization operation;

[0025] inputting into a pooling layer, an adaptive average pooling layer and a flattening and fully connected layer in sequence, outputting a feature vector .

[0026] Further, in the spatial-band domain branch network, the input is a five-dimensional tensor , the five-dimensional tensor is a tensor obtained after expanding the batch dimension and the channel dimension of ;

[0027] First, a layer-by-layer 3D convolution operation is performed:

[0028] ;

[0029] obtaining an output , wherein is the number of output channels, is the number of input channels, is the size of the convolution kernel, represents the input, represents the layer-by-layer 3D convolution, represents the convolution kernel, represents the bias, and the subscript represents the layer 3D convolution;

[0030] Then, batch normalization and activation are performed on , obtaining , ;

[0031] inputting into a pooling layer, an adaptive average pooling layer and a flattening and fully connected layer in sequence, outputting a feature vector .

[0032] Further, in the feature fusion classification network of the attention weight, two feature vectors are connected , represents the dimension of , represents the dimension of , and is input into a standard 2-layer MLP network to calculate the fusion weight of the two branches , wherein is a weight coefficient parameter, represents the number of hidden layer neurons of the MLP network, express Dimensions Weights are assigned to attention, and the final fused features are Finally, a two-layer MLP network is used to fuse the features. Mapping to imagined category classification results This yields the probability corresponding to each category of imagination.

[0033] The beneficial effects of the method described in this invention are as follows:

[0034] (1) In terms of rectangular BEAM reconstruction, a frequency band-specific adaptive variogram Kriging interpolation model was built. The optimal variogram function was automatically selected based on the spatial variation characteristics of EEG power in each frequency band, and a high-resolution BEAM sequence under multiple frequency bands and multiple time steps was constructed, taking into account both spatial continuity and CNN structural adaptability.

[0035] (2) In terms of continuous EEG time-frequency-space structure optimization, based on bandpass filter and power spectrum estimation, the spatial distribution features of multiple typical frequency bands are extracted. By stacking the frequency band maps, a three-dimensional space-frequency band tensor is formed to model the spectral band information and its spatial projection rules in the language imagination process. A space-time BEAM map sequence is constructed, and the BEAM maps corresponding to the continuous time frames in the language imagination trials are stacked into a three-dimensional tensor to capture the spatiotemporal dynamic evolution mode of the electric field.

[0036] (3) A dual-branch three-dimensional convolutional neural network was designed to extract high-dimensional features from the space-time and space-frequency tensors respectively, explore the multi-scale structure and dynamic expression ability of neural activity, dynamically weight the discriminativeness of the two feature vector branches, and improve the classifier's ability to recognize language imagination EEG patterns. Attached Figure Description

[0037] Figure 1 This is a flowchart of the method described in an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of the three-dimensional continuous space-frequency band characterization method and the three-dimensional continuous space-time characterization method in the embodiments of the present invention;

[0039] Figure 3 This is a schematic diagram of the feature fusion encoding and decoding method for attention weights in an embodiment of the present invention. Detailed Implementation

[0040] The technical solutions of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0041] Embodiment 1,

[0042] The present embodiment provides a three-dimensional convolution method for imagined language electroencephalogram topographic decoding, and the flow of the method is as shown in Figure 1 The method comprises the following steps:

[0043] S1, collecting a real imagined language electroencephalogram signal dataset, including four kinds of imagined upper, lower, left and right;

[0044] S2, using a frequency band-specific Kriging interpolation method for the electroencephalogram signal, constructing a frequency band-specific rectangular BEAM sequence, mapping the original electrode spatial layout into a fixed-size two-dimensional rectangular grid to obtain a BEAM graph of the electroencephalogram signal;

[0045] S3, using a three-dimensional continuous space-frequency band representation method, stacking the BEAM graphs of multiple frequency bands at the same time point of the electroencephalogram signal;

[0046] S4, using a three-dimensional continuous space-time representation method, stacking the BEAM graphs of continuous multiple time points of the same frequency band of the electroencephalogram signal;

[0047] S5, using a double-branch three-dimensional imagined language coding method, integrating the discriminative features under two perspectives of steps S3 and S4, using an attention weight feature fusion decoding method to realize adaptive weighted fusion of the importance of different modal features, and then using a 2-layer MLP network to map the fused features to a 4-classification result.

[0048] Embodiment 2,

[0049] The present embodiment is a further limitation of embodiment 1, and further describes step S2.

[0050] The present embodiment constructs a frequency band-specific adaptive Kriging interpolation based on which the original electrode spatial layout is mapped into a fixed-size two-dimensional rectangular grid to adapt the input structure of the electroencephalogram signal data in subsequent space-time and space-frequency band modeling.

[0051] Band-specific Kriging interpolation method, according to the spatial structure characteristics of different frequency bands, set to linear, Gaussian, exponential, spherical variation model. Adopt the strategy of Leave-One-Out Cross Validation (LOOCV), select the mean square error (MSE) as the index, select the variation model with the smallest MSE for the whole map interpolation.

[0052] First, the original electroencephalogram signal is segmented , divided into frequency points in logarithmic intervals, and frequency bands are segmented, and the Morlet wave period on each frequency band is used to balance the frequency band and time resolution.

[0053] The 6 typical frequency bands are delta: 1-4 Hz, theta: 4-8 Hz, alpha: 8-13 Hz, beta: 13-30 Hz, low-gamma: 30-55 Hz and high-gamma: 55-100 Hz.

[0054] For the signal of the i-th channel , the power spectrum of the j-th frequency band at the k-th time step is extracted , is the Morlet wave with the center frequency , is the convolution operation. Channel refers to the physical lead channel (number of electrodes) during electroencephalogram signal acquisition.

[0055] Then, the two-dimensional coordinates of each electrode in the standard scalp electrode distribution are extracted

[0056] , and a two-dimensional grid is constructed within the minimum circumscribed matrix of the electrodes (the minimum circumscribed matrix of all electrodes as a whole), and a circular mask is constructed with the maximum radius of the electrodes to exclude the external area of the scalp.

[0057] For each time step and frequency band , the average power of all channels is calculated to obtain the stable spatial distribution characteristics :

[0058] ;

[0059] ​​Considering the spatial patterns of EEG activity in different frequency bands are band-dependent, the embodiment adopts an optimal selection mechanism based on variogram to model each frequency band and estimate all channels in each frequency band The interpolation result :

[0060] ; The number of EEG channels participating in the interpolation calculation, the of each of the n channels is multiplied by the corresponding weight , and then summed to obtain the interpolation result .

[0061] The weight calculated according to the variogram model , is set according to the spatial structure characteristics of different frequency bands is a linear, Gaussian, exponential, or spherical variogram model. Using the leave-one-out cross-validation strategy, the mean squared error is selected as the indicator, and the variogram model with the smallest MSE is selected for full-image interpolation.

[0062] The size of the rectangular grid is set to 500x500, and the above frequency-specific Kriging interpolation method is applied to each point in the grid to obtain a spatially continuous power estimation map, i.e., a circular BEAM map.

[0063] Since the grid is rectangular and the electrode distribution is usually approximately circular, there are missing areas outside the grid. To improve image continuity and network adaptability, an edge compensation strategy is proposed for the spatially continuous power estimation map. The edge region with a radius of from the farthest point to the electrode , is the width of the edge region. For all points, calculate the minimum distance to the edge point, find the nearest edge value , and apply the exponential decay interpolation formula to realize a natural transition from the circular BEAM map to the rectangular BEAM map, where is the power value of the nearest boundary point, is the normalized distance from the external point to the nearest boundary point, represents the power estimation value of the external region of the circular mask in the rectangular BEAM map. is not directly represented as the entire rectangular BEAM map, but as the power estimation value of the external region of the circular mask in the rectangular BEAM map (i.e., the rectangular region not covered by the original circular BEAM map). It is a supplement to the circular BEAM map, and ultimately forms a complete rectangular BEAM map together with the power values of the circular region.

[0064] Example 3

[0065] This embodiment further defines embodiment 1 and provides further explanation of steps S3-S4.

[0066] like Figure 2 As shown, for each EEG channel signal Power spectral density was extracted for six typical frequency bands (delta: 1–4 Hz, theta: 4–8 Hz, alpha: 8–13 Hz, beta: 13–30 Hz, low-gamma: 30–55 Hz, and high-gamma: 55–100 Hz) using bandpass filters, yielding the characteristic representation of each channel in different frequency bands. , This refers to the number of frequency bands.

[0067] Then for each frequency band Frequency-specific rectangular EEG topographic maps are constructed using the interpolation method in step S2. , It is the size of the spatially interpolated image. After stacking all the frequency band images, a spatial-frequency joint tensor is formed. : .

[0068] Figure 2 In this model, each channel of the tensor corresponds to the spatial distribution of a specific frequency band, which can be viewed as a spatial projection of brain neural activity under multiple frequencies during language imagination. This representation method effectively preserves the physiological meaning of different frequency bands and enhances the spatial specificity of local brain region activity through spatial interpolation.

[0069] like Figure 2 As shown, the continuous The electrode amplitude at each time point within each time step is interpolated in step S2 to generate a rectangular EEG topography map of consecutive frames. This represents the electric field distribution at this point in time. Stacking all t frames in chronological order forms a three-dimensional continuous space-time joint tensor. , .like Figure 2 As shown, each channel of this tensor corresponds to a spatial distribution at a point in time, representing the dynamic changes of EEG signals over a continuous period of time.

[0070] Example 4

[0071] This embodiment further defines embodiment 1 and provides further explanation of step S5.

[0072] like Figure 3As shown, the embodiment designs a double-branch 3D convolution model based on three-dimensional continuous space-time-frequency band features, extracts high-dimensional dynamic representations from the space-time domain and the space-frequency domain respectively, combines the time evolution mode and the frequency band characteristics of the electroencephalogram signal, and enhances the perception ability of the model to the differences in brain activity in the language conception process.

[0073] The model includes a space-time domain branch network, a space-frequency domain branch network, and a feature fusion classification network (decoding network) of attention weights.

[0074] (1) Space-time domain branch network: The network is used to extract three-dimensional features of the space-time EEG image (spatial structure + time sequence dimension), and map it to a 128-dimensional vector.

[0075] Input five-dimensional tensor , the five-dimensional tensor is obtained after the batch dimension and channel dimension of are expanded, is the batch size, 1 is the channel C dimension, and each pixel only contains an amplitude, is the number of time frames, is the spatial size of the rectangular BEAM image, and after the layer-by-layer 3D convolution operation:

[0076] ;

[0077] is obtained, wherein is the number of output channels, is the number of input channels, is the convolution kernel size 5x3x3 (capturing 3x3 local spatial features of 5 consecutive time steps).

[0078] Batch normalization and activation are performed on to obtain , wherein is an activation function, is three-dimensional batch normalization.

[0079] is sequentially input into the pooling layer (only spatial downsampling), the adaptive average pooling layer, and the flattening and fully connected layer, and the output feature vector is obtained,

[0080] ;

[0081] wherein represents the input feature, and ​​denote the features after the first and second pooling operations, respectively, , and denote the first, second, and third layer 3D max pooling, respectively.

[0082] ;

[0083] wherein, denotes an adaptive average pooling operation, denotes a flattening operation, denotes a full connection operation.

[0084] (2) Spatial-band domain branch network: This network is used to extract the three-dimensional features of the spatial-band EEG image (spatial structure + band sequence dimension), and finally map it to a 128-dimensional vector. The input shape is a five-dimensional tensor , the five-dimensional tensor is obtained after expanding the batch dimension and channel dimension of ;

[0085] After 3D convolution operation:

[0086] ;

[0087] wherein, is the number of output channels, is the number of input channels, is the size of the convolution kernel × 3 × 3 (capture 3 × 3 local spatial features of the interaction of adjacent bands).

[0088] Batch normalization and activation are performed on to obtain :

[0089] .

[0090] Input into the pooling layer (only spatial down-sampling), the adaptive average pooling layer, and the flattening and full connection layer, respectively, to output the feature vector . The operations in the pooling layer are:

[0091] ;

[0092] wherein, denotes the input feature, and denote the features after the first and second pooling operations, respectively, , and denotes the first layer three-dimensional maximum pooling, the second layer three-dimensional maximum pooling, and the third layer three-dimensional maximum pooling in sequence.

[0093] .

[0094] (3) Attention weight feature fusion classification network: In order to make full use of the complementarity of the two modalities and avoid the redundancy or information interference caused by simple splicing, an attention weight-based feature fusion module is designed to realize adaptive weighting of the importance of different modal features, thereby enhancing the expression ability of the final feature representation.

[0095] , denotes the dimension of , denotes the dimension of , and is input into a standard 2-layer MLP network to calculate the fusion weight of the two branches , wherein is a weight coefficient parameter, denotes the number of hidden layer neurons of the MLP network, denotes the dimension of , is the attention allocation weight, and the final fused feature is ; finally, a 2-layer MLP network is used to map the fused feature to the imagination class classification result , to obtain the probability corresponding to each imagination class.

[0096] Example 5,

[0097] This embodiment is a further limitation of Examples 1-4, and the beneficial effects of the method described in the application are illustrated by comparative experiments.

[0098] The method of the application is used to classify and predict 1200 samples (each sample is composed of 50 consecutive frames) of four types of directional imagination tasks (up, down, left, right) under three task conditions (actual pronunciation Pron, language imagination Inner and visual imagination Vis) of 10 healthy subjects, and the comparison is shown in Table 1 (This paper represents the method in the application). Compared with the electroencephalogram decoding methods KNN, SVM, XGBoost, LSTM, BiLSTM and EEGNet in recent years, the 3D CNN method proposed in the application achieves higher classification accuracy than the existing methods under the Pron and Inner two task modalities, with the accuracy reaching 62.67% and 61.79% respectively, which is significantly higher than the traditional KNN and SVM models (the highest is 59.55%), and is also significantly better than the time series modeling methods such as LSTM and BiLSTM widely used in recent years (the highest is 31.30%). Even in the most challenging visual imagination task, the model in this paper still achieves an accuracy of 53.25%, which is much higher than the 29.67% of EEGNet and the performance of other deep learning methods.

[0099] Among them, Work represents the source of the method, Classifier represents the classification method corresponding to the method, and Accuracy represents the accuracy of classification.

[0100] [1] represents the literature as follows: Lopez-Bernal D, Balderas D, Ponce P, Molina A. Exploring inter-trial coherence for inner speech classification in EEG-based brain-computer interface. J Neural Eng 26, 21 (2024).

[0101] [2] represents the literature as follows: Van den Berg B, Van Donkelaar S and Alimardani M 2021 Inner speech classification using EEG signals: A deep learning approach 2021 IEEE 2nd Int. Conf. on Human-Machine Systems (ICHMS) (IEEE) 1-4 (2021).

[0102] The document indicated by [3] is specifically: Gasparini F, Cazzaniga E and Saibene A. Inner speech recognition through electroencephalographic signals, (2022).

[0103] The document indicated by [4] is specifically: Merola, N.R., Venkataswamy, N.G., Imtiaz, M.H. Can Machine Learning Algorithms Classify Inner Speech from EEG Brain Signals? 2023 IEEE World AI IoT Congress (AIIoT), Seattle, WA, USA, 7–10 June 2023, 466–470 (2023).

[0104] The document indicated by [5] is specifically: Ng, H. W., & Guan, C. Efficient representation learning for inner speech domain generalization. In International conference on computer analysis of images and patterns, 131–141 (2023).

[0105] The document indicated by [6] is specifically: Ng H W, Guan C. Subject-independent meta-learning framework towards optimal training of EEG-based classifiers. Neural Networks 172, (2024).

[0106] Table 1 Comparison of four-classification accuracy of different methods on the same dataset:

[0107]

Claims

1. A three-dimensional convolution method for imagined language electroencephalography topography decoding, characterized in that, The method comprises the following steps: S1, collecting a real imagination language electroencephalogram signal dataset by using electrodes, including an imagination category; S2, a frequency band specific Kriging interpolation method is used for the electroencephalogram signal to construct a frequency band specific rectangular BEAM sequence, and the original electrode spatial layout is mapped into a fixed size two-dimensional rectangular grid to obtain a BEAM graph of the electroencephalogram signal; When the frequency band specific Kriging interpolation method is used for the electroencephalogram signal to construct the frequency band specific rectangular BEAM sequence, the linear, Gaussian, exponential or spherical variation model is set according to the spatial structure characteristics of different frequency bands, the leave-one-out cross-validation strategy is adopted, the average mean square error is selected as an index, and the variation model with the minimum average mean square error is selected for full graph interpolation; A kind of edge compensation strategy is proposed when mapping the original electrode spatial layout into a fixed size two-dimensional rectangular grid, and the edge region with a radius of from the farthest point of the electrode is extracted , representing the point in the two-dimensional rectangular grid, is the width of the edge region, and the minimum distance from the edge point is calculated for all points, and the nearest edge value is found An exponential decay interpolation formula is applied to realize the natural transition from the circular BEAM pattern to the rectangular BEAM pattern, where is the power value from the nearest boundary point, is the normalized distance from the external point to the nearest boundary point, and represents the power estimation value of the external region of the circular mask in the rectangular BEAM pattern. S3, a three-dimensional continuous space-frequency band representation method is used to stack the BEAM graphs of multiple frequency bands at the same time point of the electroencephalogram signal; S4, a three-dimensional continuous space-time representation method is used to stack the BEAM graphs of the same frequency band at continuous multiple time points of the electroencephalogram signal; S5, adopt a double-branch three-dimensional imagination language coding method, integrate the discriminant features under two perspectives of steps S3 and S4, adopt a feature fusion decoding method of attention weight, realize adaptive weighted fusion of the importance of different modal features, and then utilize a 2-layer MLP network to map the fused features to a classification result, and obtain a probability corresponding to each classification result.

2. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 1, characterized in that, Step S3 is specifically: extracting the power spectrum density of each electrode pair corresponding to the brain electrical signal The band-pass filter is applied to extract the power spectrum density of 6 typical frequency bands, and the feature representation of each channel under different frequency bands is obtained , is the number of frequency bands; The 6 typical frequency bands are delta: 1-4 Hz, theta: 4-8 Hz, alpha: 8-13 Hz, beta: 13-30 Hz, low-gamma: 30-55 Hz and high-gamma: 55-100 Hz, corresponding to , , , , and respectively; Then for each typical frequency band The rectangular scalp topography of each frequency band is constructed by the interpolation method in step S2 , is the size of the spatial interpolation image. After stacking all the frequency band images, a spatial-frequency joint tensor is constructed : .

3. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 2, characterized in that, Step S4 is specifically: interpolate the electrode amplitudes at each time point within the continuous frame of rectangular electroencephalographic topographic maps representing the electric field distribution at this time point, representing the number of electrodes, stack all frames in chronological order to form a space-time joint tensor , .

4. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 3, characterized in that, In step S5, the double-branch three-dimensional imagined language coding method is performed through the space-time domain branch network and the space-frequency domain branch network, and the attention weight feature fusion decoding method is performed through the attention weight feature fusion classification network.

5. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 4, characterized in that, In the space-time domain branch network, the input is a five-dimensional tensor , the five-dimensional tensor is obtained after the batch dimension and the channel dimension of the tensor are expanded, is the batch size, is the channel dimension; firstly, a layer-by-layer 3D convolution operation is performed: ; obtaining an output wherein, is the number of output channels, is the number of input channels, is the size of the convolution kernel, denotes the input, denotes a layer-wise 3D convolution, denotes the convolution kernel, denotes the bias, the lower index denotes the layer 3D convolution; Then the batch normalization and activation, obtaining , wherein is an activation function, is a three-dimensional batch normalization operation; The input is sequentially inputted into a pooling layer, an adaptive average pooling layer, and a flattening and fully connected layer, and a feature vector is outputted .​ 6. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 5, characterized in that, In the space-band domain branch network, the input is a five-dimensional tensor , the five-dimensional tensor is obtained after the batch dimension and the channel dimension of the tensor are expanded ; firstly, a layer-by-layer 3D convolution operation is performed: ; obtaining an output wherein, is the number of output channels, is the number of input channels, is the size of the convolution kernel, denotes the input, denotes a layer-wise 3D convolution, denotes the convolution kernel, denotes the bias, the lower index denotes the layer 3D convolution; Then the batch normalization and activation, obtaining , ; will be described below. The feature vector is output by sequentially inputting a pooling layer, an adaptive average pooling layer, and a flattening and fully connected layer .

7. The three-dimensional convolution method for imagined language electroencephalogram topographic map decoding according to claim 6, characterized in that, The two feature vectors are connected in the attention weight feature fusion classification network , represents the dimension of , represents the dimension of , input into a standard 2-layer MLP network, and the fusion weight of the two branches is calculated , wherein is a weight coefficient parameter, represents the number of hidden layer neurons of the MLP network, represents the dimension of , is an attention distribution weight, and the final fused feature is ; and finally, a 2-layer MLP network is used to map the fused feature to an imagination category classification result , to obtain the probability corresponding to each imagination category.

Citation Information

Patent Citations

  • A convolutional neural network motor imagery electroencephalogram recognition method based on a time-frequency domain

    CN109711383A

  • Motor imagery electroencephalogram decoding method based on multi-band dual-stage feature extraction network

    CN117743942A