Brain electrolysis interpretation method of Transformer based on multi-granularity combination
Through the multi-grained Transformer method, the problem of time information loss and insufficient classification accuracy in electroencephalogram interpretation is solved, and efficient decoding of EEG signals is achieved, especially in diseases such as epilepsy and sleep disorders to improve the decoding accuracy.
Patent Information
- Application Number
- CN202510393705.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Existing electrocerebral interpretation methods have problems such as loss of time information and insufficient classification accuracy, especially in diseases such as epilepsy and sleep disorders, where machine learning decoding accuracy is limited.
The Transformer method based on multi-grained size combination is adopted to generate multi-level information by obtaining differential entropy features, spatial feature extraction, shallow temporal feature extraction, fine-grained convolution modules and self-attention and guide attention optimization, and combine SA and GA networks to achieve accurate decoding of brain activity.
It significantly improves the decoding accuracy of EEG signals, fully extracts spatial and temporal features, solves the problem of insufficient utilization of spatial features, and ensures the accuracy and efficiency of decoding.
Smart Images

Figure CN120336805A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electroencephalogram signal interpretation, and specifically relates to an electroencephalogram interpretation method based on a multi-granularity combined Transformer. Background Art
[0002] Using deep learning algorithms to decode physiological, psychological or pathological states from electroencephalogram signals plays an important role. The supervised learning method represented by deep learning can improve the accuracy of electroencephalogram signal decoding to a certain extent with its powerful feature learning ability, and this improvement in accuracy is particularly important in diseases such as epilepsy and sleep disorders.
[0003] However, electroencephalogram signals have disadvantages such as low signal-to-noise ratio and low spatial resolution, which limit the accuracy of machine learning decoding and cause many difficulties in practical applications. Summary of the Invention
[0004] Aiming at the problems of time information loss and insufficient classification accuracy in the existing electroencephalogram interpretation methods, the present invention provides an electroencephalogram interpretation method based on a multi-granularity combined Transformer.
[0005] To this end, the present invention provides an electroencephalogram interpretation method based on a multi-granularity combined Transformer, which is characterized by including the following steps:
[0006] Step 1: Obtain the differential entropy feature x from the original electroencephalogram signal e ;
[0007] Step 2: Use a spatial feature extraction Transformer to obtain the spatial features of the electroencephalogram region;
[0008] Step 3: Use a shallow time feature extraction module to extract the time features of the electroencephalogram signal;
[0009] Step 4: Use a fine-grained convolution module to extract fine-grained time features;
[0010] Step 5: Use self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the emotion label corresponding to the electroencephalogram activity.
[0011] Furthermore, in the above Step 1: The specific process of obtaining the differential entropy feature from the original electroencephalogram signal is: divide the original electroencephalogram signal into N non-overlapping segments with a length of several seconds, X = [X1,..., XN], and divide the electroencephalogram signal into different frequency bands through a filter, including Delta (1–4 Hz), Theta (4–8 Hz), Alpha (8–14 Hz), Beta (14–30 Hz) and Gamma (30–45 Hz).
[0012] DE = ∫s(x i ) log(s(x i )) dx i (1)
[0013] where s(x i ) is the probability density function of x i , and DE represents the features obtained for each time window.
[0014] Furthermore, the specific process of step 2: using the spatial feature extraction Transformer to obtain the spatial features of the EEG region is as follows:
[0015] First, convert the differential entropy feature x e into queries Q, keys K, and values V; the process of querying and the spatial operations on keys K and values V can be expressed as:
[0016] Q = LN(Linear(X e )) (2)
[0017] SP(K, V) = LN(DWConv(X e )) (3)
[0018] where LN refers to layer normalization, DWConv represents depth convolution operation, and SP represents the operation on the spatial dimension of the input sequence;
[0019] Secondly, the calculation of the attention operation is:
[0020]
[0021] where is an additional scaling factor;
[0022] The processing process of the single-head attention mechanism SHA is as follows:
[0023] SHA(Q, K, V) = Attention(QW Q , SP(K)W K , SP(V)W V ) (5)
[0024] W Q , W K , W V are linear projection parameters, and Q, SP(K), SP(V) are the query q, key k, and value V of Attention(·); the regional spatial learning transformer structure that provides the embedded feature X0 can be expressed as follows:
[0025]
[0026] where Xl Denote the output feature of the l-th layer as \(X\). l-1 Denote the output feature of the \((l - 1)\)-th layer as \(X\). \(LN\) represents layer normalization, \(SHA\) represents single-head attention mechanism, and \(MLP\) represents multi-layer perceptron. L Denote the output feature after being processed through all \(L\) layers as \(R\). s Denote the final encoder output feature.
[0027] Furthermore, the specific process of step 3: using the shallow temporal feature extraction module to extract the temporal features of the EEG signal is as follows:
[0028] After the EEG signal is input, through operations such as convolution, BN, and ELU, coarse-grained temporal features are obtained, and then the dimensions of the feature map are rearranged. The size of the temporal CNN kernel is \((1, 0.1f)\), where \(f\) represents the sampling rate of the EEG. After activation by the ELU function, max-pooling is performed on every two data points of the learned features. The shallow feature extraction module can be expressed as:
[0029] \(F=\Gamma(\text{MaxPool}(\text{ELU}(\text{BN}(\text{CNN}(X)))))+P\ (7)\)
[0030] where \(\Gamma\) is the rearrangement operation, \(\text{MaxPool}\) is max-pooling, \(\text{ELU}\) is the activation function, \(\text{BN}\) is the batch normalization layer, and \(P\) represents a learnable unknown encoding. The positional encoding can help the model distinguish signals at different time points, thus better capturing temporal dependencies. Its size is \(c*0.5^l\), where \(c\) is the number of channels and \(l\) is the number of upsampled data points on each channel.
[0031] Furthermore, the specific process of step 4: using the fine-grained convolution module to extract fine-grained temporal features is as follows:
[0032] The EEG signal first passes through a one-dimensional convolutional layer, which uses a small convolutional kernel to slide in the time dimension to capture fine-grained temporal information. After the dropout layer, the learned representation is fed into a one-dimensional CNN layer, followed by a batch normalization layer, an ELU activation layer, and a max-pooling layer. The fine-grained temporal information extraction process can be expressed as follows;
[0033]
[0034] where, Denote the fine-grained temporal feature of the \(i\)-th layer as \(F\). \(DP\) represents the Dropout layer, \(CNN\) represents convolution, \(\text{ELU}\) represents the neural network activation function, \(\text{MaxPool}\) represents max-pooling. iDenote the features input into the fine-grained convolution module. F1 is the temporal feature extracted by the Temporal Conv in the temporal feature module; F2 is the temporal feature obtained by the EEG after passing through the Temporal Conv and the first FFT module; F3 is the temporal feature obtained by the EEG after passing through the Temporal Conv and the first two FFT modules; F4 is the temporal feature obtained by the EEG after passing through the Temporal Conv and the first three FFT modules.
[0035] Furthermore, in step 5: the specific process of using self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the learnable dynamic weight γ is as follows:
[0036] Use self-attention (SA) and guided attention (GA) to optimize the generated multi-level information; these two features are superimposed to obtain the mixed information, and then a linear transformation is performed to obtain the learnable dynamic weight γ. The process can be expressed as:
[0037]
[0038] where, v g and v l respectively represent the global and local features after information interaction. and respectively represent the global and local features after being processed by the self-attention (SA) and guided attention (GA) modules. denotes element-wise addition, that is, adding corresponding elements, and σ represents the activation function sigmoid function.
[0039]
[0040] γ1, γ2 = softmax(σ(v’W α )W β ) (11)
[0041]
[0042] where, and represent the global and local features after feature interaction. W α , W β are learnable weights. v', γ1, γ2 all represent intermediate variables. The update of the W α , W β matrices is achieved through an optimization algorithm. Softmax ensures that the sum of the two weights γ1 and γ2 is 1, and v represents the fusion output of the spatial feature and the multi-granularity temporal feature.
[0043] The advantages of the present invention are as follows: The present invention provides a brain electrical signal decoding method based on a multi-granularity combined Transformer. Compared with traditional brain electrical signal decoding algorithms, this method has made significant progress in extracting fine-grained features and utilizing the spatial features of electrodes. By adding spatial feature extraction and temporal feature extraction operations, the algorithm efficiency is improved, effectively solving problems such as insufficient utilization of spatial features and ensuring the decoding accuracy. In addition, this method uses a spatial feature extraction Transformer to obtain regional spatial features of 5 brain regions, which can effectively obtain the spatial features of brain electrical signals, and this is crucial for the effect of brain electrical signal decoding. Moreover, the repeated use of the temporal convolution module enables the model to fully extract deep temporal features while maintaining computational efficiency and improving decoding accuracy. This method combines the SA and GA networks, as well as the interaction steps of spatial feature and temporal feature extraction. By combining the extracted coarse-grained spatial features and fine-grained temporal features, it is possible to generate masks for local features to filter global features, and at the same time enable global features to directly supplement local features, so that temporal features and spatial features are fully obtained and utilized.
[0044] The following will make a detailed description of the present invention in conjunction with the drawings and embodiments. Description of the Drawings
[0045] Figure 1 It is a flowchart of a brain electrical signal decoding algorithm based on a multi-granularity combined Transformer.
[0046] Figure 2 It is a structural diagram of a spatial feature extraction Transformer model.
[0047] Figure 3 It is a structural diagram of a shallow temporal feature extraction module.
[0048] Figure 4 It is a structural diagram of a fine-grained convolution.
[0049] Figure 5 It is a structural diagram of a temporal and spatial feature combination module.
[0050] Figure 6 It is a structural diagram of a self-attention module.
[0051] Figure 7 It is a structural diagram of a guided attention module.
[0052] Figure 8 It is a schematic diagram of a brain electrical signal. Specific Embodiments
[0053] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the following will make a detailed description of the specific embodiments, structural features and their effects of the present invention in conjunction with the drawings and embodiments.
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] All the features disclosed in this specification, or all the steps in any method or process disclosed, except for mutually exclusive features and / or steps, may be combined in any manner.
[0056] Any feature disclosed in this specification (including any additional claims, abstract, and drawings), unless specifically stated, may be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically stated, each feature is only an example of a series of equivalent or similar features.
[0057] Embodiment 1
[0058] To solve the problems existing in the existing electroencephalogram (EEG) decoding methods, such as loss of time information and insufficient classification accuracy.
[0059] This embodiment provides an EEG decoding method based on a multi-granularity combined Transformer as Figures 1 to 7 shown, including the following steps:
[0060] Step 1: Obtain the differential entropy feature x from the original EEG signal e ;
[0061] Step 2: Use the spatial feature extraction Transformer to obtain the spatial features of the EEG region;
[0062] Step 3: Use the shallow time feature extraction module to extract the time features of the EEG signal;
[0063] Step 4: Use the fine-grained convolution module to extract the fine-grained time features;
[0064] Step 5: Use self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the emotion labels corresponding to the EEG activities.
[0065] Further, the specific process of step 1: obtaining differential entropy features from the original EEG signals is as follows: The original EEG signals are divided into N non-overlapping segments with a length of several seconds, X = [X1,..., XN]. The EEG signals are divided into different frequency bands through a filter, including Delta (1–4 Hz), Theta (4–8 Hz), Alpha (8–14 Hz), Beta (14–30 Hz), and Gamma (30–45 Hz).
[0066] DE = ∫s(x i )log(s(x i ))dx i (1)
[0067] where s(x i ) is the probability density function of x i , x i represents the signal segment of the original EEG signal in the i-th frequency band, DE represents the features obtained for each time window, and the number of features obtained for each time window is d.
[0068] The differential entropy feature matrix x e has a size of c*d, where c is the number of EEG channels and d is the number of features obtained for each window.
[0069] Further, the specific process of step 2: using the spatial feature extraction Transformer to obtain the spatial features of the EEG regions is as follows:
[0070] The EEG signals obtained from the electrode patches in the brain regions are input in parallel. The electrode patches near the brain are roughly divided into five regions. After processing by this structure, the encoder obtains the regional spatial features of 5 brain regions. The specific steps are as follows: First, the differential entropy feature x e is converted into query Q, key K, and value V; the process of query and the spatial operations of key K and value V can be expressed as:
[0071] Q = LN(Linear(X e )) (2)
[0072] SP(K, V) = LN(DWConv(X e )) (3)
[0073] where LN refers to layer normalization, DWConv represents depth convolution operation, and SP represents the operation on the spatial dimension of the input sequence;
[0074] Secondly, the calculation of the attention operation is:
[0075]
[0076] where, The item is an additional scaling factor;
[0077] The processing process of the single-head attention mechanism SHA is as follows:
[0078] SHA(Q, K, V) = Attention(QW Q , SP(K)W K , SP(V)W V ) (5)
[0079] W Q , W K , W V W, W, W are linear projection parameters, and Q, SP(K), SP(V) are the query q, key k, and value V of Attention(·); the regional spatial learning transformer structure that provides the embedded feature X0 can be expressed as follows:
[0080]
[0081] Among them, X l represents the output feature of the l-th layer, X l-1 represents the output feature of the (l - 1)-th layer, LN represents layer normalization, SHA represents the single-head attention mechanism, MLP represents the multi-layer perceptron, X L represents the output feature after being processed by all L layers, and R s represents the final encoder output feature.
[0082] Furthermore, the specific process of step 3: using the shallow time feature extraction module to extract the time features of the EEG signal is as follows:
[0083] After the EEG signal is input, it undergoes operations such as convolution, BN, and ELU to obtain coarse-grained time features, and then the dimensions of the feature map are rearranged; among them, the size of the time CNN kernel is (1, 0.1f), where f represents the sampling rate of the EEG; after activation by the ELU function, max pooling is performed on every two data points of the learned features; the shallow feature extraction module can be expressed as:
[0084] F = Γ(MaxPool(ELU(BN(CNN(X)))))+P (7)
[0085] Among them, Γ is the rearrangement operation, MaxPool is max pooling, ELU is the activation function, BN is the batch normalization layer, P represents the learnable unknown encoding, and the positional encoding can help the model distinguish signals at different time points, thereby better capturing the time dependence, with a size of c * 0.5l, where c is the number of channels and l is the number of sampled data points on each channel.
[0086] Further, the specific process of step 4: using the fine-grained convolution module to extract fine-grained time features is as follows:
[0087] The EEG signal first passes through a one-dimensional convolutional layer, which uses a small convolutional kernel to slide in the time dimension to capture fine-grained time information; after the dropout layer, the learned representation is fed into a one-dimensional CNN layer, followed by a batch normalization layer, an ELU activation layer, and a max pooling layer; the fine-grained time information extraction process can be expressed as follows;
[0088]
[0089] where, represents the fine-grained time feature of the i-th layer, DP represents the Dropout layer, CNN represents convolution, ELU represents the neural network activation function, MaxPool represents max pooling, and F i represents the feature input to the fine-grained convolution module, F1 is the time feature extracted by the Temporal Conv of the time feature module; F2 is the time feature obtained by the EEG passing through the Temporal Conv and the first FFT module; F3 is the time feature obtained by the EEG passing through the Temporal Conv and the first two FFT modules; F4 is the time feature obtained by the EEG passing through the Temporal Conv and the first three FFT modules.
[0090] Further, the specific process of step 5: using self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the learnable dynamic weight γ is as follows:
[0091] The time features extracted by the fine-grained convolution module and the Transformer layer are combined to provide more comprehensive time and space features; self-attention (SA) and guided attention (GA) are used to optimize the generated multi-level information; these two features are stacked to obtain hybrid information, and then a linear transformation is performed to obtain the learnable dynamic weight γ, and the process can be expressed as:
[0092]
[0093] where, v g and v l respectively represent the global and local features after information interaction, and respectively represent the global and local features after being processed by the self-attention (SA) and guided attention (GA) modules, represents element-wise addition, that is, adding corresponding elements, and σ represents the sigmoid activation function.
[0094]
[0095] γ1, γ2 = soft max(σ(v’W α )W β ) (11)
[0096]
[0097] wherein, and represent the global and local features after feature interaction, W α , W β are learnable weights, v', γ1, γ2 all represent intermediate variables, and the updates of the matrices W α , W β are achieved through an optimization algorithm. Softmax ensures that the sum of the two weights γ1 and γ2 is 1, and v represents the fused output of spatial features and multi-granularity time features.
[0098] According to the value of the fused output v of spatial features and multi-granularity time features, the emotion labels corresponding to the electroencephalogram activities can be classified.
[0099] Emotion types: happy, fearful, surprised, sad, angry, and calm.
[0100] Principle of emotion label classification: It is judged according to the MLP classifier. The output of the input MLP classifier is a probability distribution, representing the probability that the sample belongs to each category. The fused output v of spatial features and multi-granularity time features is input into the MLP classifier, and according to the output probability, the category with the highest probability is selected as the final prediction result. For example, if the output is [0.1, 0.2, 0.1, 0.1, 0.1, 0.4], the emotion is judged to be calm.
[0101] In summary, compared with traditional electroencephalogram (EEG) decoding algorithms, the EEG decoding method based on multi-granularity combined Transformer provided in this embodiment has made significant progress in extracting fine-grained features and utilizing the spatial features of electrodes. By adding spatial feature extraction and temporal feature extraction operations, the algorithm efficiency is improved, the problem of insufficient utilization of spatial features is effectively solved, and the decoding accuracy is ensured. In addition, the method uses spatial feature extraction Transformer to obtain the regional spatial features of 5 brain regions, which can effectively obtain the spatial features of EEG signals, which is crucial for the effect of EEG decoding. In addition, the repeated use of the temporal convolution module enables the model to fully extract deep temporal features while maintaining computational efficiency and improving decoding accuracy. This method combines SA and GA networks, as well as the interaction steps of spatial feature and temporal feature extraction, combines the extracted coarse-grained spatial features and fine-grained temporal features, can generate a mask of local features to filter global features, and at the same time enables global features to directly supplement local features, so that temporal features and spatial features are fully obtained and utilized.
[0102] Embodiment 2
[0103] Use the EEG decoding method based on multi-granularity combined Transformer shown in Embodiment 1 to Figure 8 decode the shown EEG signals. EEG signals have different manifestations in different frequency bands at different times of emotions. When an individual is in negative emotions such as anxiety and depression, the power of theta waves in the frontal lobe region will increase. Alpha waves (8-14 Hz) are closely related to the relaxed and calm emotional state. When an individual enters a relaxed state from a tense state, the alpha wave activity in the occipital lobe region of the brain increases significantly. Beta waves (14-30 Hz) are related to alert and excited emotions. When facing a stress task or in a stress state, the power of beta waves in multiple regions of the brain increases significantly. In EEG emotion classification, different frequency bands of brain waves can be considered.
[0104] Obtain Figure 8The EEG signal data shown is divided into N non-overlapping segments of several seconds in length according to the method in step 1 of Embodiment 1. The original EEG signal is divided into different frequency bands of Delta (1–4Hz), Theta (4–8Hz), Alpha (8–14Hz), Beta (14–30Hz), and Gamma (30–45Hz) using a filter, and the differential entropy features of each frequency band are calculated according to the formula. The calculated differential entropy features are converted into query Q, key K, and value V in the manner of step 2 in Embodiment 1. Through relevant spatial operations and attention mechanism calculations, the spatial features of the EEG region are obtained using the spatial feature extraction Transformer, and this process involves operations such as layer normalization (LN) and depthwise convolution operation (DWConv). Another branch inputs the EEG signal into the shallow temporal feature extraction module, and according to step 3 in Embodiment 1, it sequentially undergoes operations such as convolution, batch normalization (BN), and exponential linear unit (ELU) activation to obtain coarse-grained temporal features, and then the dimensions of the feature map are rearranged. The size of the temporal CNN kernel is set to (1, 0.1f) (f is the EEG sampling rate). After ELU activation, max-pooling operations are performed on every two data points of the learned features. The fine-grained temporal feature extraction module enables the EEG signal to first pass through a one-dimensional convolutional layer, which uses a small convolutional kernel to slide in the temporal dimension to capture fine-grained temporal information. After passing through the dropout layer, the learned representation is input into the one-dimensional CNN layer, followed by batch normalization, ELU activation, and max-pooling operations to complete the extraction of fine-grained temporal features. The specific operations refer to step 4 in Embodiment 1. The fine-grained convolution module is combined with the temporal features extracted by the Transformer layer, and self-attention (SA) and guided attention (GA) are used to optimize the generated multi-level information. The SA and GA features are superimposed to obtain the hybrid information, and then a linear transformation is performed to obtain the learnable dynamic weights, as described in step 5 of Embodiment 1.
[0105] Through the Figure 8 interpretation practice of the EEG signal shown, the model after feature extraction is trained using the EEG signal dataset with emotion labels. Here, the cross-entropy loss function can be used to measure the difference between the model prediction result and the true label, and optimization algorithms such as stochastic gradient descent are used to continuously adjust the model parameters to gradually reduce the value of the loss function, so that the model can learn effective classification patterns. For example, the dataset is divided into a training set, a validation set, and a test set. The model is trained on the training set, and the model hyperparameters are adjusted on the validation set to avoid overfitting. Finally, the model performance is evaluated on the test set. After training is completed, the Figure 6When the EEG signal is input into the trained model, the model will output the predicted emotion category. Common emotion categories include happy, sad, angry, calm, etc. Based on the output result of the model, the emotion state corresponding to the EEG signal can be judged. It verifies again the feasibility and effectiveness of the EEG decoding method based on the multi-granularity combined Transformer when dealing with actual EEG signals. It can fully extract the spatial and temporal features of EEG signals, effectively solve problems such as insufficient utilization of spatial features, and improve the decoding accuracy.
[0106] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for brain electrophysiological interpretation based on a multi-granularity combined Transformer, characterized in that It includes the following steps: Step 1: Obtain the differential entropy feature x from the original EEG signal e ; Step 2: Use the spatial feature extraction Transformer to obtain the spatial features of the EEG region; Step 3: Use the shallow temporal feature extraction module to extract the temporal features of the EEG signal; Step 4: Use the fine-grained convolution module to extract fine-grained temporal features; Step 5: Use self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the emotion label corresponding to the EEG activity.
2. The method for decoding brain electricity based on a multi-granularity combined Transformer according to claim 1, wherein: The specific process of obtaining differential entropy features from the original EEG signal in Step 1 is as follows: The original EEG signal is divided into N non-overlapping segments of several seconds in length, X = [X1,..., XN]. The EEG signal is divided into different frequency bands through a filter, including Delta (1–4 Hz), Theta (4–8 Hz), Alpha (8–14 Hz), Beta (14–30 Hz), and Gamma (30–45 Hz). DE = ∫s(x i ) log(s(x i )) dx i (1) where s(x i ) is the probability density function of x i , and DE represents the features obtained for each time window.
3. A method for decoding brain electricity based on a multi-granularity combined Transformer according to claim 1, characterized in that: The specific process of using the spatial feature extraction Transformer to obtain the spatial features of the EEG region in Step 2 is as follows: First, convert the differential entropy feature x e into query Q, key K, and value V; the process of querying and the spatial operations on key K and value V can be expressed as: Q = LN(Linear(X e )) (2) SP(K, V) = LN(DWConv(X e )) (3) Among them, LN refers to layer normalization, DWConv represents depth convolution operation, and SP represents the operation on the spatial dimension of the input sequence; Secondly, the calculation of the attention operation is: Among them, The item is an additional scaling factor; The processing process of the single-head attention mechanism SHA is as follows: SHA(Q, K, V) = Attention(QW Q , SP(K)W K , SP(V)W V ) (5) W Q , W K , W V is a linear projection parameter, Q, SP(K), and SP(V) are the query q, key k, and value V of Attention(·); the regional spatial learning transformer structure that provides the embedded feature X0 can be expressed as follows: Among them, X l represents the output feature of the l-th layer, X l-1 represents the output feature of the (l-1)-th layer, LN represents layer normalization, SHA represents single-head attention mechanism, MLP represents multi-layer perceptron, X L represents the output feature after being processed through all L layers, R s represents the final encoder output feature.
4. The brain electrointerpretation method based on a multi-granularity combined Transformer as claimed in claim 1, wherein: The specific process of using the shallow temporal feature extraction module to extract the temporal features of the EEG signal in Step 3 is as follows: After the EEG signal is input, it undergoes operations such as convolution, BN, and ELU to obtain coarse-grained temporal features, and then the dimensions of the feature map are rearranged; the size of the temporal CNN kernel is (1, 0.1f), where f represents the sampling rate of the EEG; after activation by the ELU function, max pooling is performed on every two data points of the learned features; the shallow feature extraction module can be expressed as: F = Γ(MaxPool(ELU(BN(CNN(X)))))+P (7) Among them, Γ is the rearrangement operation, MaxPool is max pooling, ELU is the activation function, BN is the batch normalization layer, and P represents a learnable unknown encoding. The positional encoding can help the model distinguish signals at different time points, thereby better capturing temporal dependencies. Its size is c*0.5l, where c is the number of channels and l is the number of sampled data points on each channel.
5. The brain electrointerpretation method based on a multi-granularity combined Transformer according to claim 1, characterized in that: The specific process of using the fine-grained convolution module to extract fine-grained temporal features in Step 4 is as follows: The EEG signal first passes through a one-dimensional convolutional layer, which uses a small convolutional kernel to slide in the time dimension to capture fine-grained temporal information; after the dropout layer, the learned representation is fed into a one-dimensional CNN layer, followed by a batch normalization layer, an ELU activation layer, and a max pooling layer; the fine-grained temporal information extraction process can be expressed as follows; Among them, represents the fine-grained time feature of the i-th layer, DP represents the Dropout layer, CNN represents convolution, ELU represents the neural network activation function, MaxPool represents max pooling, F i represents the feature input into the fine-grained convolution module. F1 is the time feature extracted by the Temporal Conv of the time feature module; F2 is the time feature obtained by the EEG passing through the Temporal Conv and the first FFT module; F3 is the time feature obtained by the EEG passing through the Temporal Conv and the first two FFT modules; F4 is the time feature obtained by the EEG passing through the Temporal Conv and the first three FFT modules.
6. The method for electroencephalogram interpretation based on a multi-granularity combined Transformer according to claim 3, wherein: The specific process of using self-attention (SA) and guided attention (GA) to optimize the generated multi-level information to obtain the learnable dynamic weight γ in Step 5 is as follows: Self-attention (SA) and guided attention (GA) are used to optimize the generated multi-level information; these two features are superimposed to obtain hybrid information, and then a linear transformation is performed to obtain a learnable dynamic weight γ, and the process can be expressed as: Among them, v g and v l respectively represent the global and local features after information interaction, and respectively represent the global and local features after being processed by the self-attention (SA) and guided-attention (GA) modules, denotes element-wise addition, that is, adding corresponding elements, and σ represents the sigmoid activation function. γ1, γ2 = softmax(σ(v’W α )W β ) (11) Among them, and represent the global and local features after feature interaction. W α , W β are learnable weights, v', γ1, and γ2 all represent intermediate variables. W α , W β The update of the matrix is achieved through an optimization algorithm. Softmax ensures that the sum of the weights of γ1 and γ2 is 1. v represents the fusion output of the spatial features and the multi-granularity time features.