EEG classification 2D-MDAGTS model and construction method thereof
By proposing the 2D-MDAGTS model in the EEG classification, using multi-level information mining and feature dependency capture, the problems of universality and noise complexity of deep learning methods in PD diagnosis are solved, and higher diagnostic accuracy and interpretability are achieved.
Patent Information
- Application Number
- CN202311573706.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
Existing deep learning methods have problems such as insufficient universal verification, noise complexity and difficulty in extracting redundant information in PD diagnosis and classification, resulting in insufficient diagnostic accuracy and interpretability.
A 2D-MDAGTS model of EEG classification is proposed, including a first-time separable convolution module, a dual attention network and a gated loop unit. Through multi-level information mining, feature dependency capture and long-term dependency capture, more accurate feature information is extracted.
It improves the accuracy and interpretability of PD diagnosis, enhances the universality and adaptability of the model, and has high classification performance and clinical application potential.
Smart Images

Figure CN120030398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an EEG classification model, in particular to a 2D-MDAGTS model of EEG classification and a construction method thereof. Background Art
[0002] Parkinson's disease (PD) is a neurodegenerative disease with more than 10 million PD patients worldwide. Symptoms can be improved through medication, lifestyle adjustments and surgery. Although PD itself is not fatal, the complications of the disease can be serious. The Centers for Disease Control and Prevention (CDC) ranks complications of PD as the 14th leading cause of death in the United States.
[0003] As a non-invasive detection technology, electroencephalogram (EEG) has the advantages of high temporal resolution, low cost and fast computing speed, which will help to understand the pathogenesis of PD patients and assist doctors in early diagnosis. Due to the complexity and non-stationarity of EEG signals. In recent years, deep learning algorithms have received more and more attention and applications in the field of EEG signal processing. By constructing multi-layer neural networks, deep learning algorithms can learn important features from large-scale EEG data and classify them. They can effectively process the complexity and high dimensionality of EEG signals and improve the accuracy and reliability of disease detection.
[0004] Current deep learning methods are more accurate in distinguishing PD patients from healthy subjects (HC). However, there are still some limitations in deep learning-based PD diagnosis and classification methods. First, the universality of deep learning methods needs to be further verified, and their high classification accuracy in PD diagnosis and classification tasks needs to be verified on multiple datasets. Second, EEG signals have complex noise characteristics, and it is difficult to extract important information from them. Therefore, it is necessary to select suitable feature extraction methods to extract PD-related features to better reveal the potential information in EEG signals. In the PD classification problem, it is still necessary to seek the best performing model to improve diagnostic accuracy and interpretability. Therefore, although deep learning methods have made progress in PD diagnosis and classification, further research is still needed to overcome these limitations and improve their potential for assisting doctors in clinical practice. Summary of the invention
[0005] One of the objectives of the present invention is to provide a 2D-MDAGTS model for EEG classification to establish an effective EEG classification model.
[0006] The second purpose of the present invention is to provide a method for constructing a 2D-MDAGTS model for EEG classification to solve the problem that complex noise and redundant information in EEG signals make it difficult to extract important information.
[0007] One of the purposes of the present invention is achieved in that:
[0008] A 2D-MDAGTS model for EEG classification, including:
[0009] The first time-separable convolution module is used to mine multi-level information of the model input;
[0010] A second time-separable convolution module, connected to the first time-separable convolution module, for mining multi-level information of a feature map output by the first time-separable convolution module;
[0011] A third time-separable convolution module, connected to the second time-separable convolution module, for mining multi-level information of a feature map output by the second time-separable convolution module;
[0012] A dual attention network, connected to the temporal separable convolutional module, is used to capture the feature dependencies of the feature maps output by the third temporal separable convolutional module in both spatial and channel dimensions; and
[0013] The gated recurrent unit, connected to the temporal separable convolution module, is used to capture the long-term dependencies of the feature maps output by the third temporal separable convolution module.
[0014] Furthermore, it also includes:
[0015] The fully connected layer is used to forward propagate the feature maps output by the gated recurrent unit and the dual attention network after flattening and concatenating them according to the channel dimension, and then classify them using the Sigmoid activation function.
[0016] Furthermore, the number of neurons input to the fully connected layer is 7680, and the number of neurons output is 2.
[0017] Furthermore, it also includes:
[0018] The first convolutional layer is used to increase the number of channels of the model input;
[0019] a second convolution layer, connected to the first time-separable convolution module, and configured to increase the number of channels of the feature map output by the first time-separable convolution module; and
[0020] The third convolution layer is connected to the second time-separable convolution module, and is used to increase the number of channels of the feature map output by the second time-separable convolution module.
[0021] Furthermore, the input of each time-separable convolution module is divided into two branches, one of which passes through a kernel of size size1×5, step size is 1×1, padding is (0, d×(kernel size -1)) after the dilated causal convolution layer, and then through batch normalization, nonlinear activation by ReLU activation function, and discarding of the dropout unit, and through a convolution kernel kernel size The size is 1×5, the stride is 1×1, and the padding is (0, kernel size / 2) after a separable convolutional layer, and then undergoes batch normalization, activation by the ELU function, and discarding by the discard unit; the other branch undergoes shortcut convolution; the feature maps output by the two branches are nonlinearly activated using the ReLU activation function through the feature maps after residual connection.
[0022] Furthermore, the expansion coefficient d of the first time-separable convolution module is 1, the convolution kernel size is 1×5, and the number of convolution kernels is 32; the expansion coefficient d of the second time-separable convolution module is 2, the convolution kernel size is 1×5, and the number of convolution kernels is 16; the expansion coefficient d of the third time-separable convolution module is 4, the convolution kernel size is 1×5, and the number of convolution kernels is 32.
[0023] Furthermore, the dual attention network includes:
[0024] A position attention module is used to capture the feature dependencies in the spatial dimension of the feature map that has passed through the time-separable convolution module, obtain the spatial dimension feature map and output it; and
[0025] The channel attention module is used to capture the feature dependencies in the channel dimension of the feature map that has passed the time-separable convolution module, obtain the channel dimension feature map and output it;
[0026] The spatial dimension feature map output by the position attention module and the channel dimension feature map output by the channel attention module are concatenated to obtain a combined feature map. After the fourth convolutional layer performs dimensionality reduction processing on the combined feature map, the output of the dual attention network is obtained.
[0027] The second object of the present invention is achieved in this way:
[0028] A method for constructing a 2D-MDAGTS model for EEG classification comprises the following steps:
[0029] S1. Preprocessing: Perform denoising on the original EEG signal, including filtering, artifact removal and wavelet threshold denoising; divide the denoised signal into 2s-epochs of equal length to obtain samples; use a random division method to divide 90% of the samples into training sets and 10% of the samples into test sets. The samples in the training set are used for model training, and the samples in the test set are used for model performance testing;
[0030] S2. Model construction: construct 2D-MDAGTS model;
[0031] S3. Feature extraction: Calculate the fuzzy entropy features of each channel of the preprocessed samples with a scale factor of τ = 1, 2, and perform normalization processing to finally obtain feature vectors of two scales;
[0032] S4. Fusion and splicing: The feature vectors with scale factors τ = 1 and τ = 2 are fused and spliced to obtain the input of the model;
[0033] S5. Model training: input the training set samples into the 2D-MDAGTS model for training, and input the test set samples into the 2D-MDAGTS model for testing. The model training and testing are repeated 200 times, and the parameters of the model are determined after training.
[0034] S6. Model evaluation: Accuracy, precision, recall, F1-score and Kappa evaluation indicators are used to evaluate the performance of the model.
[0035] Furthermore, step S2 includes the following sub-steps:
[0036] S2-1. Each sample after preprocessing is a time series u(i) containing N sampling points, where 1≤i≤N.
[0037] S2-2 selects appropriate embedding dimension and time delay according to the time series u(i) to reconstruct the phase space and obtain the phase space phasor for:
[0038]
[0039]
[0040] Among them, u 0 (i) is the reference value, i=1,2,...,N-m+1, m is the embedding dimension, Represents m consecutive u values, subtracting the reference value u from the i-th point 0 (i) to distinguish differences;
[0041] S2-3. There are two vectors with equal elements With X j m , then we can obtain With X j m Chebyshev distance between Right now With X j mThe maximum absolute difference between the corresponding scalar components for:
[0042]
[0043] Among them, X j m ={u(j),u(j+1),...,u(j+m-1)}-u 0 (i);
[0044] S2-4. Based on the fuzzy membership function Two vectors and X j m The similarity between for:
[0045]
[0046] Where n is the fuzzy membership function The fuzzy index of r is the fuzzy membership function Similarity tolerance threshold of ;
[0047] S2-5. Based on Determine the function φ m (n,r),φ m (n,r) is:
[0048] S2-6. Vector With X j m The embedding dimension of is increased by 1, and we get With X j m+1 ,calculate With X j m+1 Chebyshev distance Right now With X j m+1 The maximum absolute difference between the corresponding scalar components for:
[0049]
[0050] According to the fuzzy membership function Two vectors and X j m+1 The similarity between for:
[0051]
[0052] based on Get φ m+1 (n,r),φ m+1 (n,r) is:
[0053] S2-7. Based on φ m (n,r) and φ m+1 (n,r) determines the fuzzy entropy function of the time series FuzzyEn(m,n,r) as: FuzzyEn(m,n,r)=lim N→∞ [lnφ m (n,r)-lnφ m+1 (n,r)];
[0054] S2-8. Since the time series is EEG data of finite length, the fuzzy entropy function is expressed as:
[0055] FuzzyEn(data,m,n,r)=lnφ m (n,r)-lnφ m+1 (n,r)
[0056] Among them, data is the input data;
[0057] S2-9. Downsample the original time series by the scale factor τ. The mathematical form of the calculation method is as follows:
[0058]
[0059] Among them, y δ (s) is the time series after downsampling, is the process of downsampling the original time series using the scale factor τ, i is any original sampling point, N is the total number of original sampling points, s is any sampling point after downsampling, and M is the total number of sampling points after downsampling;
[0060] S2-10. Thus, the multi-scale fuzzy entropy function FuzzyEn(data,m,n,r,τ) is obtained;
[0061] S2-11. Calculate the fuzzy entropy features of each channel for scale factors τ = 1 and τ = 2, and perform normalization to obtain the feature vector n τ ,in, g is the number of EEG electrodes, is the real number space.
[0062] Furthermore, in step S3, the feature vectors with scale factors τ=1 and τ=2 are fused and concatenated to obtain a two-dimensional feature matrix n in , and for n in Adjust and finally get the model input xin ,in,
[0063] In feature extraction, the present invention manually extracts features from two fuzzy entropy scales respectively and fuses and splices them. Such a choice can maximize the use of useful features in EEG signals and avoid introducing information that is regarded as noise or redundant. In model construction, the time-separable convolution module is first used to mine multi-level information in the data, and the gated recurrent unit module and the dual attention network are used in parallel to capture feature dependencies from multiple scales, so as to further extract and integrate feature information and make the extracted feature information more accurate.
[0064] The present invention evaluates the universality and effectiveness of the present invention by comparing the performance of the 2D-MDAGTS model on different data sets, verifies the adaptability of the 2D-MDAGTS model in different application scenarios, and helps to determine the application potential of the present invention in clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 Flowchart of the method for building a 2D-MDAGTS model for EEG classification.
[0066] Figure 2 Flowchart of EEG preprocessing, feature extraction, feature selection, feature fusion, and model classification.
[0067] Figure 3 It is a structural diagram of the 2D-MDAGTS model of the present invention.
[0068] Figure 4 Figure 2 is the structural diagram of the temporal separable convolution module.
[0069] Figure 5 Flowchart of EEG preprocessing and feature extraction.
[0070] Figure 6 Confusion matrix of HC and drug-naive PD patients for the 2D-MDAGTS model on the San Diego dataset.
[0071] Figure 7 Confusion matrix of HC and PD patients taking medication for the 2D-MDAGTS model on the San Diego dataset.
[0072] Figure 8 Is the accuracy of different models on the San Diego dataset. DETAILED DESCRIPTION
[0073] The present invention is further described below.
[0074] like Figure 1 and Figure 2 As shown, the 2D-MDAGTS model for extracting features and the construction method thereof provided by the present invention specifically include the following steps:
[0075] S1. Preprocessing: The original EEG signal is denoised by filtering, artifact removal and wavelet threshold denoising. Considering the limitation of sample number, the method of increasing the number of samples is used for data enhancement. The denoised signal is divided into 2s-epochs of equal length to obtain samples. 90% of the samples are randomly divided into training sets for model training and 10% of the samples are divided into test sets for model performance testing.
[0076] The samples in the embodiment of the present invention adopt the public San Diego data set, and the samples are labeled, the HC subject label is "0", and the PD patient label is "1".
[0077] S2. Model construction: construct a 2D-MDAGTS model.
[0078] like Figure 3 As shown in the figure, the 2D-MDAGTS model includes a dual attention network (DANet), a gated recurrent unit (GRU) and three temporal separable convolution modules (TSCN), and the three temporal separable convolution modules are: the first temporal separable convolution module (TSCN Block1), the second temporal separable convolution module (TSCN Block2) and the third temporal separable convolution module (TSCN Block3).
[0079] The input feature map of the 2D-MDAGTS model passes through the first convolution layer with a convolution kernel size of 1×1. This convolution layer increases the dimension according to the number of input channels to obtain the output feature map Out 1 :
[0080] Out 1 =ReLU(BatchNorm(conv 1×1 (x)))
[0081] The 2D-MDAGTS model also includes a second convolutional layer and a third convolutional layer. The feature map output by the first convolutional layer passes through a first time-separable convolutional module connected to the first convolutional layer, a second convolutional layer, a second time-separable convolutional module connected to the second convolutional layer, a third convolutional layer, and a third time-separable convolutional module connected to the third convolutional layer in sequence. The second convolutional layer is used to increase the number of channels of the feature map output by the first time-separable convolutional module; the third convolutional layer is used to increase the number of channels of the feature map output by the second time-separable convolutional module.
[0082] like Figure 4 As shown in Figure 1, the structure of each time-separable convolution module is the same. Taking the first time-separable convolution module as an example, the input feature map will pass through two branches, one branch passes through a convolution kernel kernel. size The size is 1×5, the stride is 1×1, and the padding is (0, d×(kernel size -1)) of the dilated causal convolutional layer, the output feature map is Out 2 , and then pass through a convolution kernel size The size is 1×5, the stride is 1×1, and the padding (0, kernel size / 2) of the separable convolutional layer, the output feature is Out 3 The other branch passes through the shortcut convolution module, which ensures that the tensor scale between the input and output is the same, and the feature maps output by the two branches are nonlinearly activated through residual connection and ReLU activation function to obtain the output feature map Out of the first time separable convolution module. 4 ; Finally, after the second convolution layer with a convolution kernel size of 1×1, the channel dimension of the feature map is matched, and the output feature map is Out 5 .
[0083] Among them, Out 2 、Out 3 、Out 4 and Out 5 The expression is:
[0084] Out 2 =Dropout(ReLU(BatchNorm(conv dilated causal (Out 1 ))))
[0085] Out 3 = Dropout(ELU(BatchNorm(conv separable (Out 2 ))))
[0086] Out 4= ReLU(conv Short Cut (Out 1 ) + Out 3 )
[0087] Out 5 = conv 1×1 (Out 4 )
[0088] The vector output after passing through the first depthwise separable convolution module and the second convolutional layer is then passed through the second depthwise separable convolution module, the third convolutional layer, and the third depthwise separable convolution module. The second depthwise separable convolution module and the third depthwise separable convolution module process the input feature map in the same way as the first depthwise separable convolution module, and the feature maps Out 6 and Out 7 are obtained in sequence. Finally, the feature map Out 8 output by the second depthwise separable convolution module is obtained. After passing through the third convolutional layer with a kernel size of 1×1, Out 9 is obtained to match the channel dimension of the feature map. Finally, after passing through the third depthwise separable convolution module, Out 10 and Out 11 are obtained in sequence. Finally, the feature map Out 12 output by the third depthwise separable convolution module is obtained.
[0089] Among them, the expressions of Out 6 , Out 7 , Out 8 , Out 9 , Out 10 , Out 11 and Out 12 are as follows:
[0090] Out 6 = Dropout(ReLU(BatchNorm(conv dilated causal (Out 5 ))))
[0091] Out 7 = Dropout(ELU(BatchNorm(conv separable (Out 6 ))))
[0092] Out 8 = ReLU(conv Short Cut (Out 5 ) + Out 7 )
[0093] Out 9 =conv 1×1 (Out 8 )
[0094] Out 10 =Dropout(ReLU(BatchNorm(conv dilated causal (Out 9 ))))
[0095] Out 11 = Dropout(ELU(BatchNorm(conv separable (Out 10 ))))
[0096] Out 12 =ReLU(conv Short Cut (Out 11 )+Out 9 )
[0097] Among them, as shown in Table 1, the parameters used by the three time-separable convolution modules are: the expansion coefficient d of the first time-separable convolution module is 1, the convolution kernel size is 1×5, and the number of convolution kernels is 32; the expansion coefficient d of the second time-separable convolution module is 2, the convolution kernel size is 1×5, and the number of convolution kernels is 16; the expansion coefficient d of the third time-separable convolution module is 4, the convolution kernel size is 1×5, and the number of convolution kernels is 32. The expansion coefficient can improve the receptive field, which represents the perception range of neurons for input data. The time-separable convolution module is used to mine multi-level information in time and channel dimensions, that is, through the expansion coefficient, the receptive field of different levels of the convolution kernel is effectively expanded.
[0098] Input data Out 12 Reshape the data into the gated recurrent unit input Out 13 , use the gated recurrent unit to capture the long-term dependencies of time series data and obtain the output feature Out gru , where the size of the hidden layer output feature is 32 and the number of hidden layers is 1.
[0099] Among them, Out 13 and Out gru The expression is:
[0100] Out 13 =Reshape(Out 12 )
[0101] Out gru=GRU(Out 13 )
[0102] A dual attention network is used to capture feature dependencies in spatial and channel dimensions. The dual attention network includes a position attention module and a channel attention module. The position attention module captures the feature dependencies of the feature map output by the third time-separable convolution module in the spatial dimension, and obtains the spatial dimension feature map Out PAM The channel attention module captures the feature dependency of the feature map output by the temporal separable convolution module in the channel dimension, and obtains the channel dimension feature map Out CAM , the spatial dimension feature map Out PAM And channel dimension feature map Out CAM Perform channel splicing to obtain the combined feature map Out DANet ; Then the combined feature map passes through the fourth convolution layer with a convolution kernel size of 1×1, reducing the channel dimension of the feature map to 32, reducing the computational complexity, and finally obtaining the feature map Out output by the dual attention network 14 .
[0103] Among them, Out PAM 、Out CAM 、Out DANet and Out 14 The expression is:
[0104] Out PAM =PAM(Out 12 )
[0105] Out CAM =CAM(Out 12 )
[0106] Out DANet =concat(Out PAM ,Out CAM )
[0107] Out 14 =conv 1×1 (Out DANet )
[0108] The output feature maps of the gated recurrent unit and the dual attention network are flattened and spliced according to the channel dimension. Concat Passed as input to the fully connected layer for processing.
[0109] Out Concat =Concat(Flatten(Out gru ),Flatten(Out 14 ))
[0110] Out Concat After being connected through the fully connected layer, the classification is finally performed through the Sigmoid activation function. The number of input neurons in the fully connected layer is 7680, and the number of output neurons is 2.
[0111] Table 1: Hyperparameters of the 2D-MDAGTS model structure
[0112]
[0113]
[0114] After the 2D-MDAGTS model is built, it is necessary to extract features from the samples. Due to the complex noise in the original EEG signal, if it is directly input into the model, redundant information may be learned, resulting in model overfitting. Using a good feature extraction method to extract effective features is the key to avoiding this problem. Entropy-based feature extraction methods have been widely used to identify nonlinear features in EEG signals. In addition, many studies have verified the unique frequency domain characteristics of PD patients. Power spectral density (PSD) is a commonly used frequency domain feature extraction method that can be used to analyze the power distribution of EEG signals at different frequencies; fuzzy entropy (FuzzyEn) is an algorithm for measuring the complexity of time series.
[0115] As shown in Table 2, since the structures and hyperparameters of the MDAGTS model and the 2D-MDAGTS model are the same, the input features of the MDAGTS model are one-dimensional vectors, and the input features of the 2D-MDAGTS model are two-dimensional matrices, the classification performance indicators of the model under different feature extraction methods are determined by the MDAGTS model, and the classification performance indicators of the feature extraction method using fuzzy entropy are obtained based on the scale factor τ = 1. In the present invention, the classification performance indicators of the MDAGTS model under different feature extraction methods are verified by evaluating the indicators accuracy, precision, recall, F1-score and Kappa.
[0116]
[0117]
[0118]
[0119]
[0120] Among them, TP represents the number of correctly identified positive samples, FP represents the number of incorrectly identified positive samples, TN represents the number of correctly identified negative samples, and FN represents the number of incorrectly identified negative samples.
[0121]
[0122]
[0123] Among them, P e represents accidental consistency, a 1 is the number of real HC samples; a 2 is the number of true PD samples, b 1 is the number of HC predictions; b 2 The number predicted for PD.
[0124] OFF-PD refers to PD patients who have not taken medication for at least 12 hours, and ON-PD refers to PD patients who have taken medication within 12 hours.
[0125] Table 2: Classification performance indicators of MDAGTS model for different features
[0126]
[0127] As can be seen from Table 2, when comparing power spectral density with fuzzy entropy, the classification performance index of the MDAGTS model is better when fuzzy entropy is used as the feature extraction method, that is, fuzzy entropy is more superior in describing the EEG signal characteristics of PD patients, so fuzzy entropy is used as the feature extraction method.
[0128] When extracting fuzzy entropy features, a multi-scale fuzzy entropy comparison experiment is performed. As shown in Table 3, the fuzzy entropy features of each channel are calculated for scale factors τ = 1, 2, 3, 4, and 5, and normalized to obtain the feature vectors n of five scales. τ , Where g is the number of EEG electrodes. Since the input feature of the MDAGTS model is a one-dimensional vector, the classification performance of the MDAGTS model at different scales is determined through multi-scale fuzzy entropy comparison experiments.
[0129] Table 3: Classification performance indicators of the MDAGTS model with multi-scale fuzzy entropy features
[0130]
[0131]
[0132] As shown in Table 3, when the scale factor τ = 2, the classification performance of the model is the best. The accuracy of HC vs. OFF-PD is 98.35%, and the accuracy of HC vs. ON-PD is 98.68%. At the same time, when τ = 1 and τ = 3, the model also shows good classification performance. However, after τ = 2, the classification performance of the model decreases as the scale factor increases. Therefore, in the subsequent multi-scale fusion experiments, only the feature vectors with scale factors τ = 1, 2, and 3 are retained.
[0133] The fusion of multi-scale entropy maximizes the extraction of useful features in EEG signals and avoids the introduction of what may be considered as noise or redundant information. As shown in Table 4, the feature vectors with scale factors τ = 1, 2, and 3 are concatenated into a two-dimensional feature matrix. And adjust the feature matrix to have channel dimensions to get the input of the model Where c is the number of scale factor dimensions. Since the input features of the 2D-MDAGTS model are two-dimensional matrices, the fused and spliced feature matrices are sent to the 2D-MDAGTS model for training. Through multi-scale fuzzy entropy fusion feature comparison experiments, the classification performance of the 2D-MDAGTS model under different scale factor fusion strategies is determined.
[0134] Table 4: Classification performance indicators of 2D-MDAGTS model with multi-scale fuzzy entropy fusion features
[0135]
[0136] From Table 4, we can see that the feature fusion with scale factor τ = 1 and τ = 2 achieves better results on the 2D-MDAGTS model. It can be judged that features of different scales can capture information of different levels and levels in the data. When these features are combined or spliced together, the data can be more comprehensively represented and richer and more discriminative features can be extracted. The reason why the accuracy of other multi-scale feature fusion splicing is not improved may be due to the large overlap between the fused features or the inconsistency between the features, which causes the spliced features to lose some key information, thus affecting the performance of the model.
[0137] As shown in Table 5, after determining the scale factor of the 2D-MDAGTS model, since the 2D-MDAGTS model is obtained by combining the dual attention network (DANet), the gated recurrent unit (GRU) and three temporal separable convolution modules (TSCN), it is necessary to judge the performance of the 2D-MDAGTS model through ablation experiments, that is, to compare the 2D-MDAGTS model with the models obtained by combining the three modules in different ways. Combined with the feature vectors of the scale factors τ = 1 and τ = 2, by gradually removing specific components or modules in the model, the impact of these components on the model performance is verified. Ablation experiments can provide an in-depth understanding of the contribution of different parts of the model to its overall performance, which helps to optimize the model and improve its performance.
[0138] Table 5: Fuzzy entropy feature classification performance indicators of different models
[0139]
[0140] It can be seen from Table 5 that when the three modules are combined together (i.e., the 2D-MDAGTS model), the various performance indicators are higher, so the 2D-MDAGTS model has higher classification performance.
[0141] S3. Feature extraction: The fuzzy entropy features of the preprocessed EEG signals are calculated with scale factors τ=1 and τ=2 respectively, and normalized to finally obtain feature vectors of two scales.
[0142] Each sample after preprocessing is a time series u(i) containing N sampling points, where 1≤i≤N;
[0143] Then, according to the time series u(i), the appropriate embedding dimension and time delay are selected to reconstruct the phase space and obtain the phase space phasor. The calculation formula of the phase space vector is as follows:
[0144]
[0145]
[0146] Among them, u 0 (i) is the reference value, i=1,2,...,N-m+1, m is the embedding dimension, Represents m consecutive u values, subtract u from the i-th point 0 (i) to distinguish differences;
[0147] There exist two vectors whose elements are equal and Then we can get With X j m Chebyshev distance between Right now With X j m The maximum absolute difference between the corresponding scalar components for:
[0148]
[0149] In addition, according to the fuzzy membership function Two vectors and X j m The similarity between for:
[0150]
[0151] Where n is the fuzzy membership function The fuzzy index of r is the fuzzy membership function similarity tolerance threshold.
[0152] Afterwards based on Determine the function φ m (n,r),φ m (n,r) is:
[0153]
[0154] Vector With X j m The embedding dimension of is increased by 1, and we get With X j m+1 ,calculate With X j m+1 Chebyshev distance between Right now With X j m+1 The maximum absolute difference between the corresponding scalar components for:
[0155]
[0156] According to the fuzzy membership function Two vectors and X j m+1 The similarity between for:
[0157]
[0158] based on Get φ m+1 (n,r),φm+1 (n,r) is:
[0159] Finally, based on φ m (n,r) and φ m+1 (n,r) determines the fuzzy entropy function of the time series FuzzyEn(m,n,r) as: FuzzyEn(m,n,r)=lim N→∞ [lnφ m (n,r)-lnφ m+1 (n,r)];
[0160] Since the time series is EEG data of finite length, the fuzzy entropy function can be expressed as:
[0161] FuzzyEn(data,m,n,r)=lnφ m (n,r)-lnφ m+1 (n,r)
[0162] Among them, data is the input data.
[0163] The original time series is downsampled by the scale factor τ, that is, τ-1 points are skipped from the original time series u(i) and the next point is selected as the sampling point, thereby reducing the number of sampling points. The mathematical form of the calculation method is as follows:
[0164]
[0165] Among them, y δ (s) is the time series after downsampling, is the process of downsampling the original time series using the scale factor τ, i is any original sampling point, N is the total number of original sampling points, M is the total number of sampling points after downsampling, and s is any sampling point after downsampling.
[0166] Thus, the multi-scale fuzzy entropy function FuzzyEn(data,m,n,r,τ) is obtained.
[0167] like Figure 5 As shown in the figure, the fuzzy entropy features of each channel are calculated for the scale factors τ = 1 and τ = 2, and then normalized to obtain two feature vectors n τ .in, g is the number of EEG electrodes, n is s is the number of samples to be segmented.
[0168] S4, fusion and splicing: fuse and splice the feature vectors with scale factors τ=1 and τ=2 to obtain the input of the model.
[0169] The feature vectors with scale factors τ = 1 and τ = 2 are fused and concatenated into a two-dimensional feature matrix n in , and adjust the feature matrix so that the input feature matrix has channel dimensions, and finally get the input x of the model in .in,
[0170] S5. Model training: The training set samples are input into the 2D-MDAGTS model for training, and the test set samples are input into the 2D-MDAGTS model for testing. The model training and testing are repeated 200 times, and the parameters of the model are determined after training.
[0171] S6. Model evaluation: Evaluate the performance and universality of the model.
[0172] As shown in Table 6, in order to verify the performance and universality of the method proposed in the present invention, experiments are conducted on two other public datasets: the University of Iowa (Iowa) dataset and the University of New Mexico (UNM) dataset, after training and testing.
[0173] Table 6: Total number of samples in each category for each dataset
[0174] OFF-PD\ON-PD San Diego Iowa UNM HC vs. OFF-PD 3026 - 5672 HC vs. ON-PD 3022 2598 5783
[0175] The performance of the 2D-MDAGTS model was evaluated using the evaluation indicators accuracy, precision, recall, F1-score, and Kappa. The method in this embodiment uses the Pytorch framework. The model was trained using the Adam optimizer, with the learning rate set to 0.0009, the batch size set to 64, and the loss function set to CrossEntropy.
[0176] In order to verify the performance of the method proposed in the present invention, according to Figure 6 The accuracy, precision, recall, F1-score and Kappa of the 2D-MDAGTS model in identifying HC and unmedicated PD patients in Table 7 can be calculated. Figure 7 The accuracy, precision, recall, F1-score and Kappa of the 2D-MDAGTS model in identifying HC and PD patients on medication can be obtained in Table 7. When calculating the accuracy, a positive sample is taken from each of the two categories, and then the average is calculated, and the average is used as the accuracy. When calculating the recall rate, a positive sample is taken from each of the two categories, and then the average is calculated, and the average is used as the recall rate. When calculating the F1-score, a positive sample is taken from each of the two categories, and then the average is calculated, and the average is used as the F1-score.
[0177] Table 7: Performance on three public datasets
[0178]
[0179] According to Table 7, the accuracy, precision, recall and F1-score of the 2D-MDAGTS model on the three datasets are all around 98%; the accuracy, precision, recall and F1-score of the 2D-MDAGTS model in identifying HC and unmedicated PD patients on the San Diego dataset are all around 99%; the accuracy, precision, recall and F1-score on the Iowa dataset and UNM dataset are all around 99%.
[0180] like Figure 8 As shown in the figure, in order to further compare the performance of the 2D-MDAGTS model with other models, the accuracy of the 2D-MDAGTS model, CNN GRU model, ResNet CBAM model, EEGNet model and DenseNet model on the San Diego dataset is compared. The 2D-MDAGTS model has a significant improvement in accuracy.
Claims
1. A 2D-MDAGTS model for EEG classification, It is characterized in that include: The first time-separable convolution module is used to mine multi-level information of the model input; A second time-separable convolution module, connected to the first time-separable convolution module, for mining multi-level information of a feature map output by the first time-separable convolution module; A third time-separable convolution module, connected to the second time-separable convolution module, for mining multi-level information of a feature map output by the second time-separable convolution module; A dual attention network, connected to the third temporal separable convolutional module, is used to capture the feature dependencies of the feature maps output by the third temporal separable convolutional module in spatial and channel dimensions; as well as The gated recurrent unit is connected to the third time-separable convolution module and is used to capture the long-term dependencies of the feature maps output by the third time-separable convolution module.
2. The 2D-MDAGTS model for EEG classification according to claim 1, It is characterized in that Also includes: The fully connected layer is used to forward propagate the feature maps output by the gated recurrent unit and the dual attention network after flattening and concatenating them according to the channel dimension, and then classify them using the Sigmoid activation function.
3. The 2D-MDAGTS model for EEG classification according to claim 2, It is characterized in that The number of input neurons of the fully connected layer is 7680, and the number of output neurons is 2.
4. The 2D-MDAGTS model for EEG classification according to claim 1, It is characterized in that Also includes: The first convolutional layer is used to increase the number of channels of the model input; A second convolution layer, connected to the first time-separable convolution module, is used to increase the number of channels of the feature map output by the first time-separable convolution module; as well as The third convolution layer is connected to the second time-separable convolution module, and is used to increase the number of channels of the feature map output by the second time-separable convolution module.
5. The 2D-MDAGTS model for EEG classification according to claim 1, 2 or 3, It is characterized in that The input of each time-separable convolution module is divided into two branches, one of which passes through a kernel of size size 1×5, step size is 1×1, padding is (0, d×(kernel size -1)) after the dilated causal convolution layer, and then batch normalization, nonlinear activation by ReLU activation function, and discarding of the dropout unit, and then passing through a convolution kernel kernel size The size is 1×5, the stride is 1×1, and the padding is (0, kernel size / 2) after a separable convolutional layer, and then undergoes batch normalization, activation by the ELU function, and discarding by the discard unit; the other branch undergoes shortcut convolution; the feature maps output by the two branches are nonlinearly activated using the ReLU activation function through the feature maps after residual connection.
6. The 2D-MDAGTS model for EEG classification according to claim 1, 2 or 3, It is characterized in that The expansion coefficient d of the first time-separable convolution module is 1, the convolution kernel size is 1×5, and the number of convolution kernels is 32; the expansion coefficient d of the second time-separable convolution module is 2, the convolution kernel size is 1×5, and the number of convolution kernels is 16; the expansion coefficient d of the third time-separable convolution module is 4, the convolution kernel size is 1×5, and the number of convolution kernels is 32.
7. The 2D-MDAGTS model for EEG classification according to claim 1 or 2, It is characterized in that The dual attention network includes: A position attention module is used to capture the feature dependencies in the spatial dimension of the feature map that has passed through the time-separable convolution module, obtain the spatial dimension feature map and output it; and The channel attention module is used to capture the feature dependencies in the channel dimension of the feature map that has passed the time-separable convolution module, obtain the channel dimension feature map and output it; The spatial dimension feature map output by the position attention module and the channel dimension feature map output by the channel attention module are concatenated to obtain a combined feature map. After the fourth convolutional layer performs dimensionality reduction processing on the combined feature map, the output of the dual attention network is obtained.
8. A method for constructing a 2D-MDAGTS model for EEG classification. It is characterized in that The steps include: S1. Preprocessing: Perform denoising on the original EEG signal, including filtering, artifact removal and wavelet threshold denoising; divide the denoised signal into 2s-epochs of equal length to obtain samples; use a random division method to divide 90% of the samples into training sets and 10% of the samples into test sets. The samples in the training set are used for model training, and the samples in the test set are used for model performance testing; S2. Model construction: construct 2D-MDAGTS model; S3. Feature extraction: Calculate the fuzzy entropy features of each channel of the preprocessed samples with a scale factor of τ = 1, 2, and perform normalization processing to finally obtain feature vectors of two scales; S4. Fusion and splicing: The feature vectors with scale factors τ = 1 and τ = 2 are fused and spliced to obtain the input of the model; S5. Model training: input the training set samples into the 2D-MDAGTS model for training, and input the test set samples into the 2D-MDAGTS model for testing. The model training and testing are repeated 200 times, and the parameters of the model are determined after training. S6. Model evaluation: Accuracy, precision, recall, F1-score and Kappa evaluation indicators are used to evaluate the performance of the model.
9. The method for constructing a 2D-MDAGTS model for EEG classification according to claim 8, It is characterized in that Step S2 includes the following sub-steps: S2-1. Each sample after preprocessing is a time series u(i) containing N sampling points, where 1≤i≤N; S2-2. Select appropriate embedding dimension and time delay according to the time series u(i) to reconstruct the phase space and obtain the phase space phasor for: Among them, u 0 (i) is the reference value, i=1,2,...,N-m+1, m is the embedding dimension, Represents m consecutive u values, subtracting the reference value u from the i-th point 0 (i) to distinguish differences; S2-3. There are two vectors with equal elements and Then we can get and Chebyshev distance between for: (i,j=1,2,...,Nm,j≠i) in, S2-4. Based on the fuzzy membership function Two vectors and The similarity between for: Where n is the fuzzy membership function The fuzzy index of r is the fuzzy membership function Similarity tolerance threshold of ; S2-5. Based on Determine the function φ m (n,r),φ m (n,r) is: S2-6. Vector and The embedding dimension of is increased by 1, and we get and calculate and Chebyshev distance between for: According to the fuzzy membership function Two vectors and Similarity between for: based on Get φ m+1 (n,r),φ m+1 (n,r) is: S2-7. Based on φ m (n,r) and φ m+1 (n,r) determines the fuzzy entropy function of the time series FuzzyEn(m,n,r), FuzzyEn(m,n,r) is: FuzzyEn(m,n,r)=lim N→∞ [lnφ m (n,r)-lnφ m+1 (n,r)]; S2-8. Since the time series is EEG data of finite length, the fuzzy entropy function is expressed as: FuzzyEn(data,m,n,r)=lnφ m (n,r)-lnφ m+1 (n,r) Among them, data is the input data; S2-9. Downsample the original time series by the scale factor τ. The mathematical form of the calculation method is as follows: Among them, y δ (s) is the time series after downsampling, is the process of downsampling the original time series using the scale factor τ, i is any original sampling point, N is the total number of original sampling points, s is any sampling point after downsampling, and M is the total number of sampling points after downsampling; S2-10. Thus, the multi-scale fuzzy entropy function FuzzyEn(data,m,n,r,τ) is obtained; S2-11. Calculate the fuzzy entropy features of each channel for scale factors τ = 1 and τ = 2, and perform normalization to obtain the feature vector n τ ,in, g is the number of EEG electrodes, is the real number space.
10. The method for constructing a 2D-MDAGTS model for EEG classification according to claim 8, It is characterized in that In step S3, the feature vectors of scale factors τ = 1 and τ = 2 are fused and spliced to obtain a two-dimensional feature matrix n in , and for n in Adjust and finally get the model input x in ,in,
Citation Information
Cited By
Partial discharge intelligent identification method and system based on collaborative reasoning
CN120470462A
Identification system and identification method for attention deficit hyperactivity disorder
CN120661143A