Brain wave processing method and application
By building a dual-branch network in time and time frequency domain, using multi-scale time filtering and frequency domain feature selection strategies, the time and frequency domain characteristics of EEG signals are fused, and the problem of insufficient utilization of EEG signal information in the prior art is solved, and high-precision Alzheimer's disease diagnosis is achieved.
Patent Information
- Application Number
- CN202510277234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-08-01
AI Technical Summary
Existing deep learning models fail to fully utilize the time and frequency domain information of EEG signals in Alzheimer's disease diagnosis, and the fusion of time and frequency domain features is difficult to effectively carry out, resulting in limited classification accuracy and generalization capabilities.
A dual-branch network in time and time frequency domain is built, and a multi-scale time filtering module and a multi-scale feature selection strategy based on the frequency domain are used. By splicing and fusion of time and frequency domain features, an efficient feature fusion module is designed, and the complementarity and correlation of EEG signals are used for modeling.
It improves the accuracy of Alzheimer's disease diagnosis and the robustness of the model, and achieves higher accuracy classification effects.
Smart Images

Figure CN120408258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroencephalogram processing, and particularly to a method and application for processing electroencephalograms. Background Art
[0002] Alzheimer's disease (AD) is the most common cause of dementia, especially common in the elderly population over 65 years old. According to statistics from the World Health Organization (WHO), currently more than 550 million people suffer from dementia, and dementia is the seventh leading cause of death globally, with nearly 10 million new cases reported each year. Alzheimer's disease usually presents as progressive memory loss, cognitive decline, and a decrease in the ability to perform daily activities. There is currently no known cure for Alzheimer's disease, but the latest progress in preventive drug trials and therapeutic interventions has sparked great interest in developing early detection algorithms. AD can be detected or predicted using various techniques, including neuroimaging methods (such as magnetic resonance imaging MRI and positron emission tomography PET), cognitive function tests (such as memory assessment, language ability tests), biomarker analysis (such as cerebrospinal fluid). However, these methods often suffer from problems such as high equipment costs, invasiveness, complex operations, and subjective factors, making it difficult to achieve widespread screening and accurate diagnosis. In contrast, electroencephalography (EEG) is considered a potential clinical alternative due to its non-invasiveness, cost-effectiveness, and portability. EEG signals reflect the electrical activity of the brain and can be used to evaluate neuronal degeneration related to the progression of AD. Medical research has shown that the EEG signals of Alzheimer's disease patients exhibit significant changes in the time domain and frequency domain, such as changes in brain wave activity, enhanced slow waves, etc., and these changes provide potential clues for early diagnosis.
[0003] Traditional EEG signal analysis methods mainly rely on manual feature extraction and traditional machine learning algorithms, such as support vector machines (SVM), decision trees, etc. These methods can extract some patterns in EEG signals to a certain extent, but their classification accuracy and generalization ability are often limited by the selection of manual features and the complexity of the model. With the rapid development of deep learning technology, especially the successful application of convolutional neural networks (CNN) in the field of image processing, more and more research has begun to attempt to apply deep learning methods to the automatic feature extraction and classification tasks of EEG signals.
[0004] However, most deep learning models only process features from a single time domain or time-frequency domain, failing to fully utilize the rich information contained in EEG signals in these two domains. The CNN-ViT model combines CNN and Vision Transformer components to extract time-domain and frequency-domain features from EEG signals and fuse these features through concatenation. These methods have demonstrated that joint modeling of the time domain and frequency domain is superior to a single domain in extracting time-frequency features. However, due to the non-stationary nature of EEG data, directly applying the Fast Fourier Transform (FFT) to obtain frequency-domain information is inappropriate because FFT assumes that the signal is stationary during the analysis window. Using FFT in this context can lead to the loss of local information and hinder the capture of time-varying frequency characteristics. In addition, due to the different characteristics of time-domain and frequency-domain signals, it is challenging to control their fusion and interaction. Therefore, how to simultaneously combine time-frequency features and achieve higher-precision Alzheimer's disease detection through deep learning methods remains an urgent technical problem to be solved. Summary of the Invention
[0005] The primary objective of the present invention is to overcome the problems existing in the prior art and provide a method for processing brain waves and its application. The present invention can improve the accuracy of Alzheimer's disease diagnosis.
[0006] As another objective of the present invention, there is provided a non-volatile storage medium suitable for storing a computer program implemented according to the above method.
[0007] To achieve the above objectives, the present invention provides a method for processing brain waves, which includes the following steps: Obtain brain wave data; Preprocess the brain wave data, and divide the preprocessed brain wave data into a training set, a validation set, and a test set according to a ratio. Construct an image processing model, which includes a time-domain feature extraction network, a time-frequency domain feature extraction network, a feature fusion module, and a fully connected layer classifier. Use the training set to train the image processing model, use the validation set to adjust the parameters of the trained image processing model, use the test set to test the image processing model with adjusted parameters, and obtain a trained image processing model. Input the brain wave data to be classified into the trained image processing model to obtain a processing result.
[0008] Further, before training the image processing model using the training set, splice the electroencephalogram data of the subjects in the training set in the time dimension, calculate the global mean and standard deviation for each channel of the spliced training set, standardize the training set and the test set using the global mean and the standard deviation, and use the standardized training set and test set to train and validate the image processing model.
[0009] Further, the specific calculation method for standardizing the training set and the test set using the global mean and the standard deviation is as follows:
[0010] where is the value of the th data point of the electroencephalogram signal in the standardized training set and test set, is the value of the th data point of the electroencephalogram signal in the training set and test set, is the global mean, The calculation method of
[0011] is the standard deviation, and the calculation method of
[0012] is the total number of data points of the electroencephalogram signal in the training set and test set.
[0013] Further, the preprocessing of the electroencephalogram data and the proportional division of the preprocessed electroencephalogram data into a training set, a validation set, and a test set specifically include: ]>Use a filter to denoise the electroencephalogram data to obtain denoised electroencephalogram data; Remove the artifacts in the denoised electroencephalogram data to obtain processed electroencephalogram data, where the artifacts include signals generated by eye movement and muscle activity; Segment the processed electroencephalogram data in chronological order, and record the ID and label to which the segmented electroencephalogram data belongs. The label is Alzheimer's patients and normal people, to obtain a segmented electroencephalogram data set; Divide the segmented electroencephalogram data set into a training set, a validation set, and a test set according to the ratio.
[0014] Further, the time-domain feature extraction network includes a spatial filtering convolution module and a multi-scale time selection module connected in sequence. The spatial filtering convolution module is used to obtain a spatial feature map based on the input electroencephalogram signal, and the multi-scale time selection module is used to obtain a weighted time feature map based on the spatial feature map.
[0015] Further, the multi-scale time selection module includes a multi-scale time filtering module and a frequency-domain-based scale selection module. The multi-scale time filtering module is composed of multiple depthwise dilated convolution layers, and the output end of each depthwise dilated convolution layer is connected to the input end of the frequency-domain-based scale selection module. The multi-scale time filtering module is used to obtain an initial time feature map based on the spatial feature map, and the frequency-domain-based scale selection module is used to obtain a weighted time feature map based on the initial time feature map.
[0016] Further, the time-frequency domain feature extraction network includes an initial convolution module, a second max pooling module, a multi-stage residual module, and a global average pooling module connected in sequence. The initial convolution module is used to obtain a feature map based on the time-frequency image, and the time-frequency image is obtained by performing a short-time Fourier transform on the electroencephalogram signal; the second max pooling layer is used to perform downsampling on the feature map to obtain a pooling result; the multi-stage residual module includes 4 stages, the number of channels in each stage is 64, 128, 256, and 512 in sequence, each stage includes 2 residual units, and each residual unit contains 2 convolutional layers. The multi-stage residual module is used to extract deep features of the pooling result to obtain an extraction result, and the global average pooling module is used to perform pooling on the extraction result to obtain an output.
[0017] Further, the feature fusion module includes: a time-domain feature mapping layer, a time-frequency domain feature mapping layer, a feature splicing layer, and a fused feature mapping layer.
[0018] The present invention also provides an Alzheimer's disease diagnosis method based on electroencephalogram. The method is based on the above-mentioned electroencephalogram processing method and includes: Obtain electroencephalogram data; Preprocess the electroencephalogram data, and divide the preprocessed electroencephalogram data into a training set, a validation set, and a test set according to a ratio; Construct an image processing model, and the image processing model includes a time-domain feature extraction network, a time-frequency domain feature extraction network, a feature fusion module, and a fully connected layer classifier; Use the training set to train the image processing model, use the validation set to adjust the parameters of the trained image processing model, use the test set to test the image processing model with adjusted parameters, and obtain a trained image processing model; Input the electroencephalogram data to be classified into the trained image processing model to obtain a processing result, and obtain a diagnosis result according to the processing result.
[0019] To achieve another object of the present invention, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer storage program is executed by a processor, the processing method of an electroencephalogram is implemented.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention extracts the time-frequency-space features of EEG signals by constructing a two-branch network in the time domain and the time-frequency domain. At the same time, an efficient multi-scale time filtering module is designed to fuse the time features of different scales in the time domain; and a multi-scale feature selection strategy based on the frequency domain is adopted to calculate the weights of information at different levels to further improve the representation ability and robustness of the overall model; and, the present invention uses the complementarity and correlation between different perspective information of EEG signals for modeling, and then conducts the diagnosis of AD, effectively improving the accuracy of AD diagnosis. Description of the Drawings
[0021] Figure 1 is a flowchart of a method for processing electroencephalogram according to Embodiment 1 of the present invention; Figure 2 is a structural diagram of an Alzheimer's disease classification model according to Embodiment 1 of the present invention; Figure 3 is a structural diagram of a multi-scale time selection module according to Embodiment 1 of the present invention; Figure 4 is a structural diagram of a time-frequency domain feature extraction network according to Embodiment 1 of the present invention; Figure 5 is a structural diagram of a dual-domain feature fusion module according to Embodiment 1 of the present invention. Detailed Embodiments
[0022] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0023] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0024] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0025] In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0026] Embodiment 1 As Figure 1 shown, a method for processing brain waves according to a preferred embodiment of an embodiment of the present invention includes: S1: Obtain brain wave data; In one embodiment, a dataset from the open-source neuroscience computing platform OpenNeur was used. Download the specified Alzheimer's disease dataset A dataset of EEG recordings from: Alzheimer's disease, Frontotemporal dementia and Healthy subjects on the OpenNeur official website, hereinafter referred to as ADFD. After downloading and obtaining the data, determine the detailed information of the dataset according to the published data descriptor paper "Dataset of scalp electroencephalogram recordings from Alzheimer's disease, frontotemporal dementia, and healthy subjects with conventional electroencephalograms".
[0027] S2: Preprocess the brain wave data, and divide the preprocessed brain wave data into a training set, a validation set, and a test set according to a ratio; In one embodiment, the preprocessing specifically includes: segmenting the EEG signal data in chronological order according to the sampling rate of 500 Hz and the recording length of the data set. The preprocessed EEG signals (500 Hz) of each subject are segmented into 10-second segments in chronological order, with a step size of 5 seconds between each segment, that is, the overlap rate is 50%. Starting from the beginning of the signal, segments with a length of 10 seconds (500 seconds * 10 Hz = 5000 time points) are extracted, and then the next segment is extracted by moving 5 seconds each time until the entire signal is covered. This overlapping method can increase the amount of data and maintain the continuity and dynamic characteristics of the signal. Record the subject ID and label to which the segment belongs. A total of 10,550 segments are segmented, and the shape of each segment is (19, 5000), where 19 is the number of electrodes and 5000 is the time length. Among them, the number of segments of AD patients is 5771, and the number of segments of healthy controls is 4779. The number of segments for each subject using this method is different, and the number of segments for each category also varies slightly, but the label ratio of the overall data set is relatively balanced, and the model will not be biased towards a certain class due to class imbalance.
[0028] Divide the EEG data segments into a training set, a validation set, and a test set according to the ratio of 70%:15%:15% of the subjects for subsequent model training and evaluation; Concatenate the data of all subjects in the training set in the time dimension to form a large matrix with a shape of (number of electrodes, total concatenated time length). Then, calculate the global mean and standard deviation for each channel: calculate the mean and standard deviation based on the divided training set data to standardize the training set and the test set to avoid information leakage.
[0029] The global mean and the standard deviation are used to standardize the training set and the test set, and the specific calculation method is as follows:
[0030] where, is the value of the th data point of the brain wave signal in the standardized training set and test set, is the value of the th data point of the brain wave signal in the training set and test set, is the global mean, The calculation method of
[0031] is the standard deviation, and the calculation method of
[0032] is the total number of electroencephalogram signal data points in the training set and the test set.
[0033] S3: Construct an image processing model, which includes a time-domain feature extraction network, a time-frequency domain feature extraction network, a feature fusion module, and a fully connected layer classifier; In one embodiment, the time-domain feature extraction network includes a spatially filtered convolutional module and a multi-scale time selection module connected in sequence. The spatially filtered convolutional module is used to obtain a spatial feature map according to the input electroencephalogram signal, and the multi-scale time selection module is used to obtain a weighted time feature map according to the spatial feature map. Specifically, in the time-domain branch (time-domain feature extraction network), for the time-domain branch input data (batch size, number of electrodes, sequence length), add one dimension to make its shape become (batch size, 1, number of electrodes, sequence length), so as to be consistent with the time-frequency domain data and directly use 2D convolution.
[0034] Further, the multi-scale time selection module includes a multi-scale time filtering module and a frequency-domain-based scale selection module, see Figure 3 In this embodiment, to solve the computational challenges brought by parallel convolutions, multi-layer dilated convolutions are used to capture multi-scale temporal relationships in EEG signals. These dilated convolutions simulate the sequential temporal filtering of EEG signals. Compared with using a single large convolutional kernel, layer-by-layer decomposition is more effective in terms of feature extraction and computational efficiency. To avoid overly deep architectures, the dilation rate of the deep dilated convolutions gradually increases with each layer, so that the receptive field expands rapidly. A frequency-domain attention mechanism is applied to the output of each layer of the deep dilated convolutions. This mechanism assigns weights to different temporal features based on the amplitude features of various receptive fields at different scales, thereby optimizing feature selection. SKNet introduces multiple branches with different convolutional kernels, and these branches are selectively combined along the channel dimension.
[0035] Specifically, the multi-scale time filtering module is a multi-layer dilated convolution. The output of each layer of the dilated convolution is normalized using batch normalization (BatchNorm2d), and the outputs are respectively denoted as: , where n is a positive integer representing the total number of layers of the dilated convolution; each layer of the dilated convolution uses the same small convolutional kernel size, and the dilation rate is the same as the number of convolutional layers.
[0036] The frequency-domain-based scale selection module includes a max pooling layer with a 4-fold increase in the time dimension, a 1*1 dimensionality reduction convolutional layer, a scale feature concatenation layer, a fast Fourier transform modulus layer, a fully connected layer for scale weight calculation, a weight application layer, and a 1*1 dimensionality increase convolutional layer.
[0037] The output of each layer of the multi-scale time filtering module Dimensionality reduction and within-scale information mixing are achieved using a max pooling layer and an independent 1×1 dimensionality reduction convolutional layer to obtain . The scale feature concatenation layer concatenates in the channel dimension to obtain a fused feature representation .
[0038] The EEG frequency range is closely related to the temporal receptive field. A larger temporal receptive field covers more temporal information and corresponds to the low-frequency components in the frequency domain. Based on this, we designed a frequency domain attention mechanism that dynamically selects appropriate receptive fields based on the amplitude. This allows the model to focus on the temporal context regions most relevant to the pathological changes in EEG. Our attention mechanism performs FFT on the temporal signal, transforms it to the frequency domain, and adaptively aggregates information from each layer of dilated convolutions. This mechanism models the frequency dependence between feature maps of different scales, unlike temporal operations (such as average pooling) that can lead to insufficient or lost temporal information. Different from EEG-Inception, which only connects features along the channel dimension, our method avoids feature redundancy by focusing on frequency information. The formulas for FFT and calculating the complex modulus are as follows:
[0039] The specific operation is as follows. In the fast Fourier transform modulus layer, performs fast Fourier transform and obtains the frequency domain feature representation through modulus operation
[0040] The fully connected layer for scale weight calculation includes an unbiased linear transformation layer, a Dropout layer, a rectified linear unit (ReLU), an unbiased linear transformation layer, and a Sigmoid activation function connected in sequence. Through this fully connected layer, the input feature is mapped to a weight vector with the size number . This weight vector is applied to the feature . The specific process is as follows: for each scale of the weight and the corresponding feature are multiplied element-wise, and finally a 1×1 dimensionality increase convolutional layer is used to restore to the original dimension.
[0041] The amplitude of the frequency components obtained by calculating FFT through modulus is taken, and the obtained amplitude is converted into a scale-specific attention map through two fully connected layers, which use a bottleneck architecture and a non-linear activation function. The convolutional kernel weights for each scale are calculated using the sigmoid activation function. Finally, a 1×1 convolution is applied to restore the dimension and mix the cross-scale channel information.
[0042] Furthermore, the time-frequency domain feature extraction network is a Residual Network ResNet18, as shown in Figure 4 , which includes, connected in sequence: an initial convolution module, a max pooling module, a multi-stage residual module, and a global average pooling module. The input of the initial convolution module is an image with a size of the number of electrode channels C × time T × frequency F. The initial convolution module is responsible for extracting preliminary features from the input time-frequency image. The output is a feature map after convolution operation, batch normalization, and ReLU activation function. The output size is 64 × T / 2 × F / 2. The max pooling operation is used to downsample the size of the image, reduce the computational amount, control the complexity of the model, and improve the robustness of the model to translation. The output size is 64 × T / 2 × F / 2. The input of each residual module is the output of the previous layer. The global average pooling module averages the spatial features of each channel, and the output is the feature after global average pooling, with a size of 512 × 1 × 1, that is, one average value for each channel, and there are 512 channels in total. The output of each channel is the average value of that channel.
[0043] Specifically, the multi-stage residual module is divided into 4 stages, and each stage includes 2 residual units. The first residual unit of each stage reduces the time-frequency resolution of the time-frequency map through a convolution operation with a stride of 2, so as to extract deeper features. The number of channels in each stage increases sequentially, which are 64, 128, 256, and 512 respectively. In addition, each residual unit includes 2 layers of convolution, and after each layer of convolution, batch normalization and ReLU activation function are connected, and the identity mapping realizes feature fusion through a shortcut connection.
[0044] Before inputting into the network, the torch.stft function in the PyTorch library is used to perform short-time Fourier transform (STFT) on the original EEG data. The STFT formula is as follows:
[0045] The parameters of the STFT are set based on the EEG sampling rate fs. The window size n_fft and the window length win_length of the FFT are both equal to the sampling rate fs, thus maintaining a good balance between frequency and time resolution. The step size hop_length between adjacent frames is set to one-fourth of the sampling rate fs to enhance the time resolution. The window function Hann window is selected. Due to the adaptability of the Hann window to non-periodic signals, it has been widely used in EEG signal processing. This window effectively reduces spectral leakage, maintains a reasonable frequency resolution, and avoids excessive sacrifice of time resolution. The EEG signal is converted into a time-frequency map using the STFT, which is usually represented as a two-dimensional matrix, where one dimension corresponds to time and the other dimension corresponds to frequency. The STFT produces complex values containing amplitude and phase information. The amplitude spectrum is calculated by taking the modulus of the complex number and is input into the time-frequency domain feature extraction network based on the ResNet18 model.
[0046] Furthermore, the feature fusion module includes a time-domain feature mapping layer, a time-frequency domain feature mapping layer, a feature concatenation layer, and a fused feature mapping layer. Features from different domains are effectively combined through simple linear transformation, connection, activation, and normalization steps, thereby reducing feature redundancy and fusing complementary features. Each input feature vector is independently processed through its respective linear layer and mapped to the same dimensional space. Then the transformed feature vectors are concatenated into a unified representation. The subsequent fully connected layer reduces the dimension of the concatenated features, and the projected output is activated by an activation function and regularized using Dropout. Residual connections are applied to the transformed time-domain features to facilitate the gradient flow and prevent over-smoothing of the feature representation. Finally, the output is normalized to ensure stability during training and accelerate convergence.
[0047] S4: Train the image processing model using the training set, adjust the parameters of the trained image processing model using the validation set, test the image processing model with adjusted parameters using the test set, and obtain the trained image processing model; In one embodiment, the cross-entropy loss function is used to train the network, and the Adam optimizer is used to update the parameters of the neural network. Through the backpropagation algorithm, the network parameters are continuously adjusted to optimize the classification accuracy. In the test phase, EEG signal segments of untrained subjects are input into the trained network model. The network outputs prediction results based on the characteristics of the input signal. In addition, the model can also output the confidence score of the classification, which is used to measure the credibility of the prediction result. After the classification of all test segment samples is completed, the performance of the model is evaluated using standard evaluation metrics (accuracy, precision, recall, F1-score). By comparing the prediction results with the actual labels, the classification accuracy and robustness of the model are calculated. After evaluation, the model shows a detection accuracy higher than that of traditional methods and can effectively distinguish Alzheimer's disease patients from healthy people.
[0048] S5: Input the electroencephalogram data to be classified into the trained image processing model to obtain a processing result.
[0049] Example 2 In order to verify the superiority of the method proposed in Example 1, this example compares the proposed network model with the following methods. All comparative experiments are carried out on a public dataset, and the subject-independent verification method is adopted. The same training set and test set division method is used for division to reduce the influence of the experimental results caused by the dataset division. The experimental data comes from public datasets, hereinafter referred to as ADFD and APAVA respectively, as shown in the following table: Table 1 Datasets Used
[0050] In the model classification performance evaluation, the evaluation metrics adopted include accuracy (Accuracy, ACC), precision, recall, and F1-score. These metrics are defined as:
[0051] In the formula, TP, TN, FP, and FN represent true positive value, true negative value, false positive value, and false negative value respectively.
[0052] The focus of this application is the classification of EEG (multi-channel time series) data, especially for the records of AD patients and cognitively normal (CN). To prove the significant performance of the model, this example compares the proposed method with the most excellent time series methods open-sourced in recent years. Five random seeds (41 - 45), fixed training, validation, and test sets are used to calculate the average result and standard deviation.
[0053] All models represent the most advanced time series models developed in recent years, including Autoformer, Crossformer, FEDformer, Informer, iTransformer, MTST, Nonformer, PatchTST, Reformer, and Transformer. All models use 6-layer encoders, the self-attention dimension D is set to 128, and the hidden dimension of the feed-forward network is set to 256. The optimizer used is Adam, and the learning rate is 1e-4. The batch sizes for the APAVA and ADFD datasets are set to 32 and 128 respectively. The training process spans 100 epochs, and early stopping is triggered if the F1-score on the validation set does not improve within 10 consecutive epochs. The model with the highest F1 score on the validation set is saved and then evaluated on the test set.
[0054] Tables 2 and 3 show the performance of various methods in the EEG-based AD / CN binary classification and AD / CN / FTD ternary classification tasks on two datasets, as shown in the following table: Table 2 Performance comparison between the model proposed in this application and other models on the APAVA data test set
[0055] Table 3 Performance comparison between the model proposed in this application and other models on the ADFD data test set
[0056] According to Figure 2 and Figure 3 it can be seen that our model significantly outperforms all state-of-the-art models in four metrics. On the ternary classification ADFD dataset, the model has the highest precision, ranks third in accuracy and recall, and ranks second in F1 score. Overall, the model has achieved the best average ranking among all models on the two datasets.
[0057] To verify the effectiveness of the model strategy, three aspects of ablation experiments were conducted in this embodiment using the ADFD dataset and the leave-one-subject experimental setting. As shown in Table 4: Table 4 Ablation experiment results
[0058] (1) To illustrate that the time-domain network using time-domain multi-scale time selection convolution is superior in performance to previous time-domain networks, a large-kernel time convolution equivalent to the final receptive field of multi-layer depth dilated convolution is used to replace the multi-scale time selection convolution module, so as to prove the efficient performance of multi-layer depth dilated convolution. A multi-scale convolution module (from the EEG-Inception model) equivalent to the receptive field of each layer of multi-layer depth dilated convolution is used to replace the multi-scale time selection convolution module, so as to prove the efficient performance of the frequency-domain-based scale selection mechanism. For more concise representation, the large-kernel time convolution is denoted as TimeConv, and the multi-scale convolution module of the EEG-Inception model is denoted as MultiConv.
[0059] (2) To illustrate that the dual-domain feature fusion module can effectively fuse time-domain and time-frequency domain features, it is replaced with common feature fusion methods in deep learning: concatenation and weighted fusion. For more concise representation, concatenation and weighted fusion are denoted as Cat and Weighted (3) To illustrate that the model obtained by joint time-domain and time-frequency domain feature extraction is superior in performance to the models constructed by only one of them, it is modified to use only the time-domain feature extraction network or the time-frequency domain feature extraction network. For more concise representation, the single time-domain feature extraction network and the single time-frequency domain feature extraction network are respectively denoted as "T-only" and "TF-only".
[0060] Table 4 shows that each module of the model is crucial for achieving the best performance. The results of single-domain feature extraction in the time domain and time-frequency domain are inferior to the joint time domain and time-frequency domain, emphasizing the important role of time-frequency complementary information. In the study of the time-domain branch, the multi-scale time selection convolutional module based on frequency attention can effectively extract time-domain information at different scales, which is superior to the large-kernel convolution and multi-scale parallel convolution with equivalent receptive fields. In the study of feature fusion, by comparing with common methods such as splicing and weighted fusion, it is proved that the adaptive feature fusion module proposed in this embodiment can more effectively fuse the complementary features of the two domains.
[0061] Embodiment 3 The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer storage program is executed by a processor, the processing method of an electroencephalogram is realized.
[0062] Through the embodiment of the present invention, a processing method and application of an electroencephalogram are provided. By constructing a dual-branch network in the time domain and time-frequency domain to extract the time-frequency-space features of EEG signals, at the same time, an efficient multi-scale time filtering module is designed to fuse the time features of different scales in the time domain; and a multi-scale feature selection strategy based on the frequency domain is adopted to calculate the weights of information at different levels to further improve the representation ability and robustness of the overall model; and, the present invention uses the complementarity and correlation between different perspective information of EEG signals for modeling, and then conducts the diagnosis of AD, effectively improving the accuracy of AD diagnosis.
[0063] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.
Claims
1. A method for processing brain waves, characterized in that, The method includes the following steps: Obtain electroencephalogram data; Preprocess the electroencephalogram data, and divide the preprocessed electroencephalogram data into a training set, a validation set, and a test set according to a ratio; Construct an image processing model, where the image processing model includes a time-domain feature extraction network, a time-frequency domain feature extraction network, a feature fusion module, and a fully connected layer classifier; Use the training set to train the image processing model, use the validation set to adjust the parameters of the trained image processing model, and use the test set to test the image processing model with adjusted parameters to obtain a trained image processing model; Input the electroencephalogram data to be classified into the trained image processing model to obtain a processing result.
2. The processing method of brain waves according to claim 1, characterized in that Before using the training set to train the image processing model, splice the electroencephalogram data of the subjects in the training set in the time dimension, calculate the global mean and standard deviation for each channel of the spliced training set, use the global mean and the standard deviation to standardize the training set and the test set, and use the standardized training set and test set to train and validate the image processing model.
3. The method for processing brain waves according to claim 2, wherein The specific calculation method for standardizing the training set and the test set using the global mean and the standard deviation is as follows: Among them, is the value of the th data point of the electroencephalogram signal in the standardized training set and test set, is the value of the th data point of the electroencephalogram signal in the training set and test set, is the global mean value, is calculated as follows: is the standard deviation, and its calculation method is as follows: is the total number of electroencephalogram signal data points in the training set and the test set.
4. A method for processing brain waves according to claim 1, characterized in that, The preprocessing of the electroencephalogram data and the division of the preprocessed electroencephalogram data into a training set, a validation set, and a test set according to a ratio specifically include: Use a filter to denoise the electroencephalogram data to obtain denoised electroencephalogram data; Remove the artifacts in the denoised electroencephalogram data to obtain processed electroencephalogram data, where the artifacts include signals generated by eye movement and muscle activity; Segment the processed electroencephalogram data in chronological order, and record the ID and label of the segmented electroencephalogram data. The label is Alzheimer's patients and normal people, to obtain a segmented electroencephalogram data set; Divide the segmented electroencephalogram data set according to into a training set, a validation set, and a test set.
5. A method for processing brain waves according to claim 1, characterized in that, The time-domain feature extraction network includes a spatially filtered convolutional module and a multi-scale time selection module connected in sequence. The spatially filtered convolutional module is used to obtain a spatial feature map according to the input electroencephalogram signal, and the multi-scale time selection module is used to obtain a weighted time feature map according to the spatial feature map.
6. A method for processing brain waves according to claim 5, characterized in that, The multi-scale time selection module includes a multi-scale time filtering module and a frequency-domain-based scale selection module. The multi-scale time filtering module is a plurality of depthwise dilated convolutional layers, and the output end of each depthwise dilated convolutional layer is connected to the input end of the frequency-domain-based scale selection module. The multi-scale time filtering module is used to obtain an initial time feature map according to the spatial feature map, and the frequency-domain-based scale selection module is used to obtain a weighted time feature map according to the initial time feature map.
7. A method for processing brain waves according to claim 1, characterized in that, The time-frequency domain feature extraction network includes an initial convolution module, a second max pooling module, a multi-stage residual module, and a global average pooling module connected in sequence. The initial convolution module is used to obtain a feature map based on a time-frequency image, where the time-frequency image is obtained by performing short-time Fourier transform on an electroencephalogram (EEG) signal. The second max pooling layer is used to perform downsampling on the feature map to obtain a pooling result. The multi-stage residual module includes 4 stages, and the number of channels in each stage is 64, 128, 256, and 512 in sequence. Each stage includes 2 residual units, and each residual unit contains 2 convolutional layers. The multi-stage residual module is used to extract deep features of the pooling result to obtain an extraction result, and the global average pooling module is used to perform pooling on the extraction result to obtain an output.
8. A method for processing brain waves according to claim 1, characterized in that, The feature fusion module includes: a time-domain feature mapping layer, a time-frequency domain feature mapping layer, a feature splicing layer, and a fusion feature mapping layer.
9. A method for diagnosing Alzheimer's disease based on brain waves, the method being based on a method for processing brain waves according to any one of claims 1 to 8, characterized in that, including: Obtain EEG data; Preprocess the EEG data, and divide the preprocessed EEG data into a training set, a validation set, and a test set according to a ratio. Construct an image processing model, where the image processing model includes a time-domain feature extraction network, a time-frequency domain feature extraction network, a feature fusion module, and a fully connected layer classifier. Use the training set to train the image processing model, use the validation set to adjust the parameters of the trained image processing model, use the test set to test the image processing model with adjusted parameters, and obtain a trained image processing model. Input the EEG data to be classified into the trained image processing model to obtain a processing result, and obtain a diagnosis result based on the processing result.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer storage program is executed by a processor, it implements a method for processing EEG as described in any one of claims 1 to 8.