Self-supervised electroencephalogram classification method based on comparative learning and space-time mask reconstruction
By adopting a self-supervised learning framework based on contrast learning and spatiotemporal mask reconstruction in EEG data analysis, the problems of insufficient spatiotemporal correlation modeling and insufficient information combination in the prior art are solved, and better EEG data classification performance is achieved.
Patent Information
- Application Number
- CN202510051601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively capture spatiotemporal correlation and integrate fragment-level and instance-level information in EEG data analysis, resulting in insufficient generalization performance of the model in EEG data classification tasks.
A self-supervised learning framework based on contrast learning and spatiotemporal mask reconstruction is designed, and mask reconstruction of time dimensions and channel dimensions is combined with contrast learning, the spatiotemporal relationships of EEG data are captured and modeled at the fragment and instance levels.
The performance of EEG data classification tasks is effectively improved, and the generalization ability of the model is improved by integrating spatiotemporal dimensions and combining fragment-level and instance-level information.
Smart Images

Figure BSA0000299941760000046 
Figure BSA0000299941760000048 
Figure BSA0000299941760000051
Abstract
Description
Technical Field
[0001] The embodiment of the present invention relates to the field of medical time series technology, and in particular to a self-supervised electroencephalogram classification method based on contrastive learning and spatiotemporal mask reconstruction. Background Art
[0002] Electroencephalography (EEG) is a non-invasive technology that collects electrical signals through scalp electrodes to record and measure brain electrical activity. EEG signals reflect the intrinsic neural activity of the brain, have high temporal resolution, and can capture millisecond-level dynamic changes in the brain. They are widely used in epilepsy recognition, emotion analysis, sleep research, and brain-computer interfaces, and have important clinical value.
[0003] Although EEG technology has shown broad application prospects in many fields, its analysis still faces several challenges. First, labeled samples are scarce, and obtaining large-scale labeled data is time-consuming and expensive, and requires professional medical knowledge and experience. Second, due to the complexity and diversity of EEG signals, different experts often have subjective differences in their interpretation of the same EEG signal segment, which further affects the learning effect of the model. In addition, in some cases, it may be challenging to accurately understand the thoughts or behaviors of participants in cognitive neuroscience experiments, so it is difficult to obtain accurate labels. For example, in an imagination task, the subjects may not follow the instructions, or the brain activity process being studied (such as meditation state, emotional changes, etc.) itself is difficult to accurately quantify through objective indicators. Therefore, traditional supervised learning methods face significant bottlenecks in EEG analysis and it is difficult to fully tap the potential of EEG data.
[0004] Self-supervised learning (SSL) has received widespread attention in recent years as a learning paradigm that does not require labels. SSL learns by mining the internal structure of data, thereby effectively reducing the reliance on external labels and lowering the cost and technical threshold of data preparation. Especially for the EEG field, SSL can use the time series characteristics of the data itself to train the model without clear labels, thereby improving the generalization performance of the model, making it a promising method to solve the problem of EEG data analysis. Summary of the invention
[0005] Purpose of the invention: Although SSL has made significant progress in the field of EEG, there is still much room for improvement in performance due to the failure to fully utilize the characteristics of EEG data, specifically:
[0006] 1) Insufficient modeling of spatiotemporal correlations: The core of EEG sequence analysis is to capture and utilize the complex temporal dependencies and inter-channel correlations in the observed values. Temporal dependencies can help grasp the changing trends of EEG data, and spatial dependencies can model and learn the functional dependencies between different brain regions. EEG data contains complex spatiotemporal correlations, but existing methods are often limited to modeling in a single dimension of time or space.
[0007] 2) Insufficient integration of segment-level and instance-level information: Segment-level modeling focuses on mining local changes in signals. By modeling EEG segments, the model can effectively capture instantaneous changes and short-term trends in EEG signals. By understanding and encoding this dynamic pattern, the model can achieve better generalization performance. Instance-level modeling can enhance the model's understanding of the global structure of the EEG sequence. Downstream tasks of EEG data usually involve classifying the currently observed EEG sequence. For classification tasks, the class is labeled on the entire time series (instance), so instance-level representation is required to learn user-specific features.
[0008] In view of the above problems, the purpose of the present invention is to propose improvements to the shortcomings of the prior art, and to design and implement a self-supervised learning framework based on contrastive learning and spatiotemporal mask reconstruction. Specifically, the framework introduces a channel mask strategy based on the reconstruction of the time dimension mask, thereby effectively capturing the spatiotemporal relationship of EEG data and solving the problem that the existing methods are difficult to integrate the spatiotemporal dimensions. The modeling of the fragment and instance levels is achieved through mask reconstruction tasks and contrastive learning, respectively. The reconstruction task enables the model to learn the fine-grained relationship between fragments by reconstructing the masked fragments, and contrastive learning shortens the distance between the instance and its potential positive pair, which can effectively improve the performance of downstream classification tasks when the positive pair is properly constructed.
[0009] Technical solution: To achieve the above technical effects, the technical solution proposed by the present invention is:
[0010] A self-supervised learning framework based on contrastive learning and spatiotemporal mask reconstruction, including the following steps:
[0011] Step S1: pre-process the original EEG data by filtering and downsampling, and window it into segments consisting of 256 time steps as model input;
[0012] Step S2: For each sequence, a pair of mutually positive augmented samples is generated through time clipping and frequency mixing techniques;
[0013] Step S3: applying the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences;
[0014] Step S4: input the four sequences into the encoder to encode and obtain their respective feature representations, perform unsupervised pre-training on the augmented sequence pairs using a contrastive learning function, and input the representation of the masked sequence into a multi-layer perceptron to reconstruct the masked part of the original sequence;
[0015] Step S5: After pre-training, freeze the encoder of the original model, and input the representation output by the encoder into the logistic regression classifier for training;
[0016] Step S6: During the testing phase, the data in the test data set is input into the trained model to perform the EEG classification task. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a framework diagram for overall training of a self-supervised EEG classification method based on contrast and spatiotemporal mask reconstruction proposed in the present invention;
[0018] Figure 2 for Figure 1 A flow chart of a self-supervised EEG classification method based on contrast and spatiotemporal mask reconstruction is shown in FIG.
[0019] Figure 3 Schematic diagram for data augmentation;
[0020] Figure 4 Schematic diagram of spatiotemporal mask reconstruction; DETAILED DESCRIPTION
[0021] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0022] The present invention provides a self-supervised EEG classification method based on contrast and spatiotemporal mask reconstruction, such as Figure 2 As shown, it specifically includes steps S1 to S6.
[0023] Step S1: pre-process the original EEG data by filtering and downsampling, and window it into segments consisting of 256 time steps as model input;
[0024] Step S2: For each sequence, a pair of mutually positive augmented samples is generated through time clipping and frequency mixing techniques;
[0025] Step S3: applying the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences;
[0026] Step S4: input the four sequences into the encoder to encode and obtain their respective feature representations, perform unsupervised pre-training on the augmented sequence pairs using a contrastive learning function, and input the representation of the masked sequence into a multi-layer perceptron to reconstruct the masked part of the original sequence;
[0027] Step S5: After pre-training, freeze the encoder of the original model, and input the representation output by the encoder into the logistic regression classifier for training;
[0028] Step S6: During the testing phase, the data in the test data set is input into the trained model to perform the EEG classification task.
[0029] The step S1 specifically includes: Derived from the Crowdsourced dataset, the subjects were collected in a resting state, including two phases of eyes open and eyes closed, each phase lasting 2 minutes, the data was initially recorded at 2048Hz and then downsampled to 128Hz. It was then windowed into segments consisting of 256 time steps. The processed dataset For each input sample Where T is the length of the timestamp and F is the number of feature channels.
[0030] In the step S2, for each sequence, a pair of augmented samples that are positive to each other are generated by time clipping and frequency mixing technology, which is specifically expressed as:
[0031] like Figure 3 As shown, for each instance x i In the first step, the present invention randomly selects two overlapping time periods [a1, b1] and [a2, b2] to obtain x i,1 and x i,2 , where 0<a1 a2 b1 b2 t and the constraint b1-a2 is not less than the threshold ω;
[0032] The second step is to calculate the time period x. i,2 and another random training instance x from the same batch k , use the fast Fourier transform FFT to transform it into the frequency domain, expressed as
[0033]
[0034] In the formula Respectively represent their corresponding spectrum representations in the frequency domain;
[0035] Step 3: Random selection In the figure, except for the two with the largest amplitude, half of the frequency components are replaced by the corresponding The frequency components in middle;
[0036] The fourth step is to Through the inverse Fourier transform iFFT, it is converted to the time domain and x i,1 Mutually augmented positive samples It is expressed as:
[0037]
[0038] Finally, the sample pairs generated by time clipping and frequency mixing Abbreviated as (x, x′).
[0039] The step S3 specifically includes applying the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences, such as Figure 4 As shown, it is specifically expressed as: for the augmented samples x and x′, the first step is to calculate the size of a single reserved block according to the mask rate and the number of mask blocks:
[0040]
[0041] Where ρ represents the mask rate, β represents the number of mask blocks, and l represents the total number of time steps of the sample.
[0042] The second step is to generate a mask matrix of the same size and all ones for x and x′ respectively. For each channel of each sample, randomly generate β starting time points, which should be less than lB. Set the values of the starting time point of each channel in the mask matrix and the position of the B time points thereafter to 0, and obtain the mask matrix M of the time dimension. The corresponding position value of the matrix is 1, which means that the position is masked.
[0043] The third step is to randomly set all the values of the mask matrix of a certain channel to 1, indicating masking in the spatial dimension, and obtain the spatiotemporal mask matrix M;
[0044] The fourth step is to set the positions in x and x' where the corresponding mask matrix values are 1 to 0, and obtain the masked sequence and
[0045] The step S4 specifically includes inputting the above four sequences into the encoder for encoding to obtain their respective feature representations for comparison and reconstruction, which is specifically represented by inputting the four sequences into the encoder for encoding to obtain their respective feature representations r, r′, and r and r′ are max-pooled along the time dimension and then input into a linear layer to obtain the final representations h and h′. Unsupervised pre-training is then performed using the contrastive learning function. The contrastive loss function for the input anchor sample x is defined as:
[0046]
[0047] in Constraint embedding is consistent, and v(h′) constrain the standard deviation of each variable embedded in a batch to be above a given threshold, forcing the embedding vectors of samples in the batch to be different. d represents the dimension of the embedding, and h j Represents a vector consisting of the values of all vectors in the same batch at dimension j. and c(h′) penalize the non-diagonal elements in the embedded covariance matrix to reduce the correlation between features. λ, μ, and γ are hyperparameters for adjusting these three losses.
[0048] Will and Input into a multi-layer perceptron to reconstruct the masked part of the original sequence, and the reconstruction loss function is defined as:
[0049]
[0050] Finally, the total loss function is defined as:
[0051]
[0052] Among them, λ1 and λ2 are hyperparameters for adjusting the two losses.
[0053] In step S5, the representation output by the pre-trained encoder is input into a logistic regression classifier for training using binary cross entropy, specifically including:
[0054] Freeze the encoder of the pre-trained model, and input the representation of the encoder output and its corresponding label into the logistic regression classifier for training using binary cross entropy.
[0055] The step S6 specifically includes inputting the test data set into the trained model to perform the EEG classification task during the test phase.
[0056] like Figure 1 As shown, the present invention first performs data enhancement on the original EEG sequence x to obtain an enhanced data pair (x, x′). Then, a spatiotemporal masking operation is performed on x and x′ to generate a corresponding masked sequence and Among them, x and x′ are used for comparison tasks to mine instance-level similarity information, while and Then, x, x′, and Input the backbone encoder to obtain their representations r, r′, and In order to alleviate the gradient interference problem that may be caused by optimizing the comparison and reconstruction goals at the same time, this paper constrains the two tasks to be performed at different levels. For the comparison task, r and r′ are max-pooled along the time dimension and then input into a linear layer to obtain the final representations h and h′. This process ensures that the comparison task is performed in the high-level feature space, and these features can reflect the global structure and semantic information of the data. For the reconstruction task, and Input a multi-layer perceptron (MLP) to get x R and x′ R , to reconstruct the masked parts of the original sequence. This process is carried out in the low-level feature space, which focuses more on capturing local and detailed information of the data. By performing these tasks in different feature spaces, not only can gradient interference be reduced, but the advantages of each task can also be better utilized, thereby generating more effective and robust feature representations to improve the generalization ability of the model.
[0057] In summary, the self-supervised EEG classification method based on contrastive learning and spatiotemporal mask reconstruction proposed in the present invention is of great help in studying the classification generalization of EEG data in reality.
[0058] The above specific embodiments are only for explaining the principle and technical method of the method proposed in the present invention, and are not intended to limit the implementation of the technical solution of the present invention. For those skilled in the art, the technical solution of the method proposed in the present invention can be reasonably adjusted and replaced according to needs, and these adjustments and replacements are protected by the claims of the present invention.
Claims
1. A self-supervised EEG classification method based on contrastive learning and spatiotemporal mask reconstruction, characterized in that The method at least comprises: Step S1: pre-process the original EEG data by filtering and downsampling, and window it into segments consisting of 256 time steps as model input; Step S2: For each sequence, a pair of mutually positive augmented samples is generated through time clipping and frequency mixing techniques; Step S3: applying the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences; Step S4: input the four sequences into the encoder to encode and obtain their respective feature representations, perform unsupervised pre-training on the augmented sequence pairs using a contrastive learning function, and input the representation of the masked sequence into a multi-layer perceptron to reconstruct the masked part of the original sequence; Step S5: After pre-training, freeze the encoder of the original model, and input the representation output by the encoder into the logistic regression classifier for training; Step S6: During the testing phase, the data in the test data set is input into the trained model to perform the EEG classification task.
2. The electroencephalogram signal classification process according to claim 1, characterized in that: Step S1 specifically includes: Derived from the Crowdsourced dataset, the subjects were collected in a resting state, including two phases of eyes open and eyes closed, each phase lasting 2 minutes, the data was initially recorded at 2048Hz and then downsampled to 128Hz. It was then windowed into segments consisting of 256 time steps. The processed dataset For each input sample Where T is the length of the timestamp and F is the number of feature channels.
3. The electroencephalogram signal preprocessing process according to claim 2, characterized in that: For each sequence described in step S2, a pair of mutually positive augmented samples are generated by time clipping and frequency mixing technology, specifically including: For each instance x i , in the first step, the present invention randomly extracts two overlapping time periods [a1, b1] and [a2, b2] to obtain x i,1 and x i,2 , where 0 < a1 < a2 < b1 < b2 < t and it is constrained that b1 - a2 is not less than the threshold ω; The second step is to calculate the time period x. i,2 and another random training instance x from the same batch k , use the fast Fourier transform FFT to transform it into the frequency domain, expressed as In the formula Respectively represent their corresponding spectrum representations in the frequency domain; Step 3: Random selection In the figure, replace half of the frequency components except the two with the largest amplitudes with the corresponding The frequency components in middle; The fourth step is to Through the inverse Fourier transform iFFT to convert to the time domain, we can get the same value as x i,1 The augmented positive samples It is expressed as: Finally, the sample pairs generated by time clipping and frequency mixing Abbreviated as (x, x′).
4. The method for data augmentation of sample pairs generated by time clipping and frequency mixing according to claim 3, characterized in that: Step S3 applies the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences, which specifically includes: For the augmented samples x and x′, the first step is to calculate the size of a single retained block based on the mask rate and the number of mask blocks: Where ρ represents the mask rate, β represents the number of mask blocks, and l represents the total number of time steps of the sample; The second step is to generate a mask matrix of the same size for x and x′ respectively. For each channel of each sample, β starting time points are randomly generated. The starting time point should be less than -B. The values of the starting time point of each channel in the mask matrix and the position of the B time points thereafter are set to 0 to obtain the mask matrix M of the time dimension. The corresponding position value of the matrix is 1, which means that the position is masked; The third step is to randomly set all the values of the mask matrix of a certain channel to 1, indicating masking in the spatial dimension, and obtain the spatiotemporal mask matrix M; The fourth step is to set the positions in x and x' where the corresponding mask matrix values are 1 to 0, and obtain the masked sequence and 5. According to claim 4, the method of applying the spatiotemporal masking technique to the augmented sample pairs to obtain corresponding mask sequences, characterized in that: In step S4, the four sequences are input into the encoder to be encoded to obtain their respective feature representations r, r′, and r and r′ are max-pooled along the time dimension and then input into a linear layer to obtain the final representations h and h′. Unsupervised pre-training is then performed using the contrastive learning function. The contrastive loss function for the input anchor sample x is defined as: in Constraint embedding is consistent, and v(h′) constrain the standard deviation of each variable embedded in a batch to be above a given threshold, forcing the embedding vectors of samples in the batch to be different. d represents the dimension of the embedding, and h j Represents a vector consisting of the values of all vectors in the same batch at dimension j. and c(h′) penalize the non-diagonal elements in the embedded covariance matrix to reduce the correlation between features. λ, μ, and γ are hyperparameters for adjusting these three losses. Will and Input into a multi-layer perceptron to reconstruct the masked part of the original sequence, and the reconstruction loss function is defined as: Finally, the total loss function is defined as: Among them, λ1 and λ2 are hyperparameters for adjusting the two losses.
6. The method of claim 5 wherein the respective feature representations obtained from the encoder are unsupervisedly pre-trained using a contrastive learning function and a reconstruction loss, wherein: In step S5, the encoder of the pre-trained model is frozen, and the representation output by the encoder and its corresponding label are input into the logistic regression classifier for training using binary cross entropy.
7. The method of inputting the representation into a logistic regression classifier for training using binary cross entropy according to claim 6, characterized in that: In step S6, during the testing phase, the test data set is input into the trained model to perform the EEG classification task.
Citation Information
Cited By
Self-supervised graph anomaly detection method based on global space correlation perception
CN120852818A