Electroencephalogram model training method and device and storage medium
By introducing training methods for frequency domain reconstruction loss and classification loss in the EEG model, the shortcomings of existing models in EEG signal repair and generalization capabilities are solved, and stronger cross-task adaptability and signal repair quality are achieved.
Patent Information
- Application Number
- CN202510086276.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
The potential of existing large-scale EEG models in repairing EEG signals is not fully realized, and their generalization capabilities are limited in classification tasks.
A training method for EEG model is proposed. By introducing frequency domain reconstruction loss and classification loss, the model is trained to reconstruct the frequency domain characteristics of EEG signals and learn task-related tag information.
This method not only improves the semantic information representation of the EEG model in the classification task, but also enhances the model's ability in EEG signal repair, realizing the dual optimization of frequency domain reconstruction and classification task.
Smart Images

Figure CN119940465A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of brain-computer interface technology, and in particular to a training method, device, storage medium and program product for an electroencephalogram model. Background Art
[0002] Inspired by Large Language Models (LLMs), some researchers in the industry have developed several large-scale EEG (EEG) models (LEMs) to learn general representations that can be adapted to multiple tasks.
[0003] However, there are many types of devices for collecting EEG signals, and the experimental paradigms used for different tasks vary significantly. This leads to large differences in the number of electrodes, sampling rate, and data duration settings among different datasets. Researchers are looking for a sufficiently general model that can adapt to various EEG configurations and be applicable to multiple downstream tasks. Therefore, researchers have proposed several large-scale EEG models (LEMs) based on self-supervision, such as Brant-2, LaBraM, etc.
[0004] LEMs are mainly designed to abstract general representations with rich semantics for classification tasks, but their potential in raw EEG data restoration has not been fully explored. LaBraM has difficulty reconstructing raw EEG signals due to non-convergence of the loss during the training of the neural word segmenter. Brant-2 aims to reconstruct EEG signals after data augmentation. However, these models cannot be directly used for raw data restoration. Some smaller models have used Masked Autoencoder (MAE) to generate Differential Entropy (DE) features in EEG, but these methods mainly focus on emotion recognition under subject-dependent conditions, resulting in limited generalization ability.
[0005] The industry has not yet proposed a better solution to the above problems. Summary of the invention
[0006] The present application provides a training method, device, storage medium and program product of an EEG model, which are used to at least solve the problem that traditional large-scale EEG models mainly focus on learning general representations for classification tasks, but have not yet fully realized their potential in EEG signal repair.
[0007] In a first aspect, an embodiment of the present application provides a method for training an EEG model, comprising: dividing a first original EEG data sample into a first original EEG segment set, masking some of the original EEG segments in the first original EEG segment set to obtain a corresponding visible EEG segment subset and a masked EEG segment subset; encoding each visible EEG segment in the visible EEG segment subset to determine the corresponding visible EEG simulated frequency domain coding features; extracting the original signal frequency domain features corresponding to the first original EEG data sample, and reconstructing each of the visible EEG segments according to the visible EEG simulated frequency domain coding features to calculate the frequency domain reconstruction loss; decoding the visible EEG simulated frequency domain coding features to predict the codebook base class corresponding to each masked EEG segment, and calculating the classification loss in combination with the codebook base class labels of each masked EEG segment; and training and optimizing the EEG model according to the frequency domain reconstruction loss and the classification loss.
[0008] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the EEG model training method of any embodiment of the present application.
[0009] In a third aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the EEG model training method of any embodiment of the present application are implemented.
[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the EEG model training method of any embodiment of the present application.
[0011] The beneficial effects of the embodiments of the present application are: By introducing frequency domain reconstruction loss and classification loss, the frequency domain reconstruction loss enables the model to focus on recovering the frequency domain features of the signal, while the classification loss helps the model learn task-related label information, so that the EEG model can not only obtain more semantically informative representations in classification tasks, but also reconstruct missing signals when part of the EEG signal is missing, thereby enhancing the model's ability to repair original EEG data and achieving dual optimization of frequency domain reconstruction and classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 A flowchart showing an example of a method for training an EEG model according to an embodiment of the present application is shown; Figure 2 An operation flow chart of an example of pre-training an EEG model according to an embodiment of the present application is shown; Figure 3 An operational flow chart of the first stage of training of the Gram model according to an embodiment of the present application is shown; Figure 4 An operational flow chart of the second stage training of the Gram model according to an embodiment of the present application is shown; Figure 5 A schematic diagram of an operation flow of an example of basic class quantization of an EEG model trained in one stage according to an embodiment of the present application is shown; Figure 6 A schematic diagram of an operation flow of an example of two-stage training based on a layer-fusion masked autoencoder of an EEG model according to an embodiment of the present application is shown; Figure 7 A schematic diagram of the effect simulation of an example of visualization and correlation analysis of basic classes is shown; Figure 8 A schematic diagram of an effect simulation of an example of comparing the classification results of noisy data and restored data is shown; Fig. 9 It is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0015] It should be noted that the current related technology has proposed a self-supervised large-scale EEG model, LaBraM, which consists of two stages. In stage one, the frequency domain amplitude and phase information of the EEG signal are reconstructed through NEURAL TOKENIZER (neural tokenizer) to build a codebook for the original EEG signal. In stage two, the encoder performs a mask reconstruction task to predict the codebook category to which the masked EEG signal belongs. In this way, a large EEG model that can learn universal representations is obtained through two-stage pre-training.
[0016] However, LaBraM only focuses on the performance of large models in classification tasks, and does not explore the potential of large models in data repair (EEG generation tasks). In addition, using the structure of LaBraM stage 1 to reconstruct the original EEG signal will cause the loss to fail to converge, so LaBram cannot use the codebook to perform the original signal repair task. In addition, the structure of LaBraM stage 1 only adds a convolutional structure to extract features in the encoder part, and does not add the same convolutional structure to the decoder part, resulting in the loss failing to converge when reconstructing the original signal.
[0017] It should also be noted that in the current related technologies, EEG data generation and classification tasks are not integrated in the same model. Specifically, EEG data generation usually uses models such as diffusion models or generative adversarial networks (GAN) that are specifically used for generation, but these models are weak in classification capabilities. In addition, although models with structures such as Masked Autoencoder (MAE) can take into account both generation and classification tasks, the generation effect of MAE is usually poor. Moreover, due to the difficulty of training, MAE is usually only applied to EEG feature signals, not original signals, so it is impossible to generate original EEG signals.
[0018] Although there are many EEG models based on large-scale self-supervised learning (such as Brant-2 and LaBraM) for feature abstraction and emotion recognition tasks, they have obvious deficiencies in repairing EEG signals. For example, the LaBraM model failed to successfully reconstruct the original EEG signal during the neural encoder training process because the loss function failed to converge; while Brant-2 attempted to reconstruct the EEG signal after data enhancement, but did not provide a method to directly repair the original data.
[0019] In addition, although some small models use the MAE structure for EEG differential entropy feature generation, these methods mainly focus on emotion recognition tasks and mostly rely on subject-dependent conditions, so their versatility is poor.
[0020] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of the present application, and is not to be regarded as a limitation of the present application. In addition, the technical solutions described in the above-mentioned current related art are not prior art, and they may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.
[0021] In view of the above-mentioned deficiencies in the current related technologies, a general model based on the original EEG signal classification and repair task is proposed in the embodiments of the present application, named the Gram model, which aims to solve the problem of original EEG signal repair and support stronger cross-task adaptability.
[0022] The design concept of the Gram model is derived from the discretization technology in image and natural language processing. Specifically, the EEG signal is divided into multiple small blocks (or fragments) and a corresponding code book is constructed. The elements in each code book are called "basic classes". Similar to the tokenization process in image and natural language processing, the Gram model generates EEG signals by repeatedly arranging combinations of these basic classes from the code book. In this way, the Gram model can not only repair EEG signals, but also effectively perform EEG signal classification tasks.
[0023] Figure 1 An operational flowchart of an example of a method for training an EEG model according to an embodiment of the present application is shown.
[0024] like Figure 1 As shown, in step S110, the first original EEG data sample is segmented into a first original EEG segment set, and part of the original EEG segments in the first original EEG segment set are masked to obtain corresponding visible EEG segment subsets and masked EEG segment subsets.
[0025] Specifically, the first original EEG data sample is segmented according to the time window or feature window to form a series of uniformly sized subsequences (patches, fragments). These fragments usually correspond to EEG signals within a certain time period, which can be small blocks in the original EEG data or specific fragments in the frequency domain or time domain.
[0026] Furthermore, in the first original EEG segment set, a certain proportion of segments are randomly selected for masking. These masked segments will be used as "missing information" input during the training process, and masking can be achieved by replacing the data points of the selected segments with zero, noise or some default values to simulate the absence or damage of the signal. In this way, two subsets are obtained: one is the "visible EEG segment subset" (the part that is not masked), and the other is the "masked EEG segment subset" (the part that has been masked).
[0027] In step S120, each visible EEG segment in the visible EEG segment subset is encoded to determine a corresponding visible EEG simulated frequency domain encoding feature.
[0028] It should be noted that the types of frequency domain encoders can be diverse, and various non-restrictive deep neural networks (such as convolutional neural networks CNN or transformers Transformer, etc.) can be used to encode each segment in the visible EEG segment subset. Exemplarily, the visible EEG segment can be processed through a series of convolutional layers, pooling layers, or multi-layer perceptrons to extract local time domain features, and then these time domain features are converted to the frequency domain using methods such as Fourier transform or wavelet transform to obtain frequency domain features. In the frequency domain, the model can learn the frequency components of the signal, such as different frequency bands of brain wave activity (alpha waves, beta waves, etc.), which are usually related to cognitive states and emotional changes. Therefore, by encoding the EEG signal in the frequency domain, the model can more effectively extract important frequency components in the EEG signal, thereby better understanding the patterns of brain activity.
[0029] In step S130, the frequency domain features of the original signal corresponding to the first original EEG data sample are extracted, and each visible EEG segment is reconstructed according to the visible EEG simulated frequency domain coding features to calculate the frequency domain reconstruction loss.
[0030] In some embodiments, appropriate signal processing methods (such as Fourier transform, wavelet transform, etc.) are used to extract frequency domain features of the first original EEG data sample. These features represent the intensity distribution of the EEG signal in different frequency bands and reflect the changes in different brain regions and activity states. For example, the EEG signal may contain low-frequency theta waves, alpha waves, and beta waves. These frequency bands are closely related to different cognitive and emotional states.
[0031] Then, the previously encoded visible EEG is used to simulate the frequency domain coding features and reconstructed through the decoding network. The decoding process reversely converts the frequency domain coding features into time domain signals to reconstruct each visible EEG segment. Specifically, the reconstruction process will restore the time domain signal of the segment as much as possible, thereby calculating the frequency domain reconstruction loss, and can use loss expression functions such as mean square error MSE or frequency domain feature difference to guide model optimization. As a result, the reconstruction process can accurately restore the masked part of the EEG signal, and the introduction of frequency domain features enables the model to better grasp the key frequency band information of the signal and reduce reconstruction errors.
[0032] In some examples of the embodiments of the present application, the frequency domain reconstruction loss is the reconstruction loss of the characteristic dimension corresponding to the frequency domain amplitude.
[0033] Specifically, in the frequency domain, the amplitude of the EEG signal is usually more informative than the phase information, especially in the spectral density and energy distribution of the EEG. By taking the frequency domain amplitude feature dimension as the target of reconstruction loss optimization, the model can focus on restoring the amplitude component of the EEG signal in the frequency domain, especially the amplitude changes in important frequency bands (such as alpha waves, beta waves, theta waves, etc.), thereby ensuring that the reconstructed EEG signal is not only continuous in time, but also maintains important cognitive information in the frequency domain.
[0034] In addition, the inventors of the present application also conducted relevant ablation experiments in the process of practicing the present application, which further proved that, as shown in Table IV, the frequency domain amplitude is the best simulation target.
[0035] In step S140 , the visible EEG simulated frequency domain coded features are decoded to predict the codebook basis class corresponding to each masked EEG segment, and the classification loss is calculated in combination with the codebook basis class labels of each masked EEG segment.
[0036] Here, the decoder type can be diverse, and a lightweight decoder based on the Transformer layer can be used to predict the codebook basis class of the masked EEG segment by processing the frequency domain coding features of the visible EEG simulation, thereby predicting the label category (such as brain wave pattern, emotional state, etc.) that the masked EEG segment may correspond to. In addition, the classification loss usually uses various unrestricted loss functions (for example, cross-entropy loss function) to evaluate the difference between the predicted label and the true label.
[0037] In step S150, the EEG model is trained and optimized according to the frequency domain reconstruction loss and the classification loss.
[0038] Here, by optimizing both the frequency domain reconstruction loss and the classification loss, the model can not only perform high-quality restoration of EEG signals, but also obtain more semantically informative representations in classification tasks. The frequency domain reconstruction loss enables the model to focus on restoring the frequency domain features of the signal, while the classification loss helps the model learn task-related label information. This dual loss function design effectively improves the convergence speed of training and the robustness of the model, and can better adapt to the diverse EEG data features.
[0039] Through the embodiment of the present application, the joint optimization of frequency domain reconstruction and classification loss is introduced, which breaks through the limitations of existing large-scale EEG models in terms of original signal repair and generalization capabilities, improves the quality of EEG signal repair, and effectively enhances the cross-device and cross-task adaptability of the model. At the same time, the self-supervised learning strategy based on EEG fragment frequency domain reconstruction further improves the learning efficiency and robustness of the model, and is applicable to data from a variety of EEG acquisition devices and different experimental paradigms, providing a more general and efficient solution for large-scale EEG signal processing and analysis.
[0040] Regarding the implementation details of step S120, in some embodiments, each visible EEG segment is encoded based on a layer fusion encoder, and the encoding information output by multiple layers of encoder layers in the layer fusion encoder is cross-layer fused to determine the corresponding visible EEG simulated frequency domain coding features.
[0041] It should be noted that cross-layer fusion can utilize the information complementarity between different levels. For example, the shallow layer contains more low-level information than the deep layer, thereby strengthening the model's comprehensive understanding of the multi-dimensional characteristics of the EEG signal, generating more accurate and semantically rich frequency domain coding features, and providing a more comprehensive EEG signal representation.
[0042] Regarding the details of the codebook base class labels of each masked EEG segment in step S140, in one example, they may be predefined label information; in another example, they may also be label information adaptively determined by other means.
[0043] In some embodiments, each first original EEG segment in the first original EEG segment set is encoded to determine the corresponding first original EEG segment encoding features. Then, the basic class matching the encoding features of each first original EEG segment is retrieved from the codebook to determine the codebook basic class label corresponding to each first original EEG segment. Specifically, a plurality of basic classes are recorded in the codebook, and the distance between the encoding features of the first original EEG segment and each basic class is calculated to obtain a matching basic class, which is used as the corresponding codebook basic class label to guide the model to learn the specific category or pattern to which each EEG segment belongs. It should be understood that the goal of feature encoding at this time is to achieve an accurate match of the basic class in the codebook, and its encoder structure can adopt the same or different encoding structure as the encoder for the visible EEG segment (e.g., layer fusion encoder) mentioned above.
[0044] Furthermore, in order to fully guarantee the matching accuracy of the basic classes in the codebook, the EEG model can be pre-trained in advance to ensure that the EEG model can fully learn the relationship between the codebook and the EEG segments.
[0045] Figure 2An operational flowchart of an example of pre-training an EEG model according to an embodiment of the present application is shown.
[0046] like Figure 2 As shown, in step S210, the second original EEG data samples are segmented into a second original EEG segment set, and each second original EEG segment in the second original EEG segment set is encoded to determine corresponding second original EEG segment encoding features.
[0047] In some implementations, a time-domain tokenizer encoder is used to encode each second raw EEG segment in the set of second raw EEG segments.
[0048] Exemplarily, the second original EEG data sample is segmented according to a set time window, each window corresponds to an EEG segment, which usually contains a fixed-length EEG data, and the common length is several seconds to tens of seconds. Each EEG segment is encoded by a time domain segmenter encoder to extract its high-dimensional feature representation to capture more fine-grained time domain and frequency domain information in the EEG signal.
[0049] In some cases, frequency domain analysis (such as FFT or wavelet transform) is used as input features to enhance the frequency domain representation of the signal, or self-attention mechanisms (such as Transformers) are adopted to capture long-term dependencies.
[0050] In step S220, for each second original EEG segment encoding feature, a base class matching the second original EEG segment encoding feature is retrieved from the codebook, and the base class is decoded to reconstruct the corresponding second original EEG segment to calculate the codebook-segment reconstruction loss.
[0051] In some embodiments, a time-domain tokenizer decoder is used to decode the base class to reconstruct the corresponding second original EEG segment. A codebook is a predefined set of vectors, and each base class represents a common EEG signal pattern or activity category. Each encoded EEG segment corresponds to a high-dimensional feature vector, and the model retrieves the base class that best matches the feature vector from the codebook, such as a method based on vector similarity (such as Euclidean distance, cosine similarity, etc.).
[0052] Here, the retrieved basis class is used as the input of the time domain word segmenter decoder, which is converted into the reconstructed signal of the EEG segment, and the high-dimensional basis class information is restored to a signal that is as similar as possible to the second original EEG segment. Then, the decoded signal is compared with the second original EEG segment, and the difference is calculated and optimized as the reconstruction error.
[0053] Specifically, the time-domain word segmenter encoder and the time-domain word segmenter decoder adopt a symmetrical structure and are both composed of convolution blocks and Transformer blocks. The convolution blocks are used to extract local features from the original EEG signal, and the self-attention mechanism of the Transformer blocks is used to capture the long-term dependencies in the signal.
[0054] It should be noted that the time-domain word segmenter encoder and the time-domain word segmenter decoder adopt a symmetrical structure, which is crucial for reconstructing the original EEG segment. The time-domain word segmenter encoder is responsible for converting the input EEG signal into compact high-dimensional features, while the time-domain word segmenter decoder recovers the output from these features as close to the original EEG signal as possible. This design ensures that the signal representation does not lose information during the processing process, and the encoding and decoding are synchronized, which can achieve high-quality signal reconstruction.
[0055] In step S230, the EEG model is pre-trained according to the codebook-segment reconstruction loss.
[0056] Here, during the pre-training phase of the EEG model, the goal of the model is to optimize the parameters of the network by minimizing the codebook-fragment reconstruction loss. At this stage, the model mainly focuses on learning how to recover fragments of the EEG signal from the codebook basis class. Specifically, during the training process, the weights of the encoder and decoder are updated through the backpropagation algorithm to reduce the reconstruction error. By optimizing the reconstruction loss, the model can continuously adjust its encoder and decoder, improve the feature extraction ability of the encoder, and optimize the decoder's ability to accurately recover the corresponding EEG signal from the codebook basis class.
[0057] In some examples of the embodiments of the present application, the visible EEG simulated frequency domain coding features are decoded based on a lightweight decoder to predict the codebook base class corresponding to each masked EEG segment. For example, the lightweight decoder can be composed of multiple Transformer layers, which can quickly infer the codebook base class corresponding to the masked EEG segment, ensuring the reasoning speed. In addition, the time domain word segmenter decoder is also used to decode the repaired EEG segment corresponding to the predicted codebook base class to further reconstruct the masked EEG segment, realize the repair of the masked EEG segment, and maintain the time domain consistency of the reconstructed EEG sample signal. The quality of signal repair is improved through the self-attention mechanism to ensure that the waveform characteristics and long-term dependencies of the signal are well restored.
[0058] The following will use examples to expand on the details of the general EEG model that takes into account both the original EEG signal restoration and classification tasks.
[0059] It should be noted that existing LEMs are relatively scarce and their potential in data reconstruction tasks has not been fully explored. In addition, how to efficiently integrate the time domain and frequency domain features of EEG data has always been a major challenge in research. In this paper, Gram, a general large-scale EEG model for raw EEG data classification and repair tasks, is proposed. The Gram model consists of two stages: 1) In the first stage, the raw EEG data is quantized into "basic classes" containing rich temporal information; 2) In the second stage, a masked autoencoder with multi-view layer fusion is used to model the time domain and frequency domain features of EEG data through dual training objectives: for visible segments, the frequency domain reconstruction objective is used; for masked segments, the "basic class" classification objective is used. Gram was pre-trained on 7000 hours of EEG data and achieved state-of-the-art performance in three cross-subject classification tasks (including event, emotion, and sleep stage classification). In the EEG data repair task, the model provided in the embodiment of the present application significantly improves the classification performance by repairing the damaged data, which is better than directly using noisy data.
[0060] The core advantage of the Gram model lies in its dual capabilities: classification and restoration of raw EEG signals. Here, the restoration process of EEG signals is divided into two stages. The first stage is base class quantization, which is to divide the raw EEG signal into multiple small blocks and build base classes based on these blocks. In the second stage, Gram performs data generation and high-order feature learning through an autoencoder structure with multi-view layer fusion. This process can not only help reconstruct EEG signals, but also perform effective tasks such as emotion recognition, event detection, and sleep stage classification.
[0061] Figure 3 The flowchart of the operation of the first stage of training of the Gram model according to an embodiment of the present application is shown.
[0062] like Figure 3 As shown, the first stage of training of the Gram model specifically includes the following steps: a) Data preparation. The original EEG signal is processed by 0-70Hz bandpass filtering and 50Hz or 60Hz notch processing. The original EEG signal is sampled to 200Hz. The original EEG signal is processed in segments.
[0063] b) Send it to the encoder. Send it to the encoder to get the encoder representation.
[0064] c) Get the base class. Find the "base class" in the codebook that is closest to the encoder representation.
[0065] d) Send to the decoder. Send the corresponding basic class to the decoder to realize the construction and learning of the codebook.
[0066] e) Calculate the loss. Use the decoder output and the original EEG signal to calculate the reconstruction loss and perform back propagation. If the specified number of iterations is reached, the iteration is stopped.
[0067] Figure 4 The flowchart of the operation of the second stage training of the Gram model according to an embodiment of the present application is shown.
[0068] like Figure 4 As shown, the second stage of training of the Gram model specifically includes the following steps: A) Data preparation. The original EEG signal is processed by 0-70Hz bandpass filtering and 50Hz or 60Hz notch processing. The original EEG signal is sampled to 200Hz. The original EEG signal is processed in segments. The processed EEG is randomly masked with a masking rate of 50%.
[0069] B) Feed the encoder. The visible signal is fed into the layer fusion encoder. The encoder fuses the outputs of multiple layers through learnable weights and finally obtains the visible encoder representation.
[0070] C) Calculate the frequency domain reconstruction loss. The reconstruction loss is calculated using the visible encoder representation and the original signal frequency domain amplitude.
[0071] D) Send to decoder. Send the visible encoder representation to the decoder to obtain the decoder representation of the masked EEG.
[0072] E) Calculate the classification loss. Use the decoder representation to predict the underlying class of the original EEG and calculate the classification loss.
[0073] F) Calculate the loss sum. Add the frequency domain reconstruction loss and the classification loss for back propagation. If the specified number of iterations is reached, the iteration is stopped.
[0074] Through experimental verification, different scales of the Gram model were compared and it was found that its performance in the EEG classification task improved significantly with the increase of the model scale, thus proving the positive impact of the model scale on the classification performance.
[0075] Therefore, the Gram model and corresponding training method provided in the embodiment of the present application solve the shortcomings of the current related technologies in EEG signal repair and classification tasks, provide a more general and adaptable EEG signal processing model, and verify its superior performance in different tasks through detailed experiments. This innovation will provide important technical support for the research and application in the field of EEG signal processing.
[0076] I. Introduction The embodiments of the present application propose Gram, a general large-scale EEG model for raw EEG data classification and reconstruction tasks. Gram's dual capabilities rely on the concept of discretization. A pioneering insight in image and text generation is to discretize the input signal and construct a corresponding codebook. The generation of an image or text involves repeatedly permuting and combining tokens from a codebook. Natural Language Processing (NLP) uses various word segmentation methods, and images are discretized based on pixels or fragments. Drawing on these concepts, the EEG signal is segmented into multiple fragments and a codebook is created. Each token in the codebook is called a "basis class." Similar to the methods of images and NLP, it is proposed that EEG signals can be constructed by repeated permutations and combinations of different basis classes.
[0077] Gram is pre-trained on 7,000 hours of EEG data from more than 15 datasets and is adaptable to three cross-subject downstream tasks: event classification, emotion classification, and sleep stage classification. It is suitable for raw EEG representation learning and data repair tasks. In order to study the impact of model scale on classification performance, the embodiment of the present application provides three Gram variants with different parameter scales, ranging from 6M to 251.28M. Gram consists of two stages: basis class quantization and multi-view layer fusion masked autoencoder. In the first stage, the original EEG fragments are classified into different basis classes and then reconstructed from these classes. In the second stage, the multi-view layer fusion masked autoencoder performs dual functions: predicting the basis class to facilitate data generation, and learning high-level representations for classification tasks.
[0078] II. Methodology A. Original signal segmentation It should be noted that the EEG signal can be composed of repeated permutations of segments corresponding to the basic classes. Therefore, the EEG signal needs to be segmented into segments. First, the input EEG signal is represented as ,in Indicates the number of channels, is the input time length. For each channel, the data is segmented along the time axis with length After splicing the fragments of all channels, we can get fragments, each fragment is represented by . Reorganized as Using this segmentation approach, Gram can be adapted to any EEG setting, since changes in the number of channels or duration only result in Modifications.
[0079] B. Basic Class Quantization Figure 5A schematic diagram of the operational flow of an example of basic class quantization of an EEG model trained in one stage according to an embodiment of the present application is shown.
[0080] The goal of this stage is to classify EEG segments into base classes. Here, the tokens in the codebook are called base classes. Figure 5 As shown, the time-domain tokenizer encoder and quantization module are responsible for classifying and compressing the original EEG segments into base classes, while the time-domain tokenizer decoder is responsible for reconstructing the original EEG segments from these base classes.
[0081] 1) Time-domain tokenizer encoder and decoder: A symmetrical time-domain tokenizer encoder and decoder structure is used. Both are composed of convolutional blocks and Transformer blocks.
[0082] Fragment embedding and decoding. Fragment embedding consists of two steps. The first step is to extract the time domain features of each fragment using a convolution block, which consists of a one-dimensional convolution, a group normalization layer, and a GELU activation function. The second step is to add position embedding to the fragment embedding. Here, fixed sine-cosine spatial and time domain embeddings are used. Since the channel and time domain dimensions of the input are flattened, the model requires additional position-related position information to identify the channel and time point of each fragment. Fragment decoding utilizes exactly the same convolution structure as fragment embedding. This symmetry is crucial for reconstructing the original EEG fragment. Experimental results show that fragment embedding and fragment decoding must work together. Without this collaboration, the loss function cannot converge, hindering the signal reconstruction process.
[0083] Transformer. The Transformer blocks are the same in the encoder and decoder. The Transformer encoder structure and some regularization methods are applied to stabilize the training, including modifying the calculation of the attention mechanism to , Formula (1) in, is the dimension of a single head, Set to 32.
[0084] 2) Quantification Vector quantization is used to represent the segments from the time-domain tokenizer encoder. Maps to repeated permutations of base classes .matrix from codebook The tokens are composed of. Among them, and They represent the embedding dimensions, and Represents the number of fragments and the number of base classes respectively. The index of the base class is found by searching for the same The closest token is used to determine.
[0085] , Formula (2) in, ,and From The linear projection of and performing index lookup in a low-dimensional space helps to enhance the utilization of the codebook.
[0086] In addition, and application Normalization can also produce a similar effect. The loss function of vector quantization is as follows: , Formula (3) in, Indicates stopping the gradient operation.
[0087] The time-domain tokenizer decoder aims to reconstruct the original EEG segment. The output of the decoder is represented as . Use cosine distance to measure and The similarity between them. Reconstruction loss And the loss of the first stage is as follows: , Formula (4) , Formula (5) C. Multi-view layer fusion mask autoencoder Figure 6 A schematic diagram of the operational flow of an example of two-stage training of an EEG model based on a layer-fusion masked autoencoder according to an embodiment of the present application is shown.
[0088] The training goal of the second stage is to obtain efficient EEG representations for downstream tasks and identify the underlying categories of the data that need to be repaired. To achieve these goals, two techniques are used: layer fusion encoder and multi-view training objectives. Figure 6 As shown in Figure 2, the proposed autoencoder consists of a layer-fusion encoder and a lightweight decoder consisting of only Transformer layers. The structure of the layer-fusion encoder is similar to the time-domain tokenizer encoder, with the main difference being the layer-fusion module. 50% of the EEG segments are randomly masked and the visible 50% of the EEG segments are passed to the layer-fusion encoder to obtain the representation ,in is the number of visible fragments. Different from using mask tokens in MAE, here, Connect with additional class tokens.
[0089] 1) Layer Fusion Encoder The shallow layers of MAE contain more low-level information than the deep layers. Fusion of outputs from different layers can effectively improve model performance. Here, linear projection is used and a feature fusion method that combines The selected layer in the encoder The output of the layers are fused.
[0090] , Formula (6) in, is assigned to The learnable weights of the layer, and , It is The output of the layer.
[0091] 2) Multi-view training objectives Efficiently integrating the time and frequency domain information of EEG data has always been a focus of EEG modeling research. Reconstructing the two types of targets directly after the decoder will lead to semantic conflicts. The encoder-decoder framework of MAE essentially provides two independent reconstruction paths, thus avoiding these conflicts. Therefore, the frequency domain amplitude of the visible EEG segment is simulated after the layer fusion encoder, thereby incorporating the frequency domain information into the encoder to enhance the representation learning of EEG data. As shown in Table IV, the frequency domain amplitude is the best simulated target. Then, the base class category of the masked EEG segment is predicted after the decoder, which serves as the basis for data generation.
[0092] By predicting the basis class, it is possible to integrate high-level temporal information within the autoencoder framework, because the basis class encapsulates different temporal features. It can be seen that the frequency domain amplitude of the EEG segment is represented as .
[0093] Then, the simulated loss is obtained , classification loss and the second-stage losses as follows: , Formula (7) , Formula (8) , Formula (9) in, is the number of masked fragments, is the probability predicted by the network conditioned on the seen segment.
[0094] D. Generator and Classifier In terms of generation, the autoencoder in stage 2 predicts the base class of the fragment to be restored, and the time-domain word segmenter decoder reconstructs the original fragment corresponding to the predicted base class, and both modules are frozen. It should be noted that although the base class is fixed, position embedding is still required to reconstruct the original EEG fragment through the base class. In other words, channel (spatial) and time-domain position embeddings affect the final generation results. Even if fragments from different channels and different time points are classified into the same base class, the reconstructed fragments will still have slight differences.
[0095] Representation learning is achieved by fine-tuning the layer-fusion encoder pre-trained in the second stage. Classification is performed using class tokens and a linear classification head.
[0096] III. Experiment Specifically, about 7,000 hours of electroencephalogram (EEG) data from more than 15 datasets with different settings (layout, duration, equipment, paradigm, etc.) are collected for pre-training. For fine-tuning, Table I shows the information of downstream tasks. The datasets used for these downstream tasks are different from those used for pre-training. Number of channels and input duration EEG signal from input . All datasets were resampled to 200 Hz. A test set was provided for TUEV. Further, the provided training set was randomly divided into training and validation sets by subject, with a ratio of 80% and 20%. For the other datasets, the subjects were randomly divided into training, validation and test sets, with a ratio of 60%, 20% and 20%, respectively. Here, Cohen's Kappa (Kappa) was used as the monitoring indicator, and the mean and standard deviation results were reported under 5 random seeds ([0, 10, 20, 30, 40]) (all random seeds were divided in the same way).
[0097] Table I: Fine-tuning dataset information
[0098] Here, 5 supervised learning models are used: SPaRCNet, ContraWR, FFCL, CNN-Trans, and ST-Trans; and 5 pre-trained self-supervised learning models are used: BIOT (56643 hours), MMM (27063 hours), LaBraM-Base (7000 hours), LaBraM-Huge (7000 hours), and BENDR (27063 hours). The duration of the pre-training dataset is given in parentheses. LaBraM was originally pre-trained on 2500 hours of data. For fair comparison, LaBraM has been retrained on the pre-training dataset.
[0099] IV. Results Analysis Table II: Performance of different algorithms on SEED-V and TUEV (%) Table II shows the performance of different algorithms on three downstream tasks. The results are divided into supervised models, self-supervised models, and Grams of different model sizes, where underlined, dashed underlined, and bold fonts indicate the best results. "B", "M", and "L" represent Gram's basic, medium, and large models, respectively. The results are sorted by model size.
[0100] A. Representation Learning It is clear from the third part of Table II that as Gram’s parameters increase, its performance also improves. Gram’s best results on various metrics outperform supervised and self-supervised learning methods on all datasets. If the model size is taken into account, all supervised learning methods, MMM, BIOT, LaBraM-Base, and Gram-B are roughly comparable in size. Gram-B achieves the best results on all datasets. BENDR, LaBraM-Huge, and Gram-L are also roughly comparable in size. Although Gram-L has nearly 100 million fewer parameters than LaBraM-Huge, Gram-L still outperforms LaBraM-Huge, especially on the TUEV dataset, where Gram-L outperforms LaBraM-Huge by 4.5% in monitoring metrics. In addition, the performance of training from scratch is compared with that of Gram-B, and the pre-training process brings significant improvement.
[0101] Ablation experiments were performed on all datasets using Gram-L. Table III shows the effectiveness of baseclass quantization (BCQ) (i.e., stage one), layer-fusion (LF) (i.e., step B of stage two), and spectral mimic (SM) (i.e., step C of stage two) modules. Removing the BCQ module means abandoning the use of the stage one model and directly reconstructing the original EEG signal of the masked segment in stage two. Table IV shows the impact of the simulation target selection. Here, an attempt is made to replace the simulation target from the frequency domain amplitude with other high-level representations learned by the model, including the output of LaBraM-B and the output of the stage one encoder. The parameters of these models are frozen and do not participate in training.
[0102] Table III: Ablation study of base class quantization (BCQ), layer fusion (LF) and frequency domain simulation (SM) modules Table IV: Ablation study of simulated target selection B. Raw data recovery Here, we propose that EEG signals can be composed by permuting different basis classes. We validate this in two ways: 1) visualizing and analyzing the correlation of basis classes, and 2) evaluating the data recovery effect through classification performance.
[0103] Figure 7 A schematic diagram of simulation effect showing an example of visualization and correlation analysis of basic classes.
[0104] Randomly select two clips from the TUEV dataset and Figure 4 The left side shows its reconstructed EEG data. Figure 7 The right side of the figure shows the results of the correlation analysis, which calculates the average correlation between the original EEG segments and the reconstructed EEG segments for each basis class in the three datasets. The number of base classes (outer ring) and original segments (inner ring) with average correlation ≥ 0.5, < 0.5, and unused are counted. This analysis shows that there is a high temporal correlation between the original EEG segments and the reconstructed segments of the corresponding base classes, indicating that Gram can effectively discretize the EEG into various basis classes based on temporal features.
[0105] Figure 8 A schematic diagram of an effect simulation of an example of comparing the classification results of noisy data and restored data is shown.
[0106] The effectiveness of model data repair is verified by comparing the classification performance of noisy data and repaired data. Noise is introduced in the EEG channel dimension. For example, if 10% of the channels are selected for noise or repair, all EEG segments associated with this channel need to be noised or repaired. The ratio of noisy data ranges from [0.005, 0.05, 0.1, 0.3,0.5, 0.7]. It is worth noting that the channels selected for noise addition and repair are exactly the same. In order to simulate the noise conditions in the real world, the training set, validation set, and test set are all subjected to the noise addition or data repair process.
[0107] Figure 8 shows the classification results of the Gram-B model using noisy and restored data. The data was restored using Gram-B, Gram-B without BCQ (as shown in Table III), and the spherical spline interpolation method. Figure 8 shows that while both spherical spline interpolation and Gram are effective, Gram consistently outperforms spherical spline interpolation. In addition, directly reconstructing the original EEG in the second stage does not perform as well as using BCQ, which confirms the advantage of discretizing the EEG into basis classes.
[0108] Based on the above research, the advanced nature of the Gram model provided in the embodiment of the present application is confirmed, that is, a general EEG model for raw data classification and repair tasks. Gram was pre-trained on more than 7,000 hours of EEG data and achieved state-of-the-art performance in three cross-subject tasks. However, large-scale EEG models still face many challenges, including the diversity of training data, expansion rules, and other unexplored areas, which are all planned research directions in the future.
[0109] Through the embodiment of the present application, by adopting a symmetrical convolution structure in the encoder and decoder, the original signal is successfully reconstructed in stage one, so that the stage one model can be used for original signal generation. In addition, by constructing a codebook, the original signal is discretized into different "basic classes", and the original EEG signal is regarded as a combination of different "basic classes", and finally the EEG data is generated. The data repair and classification of the original EEG signal are integrated into the same model, and better performance is obtained at the same time.
[0110] It should be noted that this study used publicly available human subject data for retrospective analysis. As the data used were open access and the accompanying license confirmed that no ethical approval was required, no ethical approval was required for this study.
[0111] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of actions combined, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0112] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned EEG model training methods of the present application.
[0113] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned EEG model training methods.
[0114] In some embodiments, the embodiments of the present application also provide an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the EEG model.
[0115] Fig. 9 is a schematic diagram of the hardware structure of an electronic device for executing a method for training an EEG model provided in another embodiment of the present application, such as Fig. 9 As shown, the device includes: One or more processors 910 and memory 920, Fig. 9 A processor 910 is taken as an example.
[0116] The device for executing the method for training an EEG model may further include: an input device 930 and an output device 940 .
[0117] The processor 910, the memory 920, the input device 930 and the output device 940 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.
[0118] The memory 920, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the training method of the EEG model in the embodiment of the present application. The processor 910 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 920, that is, the training method of the EEG model in the above method embodiment is implemented.
[0119] The memory 920 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 920 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 920 may optionally include a memory remotely arranged relative to the processor 910, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0120] The input device 930 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 940 may include a display device such as a display screen.
[0121] The one or more modules are stored in the memory 920, and when executed by the one or more processors 910, the training method of the EEG model in any of the above method embodiments is executed.
[0122] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.
[0123] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.
[0124] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.
[0125] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0126] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.
[0127] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for training an electroencephalogram model, comprising: Segmenting the first original EEG data sample into a first original EEG segment set, masking some of the original EEG segments in the first original EEG segment set to obtain corresponding visible EEG segment subsets and masked EEG segment subsets; encoding each visible EEG segment in the subset of visible EEG segments to determine a corresponding visible EEG simulated frequency domain encoding feature; Extracting the original signal frequency domain features corresponding to the first original EEG data sample, and reconstructing each of the visible EEG segments according to the visible EEG simulated frequency domain coding features to calculate the frequency domain reconstruction loss; Decoding the visible EEG simulated frequency domain coded features to predict the codebook basis class corresponding to each masked EEG segment, and calculating the classification loss in combination with the codebook basis class labels of each masked EEG segment; The EEG model is trained and optimized according to the frequency domain reconstruction loss and the classification loss.
2. The method according to claim 1, wherein: The frequency domain reconstruction loss is the reconstruction loss of the characteristic dimension corresponding to the frequency domain amplitude.
3. The method according to claim 1, wherein: The encoding of each visible EEG segment in the visible EEG segment subset to determine the corresponding visible EEG simulated frequency domain coding features includes: Each of the visible EEG segments is encoded based on a layer fusion encoder, and the encoding information output by multiple layers of encoder layers in the layer fusion encoder is cross-layer fused to determine the corresponding visible EEG simulated frequency domain encoding features.
4. The method according to claim 1, before calculating the classification loss by combining the codebook basis class labels of each of the masked EEG segments, the method further comprises: encoding each first original EEG segment in the first original EEG segment set to determine a corresponding first original EEG segment encoding feature; The base class matching the encoding features of each of the first original EEG segments is retrieved from the codebook to determine the codebook base class label corresponding to each of the first original EEG segments.
5. The method according to claim 4, further comprising: segmenting the second original EEG data sample into a second original EEG segment set, and encoding each second original EEG segment in the second original EEG segment set to determine a corresponding second original EEG segment encoding feature; For each second original EEG segment encoding feature, retrieving a base class matching the second original EEG segment encoding feature from a codebook, and decoding the base class to reconstruct the corresponding second original EEG segment to calculate a codebook-segment reconstruction loss; The EEG model is pre-trained according to the codebook-segment reconstruction loss.
6. The method according to claim 5, wherein: The step of encoding each second original EEG segment in the second original EEG segment set comprises: encoding each second original EEG segment in the second original EEG segment set using a time domain tokenizer encoder; The decoding of the base class to reconstruct the corresponding second original EEG segment comprises: Decoding the base classes using a time-domain tokenizer decoder to reconstruct a corresponding second raw EEG segment; Among them, the time-domain word segmenter encoder and the time-domain word segmenter decoder adopt a symmetrical structure, and both are composed of convolution blocks and Transformer blocks.
7. The method according to claim 6, wherein: The decoding of the visible EEG simulated frequency domain coding features to predict the codebook basis class corresponding to the first original EEG data sample includes: Decoding the visible EEG simulated frequency domain coded features based on a lightweight decoder to predict the codebook basis class corresponding to each masked EEG segment; The time-domain word segmenter decoder is also used to decode the repaired EEG segment corresponding to the predicted codebook basis class.
8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Method for predicting driving electroencephalogram driving intention based on fusion time-frequency characteristic large model
CN120837098A
Electroencephalogram signal processing method and device and computer equipment
CN121465609A
User intention recognition method and device and terminal equipment
CN121996077A
Epileptic seizure event detection method and device, equipment and storage medium
CN122440131A