Method and device for training electroencephalogram recovery model and storage medium
Through the EEG recovery model of mask processing and spatiotemporal attention module training, the quality degradation of EEG signals under noise and artifact interference is solved, and efficient signal recovery and robustness are achieved.
Patent Information
- Application Number
- CN202510122081.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively solve the problem of the quality of EEG signals degraded under noise and artifact interference, especially in non-laboratory environments.
The EEG recovery model is trained by segmenting the original EEG samples into a fragment set and masking some fragments using mask signal channels and time. The model includes an encoder and a symmetric decoder, which uses the temporal attention module and the spatial attention module to model the spatiotemporal characteristics, and uses the symmetric decoder to perform signal reconstruction.
It significantly improves the robustness and recovery quality of EEG signals, can effectively deal with noise and artifact interference, and improves the accuracy and stability of signal recovery.
Smart Images

Figure CN120012834A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of brain-computer interface technology, and in particular to a training method, device, storage medium and program product for an electroencephalogram recovery model. Background Art
[0002] Electroencephalo gram (EEG) can provide valuable insights into the temporal dynamics of brain function and cognitive processes. As an objective physiological signal, EEG has shown significant potential in diagnosing psychiatric disorders, affective computing, and motor imagery recognition.
[0003] However, there are considerable challenges in applying EEG to non-laboratory environments such as home or mobile. One of the key limitations is the high sensitivity of EEG to noise and artifacts. Since the EEG signal is inherently weak, it is easily disturbed by physiological activities, including eye movements, heart rhythms, and muscle contractions, as well as external factors such as cable movement and electromagnetic interference. In addition, body movement or improperly worn hats can cause electrode displacement, resulting in poor EEG recording quality. These factors can significantly degrade system performance when applying EEG-based systems from controlled laboratory conditions to real-world environments.
[0004] To address the challenges of low-quality EEG signals, researchers have adopted various methods to reduce the impact of excessive noise and data corruption on system performance. Some methods involve eliminating faulty channels or rejecting contaminated samples, but simply discarding bad channels can lead to a loss of usable information. Interpolation is a commonly used technique to reconstruct bad channels, but due to the complex spatiotemporal characteristics of EEG, interpolation often fails to capture detailed patterns, resulting in significant inaccuracies, especially when large parts of the signal are missing, which exacerbates the magnitude of the deviation.
[0005] The industry has not yet proposed a better solution to the above problems. Summary of the invention
[0006] The present application provides a training method, device, storage medium and program product of an electroencephalogram (EEG) restoration model, which are used to at least solve the key problem of quality degradation of EEG signals due to interference from noise and artifacts.
[0007] In a first aspect, an embodiment of the present application provides a training method for an EEG recovery model, comprising: dividing an original EEG sample into an original EEG segment set, and masking part of the original EEG segments in the original EEG segment set according to a mask signal channel and a mask signal time to obtain a corresponding masked EEG segment set; processing the masked EEG segment set based on an encoder to obtain an encoded EEG segment set for the masked part of the EEG segments; the encoder comprises a temporal attention module and a spatial attention module; processing the encoded EEG segment set and the masked EEG segment set based on a symmetric decoder to generate a reconstructed EEG sample, and calculating a corresponding sample reconstruction loss in combination with the original EEG sample; and training and optimizing the EEG recovery model based on the sample reconstruction loss.
[0008] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the training method of the EEG recovery model of any embodiment of the present application.
[0009] In a third aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the training method of the EEG recovery model of any embodiment of the present application are implemented.
[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the training method of the EEG recovery model of any embodiment of the present application.
[0011] The beneficial effects of the embodiments of the present application are: By dividing the original EEG samples into a set of segments and using masking operations, the partial missing and artifact situations of EEG signals in real environments are simulated. The masked EEG segment set is used to train the EEG recovery model, so that the model can effectively learn and recognize the spatiotemporal characteristics of EEG signals, thereby enhancing the robustness to noise and artifacts. Through the temporal attention module and spatial attention module introduced in the encoder, the spatiotemporal characteristics are jointly modeled, which can accurately restore the masked EEG signal. Furthermore, through the design of a symmetric decoder, the masked EEG segment set and the encoded EEG segment set are combined for signal reconstruction, so that the model can make full use of the information of the masked area and the unmasked area when restoring the signal, which can significantly improve the signal recovery quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 A flowchart showing an example of a method for training an EEG restoration model according to an embodiment of the present application is shown; Figure 2 Shown according to Figure 1 An exemplary operation flow chart of step S110 in FIG. Figure 3 A schematic diagram of the architecture connection of an example of a spatiotemporal autoencoder for EEG signal restoration according to an embodiment of the present application is shown; Figure 4 An example of experimental effect simulation diagram showing the classification accuracy of damaged and restored EEG data on different classifiers and data sets; Figure 5 A schematic diagram showing the effect simulation of an example of a reconstruction topology of γ-band test data at a single time point under different damage levels; Figure 6 It is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0015] It should be noted that in the current related technologies, researchers have tried many ways to restore missing EEG signals, such as interpolation methods, dynamic spatial filtering modules, neural processes and channel autoencoders, etc.
[0016] In the interpolation method, the approximate value of the missing signal is calculated by interpolation. For example, the spherical interpolation method is based on the spatial distribution of EEG electrodes, while the nearest neighbor interpolation is based on the calculation of signal values at adjacent moments. Specifically, the damaged channel or time point of the EEG signal is detected, and the interpolation result is calculated based on the signal value of the adjacent channel or time point, and the interpolation result is filled into the damaged part. However, the interpolation method has difficulty in capturing the complex spatiotemporal dynamic characteristics of the EEG signal, and can only be reconstructed based on local information, with low accuracy. When the damage range is large, the reconstruction error of the interpolation method increases significantly, which limits the practical application.
[0017] In the dynamic spatial filtering module, the EEG channel data is dynamically weighted to filter out noise and enhance useful signals. Specifically, the EEG signal is collected and abnormal channels are detected, the filter parameters are dynamically adjusted to reduce the impact of noise, and the filtered signal is extracted for emotion recognition model training. However, the dynamic spatial filtering module mainly targets signal damage in the channel dimension and lacks the ability to process damage in the time dimension. It cannot capture the global characteristics of the EEG signal and can only alleviate noise interference, and its adaptability is insufficient.
[0018] In the NP (Neural Process) and ANP (Attentive Neural Process) models, the missing signal is inferred through probability distribution. NP is based on Gaussian process, and ANP adds attention mechanism to enhance performance. Specifically, the model is trained with complete signal samples, the statistical characteristics of the signal are learned, the damaged signal is inferred, and the repair result is generated, which is then combined with the inferred signal for subsequent emotion recognition tasks. However, the NP and ANP models have high model complexity and strong dependence on training data, making it difficult to generalize to a variety of damage scenarios, and their ability to recover from damage in the time dimension is limited.
[0019] In particular, interpolation methods are limited by their local computational approach and fail to consider the global spatiotemporal dependencies of EEG signals. The limitation of the dynamic spatial filtering module is that it only focuses on the channel dimension, while the temporal dimension of EEG signals is equally important. The NP and ANP models are unstable in scenarios with different data distributions due to their high training data requirements, and their design does not consider the problem of temporal dimension corruption.
[0020] In the current related technologies, researchers are increasingly using deep learning methods to restore missing EEG. Some researchers have proposed a gated layer autoencoder to reconstruct incomplete EEG; however, this method only solves a limited range of EEG damage scenarios. Banville et al. proposed a dynamic spatial filtering module that enhances the robustness to damaged EEG channels and significantly improves performance in noisy environments. In addition, some researchers have proposed a robust neural process variant that improves the stability of EEG-based emotion recognition. In addition, some experts and scholars have developed a dynamic autoencoder that uses mask channel modeling to enhance the robustness of EEG-based emotion recognition in the face of damaged channels. Nevertheless, these studies only focus on the channel damage problem in EEG, ignoring the equally common problem of temporal damage, which often occurs in real-world scenarios and also limits the effectiveness of EEG applications.
[0021] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of the present application, and is not to be regarded as a limitation of the present application. In addition, the technical solutions described in the above-mentioned current related art are not prior art, and they may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.
[0022] In view of the above-mentioned deficiencies in the current related technologies studied, in order to restore EEG signals with different degrees of damage in the channel and time dimensions in real-world scenarios, a spatial-temporal autoencoder for EEG Restoration (STAR) for EEG signal repair is proposed in an embodiment of the present application. STAR uses dynamic mask channels and time modeling to simulate a variety of real-world signal degradation and repair processes. The spatiotemporal attention mechanism combines spatial and temporal information to capture the complex relationships in EEG signals. The embodiment of the present application uses three public EEG datasets for emotion recognition for system evaluation. Experimental results show that STAR can effectively reconstruct EEG signals affected by different degrees of channel and time distortion, significantly improving the performance of the emotion recognition model when facing damaged EEG data.
[0023] Figure 1 An operational flowchart of an example of a method for training an EEG restoration model according to an embodiment of the present application is shown.
[0024] like Figure 1 As shown, in step S110, the original EEG samples are segmented into original EEG segment sets, and part of the original EEG segments in the original EEG segment sets are masked according to mask signal channels and mask signal times to obtain corresponding mask EEG segment sets.
[0025] In some embodiments, the original EEG sample is divided into several EEG patches according to the time axis and the channel axis to form an EEG patch set. Each EEG patch contains signal fragments of several channels within a specific period of time, ensuring that the segmented fragments not only maintain the spatiotemporal characteristics of the EEG signal, but also facilitate batch processing of the model. The patch size can be adjusted according to specific task requirements, for example, the time window length can be tens of milliseconds to hundreds of milliseconds.
[0026] In addition, a certain proportion of EEG segments can be masked in a variety of ways, such as masking in a predetermined order or masking in a random order, and the model's adaptability to different degrees of loss can be controlled by setting the masking ratio (such as 20% or 50%). The masked EEG segment set is called the masked EEG segment set, which contains both masked EEG segments and unmasked visible EEG segments.
[0027] In this way, by masking the EEG signal in the time and channel dimensions, it is possible to construct a training data set that simulates real-life scenarios (including problems such as signal loss and artifacts), thereby enhancing the model's attention to the spatiotemporal correlation of EEG signals.
[0028] In step S120, the masked EEG segment set is processed based on an encoder to obtain an encoded EEG segment set for the masked partial EEG segment, wherein the encoder includes a temporal attention module and a spatial attention module.
[0029] Here, the encoder contains two core modules: the temporal attention module and the spatial attention module. The temporal attention module models the dynamic correlation of EEG signals in the time dimension and captures the characteristics of key time points. The spatial attention module learns the spatial correlation between different channels and understands the potential coupling relationship between electrodes to improve the comprehensive processing capability of multi-channel signals.
[0030] It should be understood that the dual attention mechanism can be in various forms, such as simultaneously modeling time and channel attention, or modeling them alternately.
[0031] In some embodiments, the temporal attention module and the spatial attention module are cascaded. Specifically, first, the temporal attention module encodes the time series features in each segment to generate a feature representation of the time dimension. Then, the spatial attention module is used to encode the correlation between different channels, extract the spatial features in the channel dimension, and then generate a set of encoded EEG segments for the masked segments.
[0032] It should be noted that temporal attention and spatial attention can be implemented based on self-attention mechanisms (such as Transformer). By calculating the attention weight of each position or channel, key information can be extracted. By jointly modeling the temporal and spatial characteristics of the encoder, the intrinsic temporal and spatial correlation of the EEG signal can be fully exploited.
[0033] In step S130 , the encoded EEG segment set and the masked EEG segment set are processed based on the symmetric decoder to generate reconstructed EEG samples, and the corresponding sample reconstruction loss is calculated in combination with the original EEG samples.
[0034] Here, the structural design of the symmetric decoder is mirror-symmetric to the encoder, that is, it also includes a time decoding module and a space decoding module. The time decoding module is responsible for restoring the temporal dynamic characteristics of the signal, and the space decoding module is responsible for reconstructing the spatial distribution pattern between channels. Exemplarily, the encoded EEG segment set and the masked EEG segment set are combined and input into the decoder, so that the unmasked part of the masked EEG segment set can provide supplementary information for the decoding process, and then output the reconstructed EEG sample. Therefore, by adopting a decoder with a spatiotemporal symmetric structure, the feature information extracted by the encoder can be effectively restored to a complete EEG signal, and the unmasked part of the masked EEG segment is used to enhance the signal reconstruction effect of the reconstructed EEG sample.
[0035] Then, the original EEG sample is compared with the reconstructed EEG sample to calculate the corresponding sample reconstruction loss through the reconstruction loss function. It should be understood that the types of reconstruction loss functions can be diverse, such as using the mean square error (MSE) loss to measure the overall error between the reconstructed EEG sample and the original EEG sample, using the spatiotemporal weighted loss to set weights for the time and channel dimensions respectively so that the model pays more attention to important time points and channels, and so on.
[0036] In step S140, the EEG restoration model is trained and optimized based on the sample reconstruction loss.
[0037] In some embodiments, the masked EEG segment set is used as input to reconstruct the EEG sample and calculate the sample reconstruction loss. The model parameters are optimized by the back propagation algorithm so that the loss function value is gradually reduced. This ensures that the model can efficiently learn the spatiotemporal characteristics of the EEG signal and gradually improve the signal reconstruction quality.
[0038] As a further optimization of the embodiment of the present application, regarding the implementation details of step S140, the reconstructed EEG sample is determined to be real or generated based on the discriminator, and the corresponding discriminator loss is calculated. Then, the EEG recovery model is trained and optimized according to the sample reconstruction loss and the discriminator loss.
[0039] Specifically, the discriminator is a neural network model that can distinguish between real samples and generated samples by learning the statistical distribution of EEG signals and the characteristics of real signals. The input of the discriminator is the reconstructed EEG sample and the original EEG sample, and the output is a binary classification result, indicating whether the input sample is real (from the original EEG data) or generated (from the reconstructed EEG model). Exemplarily, the discriminator.
[0040] In this way, the EEG recovery model is used as a generator, and its goal is to generate high-quality reconstructed EEG samples based on the masked EEG fragment set, so that it can deceive the discriminator through training optimization, making it impossible for the discriminator to distinguish between the generated samples and the real samples. In addition, the training of the generator and the discriminator can be performed alternately until the generator can generate high-quality EEG samples that are "indistinguishable from the real thing".
[0041] Through the embodiments of this application, the discriminator is introduced and a generative adversarial learning framework is formed, which significantly improves the ability of the EEG recovery model to reconstruct EEG signals. The discriminator helps the generator learn a more realistic EEG signal distribution, and has achieved remarkable results in signal detail recovery, complex loss processing, and global distribution modeling, providing strong technical support for the recovery of EEG signals and the expansion of practical application scenarios.
[0042] Figure 2 Shown according to Figure 1 An example operation flow chart of step S110 in FIG.
[0043] like Figure 2 As shown, in step S210, at least one mask signal channel and at least one mask signal time are determined.
[0044] In one example of an embodiment of the present application, certain specific channels or time points can be selected for masking, such as channels close to areas with more artifacts (such as frontal lobe areas that are more affected by eye movement artifacts) or low signal-to-noise ratios, or time windows related to specific tasks (such as key time windows in event-related potential analysis). In another example of an embodiment of the present application, at least one mask signal channel and at least one mask signal time are randomly determined according to the mask ratio, and the mask ratio defines the ratio between the number of mask signal channels and the total number of sample signal channels and the ratio between the length of the mask signal time and the total sample time. Specifically, a certain proportion of channels (such as 20% or 50%) are randomly selected as mask signal channels, or a number of time periods (such as 100ms or 500ms) are randomly selected to simulate diverse signal loss situations in real scenes.
[0045] In some embodiments, the mask ratio is dynamically adjusted for different batches of raw EEG samples. Thus, in different training rounds, by dynamically adjusting the mask ratio (e.g., gradually increasing from a low ratio), and dynamically adjusting it within a range of 20%-80%, the masked channels or time are dynamically adjusted in each training iteration, for example, by gradually enhancing the model's ability to recover from a large range of missing signals, ensuring that the model learns a variety of missing scenarios, and enhancing the generalization ability of the model.
[0046] In step S220, for each mask signal channel, the original EEG segments of the original EEG segment set at all time points of the mask signal channel are masked to determine a corresponding first mask EEG segment subset.
[0047] Exemplarily, for the selected mask signal channels, the EEG signals of these channels at all time points are set to zero or filled with a fixed value (such as the mean or median).
[0048] In step S230, for each mask signal time, mask marking is performed on the original EEG segments of all channels of the original EEG segment set at the mask signal time to determine a corresponding second mask EEG segment subset.
[0049] Exemplarily, for the selected mask signal time point, the EEG signals of all channels in the corresponding time are set to zero or filled with a fixed value.
[0050] In step S240 , a masked EEG segment set is determined according to the first masked EEG segment subset and the second masked EEG segment subset.
[0051] Exemplarily, the first mask EEG segment subset (channel masking) is fused with the second mask EEG segment subset (time masking) to generate a final mask EEG segment set. In addition, during the fusion process, the order of channel masking and time masking can be randomly disrupted to make the masking mode more diverse.
[0052] Through the embodiments of the present application, channel masking and time masking are integrated so that the masked EEG segment set covers a variety of signal loss modes (single loss, joint loss, etc.). The final masked EEG segment set is closer in distribution to the signal loss distribution in actual application scenarios, thereby improving the applicability and generalization ability of the model.
[0053] In order to enable the encoder to better complete the encoding and reconstruction of the masked EEG segments, the masked channels, time information and / or damage feature information can also be encoded with mask features, so as to use this mask encoding to guide the encoder to more accurately reconstruct the segments at the corresponding masked positions.
[0054] In some examples of the embodiments of the present application, the mask position information corresponding to the mask signal channel and the mask signal time is obtained. Here, the mask position information is a binary matrix (or tensor) used to mark the masked area (channel or time point), which is used to clearly tell the model which parts of the EEG signal have been masked.
[0055] In some embodiments, the two-dimensional sine-cosine position coding corresponding to the mask signal channel and the mask signal time can be determined. Therefore, the introduction of the two-dimensional sine-cosine position coding as the position information identifier of the mask signal channel and the mask signal time can provide the model with explicit spatiotemporal position information, so that the model can better capture the spatiotemporal characteristics of the EEG signal and the position information of the masked area.
[0056] Furthermore, based on the encoder processing of the masked EEG segment set and the mask position information, a set of encoded EEG segments for the masked partial EEG segments is obtained.
[0057] Here, by obtaining the mask position information corresponding to the mask signal channel and mask signal time, and combining this position information with the mask EEG segment set to input the encoder, the mask position information explicitly marks the masked channels and time points, so that the model can clearly identify which areas have missing signals, significantly enhancing the model's perception of the masked areas and reducing the uncertainty of the model when processing inputs. At the same time, it also provides effective recovery guidance for the temporal and spatial attention modules, thereby improving the encoder's recovery accuracy for the masked EEG signals.
[0058] The details of the space-time autoencoder for EEG signal recovery provided in the embodiments of the present application will be expanded with examples below.
[0059] Figure 3 A schematic diagram of the architecture connection of an example of a spatiotemporal autoencoder for EEG signal restoration according to an embodiment of the present application is shown.
[0060] like Figure 3 As shown in Figure 2, random channel or temporal masks are applied to EEG data, followed by segment embedding, where the masked segments are replaced by the masked tokens. The autoencoder reconstructs the masked data, and the discriminator guides the generation of more realistic signals.
[0061] In the embodiment of the present application, a space-time autoencoder (STAR) is proposed, which comprehensively solves the problem of EEG signal damage in the channel and time dimensions through a dynamic mask strategy and an alternating attention mechanism. Through dynamic mask modeling, damaged samples of the channel and time dimensions are generated to simulate the damage of EEG signals in real scenes and ensure the robustness of model training. Through the space-time alternating attention mechanism, the dynamic characteristics of the channel and time dimensions are alternately modeled to capture the complex global relationship of the EEG signal. Through the generative adversarial mechanism, the generative adversarial network is used to improve the authenticity of signal repair and ensure that the repair result is closer to the real signal.
[0062] I. Methodology A. Problem Description and Masking Strategy Problem Description set up is the DE (Differential Entropy) feature of the EEG signal, where , and denote the number of channels, the number of time series samples, and the dimension of DE features, respectively. In the training phase, it is assumed that the EEG recordings are collected in a controlled laboratory environment and all DE features are fully accessible. In contrast, in the inference phase, the EEG data are acquired in a real-world environment and are represented as , only some channels and time point Available. The main goal is to reconstruction .
[0063] Mask strategy Here, signal channels or time points are randomly selected for masking to generate damaged samples, and the damage ratio is set to 20%-80%.
[0064] In practical applications, the most common forms of EEG corruption are complete signal loss of a specific channel and signal degradation of all channels at certain time points, which are called channel corruption and temporal corruption, respectively. Therefore, this paper focuses on solving these two types of EEG corruption. Figure 3 As shown in the channel damage scenario, some channels are randomly selected and all their data are masked. Divided into and ,in , in the inference only Available. In the temporal corruption scenario, select specific time points randomly and mask the data of all channels at these time points. Divided into and ,in , only Available. For each training batch, it is randomly selected whether to apply the channel or time corruption strategy. With dynamic mask ratio, the mask ratio of channels or time points in EEG is selected from the set Random sampling with equal probability.
[0065] Therefore, dynamic mask modeling can effectively improve the model's adaptability to various types of damage.
[0066] B. Spatiotemporal Autoencoder for EEG Signal Restoration Spatiotemporal Autoencoder Here, signal embedding is achieved by extracting features from the damaged signal and encoding the spatial and temporal positions. Furthermore, in the encoder's space-time alternating modeling, the attention mechanism is used to model dynamic characteristics in the time dimension and model the inter-channel dependencies in the channel dimension, alternating multi-layer modeling to capture complex spatiotemporal characteristics.
[0067] like Figure 3 As shown, after masking the EEG signal, the embodiment of the present application will Perform spectrum fragment embedding transformation to obtain , according to the EEG channel and time point and The EEG data is divided into segments along the dimension. Then, the masked parts are replaced by mask markers to obtain In order to remember the position of each EEG segment and use it in subsequent reconstruction, a two-dimensional sine-cosine position encoding is added in the spatial and temporal dimensions. Rearrange to ,in ( ), for each element in the time dimension Apply the self-attention mechanism and collect the results. Then, perform spatial transformation to transform ,in ( ), apply self-attention mechanism to each element in the spatial dimension, and aggregate the results. This process is repeated times to alternately extract temporal and spatial information. Finally, the output passes through a symmetric decoder to obtain the prediction of the original complete data The reconstruction quality is measured by the mean square error (MSE) loss, defined as follows: , Formula (1) In the formula, represents a spatiotemporal autoencoder.
[0068] Therefore, through the alternating attention mechanism, the complex spatiotemporal dynamic characteristics of EEG signals can be fully captured to improve the restoration accuracy.
[0069] Discriminator The complete signal is reconstructed by using a symmetric decoder. Finally, the discriminator improves the authenticity of the reconstructed signal through adversarial training.
[0070] Specifically, a framework is developed to improve the realism of reconstructed EEG. The framework consists of two components: the generator (G) introduced in the previous section, which is used to reconstruct the masked EEG segments; and the discriminator (D) which is used to evaluate the realism of EEG. The discriminator performs a binary classification task to determine whether Is it real or generated. The architecture of the discriminator is similar to the encoder of the spatiotemporal autoencoder, with an additional linear projection head to output the authenticity label. Formally, the task of the discriminator model can be expressed as: , Formula (2) The training objective of the discriminator can be expressed as: , Formula (3) Therefore, by generating an adversarial mechanism, the EEG restoration model can be prompted to generate more realistic signals, reduce artifacts, improve the authenticity of signal restoration, and enhance emotion recognition performance.
[0071] Training Program The ultimate goal of the generator is to minimize the combined loss, defined as: , Formula (4) , Formula (5) In the formula, is a hyperparameter used to control the reconstruction loss and the discriminator loss Specifically, in each training cycle, the following two steps are iterated: (i) only use Train the generator; (ii) use Train the discriminator.
[0072] II. Experiment A. Dataset and Implementation Details Dataset Here, the performance of the framework provided by the embodiments of the present application is comprehensively evaluated on three sets of public emotion recognition datasets SEED, SEED-IV and SEED-V. These datasets use 62-channel EEG signals, covering 3, 4 and 5 emotions, respectively, and adopt differential entropy (DE) features using a 1-second window.
[0073] Evaluation Metrics For each emotion recognition dataset, the STAR model was initially trained on the training data of all subjects. Subsequently, classic and state-of-the-art EEG-based emotion recognition models, including CNN, LSTM, Transformer, and EmoGT, were trained on the data of each subject. “Subject-dependent” means that an independent classifier is trained for each individual subject. Here, missing EEG signals are simulated by replacing broken connections or removed parts with zeros. Specifically, simulations with different levels of channel and time point availability, including 20%, 40%, 60%, 80%, and full availability, were performed. Each configuration was repeated 40 times. The corrupted and STAR-repaired EEG signals were input into the trained emotion classifiers, and the average classification accuracy was used as the evaluation metric.
[0074] Implementation details Number of time series samples . Mask ratio set for dynamic mask strategy is [20%, 40%, 60%, 80%]. The fragment embedding dimension is set to 16, and the number of spatiotemporal attention blocks is The learning rate is 0.0001 and the batch size is 48. Hyperparameters Set to 0.01. Use Adam optimizer.
[0075] B. Experimental Results Figure 4 The experimental effect simulation diagram of an example of STAR classification accuracy of damaged and restored EEG data on different classifiers and data sets is shown, where the dotted line represents the performance of damaged data and the solid line represents the performance of restored data.
[0076] EEG restoration to improve classification accuracy In order to evaluate the effectiveness of the method provided in the embodiment of the present application in restoring EEG data, the average classification accuracy of the damaged data was compared with the performance of the restored data on multiple classifiers, such as Figure 4As shown. As the damage ratio in the channel and time dimensions increases, the classification performance of all classifiers decreases significantly, indicating the negative impact of damaged EEG data on classifier performance. When comparing the accuracy of damaged data and restored data at the same damage ratio, it can be clearly seen that restoring the damaged EEG data significantly improves the classification performance of all classifiers at different damage levels. It is worth noting that in the SEED dataset, when using the Transformer classifier, the channel damage ratio is 60% when it is improved by 11.93%, and the time damage ratio is 80% when it is improved by 29.64%. These results show that the method provided in the embodiments of the present application can effectively restore EEG data that has been damaged to varying degrees in the channel or time dimensions.
[0077] Comparison with other recovery methods Here, STAR is compared with several EEG restoration methods, including interpolation, NP, and ANP. For interpolation, the procedure is performed for channel corruption, and the nearest neighbor interpolation method is applied for temporal corruption. The NP and ANP models, which were originally used to address the channel corruption problem in the literature, are adjusted to handle both channel and temporal corruptions. These restoration methods are evaluated using the classification accuracy of the Transformer classifier and the relative root mean square error (RRMSE) of the restored data as performance metrics, as shown in Table I. The results show that STAR consistently outperforms other methods in all channel and temporal corruption scenarios and under different corruption ratios.
[0078] Table I: Classification accuracy of the Transformer model and the relative root mean square error (RRMSE) between the reconstructed data and the real data for different restoration methods at various damage levels on three datasets.
[0079] Visualization Figure 5 The figure shows an example of the effect simulation diagram of the reconstruction topology of the γ-band test data at a single time point under different damage levels. The black channel represents the damaged channel in the random sampling. The time damage will cover all the data at the time point to be stored, so it is not shown here.
[0080] In order to further evaluate the performance of the model provided by the embodiment of the present application in data reconstruction, the topographic map of the γ band of the SEED test set is shown, because this band plays an important role in emotion recognition. Figure 5 As shown, the restored topography at a single time point under different damage ratios is shown, indicating that the model can effectively reconstruct the original signal.
[0081] III. Conclusion In this paper, we introduce a spatiotemporal autoencoder (STAR) for EEG signal restoration to cope with channel and temporal corruption in EEG signals. This innovative framework adopts dynamic channel and temporal masks to simulate real-world EEG corruption scenarios, and applies a spatiotemporal alternating attention mechanism to fully exploit the spatiotemporal relationships in EEG to achieve efficient reconstruction of complete EEG signals from masked data. Experiments on the SEED, SEED-IV, and SEED-V datasets show that STAR can effectively restore EEG signals with varying degrees of channel and temporal corruption, significantly improving the performance of emotion recognition models when dealing with corrupted EEG data. These findings highlight the potential to enhance EEG-based applications in less controlled environments, bringing practical solutions closer to real-world use cases.
[0082] It should be noted that this study used publicly available human subject data for retrospective analysis. As the data used were open access and the accompanying license confirmed that no ethical approval was required, no ethical approval was required for this study.
[0083] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of actions combined, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0084] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned EEG recovery model training methods of the present application.
[0085] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned training methods for the EEG recovery model.
[0086] In some embodiments, the embodiments of the present application also provide an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a training method for an EEG recovery model.
[0087] Figure 6 is a schematic diagram of the hardware structure of an electronic device for executing a training method for an EEG recovery model provided in another embodiment of the present application, such as Figure 6 As shown, the device includes: One or more processors 610 and memory 620, Figure 6 A processor 610 is taken as an example.
[0088] The device for executing the training method of the EEG restoration model may further include: an input device 630 and an output device 640 .
[0089] The processor 610, the memory 620, the input device 630 and the output device 640 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0090] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the training method of the EEG recovery model in the embodiment of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 620, that is, the training method of the EEG recovery model in the above method embodiment is implemented.
[0091] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 620 may optionally include a memory remotely arranged relative to the processor 610, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0092] The input device 630 can receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 640 can include a display device such as a display screen.
[0093] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, the training method of the EEG recovery model in any of the above method embodiments is executed.
[0094] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.
[0095] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.
[0096] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.
[0097] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0098] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.
[0099] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for an electroencephalogram recovery model, comprising: The original EEG samples are divided into an original EEG segment set, and part of the original EEG segments in the original EEG segment set are masked according to the mask signal channel and the mask signal time to obtain a corresponding mask EEG segment set; Processing the masked EEG segment set based on an encoder to obtain an encoded EEG segment set for the masked partial EEG segment; the encoder comprises a temporal attention module and a spatial attention module; Processing the encoded EEG segment set and the masked EEG segment set based on a symmetric decoder to generate reconstructed EEG samples, and calculating corresponding sample reconstruction losses in combination with the original EEG samples; Based on the sample reconstruction loss, the EEG restoration model is trained and optimized.
2. The method according to claim 1, further comprising: Determine whether the reconstructed EEG sample is real or generated based on the discriminator, and calculate the corresponding discriminator loss; The step of training and optimizing the EEG recovery model based on the sample reconstruction loss includes: The EEG restoration model is trained and optimized according to the sample reconstruction loss and the discriminator loss.
3. The method according to claim 1, wherein: The method of masking part of the original EEG segments in the original EEG segment set according to the mask signal channel and the mask signal time to obtain a corresponding mask EEG segment set comprises: determining at least one mask signal channel and at least one mask signal time; For each of the mask signal channels, mask-marking the original EEG segments of the original EEG segment set at all time points of the mask signal channel to determine a corresponding first mask EEG segment subset; For each of the mask signal times, mask-marking the original EEG segments of all channels of the original EEG segment set at the mask signal time to determine a corresponding second mask EEG segment subset; A set of masked EEG segments is determined according to the first subset of masked EEG segments and the second subset of masked EEG segments.
4. The method according to claim 3, wherein: The determining of at least one mask signal channel and at least one mask signal time comprises: At least one mask signal channel and at least one mask signal time are randomly determined according to the mask ratio; the mask ratio defines the ratio between the number of mask signal channels and the total number of sample signal channels and the ratio between the length of the mask signal time and the total sample sampling time.
5. The method according to claim 4, wherein: The mask ratio is dynamically adjusted for different batches of raw EEG samples.
6. The method according to claim 1, further comprising: Obtaining mask position information corresponding to the mask signal channel and mask signal time; The step of processing the masked EEG segment set based on the encoder to obtain a set of encoded EEG segments for the masked partial EEG segments includes: The masked EEG segment set and the mask position information are processed based on an encoder to obtain an encoded EEG segment set for the masked partial EEG segment.
7. The method according to claim 6, wherein: The obtaining of the mask position information corresponding to the mask signal channel and the mask signal time includes: Determine the two-dimensional sine-cosine position code corresponding to the mask signal channel and the mask signal time.
8. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Pilot attention state data acquisition and processing method and device and related equipment
CN120616532A
Electroencephalogram signal processing method and device and computer equipment
CN121465609A