Electroencephalogram signal classification method based on domain adaptive data synthesis contrast diffusion model
Through the domain adaptive data synthesis contrast diffusion model, the problems of noise and individual differences in EEG signal classification are solved, high-quality data generation and feature extraction are achieved, and the robustness and adaptability of the classification model are improved.
Patent Information
- Application Number
- CN202510125527.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-09-12
AI Technical Summary
Existing EEG signal classification methods suffer from problems of training instability and insufficient feature extraction when dealing with noise and individual differences, making it difficult to achieve effective adaptation between different tasks and individuals.
A contrastive diffusion model based on domain-adaptive data synthesis is adopted. Through forward diffusion and backward diffusion processes, combined with a contrast denoising module, domain-adaptive EEG data is generated. The unit progressive concave network, continuous progressive concave network and domain noise adaptive extraction module are used to improve the accuracy of feature extraction and classification.
It improves the robustness and generalization ability of EEG signal classification, can effectively capture subtle feature changes, adapt to different tasks and individual differences, and generate high-quality EEG data.
Smart Images

Figure CN120632665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological signal processing technology, and in particular to an electroencephalogram (EEG) signal classification method based on a domain adaptive data synthesis contrast diffusion model. Background Art
[0002] Brain-computer interface technology converts brain activity into commands, allowing devices to be controlled without relying on muscle movements or gestures. In this process, EEG signals serve as a key intermediate variable and are widely used in fields such as motor imagery and emotion recognition. In motor imagery tasks, EEG signals can be used to detect hand movements, such as grasping and lifting, and to control devices such as computer mice and wheelchairs. In emotion recognition tasks, EEG signals reflect psychological and physiological states related to emotions, which are crucial for daily interactions and physical functions. These applications demonstrate the importance of EEG signals in the development of advanced brain-computer interface technologies that can replace, restore, enhance or improve neural functions.
[0003] Although deep learning methods have advanced motor imagery and emotion recognition, the network structure design of existing methods relies on extensive expertise and time. Furthermore, EEG signals are susceptible to environmental noise and individual differences, and existing denoising methods do not fully account for these complexities. Restrictions on experimental environments also lead to limited training data, making synthetic EEG data an important approach to addressing data insufficiency. Current synthesis methods based on generative adversarial networks help neural networks relearn task features, but they suffer from training instability and mode collapse, hindering practical applications. In contrast, diffusion denoising probabilistic models show greater potential in generating high-quality EEG data. However, existing research has not fully considered task and individual differences, necessitating the design of new data synthesis and adaptation frameworks to achieve more effective denoising and high-quality feature extraction. Summary of the Invention
[0004] In view of this, the present invention discloses an EEG signal classification method based on a domain-adaptive data synthesis contrast diffusion model to generate domain-adaptive EEG data.
[0005] The technical solution of the present invention is: an EEG signal classification method based on a domain adaptive data synthesis contrast diffusion model, comprising: constructing a domain adaptive data synthesis contrast diffusion model, including a forward diffusion process and a backward diffusion process;
[0006] The forward diffusion process is used to train noisy data. The original signal input passes through the forward diffusion process until it reaches the data x at time T. T ;
[0007] The reverse diffusion process is provided with a contrast denoising module, which is used to effectively learn clean signals and domain noise in the reverse diffusion process to generate clean data; the data x at time T T is a noise signal, the reverse diffusion process converts the noise signal x T By gradually removing noise and finally restoring the clean signal
[0008] Finally, the clean signal Send it to the standard classification network EEGNET for classification.
[0009] Preferably, the forward diffusion process is used to train noisy data, and the original signal input passes through the forward diffusion process until it reaches the data x at time T. T ,include:
[0010] Preprocessing: 2D EEG data Converted to stacked 3D data using sliding windows Where C represents the number of channels, W represents the window length, and N represents the number of windows;
[0011] The pre-processed two-dimensional EEG signal data is input into the forward diffusion process, and the signal passes through the intermediate state in turn by iteratively adding noise. Until the final noise state x1 is reached, the iterative process uses the sequence q(x T |x T-1 )1 means that a small amount of noise is added to the previous state at each step.
[0012] Preferably, the data x at time T T is a noise signal, the reverse diffusion process converts the noise signal x T By gradually removing noise and finally restoring the clean signal Back diffusion process using sequence express;
[0013] In the reverse diffusion process, x T First, it is sent to the clean signal extraction module and the domain noise adaptive extraction module at the same time. The contrast denoising module separates the clean EEG signal from the domain noise. The clean EEG signal is obtained by the clean signal extraction module, and the domain noise is captured by the domain noise adaptive extraction module. At the end of back diffusion, a clean signal can finally be obtained.
[0014] Preferably, the clean signal extraction module is composed of four unit progressive concave networks connected in series, and the module adopts unit progressive concave networks and residual connections, and a residual connection is added between every two consecutive unit progressive concave networks, wherein the input of each unit progressive concave network is composed of the input and output of the previous unit progressive concave network;
[0015] Each unit in the progressive concave network is called a unit, denoted as C, where i represents the index of the corresponding unit, as shown in formula (1), and Represent the input and output of the unit indexed by i respectively;
[0016]
[0017] According to the Gaussian distribution assumption in the diffusion denoising probability model, the variance is set to ε T , where ε T represents the Gaussian noise parameter; the parameter extracted from the clean signal is defined as ε C Indicates the use to restore clean signals The noise from x T arrive The time step training loss follows formula (2);
[0018] L c =MSE Loss=‖ε T -ε C ‖ 2 (2).
[0019] Preferably, the back diffusion process of the diffusion denoising probability model is composed of a continuous progressive concave network module, which is used to enhance the learning ability of high-dimensional EEG signal features;
[0020] The continuous progressive concave network module consists of two stages: downsampling and upsampling. The first stage is the downsampling process. Each downsampling block in the downsampling process contains two submodules: a fine downsampling block and a progressive downsampling block. At the same time, the upsampling process follows the processing method of the original concave network.
[0021] The downsampling process of the continuous progressive concave network module is shown in formula (3) and formula (4);
[0022]
[0023] Formula (3) represents the sequence calculation process, where each fine dimensionality reduction block is represented by F j Indicates that each progressive dimensionality reduction block is represented by P j Indicates that j is the index of the dimension reduction block; and They represent the outputs of the fine dimensionality reduction block and the progressive dimensionality reduction block with index j respectively;
[0024] The fine dimensionality reduction block contains an improved residual block in which the original two-dimensional convolution is replaced by depth-wise separable convolution and point-wise convolution; in addition, a pixel normalization layer is added to the residual block to standardize the feature distribution.
[0025] Preferably, the domain noise adaptive extraction module comprises a subject feature learning block, a cross-task feature learning block and a self-attention layer arranged in sequence;
[0026] The domain noise adaptive extraction module uses the EEG network classifier EEGNet to extract the output of the module Extract features From x T arrive The loss function Ls of the time step is based on the improved Arc-Margin Loss function, as shown in formula (5);
[0027]
[0028] The parameter m is used to increase the closeness of features within the same subject during classification, while increasing the discrimination between different subjects, and j is used to index all other subjects except the current subject; in addition, N represents the batch size used to calculate the average loss. represents the label of the target subject; Formula (6) shows The calculation process of
[0029]
[0030] In formula (6), and Represents the weight of the classifier.
[0031] Preferably, the subject feature learning module is composed of a series of stacked modules for learning subject-specific unique features in EEG data;
[0032] At the initial time step, from x T arrive Input x T is fed into the network; a two-dimensional convolutional layer is used to capture the spatial patterns in the EEG signal, thereby retaining key features, while further complex features are extracted through the residual block;
[0033] The calculation of the subject's attention is shown in formula (7), where Q Subject represents the subject embedding calculated by Q, and the inputs of K and V come from the output of the previous layer;
[0034]
[0035] Preferably, the cross-task feature learning module is used to cope with task differences and enhance the extraction of domain features in different EEG tasks, and includes three parts: Part A is region division, Part B is feature extraction, and Part C is cross-task domain calculation;
[0036] In part A, the input data first passes through the domain partitioning layer;
[0037] In Part B, the input data are domain, and emotion recognition tasks field;
[0038] In part C, the cross-domain integration module is used to solve the problem of capturing subtle relationships between features in different domains of EEG signals. The input of the cross-domain integration module is the domain features extracted in part B. Output D i is fed into the cross-domain attention layer, which Figure 8 The calculation process of cross-domain attention first transforms input 1 and input 2 according to formula (8) and formula (9);
[0039]
[0040] D i represents the input from input 1, D j represents the input from input 2; Q is the query matrix, K is the key matrix;
[0041] Subsequently, the relationship between the input fields is calculated using formula (10) and formula (11);
[0042]
[0043] in, and represents the dimension of the input query. The values of i and j in different tasks are shown in formulas (12) and (13);
[0044]
[0045] Finally, the concatenated output is passed through a random dropout layer as the final output O, where n represents the index of the attention output, as shown in formula (14);
[0046] O=Dropout(Concatenate(Attention n ))for n=1to2 (14).
[0047] Preferably, the contrast denoising module is used to separate the clean module from the noise module, thereby achieving denoising; the output of the module is represented as a negative sample The positive samples generated by the module are expressed as The contrast loss function at the current time step is defined as shown in formula (15).
[0048]
[0049] Loss function L CL It consists of two components: the first is the s and x c The similarity matrix is regularized to prevent the model from overfitting features; the second term maximizes x S and x C The cosine similarity between them is used to enhance feature alignment; through these two aspects, L CL Ensures effective alignment and separation of features;
[0050] The total loss function is shown in formula (16); the hyperparameter λ cl Controls the weight of the contrast loss, λ s Adjust the loss of the domain noise adaptive extraction module in the reverse process, λ c The reverse loss in the clean signal extraction process is crucial; ultimately, the total loss function is expressed as formula (16):
[0051] L=λ cl L CL +λ s L S +λ c L C (16).
[0052] This invention provides an EEG signal classification method based on a domain-adaptive data synthesis contrastive diffusion model, which is used to generate domain-adaptive EEG data. It achieves superior performance in EEG classification tasks, highlighting its efficiency in improving the reliability of data synthesis. The proposed classification method can effectively capture subtle feature changes in EEG signals and learn personalized noise based on different tasks. Its domain-adaptive data synthesis contrastive diffusion model outperforms the currently best recognition models.
[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0056] Figure 1 A schematic diagram of the overall architecture of data synthesis based on the adaptive contrast diffusion model provided in the embodiments disclosed in the present invention;
[0057] Figure 2 Schematic diagram of a clean signal extraction module based on a unit progressive concave network according to an embodiment of the present invention (time step ( Clean signal extraction model framework);
[0058] Figure 3 An architectural diagram of a continuous progressive concave network module according to an embodiment of the present invention;
[0059] Figure 4 A diagram of the improved residual network architecture provided by the disclosed embodiment of the present invention;
[0060] Figure 5 This is a structural diagram of the domain adaptive noise extraction module provided in the embodiment disclosed in the present invention;
[0061] Figure 6 A structural diagram of a subject feature learning module provided in an embodiment of the present invention;
[0062] Figure 7 A structural diagram of a subject feature learning module for a motor imagery task according to an embodiment of the present invention;
[0063] Figure 8 A structural diagram of a subject feature learning module for an emotion recognition task provided by an embodiment of the present invention;
[0064] Figure 9 The domain division of the motor imagery task provided by the embodiment disclosed in the present invention is: Figure 8 (a)BCI-IV-2A Figure 8 (b) BCI-IV-2B;
[0065] Figure 10 A diagram illustrating the domain division of the emotion recognition task provided by the disclosed embodiment of the present invention;
[0066] Figure 11 A structural diagram of cross-task domain computing provided by the disclosed embodiment of the present invention;
[0067] Figure 12A framework diagram of the reverse process of data synthesis based on the comparative diffusion model provided in the disclosed embodiment of the present invention;
[0068] Figure 13 This is a diagram of the internal framework of the contrast denoising model provided in the disclosed embodiment of the present invention. DETAILED DESCRIPTION
[0069] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of systems consistent with certain aspects of the present invention, as detailed in the appended claims.
[0070] In EEG research, the collection of training data is often limited due to the specific experimental environments. To address data scarcity, generating synthetic EEG data has become an important approach. Currently, mainstream synthetic data methods are based on generative adversarial networks (GANs). These methods not only alleviate the data sparsity problem but also help neural networks relearn features through data reconstruction and redefinition, thereby improving feature extraction capabilities in motor imagery and emotion recognition tasks. However, GANs suffer from instability and mode collapse during training, which can lead to reduced quality of generated data and limited application. In contrast, the diffusion denoising probabilistic model has shown great potential in generating high-quality, realistic EEG data, and recent studies have demonstrated its superior data quality compared to GANs. However, these methods fail to fully account for differences in task and individual characteristics. Therefore, designing a novel data synthesis and adaptation framework that can simultaneously accommodate both task and individual differences is crucial for optimizing EEG analysis and denoising.
[0071] Existing synthetic network architectures face challenges in the following aspects: First, the U-shaped network structure commonly used in diffusion models may lead to unstable training during the reverse process, resulting in the loss of key information; second, these architectures do not fully consider task diversity and individual differences in their design, which limits their adaptability in different scenarios; third, in complex environments, existing methods have difficulty in achieving efficient separation of clean signals and noise. These limitations lead to a lack of universality and adaptability of classification models across different tasks and individuals. Therefore, developing an innovative network architecture that can effectively address the above problems can not only improve the quality of synthetic EEG signal data, but also enhance the robustness and generalization ability of the classification model.
[0072] To this end, this embodiment provides an EEG signal classification method based on domain adaptive data synthesis contrast diffusion model. The overall process is as follows: Figure 1 As shown, Figure 1The figure above shows the forward process of the model. Figure 1 The figure below illustrates the model's reverse process. The domain-adaptive data synthesis contrastive diffusion model includes forward and backward diffusion processes. Furthermore, the model incorporates a contrastive denoising module into the backward diffusion process. This module effectively learns clean signals and domain noise during the backward diffusion process, ultimately generating clean data.
[0073] The forward diffusion process of the domain adaptive data synthesis contrast diffusion model is a standard processing flow in the diffusion denoising probability model, which is used to train noisy data, such as Figure 1 As shown in the figure above, two-dimensional EEG signal data Converted to stacked 3D data using sliding windows To enhance the utilization of temporal context information. Where C represents the number of channels, W represents the window length, and N represents the number of windows. After preprocessing, these data are input into the forward diffusion process. In this process, by iteratively adding noise, the signal passes through the intermediate state in turn. Until the final noise state x1 is reached. This process is represented by the sequence q(x T |x T-1 )1 means that a small amount of noise is added to the previous state at each step.
[0074] like Figure 1 As shown in the figure below, the reverse diffusion process of the model starts from the noise signal X T Initially, the goal is to gradually remove noise and eventually restore the clean signal This reverse process uses the sequence The figure shows a time state diagram that contains a series of contrasting diffusion models. In this figure, the starting point of the time step and end point The detailed process is represented by the green background module, and the remaining time steps are shown in dot form.
[0075] In order to highlight the difference between input and output at different time steps, the figure describes two specific time steps (positions) in detail. The first time step is from T to T-1, and the input is the signal at time T (denoted as x T ), the output is the clean signal at time T-1 (denoted as ). The second time step is from time 1 to time 0, and the input is the signal at time 1 (denoted as The output is a clean signal at time 0 (denoted as This iterative denoising process continues until the signal is finally restored to a clean signal.
[0076] In the reverse diffusion process, x TFirst, the signal is fed simultaneously into the clean signal extraction module and the domain noise adaptive extraction module. Subsequently, by designing a contrastive loss function, the contrastive denoising module effectively separates the clean EEG signal from the domain noise. The clean signal is obtained by the clean signal extraction module, while the domain noise is captured by the domain noise adaptive extraction module. At the end of back-diffusion, the clean signal is finally obtained.
[0077] like Figure 2 As shown, the time step in the reverse process is from arrive This paper uses the process of extracting clean signals as an example to illustrate the specific process of extracting clean signals. The basic diffusion denoising probabilistic model architecture is used as the baseline diffusion denoising model. However, each time step only contains a single concave network, which is well aligned with the sequential nature of EEG signal processing. Therefore, this paper designs a novel method to further enhance feature extraction capabilities.
[0078] like Figure 2 As shown in , the clean signal extraction module consists of four serially connected unit progressive concave networks. In order to gradually expand the feature space, two technologies are adopted in the module: unit progressive concave networks and residual connections. Specifically, Figure 2 As shown by the red arrow line in , a residual connection is added between every two consecutive unit progressive concave networks, where the input of each unit progressive concave network is composed of the input and output of the previous unit progressive concave network.
[0079] Each unit in the progressive concave network is called a unit, denoted as C, where i represents the index of the corresponding unit. As shown in formula (1), and Represent the input and output of the unit with index i respectively.
[0080]
[0081] In addition, according to the Gaussian distribution assumption in the diffusion denoising probability model, the variance is set to ε T , where ε T represents the Gaussian noise parameter. The parameter extracted from the clean signal is defined as ε C Indicates the use to restore clean signals Noise. From x T arrive The training loss for the time step of follows formula (2).
[0082] L c =MSE Loss=‖ε T -ε C ‖ 2 (2)
[0083] The back-diffusion process of the diffusion denoising probabilistic model consists of a concave network. The concave network is able to recover clean EEG signals and suppress artifacts. However, although the concave network architecture can be viewed as a special form of residual network, characterized by residual connections only between layers whose feature dimensions match between the encoder and decoder, it suffers from the problem of vanishing gradients as the decoder depth increases. Inspired by the architectural principle of progressive generative adversarial networks, which states that features learned at a low-resolution stage can enhance learning at a high-resolution stage, this paper proposes a novel module called the continuous progressive concave network. This module is designed to enhance the learning of high-dimensional EEG signal features.
[0084] Figure 3 The overall architecture of the continuous progressive concave network is presented, consisting of two stages: downsampling and upsampling. The first stage is the downsampling process. In this stage, each downsampling block contains two submodules: a fine downsampling block and a progressive downsampling block. Meanwhile, the upsampling process follows the original concave network.
[0085] exist Figure 3 In
[15] , the downsampling process of the continuous progressive concave network module is shown in formulas (3) and (4).
[0086]
[0087] Formula 3 represents the sequence calculation process, where each fine dimensionality reduction block is represented by F j Indicates that each progressive dimensionality reduction block is represented by P j Indicates that j is the index of the dimension reduction block. and They represent the outputs of the fine dimensionality reduction block and the progressive dimensionality reduction block with index j, respectively. According to our experiments, α is a weight with a value of 1×10 -8 The downsampling process based on formula (3) and formula (4) ensures the stability of training.
[0088] The detail dimensionality reduction block contains an improved residual block such as Figure 4 As shown in Figure 2, in the improved residual block, the original two-dimensional convolution is replaced with depthwise separable convolution and pointwise convolution. In addition, a pixel normalization layer is added to the residual block to standardize the feature distribution. The improved residual block enhances the model's ability to accurately capture complex features.
[0089] The structure of the domain noise adaptive extraction module is as follows Figure 5 The module contains a subject feature learning block, a cross-task feature learning block, and a self-attention layer, which are arranged in sequence in the module.
[0090] In order to enhance the ability to distinguish between different subjects, the domain noise adaptive extraction module uses a simple EEG network classifier (EEGNet) to extract the output of the module. Extract features From x T arrive The loss function Ls of the time step is based on the improved Arc-Margin Loss function (Arc-MarginLoss), as shown in formula (5).
[0091]
[0092] The parameter m is used to increase the closeness of features within the same subject during classification while increasing the discrimination between different subjects, while j is used to index all other subjects except the current subject. In addition, N represents the batch size used to calculate the average loss. represents the label of the target subject. Formula (6) shows The computational process of , which aims to enhance the subject separability of noise.
[0093]
[0094] In formula (6), and represents the weight of the classifier. This formula is designed to encourage the model to maximize the probability of the correct class label, thereby improving classification accuracy at all time steps.
[0095] The design of this module is based on the understanding of the complexity of EEG signals and takes into account the differences between subjects in order to more effectively extract subject-specific noise. The block network consists of a series of stacked modules such as Figure 6 As shown,
[0096] At the initial time step, from x T arrive Input x T is fed into the network. The two-dimensional convolutional layer is used to capture the spatial patterns in the EEG signal, thereby retaining the key features, while the residual block (such as Figure 5 The calculation of the subject’s attention is shown in formula (7), where Q Subject It represents the subject embedding calculated by Q, and the inputs of K and V come from the output of the previous layer.
[0097]
[0098] Subsequently, a self-attention layer facilitates the fusion of information from different channels and captures inter-channel dependencies. Finally, a group normalization layer and a convolutional layer ensure that the processed features are standardized. As a result, the subject feature learning module is able to learn unique subject-specific features in EEG data.
[0099] The cross-task feature learning module is used to deal with task differences and enhance the extraction of domain features in different EEG tasks. The framework of the cross-task feature learning module is as follows: Figure 7 shown.
[0100] The cortex in different parts of the brain has specific effects on the corresponding electrode areas in various EEG tasks. In motor imagery tasks, the best electrodes are mainly distributed in the motor cortex, covering the parietal and frontal regions, and a small amount is distributed in the occipital region. In emotion recognition tasks, the activities of different brain regions are intrinsically linked to the recognition of specific emotional states. The frontal lobe regulates high-level emotions by processing brain nerve signals and is associated with thinking and consciousness. The occipital lobe is mainly responsible for visual processing and can respond to complex stimuli such as faces, scenes, smells, and sounds. Therefore, selecting electrodes in specific areas can help extract key information related to different brain-computer interface applications.
[0101] However, current synthetic data models and EEG classification models do not fully consider the differences in electrode activity across brain regions during different tasks, nor do they explore the correlations between brain regions. To address these issues, this paper proposes a cross-task feature learning module that dynamically focuses on specific brain regions, thereby improving feature extraction capabilities across different EEG tasks.
[0102] Figure 6 and Figure 7 The framework of the cross-task feature learning module for two tasks is presented separately. This module consists of three parts: Part A is region segmentation, Part B is feature extraction, and Part C is cross-task domain computation. Each part will be described in detail below.
[0103] 1) Regional division
[0104] exist Figure 9 In part A, the input data first passes through the domain partitioning layer. The domain partitioning for different tasks is as follows Figure 9 a and Figure 9 As shown in b, the temporal lobe regions (T7, T8, TP7, TP8) with low relevance to the emotion recognition task are excluded. This method effectively considers the importance of each area in different tasks and helps to extract task-specific features.
[0105] Figure 10Figure 3. Domain division of the motor imagery task: D1 (frontocentral area) primarily targets the motor cortex, an area crucial for movement execution, marked in blue; D2 (temporal area) and D3 (parietal area) support spatial awareness and higher-level processing, represented in yellow and green, respectively.
[0106] Domain division of emotion recognition tasks: D1 (frontal lobe area) regulates emotions and consciousness, marked in blue; D2 (temporal lobe area) processes complex auditory and sensory stimuli, marked in yellow; D3 (parietal lobe area) integrates sensory information, marked in green; D4 (occipital lobe area) is crucial for visual processing, marked in orange.
[0107] exist Figure 8 In part B, the input data are the motor imagery task domain, and emotion recognition tasks Field. The fields of different tasks are as follows Figure 9 and Figure 10 As shown in Figure 2. Subsequently, the data from each domain is fed into the corresponding branch network, which extracts domain-specific features. Finally, the outputs from each domain are integrated for subsequent calculations.
[0108] In the internal structure of the branch network, domain-specific EEG data is first processed through a batch normalization layer to stabilize learning during network training. Next, a convolutional layer is used to capture the spatial features unique to each domain, thereby extracting local patterns and understanding domain activities. Subsequently, the network's ability to model complex patterns is enhanced through nonlinear processing with activation functions. On this basis, the convolutional layer is combined with the activation function layer to focus on important features. At the same time, group normalization is applied to ensure consistent feature scaling between different feature groups, while the maximum pooling layer reduces the computational burden while effectively extracting key information from the original signal. Finally, all extracted features are integrated through the convolutional layer and prepared for processing of subsequent tasks.
[0109] Cross-task domain computing: In part C, the cross-domain integration module is used to solve the problem of capturing subtle relationships between features in different domains of EEG signals. The input of the cross-domain integration module is the domain features extracted in part B. Output D i is fed into the cross-domain attention layer, which Figure 8 The calculation process of cross-domain attention first transforms input 1 and input 2 according to formula (8) and formula (9).
[0110]
[0111] D i represents the input from input 1, D jrepresents the input from input 2. Q is the query matrix and K is the key matrix.
[0112] Then, the relationship between the input fields is calculated by formula (10) and formula (11).
[0113]
[0114] in, and represents the dimension of the input query. The values of i and j in different tasks are shown in formulas (12) and (13).
[0115]
[0116] Finally, the concatenated output is passed through a random dropout layer as the final output O, where n represents the index of the attention output, as shown in formula (14).
[0117] O=Dropout(Concatenate(Attention n ))for n=1to 2 (14)
[0118] A contrast denoising module is added during the back diffusion process to separate the clean module from the noise module, thus achieving denoising. Figure 1 In the example, the output of the pink module is represented as a negative sample. The positive samples generated by the yellow module are expressed as The contrast loss function at the current time step is defined as shown in formula (15).
[0119]
[0120] Loss function L CL It consists of two components: the first is the s and x c The similarity matrix is regularized to prevent the model from overfitting features; the second term maximizes x S and x C The cosine similarity between them is used to enhance feature alignment. Through these two aspects, L CL This ensures effective alignment and separation of features. This combination effectively improves the performance of the model in processing complex EEG signals, enabling it to capture the diversity of data from different perspectives.
[0121] The total loss function is shown in Equation 16. The hyperparameter λ cl Controls the weight of the contrastive loss, ensuring that the model appropriately focuses on minimizing the difference between the clean signal and the noise. s Adjust the loss of the domain noise adaptive extraction module in the reverse process, λ cThe reverse loss in the clean signal extraction process is crucial. Finally, the total loss function is expressed as formula (16):
[0122] L=λ cl L CL +λ s L S +λ c L C (16)
[0123] The domain adaptive data synthesis contrast diffusion model proposed in the present invention reduces parameters and time, and achieves a balance between complexity and performance.
[0124] Table 1 shows a parameter comparison of different methods. The Diffusion Denoising Probabilistic Model (DDPM) contains approximately 64.45M parameters, while the Continuous Progressive Concave Network module reduces this number to 16.10M, making it more suitable for the diffusion process. Furthermore, the introduction of the Domain Noise Adaptive Extraction module increases the number of parameters in CD3Net to 24.23M. Despite this, CD3Net remains significantly more efficient than the original DDPM, achieving a balance between complexity and performance.
[0125] Table 1: Comparison of the number of parameters of the proposed methods (in trillions).
[0126] method Number of parameters Diffusion Model 64.45 Domain-Adaptive Data Synthesis Contrastive Diffusion Model 24.23 Unit progressive concave network 16.10 Domain Adaptive Noise Extraction Module 8.13
[0127] Table 2 shows a comparison of training times. The basic DDPM model takes 61.78 seconds to train per time step, while the proposed model significantly reduces this time to just 32.81 seconds. This reduction in training time not only improves efficiency but also enables faster experimentation and model optimization.
[0128] Table 2: Time comparison of the proposed method (seconds).
[0129] method Training time Diffusion Model 61.78 Domain-Adaptive Data Synthesis Contrastive Diffusion Model 32.81
[0130] Compared with traditional models, the domain-adaptive data synthesis contrast diffusion model significantly reduces computational costs and improves the ability to handle complex operations, achieving more efficient model training and better performance through rational resource utilization.
[0131] Table 3 summarizes the computational cost (FLOPS) comparison. The DDPM model has a computational cost of 228.655 FLOPS, while the proposed CD3Net model significantly reduces this to only 135.537 FLOPS. This significant reduction highlights the computational efficiency of CD3Net, which not only saves resources but also speeds up the overall processing speed of the model.
[0132] Table 3: Comparison of computational costs (FLOPS) of the proposed methods.
[0133] method Calculate costs Diffusion Model 228.655 Domain-Adaptive Data Synthesis Contrastive Diffusion Model 135.537
[0134] Table 4 provides a comparison of the memory usage (in megabytes) of the evaluated methods. The DDPM model consumes 2215.74MB of memory, while the proposed domain-adaptive data synthesis contrastive diffusion model increases its memory usage to 6265.61MB. This increase in memory usage demonstrates the advanced computational framework of the domain-adaptive data synthesis contrastive diffusion model, which enables it to handle more complex operations and achieve superior performance.
[0135] Table 4: Comparison of video memory usage (M) of the proposed methods.
[0136] method Video memory usage Diffusion Model 2215.74 Domain-Adaptive Data Synthesis Contrastive Diffusion Model 6265.61
[0137] The cross-task feature learning module effectively improves the utilization efficiency of EEG field information and the accuracy of subject classification by optimizing domain combination and attention mechanism.
[0138] Tables 5, 6, and 7 compare the motor imagery and emotion recognition task datasets using two evaluation methods. Specifically, experimental results for the BCI-IV-2A and BCI-IV-2B datasets were obtained using a subject-dependent evaluation method. For example, the domain combination [(D1, D2), (D3)] [(D1, D2), (D3)] [(D1, D2), (D3)] indicates that domains D1D1D1 and D2D2D2 are sent to the cross-domain attention module for processing, and then D3D3D3 is sent to the cross-domain attention module. "No combination" indicates a combination without domain division.
[0139] Table 5: Pre-training classification results of the cross-task feature learning module on the motor imagery task BCI-IV-2A dataset.
[0140]
[0141] Table 6: Pre-training classification results of the cross-task feature learning module on the motor imagery task BCI-IV-2B dataset.
[0142]
[0143] Table 7: Pre-training classification results of the cross-task feature learning module on the motor imagery task SEED-IV and SEED-V datasets.
[0144]
[0145] In the BCI-IV-2A dataset, the combination [(D1,D2),(D3)][(D1,D2),(D3)][(D1,D2),(D3)], in the BCI-IV-2B dataset, the combination [(D1,D2)][(D1,D2)][(D1,D2)], and in the SEED-V dataset, the combination [(D1,D3),(D2,D4)][(D1,D3),(D2,D4)][(D1,D3),(D2,D4)] shows the highest classification accuracy, which indicates that these domain combinations enable the cross-domain attention module to effectively exploit the interactions between domains.
[0146] Therefore, this improvement demonstrates that the designed cross-task feature learning module can effectively utilize the attention mechanism to better identify and utilize important EEG field information.
[0147] Example 1
[0148] as follows Figure 12 ,13 shows the specific structure and processing steps of the domain adaptive data synthesis contrast diffusion model for motion imagery and emotion classification proposed in this embodiment:
[0149] Step 1: Collect the original EEG signal and preprocess it;
[0150] The preprocessing is as follows: using a window with a length of 225 and a step size of 75 to divide the data into overlapping original signal inputs;
[0151] Step 2: In addition, no additional data processing methods are used, and the raw data is directly used as input;
[0152] Step 3: Pre-train the cross-task feature learning module to obtain the specified domain combinations under different tasks;
[0153] Step 4: The original signal input goes through the forward diffusion process until it reaches the data x at time T T ;
[0154] Step 5: Perform the reverse process and feed it into the contrast denoising model. The data of each branch comes from x T
[0155] Step 6: x T The data stream enters the clean signal extraction module, first passes through the four-unit progressive concave network, and finally passes through the convolution layer to adjust the dimension and send it to the mean square error calculation to output the clean data at the previous time step.
[0156] Step 7: At the same time, x TThe data stream enters the domain adaptation noise extraction module, which first embeds the subject code into the data through the subject feature learning block to capture the subject differences. Then the data stream enters the cross-task feature learning block to learn the differences under different tasks. Finally, it is sent to the self-attention layer to adaptively integrate the subject information and task information to output the domain adaptation noise. Then the subjects are classified into different characters through the classification layer.
[0157] Step 8: Repeat steps 6-8 for each time step from time T to time 0 until time 0 is reached, and finally the denoised data is obtained.
[0158] Step 9: The final denoised data will be obtained Send it to the standard classification network EEGNET;
[0159] The following experiments and evaluations are carried out:
[0160] The experiment uses two types of task datasets:
[0161] 1) Public datasets for motor imagery: BCI-IV-2A and BCI-IV-2B.
[0162] The BCI-IV-2A dataset contains EEG recordings from nine participants, each performing four motor imagery tasks: left hand, right hand, foot, and tongue movements. EEG data were recorded using 22 channels at a sampling rate of 250 Hz. Each participant's data consisted of two sessions, each containing 72 trials per task type, for a total of 288 trials. The 2B dataset from the BCI Competition IV contains EEG recordings from nine participants and is a valuable resource for motor imagery research. Participants completed five sessions totaling 720 trials, focusing on left and right hand motor imagery tasks. Data were acquired using three precisely placed bipolar EEG channels to ensure accurate spatial resolution. These channels are located at C3, Cz, and C4, and signals were recorded at 250 Hz.
[0163] 2) Public datasets for emotion recognition: SEED-IV and SEED-V
[0164] The SEED-IV dataset contains data from 15 participants. These participants watched 72 film clips carefully selected to elicit four specific emotions: happiness, sadness, fear, and neutrality. The dataset consists of three sessions, each involving watching 24 film clips. EEG data were collected using the international 10-20 system with 62 channels and a sampling rate of 1000 Hz.
[0165] SEED-V is a video-guided EEG dataset for emotional tasks. The experiment involved 16 participants who were exposed to 15 different stimuli across three different trials to elicit emotional responses. These 15 stimuli were designed to elicit five emotion categories: disgust, fear, sadness, neutrality, and happiness. Physiological signal data were recorded using 62 electrodes in a 10-20 system, also at a sampling rate of 1000 Hz.
[0166] The experimental environment was Python 3.8 and an NVIDIA GeForce GTX 4090 GPU. The entire network was implemented using the PyTorch framework. The dataset was trained using 10-fold cross-validation for 400 epochs, with a batch size of 128 samples and a learning rate of 0.001.
[0167] In order to verify the performance of the present invention, it was compared with different models on the motor imagery task, and the comparison results are shown in Tables 8 and 9 below. Tables 8 and 9 compare the subject-dependent average accuracy of different model types on the BCI-IV-2A and BCI-IV-2B datasets. In the BCI-IV-2A dataset, the performance of the model based on the attention mechanism varies greatly between different subjects. In contrast, the hybrid network exhibits stable performance and always maintains a high average accuracy, proving its effectiveness in dealing with subject differences. The specific comparison results are shown in Tables 8 and 9 below.
[0168] Table 8: Comparison of different models on the BCI-IV-2A dataset under the motor imagery task
[0169]
[0170] Table 9: Comparison of results of different models on the BCI-IV-2B dataset under the motor imagery task
[0171]
[0172] Experimental results show that the contrastive diffusion model based on domain-adaptive data synthesis performs well. Compared with the latest methods, the model significantly improves the classification effect, further demonstrating its efficiency and reliability in processing EEG emotion data.
[0173] Example 2:
[0174] This example uses the proposed domain-adaptive data synthesis contrast diffusion model on the emotion recognition task:
[0175] Step 1: Collect the original EEG signal and preprocess it;
[0176] The preprocessing steps are: resampling the data points of the emotion recognition dataset to the motor imagery baseline sampling rate of 250 Hz. Using a window with a length of 225 and a step size of 75, the data is divided into overlapping raw signal inputs;
[0177] The subsequent steps 2 to 10 and the experimental environment are the same as those in Example 1.
[0178] To verify the performance of the present invention on different datasets, we compared it with the comparison method in Example 1, and the comparison results are shown in Tables 10-11 below. Tables 10 and 11 compare the subject-dependent average accuracy of different methods on the SEED-IV and SEED-V datasets. In the SEED-IV dataset, the domain-adaptive data synthesis contrastive diffusion model achieved the highest accuracy of 93.87%, surpassing other hybrid network models. Similarly, in the SEED-V dataset, the domain-adaptive data synthesis contrastive diffusion model further improved, reaching an average accuracy of 96.21%. These results fully demonstrate the robustness and efficiency of the domain-adaptive data synthesis contrastive diffusion model in emotion recognition tasks on both datasets.
[0179] Table 10: Comparison of SEED-IV dataset results for emotion recognition tasks
[0180]
[0181] Table 11: Comparison of different models on the SEED-V dataset for emotion recognition tasks
[0182]
[0183] The experimental results show that the accuracy of the emotion recognition method using the domain adaptive data synthesis contrast diffusion model is 93.83% and 82.11% on the two datasets respectively, which is significantly improved compared with other comparison methods, effectively verifying that the domain adaptive data synthesis contrast diffusion model emotion recognition method can achieve good recognition accuracy.
[0184] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.
Claims
1. An EEG signal classification method based on a domain adaptive data synthesis contrast diffusion model, characterized in that: include: Construct a domain-adaptive data synthesis contrast diffusion model, including forward diffusion and backward diffusion processes; The forward diffusion process is used to train noisy data. The original signal input passes through the forward diffusion process until it reaches the data x at time T. T ; The reverse diffusion process is provided with a contrast denoising module, which is used to effectively learn clean signals and domain noise in the reverse diffusion process to generate clean data; the data x at time T T is a noise signal, the reverse diffusion process converts the noise signal x T By gradually removing noise and finally restoring it to a clean signal Finally, the clean signal Send it to the standard classification network EEGNET for classification.
2. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 1 is characterized in that: The forward diffusion process is used to train noisy data. The original signal input passes through the forward diffusion process until it reaches the data x at time T. T ,include: Preprocessing: 2D EEG data Converted to stacked 3D data using sliding windows Where C represents the number of channels, W represents the window length, and N represents the number of windows; The pre-processed two-dimensional EEG signal data is input into the forward diffusion process. By iteratively adding noise, the signal passes through the intermediate states x1, x2, ..., x T-11 , until the final noise state x1 is reached, the iterative process uses the sequence q(x T |x T-1 )1 means that a small amount of noise is added to the previous state at each step.
3. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 1, characterized in that: The data x at time T T is a noise signal, the reverse diffusion process converts the noise signal x T By gradually removing noise and finally restoring it to a clean signal Back diffusion process using sequence express; In the reverse diffusion process, x T First, it is sent to the clean signal extraction module and the domain noise adaptive extraction module at the same time. The contrast denoising module separates the clean EEG signal from the domain noise. The clean EEG signal is obtained by the clean signal extraction module, and the domain noise is captured by the domain noise adaptive extraction module. At the end of back diffusion, a clean signal can finally be obtained.
4. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 3 is characterized in that: The clean signal extraction module consists of four unit progressive concave networks connected in series. The module uses unit progressive concave networks and residual connections. A residual connection is added between every two consecutive unit progressive concave networks, where the input of each unit progressive concave network is composed of the input and output of the previous unit progressive concave network. Each unit in the progressive concave network is called a unit, denoted as C, where i represents the index of the corresponding unit, as shown in formula (1), and Represent the input and output of the unit indexed by i respectively; According to the Gaussian distribution assumption in the diffusion denoising probability model, the variance is set to ε T , where ε T represents the Gaussian noise parameter; the parameter extracted from the clean signal is defined as ε C Indicates the use to restore clean signals The noise from x T arrive The time step training loss follows formula (2); L c =MSELoss=‖e T -e C ‖ 2 (2)。 5. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 4 is characterized in that: The back-diffusion process of the diffusion denoising probabilistic model consists of a continuous progressive concave network module, which is used to enhance the learning ability of high-dimensional EEG signal features; The continuous progressive concave network module consists of two stages: downsampling and upsampling. The first stage is the downsampling process. Each downsampling block in the downsampling process contains two submodules: a fine downsampling block and a progressive downsampling block. At the same time, the upsampling process follows the processing method of the original concave network. The downsampling process of the continuous progressive concave network module is shown in formula (3) and formula (4); Formula (3) represents the sequence calculation process, where each fine dimensionality reduction block is represented by F j Indicates that each progressive dimensionality reduction block is represented by P j Indicates that j is the index of the dimension reduction block; and They represent the outputs of the fine dimensionality reduction block and the progressive dimensionality reduction block with index j respectively; The fine dimensionality reduction block contains an improved residual block in which the original two-dimensional convolution is replaced by depth-wise separable convolution and point-wise convolution; in addition, a pixel normalization layer is added to the residual block to standardize the feature distribution.
6. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 3, characterized in that: The domain noise adaptive extraction module contains a subject feature learning block, a cross-task feature learning block and a self-attention layer arranged in sequence; The domain noise adaptive extraction module uses the EEG network classifier EEGNet to extract the output of the module Extract features From x T arrive The loss function Ls of the time step is based on the improved Arc-Margin Loss function, as shown in formula (5); The parameter m is used to increase the closeness of features within the same subject during classification, while increasing the discrimination between different subjects, and j is used to index all other subjects except the current subject; in addition, N represents the batch size used to calculate the average loss. represents the label of the target subject; Formula (6) shows The calculation process of In formula (6), and Represents the weight of the classifier.
7. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 6, characterized in that: The subject feature learning module consists of a series of stacked modules for learning subject-specific unique features in EEG data; At the initial time step, from x T arrive Input x T is fed into the network; a two-dimensional convolutional layer is used to capture the spatial patterns in the EEG signal, thereby retaining key features, while further complex features are extracted through the residual block; The calculation of the subject's attention is shown in formula (7), where Q Subject represents the subject embedding calculated by Q, and the inputs of K and V come from the output of the previous layer; 8. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 6, characterized in that: The cross-task feature learning module is used to address task differences and enhance the extraction of domain features in different EEG tasks. It consists of three parts: Part A is region division, Part B is feature extraction, and Part C is cross-task domain calculation; In part A, the input data first passes through the domain partitioning layer; In Part B, the input data are domain, and emotion recognition tasks field; In part C, the cross-domain integration module is used to solve the problem of capturing subtle relationships between features in different domains of EEG signals. The input of the cross-domain integration module is the domain features extracted in part B. Output D i It is fed into the cross-domain attention layer, which is shown in the pink module in Figure 8. The calculation process of cross-domain attention first transforms input 1 and input 2 according to formula (8) and formula (9); D i Indicates the input from input 1, D j represents the input from input 2; Q is the query matrix, K is the key matrix; Subsequently, the relationship between the input fields is calculated using formula (10) and formula (11); in, and represents the dimension of the input query. The values of i and j in different tasks are shown in formulas (12) and (13); Finally, the concatenated output is passed through a random dropout layer as the final output O, where n represents the index of the attention output, as shown in formula (14); O=Dropout(Concatenate(Attention n ))for n=1 to 2 (14)。 9. The EEG signal classification method based on domain adaptive data synthesis contrast diffusion model according to claim 6, characterized in that: The contrast denoising module is used to separate the clean module from the noise module, thereby achieving denoising; the output of the module is represented as a negative sample The positive samples generated by the module are expressed as The contrast loss function at the current time step is defined as shown in formula (15). Loss function L CL It consists of two components: the first is the s and x c Regularize the similarity matrix to prevent the model from overfitting features; The second term maximizes x S and x C cosine similarity between to enhance feature alignment; Through these two aspects, L CL Ensures effective alignment and separation of features; The total loss function is shown in formula (16); the hyperparameter λ cl Controls the weight of the contrast loss, λ s Adjust the loss of the domain noise adaptive extraction module in the reverse process, λ c The reverse loss in the clean signal extraction process is crucial; ultimately, the total loss function is expressed as formula (16): L=λ cl L CL +λ s L S +λ c L C (16)。