Electroencephalogram recognition method based on 4D pulse neural network
By constructing the EEGSNet model and combining 4D time-frequency spatial representation and adaptive threshold strategy, the problem of limited recognition ability of existing SNN models in EEG cognitive recognition tasks is solved, achieving stronger recognition performance and generalization ability, and promoting the development of the field of brain-computer interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing spiking neural network models lack specificity in EEG cognitive recognition tasks and cannot effectively take into account the multidimensional characteristics of EEG and individual differences, resulting in limited recognition capabilities.
We construct the EEGSNet model based on a 4D spiking neural network, and combine 4D time-frequency spatial representation, dilated pulse encoder and adaptive threshold strategy to design an SNN structure specific to EEG. We enhance the feature representation and recognition capabilities of EEG signals through dilated convolution and PLIF model.
It improves the recognition performance and generalization ability of EEG cognitive recognition tasks, is applicable to various EEG datasets, and promotes the development of the field of brain-computer interaction.
Smart Images

Figure CN116369945B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of brain cognitive computing, and particularly to a neural electrophysiological signal analysis technology and a third-generation artificial neural network model construction method. The present application is a 4D spiking neural network-based electroencephalogram (EEG) cognitive state recognition method, which can effectively solve the problem of unstable time series signals, comprehensively consider the multi-dimensional characteristics of EEG, and be suitable for different EEG recognition tasks. BACKGROUND
[0002] Spiking Neural Networks (SNN) is the third-generation artificial neural network, which adopts a bionic mechanism and is closer to natural neural networks, and has rich spatiotemporal dynamics. Electroencephalogram (EEG) is a time-series physiological electrical signal collected from the cerebral cortex by a non-invasive method, which can objectively reflect the neural activity of the brain and has been widely used to assess the cognitive state of the brain. Since EEG data is a typical neuro-morphological data with time-frequency-space multi-dimensional information, the use of SNN or its variant neural networks for EEG recognition has attracted more and more attention from researchers. However, existing EEG-based researches only use SNN as a classifier to perform cognitive recognition tasks, and lack specific SNN models for EEG analysis. In the field of SNN, the demand for designing a specific SNN model for EEG analysis is increasing in order to further improve the EEG cognitive recognition ability of the model.
[0003] In theory, SNN uses a series of spike sequences to transmit information, which is a discrete event occurring at a time point. The purpose is to construct a spike-based neural network by simulating interpretable biological neuron mechanisms. In SNN, existing more complex and biology-related neuron models mainly include: Integrate-and-Fire (LIF), Izhikevich, Hodgkin-Huxley, and Spike Response Model (SRM) neuron models. In the learning process of SNN, the most important thing is the membrane potential of neurons. Essentially, once a neuron reaches a certain potential (i.e., threshold voltage), a spike will occur, and then the neuron that reaches the potential will be reset. For this, the most commonly used model is the LIF neuron model, whose input signal directly affects the state (membrane potential) of the neuron, and only when the membrane potential rises to the threshold potential, an output spike signal will be generated. Based on this model, some variant methods have been proposed to improve the performance of the model. For example, the PLIF model, a training algorithm that can train the synaptic weights and membrane potential time parameters of the SNN model simultaneously, improves the expression ability of the SNN.
[0004] To further optimize the SNN model, on the one hand, the existing SNN variant model mainly explores the pulse neuron and layer structure of the SNN model, such as regularization and parameterization. Among them, the regularization mainly includes the regularization optimization of SNN from the aspects of layer structure, weight and time. Parameterization mainly aims at enhancing the performance of SNN by learning synaptic weight and membrane potential time delay. On the other hand, relying on mature and effective artificial neural network (ANN) training algorithm, the SNN model guided by ANN is designed to achieve better performance. For example, by using the firing rate of each pulse neuron to approximate the corresponding ReLU activation, the pre-trained ANN is converted into SNN. In addition, due to the unique time domain information processing capability of SNN, these networks have been widely used to process neuromorphic data sets such as complex image recognition and natural language processing tasks, and have made certain progress. However, for EEG data with unstable characteristics, the performance of these networks is greatly limited.
[0005] In the field of EEG analysis, current research focuses on applying SNN models to simple classification tasks based on EEG, such as motor imagery, disease detection, and emotion recognition. Existing research characterizes electroencephalogram data from the time and space dimensions, and uses SNN as a classifier to complete the recognition of given cognitive tasks. Although these studies have certain good results, they lack comprehensive analysis of global and local fine-grained information of EEG multi-dimensional information, ignore significant individual differences and non-stationary characteristics, and limit the ability to further enhance the existing SNN model to perform EEG-based cognitive recognition tasks. Therefore, developing a reliable EEG-specific SNN model needs to design SNN structure from the nature of EEG data to improve the performance of SNN model in EEG recognition tasks.
[0006] In view of the above challenges, the present application is based on the third generation of pulse neural network, and constructs a pulse neural network specific to EEG, and a cognitive recognition method considering global and local fine-grained features of EEG, which is applicable to various cognitive recognition tasks based on EEG and suitable for various EEG data sets. SUMMARY
[0007] The purpose of the present application is to overcome the shortcomings of the prior art and provide a 4D pulse neural network-based electroencephalogram cognitive recognition method. The method of the present application proposes an EEGSNet model, a new 4D pulse neural network specific to EEG cognitive recognition tasks. According to the characteristics of EEG, a 4D time-frequency-space representation format based on time-frequency-space is designed, and an EEG-specific SNN network structure with adaptive strategy is introduced, and the global and local fine-grained features of multi-dimensional information are explored to enhance the recognition ability of SNN in EEG cognitive tasks.
[0008] The 4D-based EEGSNet provided by the application is a pulse neural network specific to EEG, and the framework mainly includes two modules: a 4D time-frequency-space representation module and an EEG-based SNN module.
[0009] The technical solution adopted by the application to solve its technical problems specifically includes the following steps:
[0010] Step 1: Data acquisition
[0011] The application is verified on three data sets, including SEED (public emotion data set), FAAD (fatigue driving data set) and ADMH (cognitive impairment cognitive data set). The SEED data set is a three-class emotion data set containing 15 subjects (positive, neutral and negative), and 62 channels of EEG data are collected. The FAAD data set collects 61 channels of EEG data of 15 subjects, including two cognitive states (awake and fatigue). The ADMH data set is an EEG data set for 33 subjects of cognitive impairment groups and control groups, collecting 32 channels of data following the international 10-20 system, and obtaining EEG data of subjects in different emotional stimuli in three cognitive states (cognitive impairment, mild cognitive impairment and normal elderly).
[0012] Step 2: Feature extraction
[0013] Firstly, the three data sets are preprocessed to filter noise and remove artifacts.
[0014] The application uses a band-pass filter to eliminate high-frequency noise and power frequency interference signals other than spontaneous brain electrical signals, and then uses an independent component analysis (ICA) method to remove artifact signals. For the preprocessed EEG data, differential entropy (DE) is used to extract signal features. Among them, the SEED and FAAD data sets both contain DE features of 5 frequency bands (delta: 1-3Hz, theta: 4-7Hz, alpha: 8-13Hz, beta: 14-30Hz, gamma: 31-50Hz), and the ADMH data set has 4 frequency bands (delta: 0.5-3Hz, theta: 4-7Hz, alpha: 8-13Hz, beta: 14-30Hz). The preprocessing and feature extraction operations provide stable EEG signal features for subsequent model construction.
[0015] Step 3: EEG cognitive recognition method based on 4D pulse neural network
[0016] 3-1. Input: DE feature sample data with cognitive state labels Where N is the total number of samples, the time step t, and the maximum training period T.
[0017] 3-2.4D time-frequency-space representation module: In order to make full use of the time-frequency-space complementary characteristics of the EEG signal, the present application converts it into a 4D time-frequency-space representation. An EEG segment containing T time steps is defined as where C represents the number of channels, and F represents the number of frequency bands extracted from the EEG signal. is the 3D space-frequency representation of the EEG feature at the t time step. Among them, the fth frequency band will be converted into a 2D space map H and W are defined as the length and width of the 2D space map of all channels. Finally, N 4D time-frequency-space representations of the EEG signal are obtained N is the total number of samples.
[0018] 3-3. Hollow pulse encoder, the 4D time-frequency-space representation of the EEG is directly used as the input of the hollow pulse encoder to retain as much information as possible. The designed hollow pulse encoder is used to encode the pulse representation based on the EEG, which can effectively and efficiently integrate and extract the comprehensive features of time-frequency-space. The hollow pulse encoder mainly includes a hollow convolution layer and a pulse neuron layer, in which the dynamics of the pulse neuron is described as charging, discharging and resetting process.
[0019] 3-3-1. Hollow convolution layer: Considering the strong correlation between cognitive state and brain functional connectivity, the internal correlation between channels in the EEG space is further explored. The introduction of hollow convolution changes the convolution receptive field, thereby extracting channel correlation information in different ranges to enhance the representation of spatial fine-grained features. An alternating strategy of convolution and hollow convolution is proposed to capture adjacent and nearby spatial fine-grained features at the same time, and to preserve the continuity between samples. Structurally, a 3x3 convolution kernel with a hollow rate of 1 is used to learn the adjacent features in the spatial domain. Then the hollow rate is expanded from 1 to 2, which aims to further obtain nearby features. By continuously iterating different sizes of hollow rate to learn the fine-grained features between adjacent and nearby channels. The extracted global and local fine-grained features are used as the input of the pulse neuron for subsequent pulse coding.
[0020] 3-3-2. Trainable membrane potential parameters: In the pulse neuron layer, the pulse neuron receives synaptic input from the hollow convolution to perform pulse training. In the SNN module, the presynaptic neuron activity affects the membrane potential of the postsynaptic neuron, and when the membrane potential exceeds a certain threshold value, a pulse will be activated. The LIF model is the simplest and most commonly used pulse neuron model, and its subthreshold (when the membrane potential does not exceed the threshold voltage) neural dynamic equation is described as:
[0021]
[0022] where, V tis the membrane potential of a neuron at time t, reset is the reset voltage after an activation spike, is the control input of the neuron, t is the decayed membrane time constant, t denotes the presynaptic input of a neuron at time t. The neurons in the SNN module use discrete difference equations to approximate the continuous differential equations, from the perspective of the difference equation, (1) can be rewritten as:
[0023]
[0024] Due to the instability of the EEG signal, the present application combines a parameterized LIF model with a learnable membrane time constant, i.e., a PLIF model, to enhance the expressiveness of the neurons in the SNN module on the EEG. In the PLIF model, the manually adjusted membrane-related parameters are replaced by a sigmoid function with trainable parameters. Further, (2) can be re-expressed as the following equation:
[0025]
[0026] where f(·) is the spiking neuron model function.
[0027] 3-3-3. Sample-based adaptive threshold: for the PLIF model, the dynamic process of all spiking neurons is described by three discrete equations, including charging, discharging, and resetting, as follows:
[0028]
[0029]
[0030]
[0031] where l is the index of the layer. H t and V t denote the membrane voltage of a neuron before and after activation at time t, respectively. O t is the output spike at time t, determined by the Heaviside step function Θ(·), i.e., there is a spike output of 1 otherwise 0. V th is the threshold voltage. The arctangent function is used as a gradient replacement method to calculate the gradient of the spiking function in the backpropagation process. Formula (4) calculates the neural dynamics of the PLIF neuron through formula (3), where w i is the synaptic weight, defined as the binarized spike output of the (l-1)-th layer at time t.
[0032] Considering the unstable characteristics of EEG signals, the present application proposes a sample-based adaptive threshold strategy to enhance the expressiveness of features and eliminate the negative effects of individual differences in EEG signals as much as possible. The adaptive threshold is determined by the mean value of the membrane voltage of the sample within T time lengths, which can be described as:
[0033]
[0034] wherein E[Z n ] is the mean value of the dynamic membrane voltage of the nth sample, is the adaptive activation threshold statistical estimation value of a small batch of samples in a layer. During the training process, the sample adaptive threshold of each layer is updated according to the membrane voltage of the neuron. The adaptive threshold strategy can keep the presynaptic stimulus intensity within a relatively stable range, effectively avoiding the disappearance or explosion of pulse activity during the training process.
[0035] 3-4. Pulse classifier: The output of the hollow pulse encoder is taken as the input of the pulse classifier. Structurally, the pulse classifier is composed of multiple layers of fully connected pulse neural layers, i.e., multiple fully connected-PLIF network structures. The full connection is realized by the Linear() linear transformation function, which maps the input to the sample label space. The PLIF structure is consistent with that in the hollow pulse encoder, i.e., the adaptive threshold is introduced in the discharge process of the pulse neuron. Finally, a binary tensor is output.
[0036] 3-5. Output layer: In the SNN module, the output of the pulse neuron is binary, and the average pulse firing frequency within T time steps is taken as the final classification decision. The output of the EEG-based SNN module is which is a CxT tensor, where C is the number of categories and T is the time step. The average pulse firing frequency of T time steps is calculated as The high and low firing frequency represents the degree of response to different categories. The present application uses mean square error (MSE) as the loss function L to encourage the highest firing frequency of the neuron to respond to the correct label, which is defined as follows:
[0037]
[0038] wherein Y i is the true one-hot label of the sample. It is worth noting that after each network parameter optimization, the state of the pulse neuron needs to be reset, because the neuron of the SNN is stateful.
[0039] The present application has the following beneficial effects:
[0040] The electroencephalogram recognition model provided by the application is a specific model for electroencephalogram tasks, is suitable for various EEG signal-based recognition tasks, has optimal performance and strong generalization ability, and further promotes the development of the field of brain-computer interaction. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 is a structural diagram of the application. DETAILED DESCRIPTION
[0042] The application will be further described below in combination with the drawings and examples.
[0043] The application first converts the extracted EEG signal features into 4D representation as the input of the spiking neural network to explore the time-frequency-space global features of EEG. The alternating transformation strategy of convolution and dilated convolution is introduced to explore the potential relationship between adjacent and similar channels of EEG signals to obtain the local fine-grained features in the spatial domain. In addition, the parameterized Integrate-and-Fire (PLIF) model is used as the benchmark network framework. Considering the unstable characteristics of EEG signals, a sample-based adaptive threshold strategy is proposed to eliminate the negative effects of the characteristics as much as possible and improve the generalization ability of the 4D spiking neural network in EEG recognition tasks. The application has achieved optimal performance on various recognition data sets, and the model effectiveness has been verified in the inter-subject and cross-subject recognition tasks. The model generalizes to various electroencephalogram recognition tasks, which greatly promotes the development of the field of brain-computer interaction.
[0044] The main contributions are described as follows: 1) based on the multi-dimensional EEG signal of time-frequency-space, a 4D spiking neural network (EEGSNet) is proposed to dynamically capture the global information of time-frequency-space domain. In order to further explore the internal relationship between adjacent and similar channels in the EEG space, dilated convolution is introduced to extract fine-grained features. 2) The designed EEGSNet is a SNN model specific to EEG signals, and its structure is based on the parameterized LIF model (PLIF) with a trainable time delay constant, and a sample-based adaptive threshold strategy is introduced to improve the ability of neurons to adapt to time-stable EEG signals.
[0045] The design of a spiking neural network (SNN) is more like a biological neuron of the brain, which can process timing information. Although SNN has made some progress in EEG recognition tasks, existing research focuses on the classification performance brought by the SNN model, ignores the unique time-frequency-space multidimensional characteristics and instability of EEG signals, and leads to great challenges for the SNN model to improve the robustness and generalization ability of EEG recognition. In order to solve this challenge, the present application proposes a 4D specific EEG recognition task based on a spiking neural network, named EEGSNet, which can enhance the recognition ability of EEG cognitive state. The present application explores the time-frequency-space multidimensional features from the global and local perspectives, and designs a new sample adaptive threshold strategy based on the threshold strategy, which can improve the generalization ability of EEGSNet in the EEG recognition task. Through feature characterization and analysis of multidimensional EEG data, cognitive state classification is realized in inter-subject and cross-subject tasks, and verification is completed on different task EEG data sets, including emotion recognition (positive, neutral and negative emotion cognitive state), fatigue detection (awake and fatigue cognitive state), cognitive impairment diagnosis (cognitive impairment, mild cognitive impairment and normal elderly cognitive state under different emotions), etc. The present application has the best performance and generalization ability, and can further promote the development of brain-computer interaction (BCI) field.
[0046] The new EEG recognition model structure proposed is shown in Figure 1 EEGSNet is a cognitive state recognition method specific to EEG based on a spiking neural network, which includes two modules: 4D time-frequency-space representation and EEG-based spiking neural network. The EEGSNet model mainly includes the following steps when performing the EEG recognition task:
[0047] Step 1: Data acquisition
[0048] The present application is verified on three data sets, including SEED (public emotion data set), FAAD (fatigue driving data set) and ADMH (cognitive impairment cognitive data set). The SEED data set is a three-class emotion data set containing 15 subjects (positive, neutral and negative), which collects 62-channel EEG data. The FAAD data set collects 61-channel EEG data of 15 subjects, including two cognitive states (awake and fatigue). The ADMH data set is an EEG data set for 33 subjects of cognitive impairment groups and control groups, which collects 32 channels of data following the international 10-20 system, and obtains EEG data of subjects in different emotional stimuli under three cognitive states (cognitive impairment, mild cognitive impairment and normal elderly).
[0049] Step 2: Feature extraction
[0050] Firstly, the three datasets are preprocessed to filter noise and remove artifacts.
[0051] The present application uses a band-pass filter to eliminate high-frequency noise and power frequency interference signals other than spontaneous brain electrical signals for raw EEG data, and then uses the independent component analysis (ICA) method to remove artifact signals. For the preprocessed EEG data, the differential entropy (DE) is used to extract signal features. Among them, the SEED and FAAD datasets contain DE features of 5 frequency bands (delta: 1-3Hz, theta: 4-7Hz, alpha: 8-13Hz, beta: 14-30Hz, gamma: 31-50Hz), and the ADMH dataset has 4 frequency bands (delta: 0.5-3Hz, theta: 4-7Hz, alpha: 8-13Hz, beta: 14-30Hz). The preprocessing and feature extraction operations provide stable EEG signal features for subsequent model construction.
[0052] Step 3: EEG cognitive recognition method based on 4D pulse neural network
[0053] 3-1. Input: DE feature sample data with cognitive state labels Where N is the total number of samples, time step t, and maximum training period T.
[0054] 3-2. 4D time-frequency-space representation module: In order to make full use of the time-frequency-space complementary features of EEG signals, the present application converts them into 4D time-frequency-space representation. An EEG segment containing T time steps is defined as Where C represents the number of channels, and F represents the number of frequency bands extracted from the EEG signal. is the 3D space-frequency representation of the EEG feature at time step t. Among them, the fth frequency band will be converted to a 2D space mapping H and W are defined as the length and width of the 2D space mapping of all channels. Finally, N 4D time-frequency-space representations of the EEG signal are obtained N is the total number of samples.
[0055] 3-3. Hollow pulse encoder, the 4D time-frequency-space representation of the EEG is directly used as the input of the hollow pulse encoder to preserve as much information as possible. The designed hollow pulse encoder is used to encode the pulse representation based on the EEG, which can effectively and efficiently integrate and extract the comprehensive features of time-frequency-space. The hollow pulse encoder mainly includes a hollow convolution layer and a pulse neuron layer, wherein the dynamics of the pulse neuron is described as a charging, discharging and resetting process.
[0056] 3-3-1. Considering the strong correlation between cognitive states and brain functional connectivity, further explore the intrinsic correlation between channels in the EEG spatial domain. Introduce a hollow convolution to change the convolution receptive field, so as to extract channel correlation information in different ranges, and enhance the representation of spatial fine-grained features. An alternating strategy of convolution and hollow convolution is proposed to capture adjacent and nearby spatial fine-grained features at the same time, and to preserve the continuity between samples. Structurally, a 3x3 convolution kernel with a hollow rate of 1 is used to learn the adjacent features in the spatial domain. Then the hollow rate is expanded from 1 to 2, the purpose is to further obtain the nearby features. By continuously iterating different sizes of hollow rate to learn the fine-grained features between adjacent and nearby channels. The extracted global and local fine-grained features are taken as the input of the pulse neuron for subsequent pulse coding.
[0057] 3-3-2. Trainable membrane potential parameters: in the pulse neuron layer, the pulse neuron receives synaptic input from the hollow convolution to perform pulse training. In the SNN module, presynaptic neuron activity affects the membrane potential of postsynaptic neurons, and when the membrane potential exceeds a certain threshold, a pulse will be activated. The LIF model is the simplest and most commonly used pulse neuron model, and its subthreshold (when the membrane potential does not exceed the threshold voltage) neural dynamic equation is described as:
[0058]
[0059] Where V t is the membrane potential of the neuron at time t, V reset is the reset voltage after the pulse is activated, is the membrane time constant that controls the decay of V t , and Z t represents the presynaptic input of the neuron at time t. The neurons in the SNN module use discrete difference equations to approximate continuous differential equations. From the perspective of difference equations, (1) can be rewritten as:
[0060]
[0061] Due to the instability of EEG signals, the present application combines a parameterized LIF model with a learnable membrane time constant, namely a PLIF model, to enhance the expressiveness of the neurons in the SNN module on EEG. In the PLIF model, the manually adjusted membrane-related parameters are replaced by a sigmoid function with trainable parameters. Further, (2) can be represented as the following equation:
[0062]
[0063] Where f(·) is the pulse neuron model function.
[0064] 3-3-3. Sample-based adaptive threshold: For the PLIF model, the dynamics of all spiking neurons are described by three discrete equations, including charge, fire and reset, as follows:
[0065]
[0066]
[0067]
[0068] where l is the index of the layer. H t and V t represent the membrane voltage before and after the activation of a neuron at the t-th time, respectively. O t is the output spike at the t-th time, determined by the Heaviside step function Θ(·), i.e., 1 if there is a spike output, otherwise 0. V th is the threshold voltage. The arctangent function is used as a gradient replacement method to calculate the gradient of the spike function in the backpropagation process. Formula (4) calculates the neural dynamics of the PLIF neuron through formula (3), where w i is the synaptic weight, defined as the binarized spike output of the (l-1)-th layer at the t-th time.
[0069] Considering the unstable characteristics of EEG signals, the present application proposes a sample-based adaptive threshold strategy to enhance the expressiveness of features and eliminate the negative effects of individual differences in EEG signals as much as possible. The adaptive threshold is determined by the mean of the sample membrane voltage within T time lengths, which can be described as:
[0070]
[0071] where E[Z n ] is the mean of the dynamic membrane voltage of the n-th sample, is the adaptive activation threshold statistical estimate value of a small batch of samples in a layer. During the training process, the sample adaptive threshold of each layer is updated according to the membrane voltage of the neuron. The adaptive threshold strategy can keep the presynaptic stimulus intensity within a relatively stable range, effectively avoiding the disappearance or explosion of pulse activity during the training process.
[0072] 3-4. Pulse classifier: The output of the hollow pulse encoder serves as the input of the pulse classifier. Structurally, the pulse classifier consists of multiple layers of fully connected pulse neural layers, i.e., multiple fully connected-PLIF network structures. Its full connection is realized by the Linear() linear transformation function, which maps the input to the sample label space. The PLIF structure is consistent with the PLIF in the hollow pulse encoder, i.e., the adaptive threshold is introduced during the discharge process of the pulse neuron. Finally, a binary tensor is output.
[0073] 3-5. Output layer: In the SNN module, the output of the pulse neuron is binary, and the average pulse firing frequency within T time steps is taken as the final classification decision. The output of the EEG-based SNN module is O = f (X) = Linear (X) is a CxT tensor, where C is the number of categories and T is the time step. The average pulse firing frequency of T time steps is calculated as The high and low firing frequencies represent the degree of response to different categories. The present application uses mean square error (MSE) as the loss function L to encourage the highest firing frequency of the neuron to respond to the correct label, which is defined as follows:
[0074]
[0075] where Y i is the true one-hot label of the sample. It is worth noting that after each network parameter optimization, the state of the pulse neuron needs to be reset because the neuron of the SNN is stateful.
[0076] The EEGSNet model specific to EEG signals proposed by the present application can be applied to any EEG-based cognitive state recognition task, fully utilizes the multi-dimensional features of EEG, to a certain extent, solves the negative effects of the instability and individual differences of EEG signals, has the advantages of high recognition performance and strong generalization ability, provides technical support for clinical diagnosis, and promotes the development of the field of brain-computer interaction.
Claims
1. A method for electroencephalic recognition of knowledge based on 4D pulse neural network, characterized in that, Comprising the following steps: Step 1, data acquisition; Step 2, feature extraction; Step 3, 4D-specific EEG recognition task based on pulse neural network, the framework mainly includes two modules: 4D time-frequency space representation module and EEG-based SNN module; wherein, the EEG-based SNN module includes a hollow pulse encoder, a pulse classifier and a pulse output; The 4D time-frequency-space representation module: in order to make full use of the time-frequency-space complementary characteristics of the EEG signal, an EEG segment containing Time steps is defined as Wherein Indicates the number of channels, Indicates the number of frequency bands extracted from the EEG signal; EEG characteristics in Time steps 3D space-frequency representation; wherein the The frequency band Is converted into 2D space mapping , ; and And Respectively defined as the length and width of the 2D space mapping of all channels; finally, the 4D time-frequency-space representation , Of the EEG signal is obtained; total sample number; The hollow pulse encoder: the 4D time-frequency space representation of EEG is directly used as the input of the hollow pulse encoder; the designed hollow pulse encoder mainly includes a hollow convolution layer and a pulse neuron layer, which is used for encoding the pulse representation based on EEG; wherein the dynamics of the pulse neuron is described as charging, discharging and resetting process; The pulse classifier includes: the output of the hollow pulse encoder is used as the input of the pulse classifier, and the structure of the pulse classifier is composed of multiple layers of fully connected pulse neural layers, and the full connection is realized by Linear() linear transformation function, which maps the input to the sample label space; in the discharging process of the pulse neuron, an adaptive threshold is introduced; finally, a binary tensor is outputted; The pulse output comprises an output layer, and the output layer: in the SNN module, the output of the pulse neuron is binarized, and an average pulse firing frequency within T time steps is taken as a final classification decision; the SNN module outputs based on the EEG is a tensor, wherein the number of categories, is a time step; the average pulse firing frequency of the calculated time step , and the high and low of the firing frequency represent the response degree to different categories; a mean square error is used as a loss function to encourage the highest firing frequency of the neuron to respond to the correct label, and is defined as shown in the following formula: ; wherein, is the true one-hot label of the sample.
2. The 4D pulse neural network based electroencephalic cognition recognition method according to claim 1, characterized in that The input of the specific implementation of step 3 includes: 3-1. Input: DE feature sample data with cognitive state labels where is the total number of samples.
3. The 4D pulse neural network based electroencephalic cognition recognition method according to claim 2, characterized in that, The pulse classifier: the output of the hollow pulse encoder is used as the input of the pulse classifier, and the structure of the pulse classifier is composed of multiple layers of fully connected pulse neural layers, and the full connection is realized by Linear() linear transformation function, which maps the input to the sample label space; The PLIF in the hollow pulse encoder is consistent with the PLIF structure, that is, an adaptive threshold is introduced in the discharging process of the pulse neuron; Finally, a binary tensor is outputted.
4. The 4D pulse neural network based electroencephalic cognition recognition method according to claim 3, characterized in that, Trainable membrane potential parameters: in the pulse neuron layer, the pulse neuron receives synaptic input from the hollow convolution to perform pulse training; the PLIF model is used as the pulse neuron model, and the neural dynamic equation is described as: ; in, yes The membrane voltage of the neuron at any given time. It is the reset voltage after the activation pulse. It is control The decaying membrane time constant, Indicates that neurons are The presynaptic input at time t; neurons in the SNN module use discrete difference equations to approximate continuous differential equations. From the perspective of difference equations, formula (2) can be rewritten as: ; In the PLIF model, the manually adjusted membrane-related parameters are replaced by a sigmoid function with trainable parameters; further, formula (3) is re-expressed as the following equation: ; wherein is Impulse neuron model function.
5. The 4D pulse neural network based electroencephalic cognition recognition method according to claim 3, characterized in that, The hollow pulse encoder specifically includes: Atrous convolution layer: an alternating strategy of convolution and atrous convolution is proposed to capture adjacent and nearby spatial fine-grained features simultaneously, so a 3 3convolution kernel learns the adjacent features in the spatial domain, then expands the atrous rate from 1 to 2 to further obtain nearby features, and finally iterates different sizes of atrous rate to learn the fine-grained features between adjacent and nearby channels. The extracted global and local fine-grained features are taken as the input of the pulse neuron for subsequent pulse coding.