Epilepsy prediction method based on self-supervised pulse graph topology collaborative network
Patent Information
- Application Number
- CN202611261282.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
传统的连续数值计算无法有效表征大脑神经元在面临癫痫突发异常放电时的离散脉冲时序动力学特征,导致模型对异常放电前兆中微弱、缓变的时序动态特征不敏感
[0093](1)、本发明针对每个脑电时间窗独立计算导联间相关关系,并融合国际10-20导联物理拓扑先验,构建与脑电时间窗一一对应的样本特异空间图。因此,所构建的图结构能够随脑电状态变化而动态表征导联间功能连接及异常放电传播关系,克服采用固定邻接矩阵难以描述癫痫脑网络动态变化的问题。
Smart Images

Figure CN122822305A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical signal processing and image detection and recognition technology, specifically relating to an epilepsy prediction method based on a self-supervised pulse graph topology cooperative network. Background Technology
[0002] Epilepsy is a chronic brain disorder caused by sudden abnormal electrical discharges in the brain. Long-term electroencephalography (EEG), with its extremely high millisecond-level temporal resolution, can record the electrophysiological activity of the cerebral cortex in real time, making it the primary means of diagnosing epilepsy and detecting abnormal discharges in clinical practice. The core task of epilepsy prediction lies in accurately determining the patient's current brain functional state through continuous monitoring of multi-channel long-term EEG signals, typically including the pre-ictal and interictal phases. Accurate prediction of the pre-ictal phase can provide valuable window time for clinical intervention, drug administration, or closed-loop stimulation, which has significant clinical value in reducing the risk of accidental injury and improving the patient's quality of life.
[0003] In recent years, a large amount of academic and engineering work has revolved around "automatic identification and prediction of epileptic states based on electroencephalogram (EEG) signals." Early research relied heavily on traditional machine learning methods, extracting statistical features (such as spectral power, approximate entropy, and coherence) of multi-channel EEG signals in the time domain, frequency domain, or nonlinear dynamics, and combining them with traditional classifiers such as support vector machines, k-nearest neighbors, and random forests to achieve state discrimination. Subsequently, traditional deep learning models such as one-dimensional convolutional neural networks were introduced, using multi-layer convolutional structures to replace manual feature engineering, thus achieving automatic learning of EEG temporal features.
[0004] As clinical demands for predictive accuracy increase, more and more studies are attempting end-to-end modeling of EEG signals from a spatiotemporal perspective. On the one hand, researchers are using multi-scale convolutional networks or recurrent neural networks (such as LSTM and GRU) to capture the dynamic evolution patterns of EEG signals over time. On the other hand, given the highly spatial topological structure of brain activity, models such as graph convolutional networks (GCN) and spatial-temporal graph networks attempt to explicitly encode each EEG lead as nodes in a graph structure, transforming the functional connections or physical distances between channels into adjacency matrices, thereby introducing a spatial perspective on brain network representation.
[0005] However, existing classification and discrimination methods for predicting epilepsy using electroencephalography (EEG) still have the following significant limitations in actual clinical monitoring and deployment scenarios of portable wearable devices:
[0006] 1. Insufficient modeling of real biological ultra-long-term time dependence and dynamic evolution.
[0007] Epilepsy is a complex, nonlinear, and gradually evolving process, with the evolution from the preictal phase to the interictal phase often lasting several minutes or even hours. Existing epilepsy prediction methods mostly employ shallow convolutional networks or standard recurrent networks, which have limited receptive fields and lack biologically heuristic modeling of transient impulse discharges (such as spikes and sharp waves) in EEG waveforms. Traditional continuous numerical calculations cannot effectively characterize the discrete-pulse temporal dynamics of brain neurons in the face of sudden abnormal discharges in epilepsy, resulting in models being insensitive to the weak, slowly changing temporal dynamics in the precursors of abnormal discharges.
[0008] 2. The contradiction between long-sequence EEG modeling and real-time clinical deployment with low power consumption remains unresolved.
[0009] In long-term EEG monitoring, the high sampling rate and multi-channel ultra-long temporal input place a significant burden on computational resources. Networks such as LSTM / GRU face severe gradient problems and inference latency when processing long sequences; while Transformer and its variants, although possessing global attention, have computational and memory overhead that is quadratically related to sequence length, making them difficult to directly apply to bedside monitoring in hospitals or portable wearable devices. How to balance prediction accuracy, high real-time performance, and ultra-low power consumption in long-sequence EEG modeling remains a challenge for existing traditional deep learning architectures.
[0010] 3. High reliance on precise expert annotation severely limits the model's generalization ability.
[0011] Constructing high-quality epilepsy EEG datasets places extremely high demands on clinical experts and annotation environments. Due to the sudden and unpredictable nature of epileptic seizures, labeled data for the "pre-ictal" and "interictal" phases collected in real clinical scenarios are extremely scarce, and the sample class distribution is severely imbalanced. Existing epilepsy prediction models mostly employ fully supervised learning, which is prone to overfitting to specific subjects or specific acquisition devices on limited labeled samples, resulting in extremely poor generalization and transfer capabilities across subjects and datasets. Although self-supervised learning has shown potential in other general EEG tasks (such as sleep staging), in the field of epilepsy prediction, self-supervised frameworks designed for the spatiotemporal evolutionary characteristics of epilepsy are still relatively scarce.
[0012] 4. Insufficient mining of heterogeneous spatiotemporal information and fragmentation of feature space
[0013] The occurrence and development of epilepsy are simultaneously reflected in EEG signals through single-channel microscopic temporal dynamic pulses and multi-channel macroscopic spatial brain network collaborative diffusion. Existing methods typically use a single data form (purely continuous numerical values or pure pulse events) for modeling, failing to recognize the efficiency of discrete pulse streams in capturing transient discharges and the complementarity of continuous manifolds in capturing macroscopic whole-brain topological networks. Due to the lack of an effective heterogeneous network collaborative architecture, traditional methods cannot achieve interactive calibration of microscopic temporal pulse features and macroscopic spatial continuous graph topological features under a unified contrastive learning paradigm, resulting in a deep lack of cross-modal collaborative coupling relationships between "space and time".
[0014] 5. Robustness and interpretability in the face of complex clinical artifacts need to be improved.
[0015] In actual long-term clinical EEG recordings, serious non-neural artifacts such as eye movements, electromyography, and head movements are inevitably mixed in, and there are significant individual differences in the baseline EEG data among different patients. Existing models show insufficient robustness in the face of these noises, distributional drift, and pathological individual differences. At the same time, due to the lack of a clear architectural design for spatiotemporal correlation, the model prediction results often lack good human readability and calibration ability.
[0016] In summary, existing EEG epilepsy prediction methods still have significant shortcomings in long-term sequence dependency modeling, multi-perspective spatiotemporal deep interaction mining, clinical annotation efficiency, low-power deployment, and cross-subject generalization. Summary of the Invention
[0017] Purpose of the invention: The purpose of this invention is to address the shortcomings of existing technologies and provide an epilepsy prediction method based on a self-supervised pulse graph topology cooperative network.
[0018] For the typical epileptic state prediction scenario of the pre-ictal / interictal phase, there is an urgent need for a method that can fully exploit large-scale unlabeled long-term EEG data using self-supervised pre-training and reduce reliance on manual labels. Furthermore, this method needs to operate within a unified framework, efficiently extracting biologically accurate temporal features through spiking neural networks and accurately capturing multi-lead spatial topological features through graph neural networks, while also considering both local pulse details and global network synergy. This is precisely the technical problem and application requirement that the proposed "Epilepsy Prediction Method and System Based on Spatiotemporal Synergy of Self-Supervised Spiking and Graph Neural Networks" aims to solve.
[0019] Technical solution: The present invention provides an epilepsy prediction method based on a self-supervised pulse graph topology cooperative network, comprising the following steps:
[0020] Step 1: Acquire multi-channel scalp EEG signal samples, and after preprocessing, divide the data into N independent EEG time windows using a sliding window. For each time window, the Pearson correlation coefficient is calculated lead by lead, and the specific adjacency matrix of the samples is obtained. Construct a sample-specific spatial map corresponding to the EEG time window. ; Initial features for nodes; final output is a pair of dual-input samples. This yields a dual-view sample.
[0021] Step 2: Dataset partitioning, which involves dividing all EEG signal samples into an unlabeled pre-training dataset, a labeled fine-tuning dataset, and a test dataset; the labeled samples include two types of labels: pre-ictal and interictal.
[0022] Step 3: Perform bidirectional cross-view pre-training on unlabeled data and conduct self-supervised pre-training on unlabeled training batches. Using any sample within a batch as the anchor sample and the remaining samples as negative samples, temporal masking enhancement is performed on the EEG sequence of the anchor sample to obtain a temporally enhanced view. Random edge pruning and normalization enhancement are performed on the specific adjacency matrix of the anchor sample to obtain a spatially enhanced map. The EEG sequence of the k-th negative sample within the batch is then used to perform the following: Gaussian noise is superimposed to obtain a negative sample time-enhanced view. The normalized adjacency matrix of the negative sample By adding random perturbations and re-adding self-connections and progress degree normalization while maintaining matrix symmetry, an enhanced view of the negative sample space is obtained. ;
[0023] The dual-view enhancement methods here include: anchor sample temporal enhancement is temporal masking, and spatial enhancement is adjacency matrix edge pruning; other samples in the batch temporal enhancement is Gaussian noise superposition, and spatial enhancement is adjacency matrix numerical noise addition; all spatial graphs are enhanced while maintaining adjacency matrix symmetry, and self-connection and degree normalization are re-executed.
[0024] Construct a parameter-shared encoder, which includes an SNN pulse temporal encoder to process the temporal augmentation view and outputs the temporal feature matrix before pooling. and the sample-level time representation after pooling ;
[0025] The GNN spatial topology encoder processes the spatial augmentation graph and outputs the spatial feature matrix of the nodes before pooling. Graph-level spatial representation after pooling ;
[0026] The temporal and spatial features of the anchor samples are projected using temporal and spatial projection heads respectively, and the homologous similarity is calculated. Negative samples within a batch are represented using the same set of shared projection heads, forming a heterologous similarity set with the anchor samples. Cosine similarity is calculated for the normalized projection vectors.
[0027] We construct a bidirectional InfoNCE loss from time to space and from space to time, minimize the batch average bidirectional cross-view loss, and simultaneously optimize the projection head. The projection head here is only used in the self-supervised pre-training stage and is discarded directly after pre-training. If there are false negative samples in the batch, we reduce the interference by increasing the batch size, similarity threshold masking, and supervised contrast loss.
[0028] The heterogeneous similarity set constructed here is the negative sample set in the subsequent InfoNCE loss, enabling the model to distinguish between "temporal-spatial representations of the same EEG time window" and "temporal-spatial representations of different EEG time windows." Specifically, the temporal and spatial projection representations of negative samples within a batch are paired with the hetero-view projection representations of anchor samples to form a heterogeneous negative sample similarity set. This also reduces the similarity between representations of different EEG time windows, thus preventing all sample representations from converging to the same features, and prompting the SNN temporal encoder and GNN spatial topological encoder to learn sample-discriminative and cross-view characteristics under unlabeled conditions. Figure 1 Characteristics of homogeneity;
[0029] Step 4: Fine-tune the labeled data and perform spatiotemporal feature fusion, i.e., load the pre-trained SNN impulse temporal encoder and GNN spatial topological encoder, and retain the temporal feature matrix before pooling. Node space feature matrix Cross-pulse attention bidirectional interaction generates unified spatiotemporal fusion features, which are then fed into the MLP classifier head to output binary prediction probabilities for pre-ictal and interictal periods.
[0030] Furthermore, after obtaining continuous multi-channel EEG signals in step 1, the EEG sequences are sequentially processed by unifying the lead channels, resampling, filtering, and Z-score normalization. Then, the sequences are divided into N independent EEG time windows by sliding windows with fixed window length and overlap rate. , , Represents the real number field. This refers to the number of EEG leads; The number of sampling points within a single window; where, sequence The OK Indicates the first The first time window The length of each lead is The EEG sequence; the window length can be set from 1 to 5 seconds according to the sampling rate and the duration of abnormal discharge, and the overlap rate between windows can be set from 0 to 75%;
[0031] No. The first window , The edge weights between leads are , ;
[0032] ; and These are the EEG sequences of the p-th and q-th leads in the i-th time window, respectively. This represents the Pearson correlation coefficient. Describing covariance, Indicates standard deviation;
[0033] For each time window, the Pearson correlation coefficient is calculated lead-by-lead to generate the instantaneous functional connectivity matrix. Integrating the international 10-20 lead fixed physical topology matrix Functional connection matrix With fixed physical topology matrix The fused adjacency matrix is obtained after fusion. ;right Perform threshold filtering or k-nearest neighbor retention, and add a C-order identity matrix. Supplementing lead self-connection and longitude normalization yields the sample-specific adjacency matrix. ; for The degree matrix is expressed as follows;
[0034] ;
[0035] ;
[0036] In the above formula, for The degree matrix, This represents the fusion weight for the two types of adjacency relationships. The result obtained here... As the input matrix for message passing and spatial augmentation in the GNN module, normalization is used to limit scale differences caused by different node degrees. When constructing the sample-specific adjacency matrix, this invention can use coherence, phase-locked values, and mutual information to replace the Pearson correlation coefficient to calculate lead edge weights. Negative correlation coefficient values can be handled using three methods: absolute value, positive / negative edge separation, and sign preservation. Sample-specific spatial graph. Each node corresponds one-to-one with an EEG lead signal, and the initial features of a node consist of the original sequence, frequency band energy, statistics, or a combination thereof for the current time window.
[0037] Furthermore, the detailed method for step 3 is as follows:
[0038] Step 3.1, Temporal-Spatial Dual-View Enhancement
[0039] Obtain unlabeled training batches Using any sample within the batch as the anchor sample The remaining samples are negative samples. , N is the number of samples in the batch;
[0040] For anchor samples Temporally enhanced views were obtained by performing temporal masking enhancement on the EEG sequences. , its special adjacency matrix Perform random edge pruning and normalized augmentation to obtain the spatial augmentation graph. ;in, and These represent the random temporal augmentation operator and the random spatial graph augmentation operator, respectively; the spatial augmentation graph contains both the augmented node features and the adjacency matrix, while maintaining the same sample index as the original EEG time window;
[0041] negative samples within a batch Negative sample time-enhanced view obtained by superimposing Gaussian noise on EEG sequences For negative samples Adding noise to the adjacency matrix yields a negative sample space augmented view. This allows the adjacency matrix to remain symmetric and enables re-self-connection and degree normalization.
[0042] The two representations obtained after temporal and spatial enhancement of the anchor sample have the same sample index and form a positive sample pair; the representations obtained after enhancement of the negative sample have different indices from the anchor sample and form a negative sample set; the positive and negative relationship is determined entirely based on the sample index within the batch, without using pre-ictal or interictal labels;
[0043] Step 3.2: Parallel Feature Extraction with Dual Encoders
[0044] Construct a parameter-shared encoder, which includes an SNN spiking temporal encoder composed of multiple layers of LIF spiking neurons. GNN spatial topology encoder composed of graph convolutions ;
[0045] SNN Pulse Timing Encoder Enhanced View of Processing Time The pulse time encoding is completed by LIF spiking neurons (first converted into a 0 / 1 pulse sequence), and the output is the time feature matrix before pooling. and the sample-level time representation after pooling The resulting representation includes the temporal pulse characteristics of EEG signals, such as EEG spikes, spiking waves, and rhythmic abrupt changes.
[0046] GNN spatial topology encoder Processing spatial augmentation maps After graph convolution message aggregation, the output node spatial feature moments before pooling are obtained. Graph-level spatial representation after pooling ;
[0047] Step 3.3, Calculation of projection and two-way contrast loss; [This involves] calculating the anchor samples within the batch. Sample-level time representation and graph-level spatial representation Input the time projection head separately and space projection head Each sample's projection representation is obtained, and the homology similarity of the anchor samples is calculated. Negative samples within the batch Time representation and spatial representation Input the time projection head separately and space projection head Each sample obtains its own projected representation, which, together with the anchor sample, forms a heterogeneous similarity set. ;
[0048] ;
[0049] and The first The temporal and spatial projection representations of each sample, both belonging to... ; To compare the dimensions of the projection space; time projection head and space projection head The structures are the same but the parameters are not shared; time projection head Parameters are shared between anchor samples and negative samples within the same batch; spatial projection head. Parameters are shared between anchor samples and negative samples within the same batch; and both projection heads can use a two-layer MLP and output representations of the same dimension.
[0050] Calculate the cosine similarity (sim) of the normalized projected vectors; after temperature coefficient... After scaling, calculate the N×N cosine similarity matrix within the batch. ;
[0051] ;
[0052] when hour, The similarity of positive samples between two views within the same EEG time window is considered a homologous positive sample pair; when hour, The similarity of heterogeneous negative samples within a batch is the same as the heterogeneous negative sample pairs.
[0053] Two views of the anchor sample are used to form the relationship between homologous positive samples, and an enhanced view of other samples within the batch is used to form the set of heterologous negative samples, enabling the temporal encoder and spatial encoder to learn sample-level consistency without using epilepsy category labels.
[0054] Then, we construct a two-way loss mechanism: time-to-space and space-to-time, and calculate the batch-average total pre-training loss. The expression is as follows:
[0055] ;
[0056] In the above formula, For time-to-space loss, The space-to-time loss is expressed as follows: ;
[0057] ;
[0058] Indicates the first The time representation is the loss of identifying the same index space representation in the entire space representation; Indicates loss in the opposite direction; This represents the batch sample size. This represents the average contrastive loss across the entire batch in both directions; the denominators in the loss formulas for both directions each include a positive sample with the same index. A heterogeneous index negative sample.
[0059] By minimizing improve and relatively reduced Errors are used to jointly update the SNN, GNN, temporal projector, and spatial projector; backpropagation synchronously updates the SNN, GNN, and dual projector; after pre-training, the projector is removed, and the SNN and GNN encoder weights are saved.
[0060] Furthermore, GNN spatial topology encoder Upon receiving the spatial augmentation view, perform the following operations:
[0061] First, the sample specific space map is... The initial features of the nodes are mapped to the hidden dimension, and then inter-lead message aggregation is performed based on the normalized adjacency matrix. The layer propagation process is represented as:
[0062] ;
[0063] in, Numbering of layers in a graph neural network; For the first The node feature matrix of the layer; The initial features of the input nodes; For the first Layer-learnable weight matrix; It is a non-linear activation function; For the first The normalized adjacency matrix of each sample; here, for the sample specific space graph corresponding to the i-th EEG time window, its node initial feature matrix is first... Mapping to a preset hidden dimension yields the input feature matrix of the graph neural network. Then, based on the normalized adjacency matrix Execute message aggregation between EEG lead nodes; the i-th sample refers to the sample generated by the i-th EEG time window. and its corresponding sample specific space diagram Composition of sample pairs ;
[0064] Simultaneously retain before level pooling The characteristics of each lead node form a spatial feature matrix. Here, C represents the number of EEG leads retained after unified lead processing, and is equal to the number of nodes in the sample specific space map, which is generally 16.
[0065] During self-supervised projection, graph-level representations are obtained through global average pooling or learnable readout layers, while node-level spatial features before pooling are used in the cross-attention stage.
[0066] ;
[0067] For the first The sample-level time representation obtained by time-encoding and pooling of a time-enhanced view; For the first The sample-level spatial representation obtained by graph encoding and reading out a spatial augmentation map;
[0068] Temporal and spatial features can have different sequence lengths or number of nodes, but are mapped to a preset hidden dimension through a linear layer before entering the projection head or cross attention; the two encoders share parameters across all batches of samples;
[0069] During the training phase, GNN and SNN run in parallel and are jointly optimized by cross-view loss; during the fine-tuning phase, their pre-trained parameters are loaded and connected to the cross-pulse attention layer; during deployment, 10-20 topological priors and normalization operators can be pre-cached, and only window-by-window related edge weights are updated in real time; depending on the device's computing power, graph convolution, graph attention, or lightweight message passing layers can be selected, and the number of layers and hidden dimensions are limited.
[0070] Furthermore, step 4 involves fine-tuning the labeled data and performing spatiotemporal feature fusion, as detailed below:
[0071] Step 4.1: Load the pre-trained SNN impulse temporal encoder and GNN spatial topological encoder, and retain the temporal feature matrix before pooling. Node space feature matrix ; Output the time steps for the SNN. For the number of lead nodes, To unify hidden dimensions;
[0072] Step 4.2: Cross-pulse attention bidirectional interactive fusion, calculating the query matrix, key matrix, and value matrix respectively, as shown in the following formulas:
[0073] ;
[0074] ;
[0075] in, , and These are the query matrix, key matrix, and value matrix of the time view, respectively. , , This is the corresponding learnable linear projection matrix; , and These are the query matrix, key matrix, and value matrix of the spatial view, respectively; , , This is the corresponding learnable linear projection matrix; and These are the temporal feature matrix and the node spatial feature matrix, respectively.
[0076] The two-way interaction between time and space, and between space and time, is represented as follows:
[0077] ;
[0078] ;
[0079] in, This is the output obtained by aggregating spatial nodes for a time-based query. The output obtained by aggregating spatial queries from time steps; Scaling factor for feature dimensions of query and key. To suppress the numerical amplification of the dot product as the dimension increases, Softmax is calculated along the time step or node dimension of the view of interest, keeping the two outputs consistent with each other. and The same first dimension length; T represents the transpose operation.
[0080] Finally, the bidirectional attention outputs are subjected to residual connections and layer normalization, and then spatiotemporal fusion features are formed by concatenation or weighted summation:
[0081] ;
[0082] ;
[0083] ;
[0084] in, Representation layer normalization; and These are the temporal and spatial characteristics after adding residual connections, respectively; Indicates feature splicing; due to and , and Since they have the same dimension, they can perform element-wise residual addition. Before concatenation, pooling can be performed on the time dimension and the node dimension separately, or the two types of pooling representations can be mapped to a unified dimension to form the final fused vector. ;
[0085] Step 4.3: Fine-tuning of the epilepsy binary classification, incorporating fusion features. The data is fed into the MLP classification head, and Softmax outputs the binary classification prediction probabilities for the pre-ictal and interictal periods. ; ;
[0086] and These are the weight matrix and bias vector of the classification layer, respectively; and Let represent the predicted probabilities of the preic and interictal periods, respectively, and their sum is 1.
[0087] Then, supervised fine-tuning is employed using cross-entropy loss. Fine-tune the encoder, cross-pulse attention module, and classification head, and save the optimal model based on the validation set;
[0088] ;
[0089] in, Adjust the number of samples in the labeled batch; Indexed by category; For the first The sample at the th Unique and genuine tags on the category; For the corresponding predicted probability; The batch average classification loss is used. This loss is minimized to jointly update the SNN, GNN, cross-pulse attention module, and MLP classification head;
[0090] The two-dimensional MLP classification head here can consist of one or more fully connected layers, nonlinear activation, and Dropout, with an output dimension of 2.
[0091] Furthermore, the present invention also includes an epilepsy prediction reasoning process, the method of which is as follows: input the EEG signal to be detected, repeat step 1 to complete the preprocessing and spatial graph construction, send it into the optimal model to output the classification probability, and determine the epilepsy state corresponding to the EEG.
[0092] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0093] (1) This invention independently calculates the correlation between leads for each EEG time window and integrates the international 10-20 lead physical topology prior to construct a sample-specific spatial graph that corresponds one-to-one with the EEG time window. Therefore, the constructed graph structure can dynamically characterize the functional connectivity and abnormal discharge propagation relationship between leads as the EEG state changes, overcoming the problem that it is difficult to describe the dynamic changes of the epileptic brain network using a fixed adjacency matrix.
[0094] (2) The present invention uses an SNN pulse timing encoder composed of LIF spiking neurons to convert continuous EEG signals into pulse activity and extract time features, which is beneficial for characterizing epilepsy-related transient discharge patterns such as spikes, spiking waves and rhythm mutations. At the same time, the pulse event-driven computing form is beneficial for reducing ineffective continuous activation, providing a basis for real-time prediction on bedside monitoring and resource-constrained equipment.
[0095] (3) This invention uses the temporal and spatial augmented views of the same EEG time window as positive sample pairs, and combines cross-views of other EEG time windows within the batch as negative sample pairs. Self-supervised pre-training is performed using bidirectional InfoNCE loss from time to space and space to time. This allows for the alignment of temporal and spatial representations using a large amount of unlabeled EEG data, reducing the model's dependence on manually labeled data and improving feature generalization ability under limited labeling and cross-patient conditions.
[0096] (4) This invention retains the time step features of the SNN pulse temporal encoder before pooling and the lead node features of the GNN spatial topology encoder before pooling, and uses bidirectional cross-pulse attention to achieve selective aggregation of temporal features to spatial nodes and reverse aggregation of spatial nodes to key time steps. Compared with simple splicing or unidirectional fusion, this method can enhance the synergistic representation ability between the dynamics of abnormal discharge time and changes in brain region functional connectivity.
[0097] (5) Under the leave-one-out-of-patient validation condition on the CHB-MIT dataset, the present invention achieved an accuracy of 85.02%, a sensitivity of 85.56%, a specificity of 80.40%, an F1 score of 82.71%, and an AUC of 92.00%. In the ablation experiment, the accuracy decreased by 4.56 percentage points after removing bidirectional cross-view pre-training, by 5.18 percentage points after removing windowed spatial graphs and GNNs, and by 3.34 percentage points after removing cross-pulse attention, indicating that the above modules have positive technical effects on unlabeled representation learning, dynamic spatial topology modeling, and spatiotemporal feature fusion, respectively. (6) In the inference stage, the present invention only needs to update the relevant edge weights of the current EEG time window and can pre-cache the fixed lead topological priors and related normalization operators, thus reducing the amount of repetitive computation in dynamic graph construction, and is suitable for continuous EEG monitoring and early warning scenarios for epileptic seizures. Attached Figure Description
[0098] Figure 1 This is a schematic diagram of the overall network structure of the present invention;
[0099] Figure 2 This is a schematic diagram of the overall prediction process of the present invention;
[0100] Figure 3 This is a schematic diagram of the label-free self-supervised training process in this invention;
[0101] Figure 4 This is a schematic diagram of the electroencephalogram (EEG) signals collected in the embodiment.
[0102] Figure 5 The diagram shows the processing results at each stage in the embodiment. Detailed Implementation
[0103] The technical solution of the present invention will be described in detail below, but the scope of protection of the present invention is not limited to the embodiments described.
[0104] like Figures 1 to 3 As shown, the epilepsy prediction method based on self-supervised pulse graph topology cooperative network of the present invention includes the following steps:
[0105] Step 1: Acquire and process multi-channel EEG signal samples, and divide the EEG into N independent EEG time windows using a sliding window; for the i-th time window... The specific adjacency matrix is obtained by calculating and processing each lead. Construct a sample-specific spatial map corresponding to the EEG time window. ; Initial features for nodes; forming dual-input sample pairs ;
[0106] Step 2: Dataset partitioning, which involves dividing all EEG signal samples into an unlabeled pre-training dataset, a labeled fine-tuning dataset, and a test dataset; the labeled samples include two types of labels: pre-ictal and interictal.
[0107] Step 3: Perform bidirectional cross-view pre-training on unlabeled data; first, pre-train on unlabeled training batches. Using any sample within a batch as the anchor sample and the remaining samples as negative samples, temporal masking enhancement is performed on the EEG sequence of the anchor sample to obtain a temporally enhanced view, and edge pruning enhancement is performed on the specific adjacency matrix of the anchor sample to obtain a spatially enhanced map; simultaneously, the EEG sequence of the k-th negative sample within the batch... Gaussian noise is superimposed to obtain a negative sample time-enhanced view. The normalized adjacency matrix of the negative sample By adding random perturbations and re-adding self-connections and progress degree normalization while maintaining matrix symmetry, an enhanced view of the negative sample space is obtained. ;
[0108] Then, the parameter-shared pulse timing encoder processes the time-enhanced view and outputs the time feature matrix before pooling. and the sample-level time representation after pooling The spatial topology encoder processes the spatial augmentation graph and outputs the spatial feature matrix of the nodes before pooling. Graph-level spatial representation after pooling ;
[0109] Next, the temporal and spatial features of the anchor samples are projected using the temporal and spatial projection heads respectively, and the homologous similarity is calculated. Negative samples within a batch are projected using the same set of projection heads, forming a heterologous similarity set with the anchor samples. The cosine similarity of the normalized projection vectors is calculated.
[0110] Finally, we construct the time-to-space and space-to-time bidirectional InfoNCE loss and optimize the projection head by minimizing the batch average bidirectional cross-view loss.
[0111] Step 4: Fine-tune the labeled data and perform spatiotemporal feature fusion, i.e., load the pre-trained SNN impulse temporal encoder and GNN spatial topological encoder, and retain the temporal feature matrix before pooling. Node space feature matrix Cross-pulse attention bidirectional interaction generates unified spatiotemporal fusion features, which are then fed into the MLP classification head to output binary classification prediction probabilities for the preic and interictal phases. This invention uses a multi-channel EEG time window and a sample-specific spatial map independently calculated from that time window as dual inputs, and outputs binary classification probabilities for the preic and interictal phases.
[0112] Figure 1 The network structure generally includes a data input and window-by-window image construction module, a time-space dual-view enhancement module, an SNN pulse temporal encoder, a GNN spatial topology encoder, a temporal projection head, a spatial projection head, a bidirectional cross-view loss module, a cross-pulse attention module, and an MLP classification head. It achieves functions such as EEG signal segmentation, window-by-window correlation matrix construction, pulse temporal feature extraction, spatial topology feature extraction, cross-attention fusion, and epilepsy prediction and classification. The epilepsy prediction and classification results of this invention only have a binary classification, namely, pre-epileptic and inter-epileptic phases.
[0113] To further verify the technical feasibility and effectiveness of this invention, this embodiment uses cross-patient EEG signal data on the CHB-MIT dataset. On the epilepsy EEG dataset CHB-MIT, accuracy, sensitivity, specificity, F1 score, and AUC-ROC are reported using either patient-only or patient-only cross-validation. To comprehensively evaluate the model's ability to distinguish between preic and interictal EEG windows, accuracy (ACC), sensitivity (SEN), specificity (SPE), F1 score, and area under the ROC curve (AUC) are used as evaluation metrics. , , and Let represent true positives, true negatives, false positives, and false negatives, respectively. The calculations for each indicator are as follows:
[0114] ;
[0115] ;
[0116] ;
[0117] ;
[0118] ;
[0119] AUC is calculated using the trapezoidal integral of adjacent sampling points on the ROC curve, where Equals sensitivity. The AUC represents the false positive rate; the closer the AUC is to 1, the stronger the model's overall discrimination ability under different classification thresholds.
[0120] Table 1. Data from the Leave-One-Out Verification Experiment
[0121]
[0122] The results of the leave-one-out validation experiment are shown in Table 1. The technical solution of this invention achieved an accuracy of 85.02%, a sensitivity of 85.56%, a specificity of 80.40%, an F1 score of 82.71%, and an AUC of 92.00% under leave-one-out validation across patients. The high sensitivity indicates that the model can effectively identify seizure-related EEG states, and the 80.40% specificity shows that the model still maintains a good ability to exclude non-target states. The F1 score and AUC further demonstrate that the network model of this invention performs stably in terms of class discrimination balance and overall discrimination ability under different thresholds.
[0123] Further comparisons were made between this invention and existing technologies. The comparative experimental results show that this invention achieves relatively balanced performance in the epilepsy prediction task, with an accuracy of 85.02%, sensitivity of 85.56%, specificity of 80.40%, F1 score of 82.71%, and AUC of 92.00%. Compared with other prediction methods, the sensitivity and AUC of this invention are outstanding, indicating that it can effectively identify epilepsy-related states while maintaining overall classification stability. Specific performance data are shown in Table 2.
[0124] Table 2 Performance comparison data with existing technologies
[0125]
[0126] Further ablation experiments were conducted on the network model of this invention, removing the self-supervised cross-view pre-training, SNN temporal branch, GNN spatial branch, cross-pulse attention module, and windowed correlation map respectively, and observing the changes in evaluation indicators to verify the independent contribution of each innovative module.
[0127] As shown in Table 3, the complete model achieved accuracy of 85.02%, sensitivity of 85.56%, specificity of 80.40%, F1 score of 82.71%, and AUC of 92.00%, respectively. Its overall performance is better than that of the ablation models, indicating that the proposed cross-view pre-training, SNN temporal encoding, windowed spatial graph modeling, and cross-pulse attention all have a positive effect on epilepsy prediction.
[0128] After removing cross-view pre-training, the model accuracy and AUC decreased by 4.56 and 2.82 percentage points, respectively, and the F1 score dropped from 82.71% to 80.82%. This indicates that bidirectional cross-view contrastive pre-training can align the temporal and spatial representations of the same time window using unlabeled EEG data, providing more discriminative and generalizable initialization parameters for downstream classification. Replacing the SNN with a conventional temporal encoder resulted in a decrease in accuracy and AUC of 2.65 and 0.79 percentage points, respectively, indicating that the spiking neural network has certain modeling advantages for discrete abnormal discharge patterns such as spikes, spikes, and rhythmic abrupt changes, but its impact is relatively smaller than that of spatial topology modeling and cross-view pre-training.
[0129] Removing the windowed spatial map and GNN resulted in the most significant performance degradation, with accuracy decreasing from 85.02% to 79.84% and AUC from 92.00% to 88.73%, representing decreases of 5.18 and 3.27 percentage points, respectively. Specificity also decreased to 74.02%. This indicates that correlation maps independently constructed for each EEG time window can effectively describe the dynamic functional connectivity and abnormal discharge propagation relationships between different leads, and are an important component in improving cross-patient epilepsy prediction capabilities. Removing cross-pulse attention further reduced accuracy and AUC by 3.34 and 1.54 percentage points, respectively, suggesting that independently encoding only temporal and spatial features is insufficient to fully utilize both types of information; cross-pulse attention can enhance the complementarity between spatiotemporal features.
[0130] In summary, the ablation results show that all modules of this invention contribute positively to the final performance. Among them, the window-by-window spatial graph and the GNN spatial branch contribute the most, followed by bidirectional cross-view pre-training and cross-pulse attention. The SNN temporal encoder further enhances the model's ability to express anomalous temporal discharge patterns.
[0131] Table 3 Comparison of ablation results
[0132]
[0133] Finally, a cross-patient experimental protocol with one target patient was adopted. For each round of the experiment, one subject was selected from CHB-MIT as the target patient, and the remaining subjects were used as source patients. The EEG data of the source patients were used for unlabeled self-supervised pre-training after the seizure state labels were removed. All data of the target patient were completely invisible at this stage, thereby verifying the model's ability to transfer from other patients to new patients and preventing the target patient's data from entering the pre-training process prematurely.
[0134] During the self-supervised pre-training phase, unlabeled EEG windows and their window-by-window spatial maps of the source patients are used to construct temporally and spatially augmented views. A general spatiotemporal representation is learned through an SNN temporal encoder, a GNN spatial encoder, dual projectors, and bidirectional cross-view InfoNCE loss. After pre-training, the two projectors are removed, and the weights of the SNN and GNN encoders are retained as initialization parameters for the individualized model of the target patient.
[0135] During the individualized adaptation phase, only labeled data from the target patient is used. The target patient's seizure events or continuous records are divided chronologically into a fine-tuning set, a validation set, and a test set. Different subsets must not contain overlapping windows, nor should adjacent windows from the same seizure be randomly distributed across different subsets. The fine-tuning set is used to update model parameters, the validation set is used for early termination, hyperparameter selection, and classification threshold determination, and the test set is used only for final performance evaluation.
[0136] To facilitate the demonstration of the complete implementation chain from raw EEG to classification results, intermediate results can be saved simultaneously during the target patient testing phase. For example... Figure 4 As shown, the original input is a multi-lead EEG time window. After preprocessing, standardized EEG windows and window-by-window correlation matrices were obtained. Then, combined with the prior of 10-20 leads, the sample-specific spatial map is obtained. The self-supervised training phase can output a time-augmented view. and spatial augmented view The fine-tuned classification network outputs a probability vector for the test window. .when When a preset threshold is reached, it is determined to be the pre-seizure period; otherwise, it is determined to be the interictal period.
[0137] Effects of epilepsy sample processing Figure 5 As shown, Figure 5 The boxes in the middle illustrate the preic seizure and interictal periods, respectively. During fine-tuning, SNN and GNN parameters pre-trained across patients are loaded, followed by a cross-pulse attention module and an MLP classification head. A smaller learning rate is used for the encoder, while larger learning rates are used for newly added attention layers and the classification head; alternatively, the encoder can be frozen initially and then gradually unfrozen. Preic seizure and interictal samples are kept roughly balanced in the fine-tuning batches. During validation and testing, random data augmentation, contrastive projection heads, and InfoNCE loss are disabled; only forward inference of the complete classification network is performed.
[0138] Accuracy, sensitivity, specificity, F1 score, AUC-ROC, and hourly false alarm rate were reported for each target patient, and the mean and standard deviation of the results for all target patients were calculated. The practical benefits of this invention are: window-by-window sample-specific spatial maps enhance the expressive ability of abnormal brain networks; SNNs enhance sensitivity to spikes, spiking waves, and rhythmic abrupt changes; bidirectional cross-view contrastive pre-training improves generalization performance under limited annotation conditions; and cross-pulse attention improves classification performance after the fusion of temporal and spatial features.
Claims
1. An epilepsy prediction method based on a self-supervised pulse graph topological cooperative network, characterized in that, Includes the following steps: Step 1: Acquire and process multi-channel EEG signals, segment the EEG into N independent EEG time windows using a sliding window, and then process the time windows. The specific adjacency matrix is obtained by per-lead processing. and sample specific space diagram ; Step 2: Dataset partitioning, which involves dividing all EEG signal samples into an unlabeled pre-training dataset, a labeled fine-tuning dataset, and a test dataset; the labeled samples include two types of labels: pre-ictal and interictal. Step 3: Perform bidirectional cross-view pre-training on unlabeled data; For unlabeled training batches, any sample within the batch is used as the anchor sample, and the remaining samples are used as negative samples. The EEG sequence and specific adjacency matrix of the anchor sample are enhanced respectively to obtain temporal and spatial enhanced views. The EEG sequences and normalized adjacency matrices of negative samples within a batch are enhanced to obtain temporal and spatial enhancement views of negative samples. Two types of views are processed using a pulse timing encoder and a spatial encoder with shared parameters, respectively, and their respective feature matrices and graph-level representations are output; anchor samples and negative samples obtain projection features through two sets of projection heads; Calculate the similarity of anchor samples to the same source view and the cosine similarity of anchor-negative samples to the different source view; construct a bidirectional loss and optimize the projection head by minimizing the batch average bidirectional cross-view loss; Step 4: Fine-tune the labeled data, perform spatiotemporal feature fusion, and feed the fused features into the MLP classification head to output the binary prediction probabilities of the pre-ictal and interictal periods.
2. The epilepsy prediction method based on self-supervised pulse graph topological cooperative network according to claim 1, characterized in that, After obtaining continuous multi-channel EEG signals in step 1, the following steps are performed sequentially: unified lead channel, resampling, filtering, Z-score normalization, and then the EEG sequences are divided into N independent EEG time windows by sliding windows with fixed window length and overlap rate. , , Represents the real number field. This refers to the number of EEG leads; This refers to the number of sampling points within a single window. For each time window, the Pearson correlation coefficient is calculated lead-by-lead to generate the instantaneous functional connectivity matrix. Integrating the international 10-20 lead fixed physical topology matrix Functional connection matrix With fixed physical topology matrix The fused adjacency matrix is obtained after fusion. ;right Perform threshold filtering or k-nearest neighbor retention, and add a C-order identity matrix. Supplementing lead self-connection and longitude normalization yields the sample-specific adjacency matrix. The expression is as follows; ; ; In the above formula, for The degree matrix, The fusion weight for the two types of adjacency relationships; Forming the i-th EEG time window Dual-input sample pairs , This is the specific space diagram of the corresponding sample.
3. The epilepsy prediction method based on self-supervised pulse graph topological cooperative network according to claim 1, characterized in that, The detailed method for step 3 is as follows: Step 3.1, Temporal-Spatial Dual-View Enhancement Obtain unlabeled training batches Using any sample within the batch as the anchor sample The remaining samples are negative samples. , N is the number of samples in the batch; For anchor samples Temporally enhanced views were obtained by performing temporal masking enhancement on the EEG sequences. , its special adjacency matrix Perform random edge pruning and normalized augmentation to obtain the spatial augmentation graph. ;in, and These represent the stochastic time augmentation operator and the stochastic spatial graph augmentation operator, respectively; For the k-th negative sample in the batch Negative sample time-enhanced view obtained by superimposing Gaussian noise on EEG sequences For negative samples Adding noise to the adjacency matrix yields a negative sample space augmented view. ; Step 3.2: Parallel Feature Extraction with Dual Encoders Construct a parameter-shared encoder, which includes an SNN spiking temporal encoder composed of multiple layers of LIF spiking neurons. GNN spatial topology encoder composed of graph convolutions ; SNN Pulse Timing Encoder Enhanced View of Processing Time The pulse timing is encoded by LIF spiking neurons, and the output is the temporal feature matrix before pooling. and the sample-level time representation after pooling ; GNN Spatial Topology Encoder Processing spatial augmentation maps After graph convolution message aggregation, the output node spatial feature moments before pooling are obtained. Graph-level spatial representation after pooling ; Step 3.3, Calculation of projection and two-way contrast loss; [This involves] calculating the anchor samples within the batch. Sample-level time representation and graph-level spatial representation Input the time projection head separately and space projection head Each sample's projection representation is obtained, and the homology similarity of the anchor samples is calculated. Negative samples within the batch Time representation and spatial representation Input the time projection head separately and space projection head Each sample obtains its own projected representation, which, together with the anchor sample, forms a heterogeneous similarity set. ; Calculate the cosine similarity of the normalized projected vectors, after temperature coefficient adjustment. After scaling, calculate the N×N cosine similarity matrix within the batch. ; when hour, The similarity of positive samples between two views within the same EEG time window is considered a homologous positive sample pair; when hour, The similarity of heterogeneous negative samples within a batch is the same as the heterogeneous negative sample pairs. Then, we construct time-to-space loss, space-to-time bidirectional loss, and batch-average total pre-training loss. The expression is as follows: ; In the above formula, For time-to-space loss, The space-to-time loss is expressed as follows: ; ; Indicates the first The time representation is the loss of identifying the same index space representation in the entire space representation; Indicates loss in the opposite direction; This represents the batch sample size. This represents the average contrast loss across the entire batch in both directions.
4. The epilepsy prediction method based on self-supervised pulse graph topology cooperative network according to claim 1 or 3, characterized in that, GNN Spatial Topology Encoder Upon receiving the spatial augmentation view, perform the following operations: First, the sample specific space map is... The initial features of the nodes are mapped to the hidden dimension, and then inter-lead message aggregation is performed based on the normalized adjacency matrix. The layer propagation process is represented as: ; in, Numbering of layers in a graph neural network; For the first The node feature matrix of the layer; The initial features of the input nodes; For the first Layer-learnable weight matrix; It is a non-linear activation function; For the first The normalized adjacency matrix of each sample; the i-th sample refers to the sample generated by the i-th EEG time window. and its corresponding sample specific space diagram Composition of sample pairs ; Simultaneously retain before level pooling The characteristics of each lead node form a spatial feature matrix. During self-supervised projection, a graph-level representation is obtained through global average pooling or a learnable readout layer, while the node-level spatial features before pooling are used in the cross-attention stage. ; For the first A time-enhanced view The sample-level time representation obtained after time encoding and pooling; For the first Space Enhancement Map Image encoding and sample-level spatial representation.
5. The epilepsy prediction method based on self-supervised pulse graph topological cooperative network according to claim 1, characterized in that, Step 4 involves fine-tuning the labeled data and performing spatiotemporal feature fusion, as detailed below: Step 4.1: Load the pre-trained SNN impulse temporal encoder and GNN spatial topological encoder, and retain the temporal feature matrix before pooling. Node space feature matrix ; Output the time steps for the SNN. For the number of lead nodes, To unify hidden dimensions; Step 4.2: Cross-pulse attention bidirectional interactive fusion, calculating the query matrix, key matrix, and value matrix respectively, as shown in the following formulas: ; ; in, , and These are the query matrix, key matrix, and value matrix of the time view, respectively. , , This is the corresponding learnable linear projection matrix; , and These are the query matrix, key matrix, and value matrix of the spatial view, respectively; , , This is the corresponding learnable linear projection matrix; and These are the temporal feature matrix and the node spatial feature matrix, respectively. The two-way interaction between time and space, and between space and time, is represented as follows: ; ; in, This is the output obtained by aggregating spatial nodes for a time-based query. The output obtained by aggregating spatial queries from time steps; Scaling factor for feature dimensions of query and key. Used to suppress the numerical amplification of the dot product as the dimension increases; T represents the transpose operation. Finally, the bidirectional attention outputs are subjected to residual connections and layer normalization, and then spatiotemporal fusion features are formed by concatenation or weighted summation. : ; ; ; in, Representation layer normalization; and These are the temporal and spatial characteristics after adding residual connections, respectively; Indicates feature splicing; Step 4.3: Fine-tuning of epilepsy binary classification. The fused feature h is fed into the MLP classification head, and Softmax outputs the binary classification prediction probabilities for the preic and interictal phases. ; ; and These are the weight matrix and bias vector of the classification layer, respectively; and Let represent the predicted probabilities of the preic and interictal periods, respectively, and their sum is 1. Then, supervised fine-tuning is employed using cross-entropy loss. Fine-tune the encoder, cross-pulse attention module, and classification head, and save the optimal model based on the validation set; ; in, Adjust the number of samples in the labeled batch; Indexed by category; For the first The sample at the th Unique and genuine tags on the category; For the corresponding predicted probability; This represents the average batch classification loss.
6. The epilepsy prediction method based on self-supervised pulse graph topological cooperative network according to claim 1, characterized in that, It also includes the epilepsy prediction reasoning process, which is as follows: input the EEG signal to be detected, repeat step 1 to complete the preprocessing and spatial graph construction, send it into the optimal model to output the classification probability, and determine the epilepsy state corresponding to the EEG.