Seizure monitoring method and system based on cross-modal timing and individual response

CN122805205APending Publication Date: 2026-09-25HEBEI MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611019175.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的之一是提供一种基于跨模态时序和个体响应的癫痫发作监测方法,以解决对癫痫发作事件检测不准确和功耗高的问题

Benefits of technology

[0136]本发明首次揭示癫痫发作外周响应的非对称异步耦合机制,将多模态响应视为非线性、异步触发的动力学系统;通过构建跨模态时序响应图谱,首次系统量化了癫痫发作过程中不同外周生理模态间的异步响应关系,解决了现有技术中简单同步融合导致的模型对噪声高度敏感的问题。本发明构建基于稳健响应表型聚合的跨个体异质性建模体系,实现从个体化碎片数据到群体响应表型规律的跨越;通过无监督聚类识别患者生理响应表型,实现了对“同病异质”问题的结构化建模,为跨个体泛化提供了新的方法路径。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805205A_ABST
    Figure CN122805205A_ABST
Patent Text Reader

Abstract

The application provides a seizure monitoring method and system based on cross-modal timing and individual response, comprising the following steps: S1. Preprocessing the signals in the data set; S2. Screening candidate events; S3. Building a map for the candidate events; S4. Training an unsupervised clustering model to obtain the response phenotype of each patient and adjusting the parameters of the unsupervised clustering model; S5. Constructing a seizure event detection model and training the seizure event detection model; S6. Preprocessing the multi-modal peripheral physiological signals of the monitored person collected in real time; S7. Determining the detected event according to the signals of the monitored person after preprocessing; S8. Building a map for the detected event according to the detected event; S9. Determining the response phenotype of the monitored person; S10. Inputting the map of the detected event and the response phenotype of the monitored person into the trained seizure event detection model, and outputting the probability of the detected event of epilepsy. The application improves the accuracy of monitoring epilepsy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for monitoring epileptic seizures, specifically a method and system for monitoring epileptic seizures based on cross-modal temporal and individual responses. Background Technology

[0002] Epilepsy is a chronic brain disorder caused by abnormal electrical activity in the brain, with a global prevalence of 4-10 per 1,000. In my country, the number of epilepsy patients exceeds 9 million, of whom approximately 30% suffer from drug-resistant epilepsy. Epileptic seizures are characterized by their sudden onset and high unpredictability, exposing patients to long-term, uncontrollable risks. In particular, patients with generalized tonic-clonic seizures (GTCS) face a 2-4 times higher risk of sudden epileptic death (SUDEP) than the general population. Therefore, timely identification of epileptic seizures in real-life environments has become a crucial starting point for reducing the risk of SUDEP.

[0003] Electroencephalography (EEG) is the gold standard for epilepsy diagnosis and seizure interpretation, but its widespread application in real-life scenarios is limited by device form factor and wearability compliance. With the development of wearable technology, researchers have gradually expanded their focus to easily acquired peripheral physiological signals such as electrocardiogram (ECG), electromyography (EMG), and acceleration (ACM), exploring their potential applications in epileptic seizure monitoring. Multiple clinical and real-world studies have shown that during epileptic seizures, especially during GTCS (Gastroesophageal Tract Surgery), detectable dynamic changes occur in peripheral autonomic nerve-related signals, making seizure detection based on wearable peripheral physiological signals an important supplementary direction to EEG.

[0004] Currently, existing epilepsy monitoring technologies suffer from the following technical limitations: During an epileptic seizure, there is an intrinsic time delay in the transmission of central nervous system discharges to different peripheral effectors. Studies have shown that heart rate changes can occur approximately 100 seconds earlier than body movement. Existing technologies generally assume that multimodal signals can be synchronized, treating the fusion of multimodal signals as equivalent to the parallel superposition or static splicing of features within the same time window. This approach implicitly assumes that the responses of different physiological modalities to the same seizure event are approximately synchronous in time, which is not entirely consistent with physiological reality. Simple synchronization alignment leads to high model sensitivity to noise, significant performance fluctuations in cross-device and cross-scenario applications, resulting in inaccurate detection of epileptic seizure events. The physiological responses of different patients vary greatly. For example, some patients' seizures are dominated by heart rate changes, while others are dominated by motor responses, leading to a high false alarm rate and poor generalization of a uniform model in real-world scenarios. While existing research attempts to mitigate this through individualized training, threshold adjustment, or transfer learning, these methods largely remain at the empirical level of parameter adjustments, resulting in inaccurate detection of epileptic seizure events. In addition, the continuous operation of multimodal deep learning models results in extremely short battery life for wearable devices, limiting the practical application of long-term monitoring. Summary of the Invention

[0005] One of the objectives of this invention is to provide a method for monitoring epileptic seizures based on cross-modal timing and individual response, in order to solve the problems of inaccurate detection of epileptic seizure events and high power consumption.

[0006] The second objective of this invention is to provide an epileptic seizure monitoring system based on cross-modal timing and individual response, so as to provide an effective system for monitoring the timing of epileptic seizures.

[0007] One of the objectives of this invention is achieved as follows:

[0008] A method for monitoring epileptic seizures based on cross-modal temporal and individual responses, comprising:

[0009] S1. Perform preprocessing on the signals in the dataset, including denoising and alignment;

[0010] S2. Filter candidate events based on the signals of the preprocessed dataset and determine the labels of the candidate events, the labels including epileptic seizures and non-epileptic seizures;

[0011] S3. Build a graph of candidate events;

[0012] S4. Train the unsupervised clustering model based on the map of candidate events labeled as epileptic seizures to obtain the response phenotype of each patient in the dataset, and adjust the parameters of the unsupervised clustering model.

[0013] S5. Construct an episodic event detection model and train the model using the candidate event map and the response phenotypes of the patients corresponding to the candidate events;

[0014] S6. Perform preprocessing, including denoising and alignment, on the multimodal peripheral physiological signals of the monitored subjects collected in real time;

[0015] S7. Determine the monitored events based on the pre-processed multimodal peripheral physiological signals of the monitored subjects;

[0016] S8. Establish a map of the monitored events based on the monitored events;

[0017] S9. Based on the graph of the detected events and the trained unsupervised clustering model, determine the response phenotype of the monitored individuals;

[0018] S10. Input the spectrum of the event to be detected and the response phenotype of the monitored person into the trained seizure event detection model, and output the epileptic seizure probability of the detected event; when the epileptic seizure probability is greater than the preset probability threshold, the detected event is determined to be an epileptic seizure and an early warning signal is issued.

[0019] Furthermore, the specific method for building the graph of candidate events in step S3 is as follows:

[0020] S3-1. Calculate the maximum correlation delay corresponding to different modal signals for a single candidate event:

[0021] τ ij (w)=argmax {τ∈[-T,T]} {corr(x i (t),x j (t+τ))}

[0022] Where T is the maximum search delay, x i (t) is the signal corresponding to mode i at time t, x j (t+τ) is the signal corresponding to mode j at time t+τ;

[0023] S3-2. Calculate the DTW distance between signals of different modes;

[0024] S3-3. Calculate the transfer entropy between different modes. :

[0025]

[0026] Where X is the source signal, Y is the target signal, and y t+1 Let Y be the state at the next moment. Let Y be the historical state composed of u points in the past. For X in the past The historical state consists of _ ...

[0027] S3-4. Calculate the coupling strength between signals of different modes;

[0028] S3-5. Use the multimodal signal of each candidate event as the node feature matrix, and the maximum correlation delay, DTW distance, propagation entropy and coupling strength of each candidate event as the adjacency matrix to obtain the event-level time series map of each candidate event.

[0029] Furthermore, the specific method for training the unsupervised clustering model in step S4 is as follows:

[0030] S4-1. All candidate events labeled as epileptic seizures from the same patient are integrated using median and interquartile range to obtain the integrated feature P for each patient. k :

[0031] P k =(median{E k,1 E k,2 ,…,E k,n ,…,E k,N ),IQR{E k,1 E k,2,…,E k,n ,…,E k,N}), n=1,...,N

[0032] Among them, E k,n This is the atlas corresponding to the nth seizure event of the kth patient;

[0033] S4-2. Input the ensemble features of each patient into the unsupervised clustering model and adjust the parameters of the unsupervised clustering model.

[0034] Furthermore, for patients whose number of seizures is less than a preset seizure frequency threshold, the integrated characteristics of the patients are obtained. The specific method is:

[0035] By using a weighted approach to fuse individual characteristics and prior group characteristics, the integrated characteristics of patients are obtained. :

[0036]

[0037] in, This is the patient-level response feature vector obtained from the existing epidemiological events of the k-th patient. To train the average response feature vector of the total number of patients in the set or patients with similar response phenotypes, , The number of seizures for the kth patient is given, and c is the smoothing coefficient.

[0038] Furthermore, the specific method by which the episodic event detection model processes the candidate event map and the response phenotype of the corresponding patient is as follows:

[0039] Based on the adjacency matrix of the candidate event spectrum, feature extraction is performed on the multimodal signal segments of the candidate events, and temporal features are output. Finally, the output temporal features are conditionally fused or conditionally modulated using the response phenotype and input into the classification head to obtain the epileptic seizure probability.

[0040] Furthermore, the event detection model is a CNN-LSTM or multimodal Transformer architecture;

[0041] When the event detection model is a CNN-LSTM, the specific method by which the event detection model extracts features from multimodal signal segments is as follows:

[0042] The input layer of CNN-LSTM receives multimodal signal segments, with each modality using an independent channel input. After that, it passes through 3 sets of convolutional-pooling blocks, batch normalization layers, ReLU activation functions, and max pooling layers. After the first set of convolutional-pooling blocks, a channel attention mechanism is introduced. The channel attention mechanism uses the adjacency matrix of the graph to perform attention calculation on the features extracted by the first set of convolutional-pooling blocks.

[0043] When the epileptic event detection model is a multimodal Transformer architecture, the specific method by which the epileptic event detection model extracts features from multimodal signal segments is as follows:

[0044] The input embedding layer of the multimodal Transformer architecture divides the multimodal signal segment into fixed-length patches. Each patch is mapped to a d-dimensional embedding vector through a linear projection layer, and position encoding and modality type embedding are added. Then, it is input to several stacked multi-head attention layers. Attention is calculated in the multi-head attention layers using the adjacency matrix of the graph.

[0045] Furthermore, the specific method by which the event detection model uses the adjacency matrix of the graph to calculate attention is as follows:

[0046] S5a-1. Normalize the adjacency matrix and perform weighted fusion to obtain the prior relation scores between modes. :

[0047]

[0048] in, and For preset or learnable weights, The maximum correlation delay between mode i and mode j. The DTW distance between mode i and mode j The coupling strength between mode i and mode j is... Let i be the transfer entropy from mode i to mode j. , , and All calculations are normalized.

[0049] S5a-2. Score all intermodal prior relations. Construct a modality-level prior adjacency matrix, and expand it into a prior matrix with the same dimension as the attention matrix based on the modality to which the input token belongs. :

[0050]

[0051] S5a-3. In attention calculation, the original attention score matrix is ​​added to the prior matrix, and then the attention weights are obtained by softmax normalization:

[0052]

[0053] Where α is the prior strength hyperparameter, Q represents the Query matrix, K represents the Key matrix, and V represents the Value matrix. Q, K, and V are all obtained by projecting the features from the previous layer through different weight matrices.

[0054] Furthermore, the specific method by which the event detection model utilizes response phenotypes for conditional fusion is as follows:

[0055] The patient's response phenotype is concatenated with the temporal features output by the episodic event detection model;

[0056] The specific method by which the episodic event detection model uses response phenotypes for conditional modulation is as follows:

[0057] Channel weighting is performed by generating modal weight vectors through fully connected layers.

[0058] Furthermore, the specific method by which the episodic event detection model concatenates the patient's response phenotype with the temporal features output by the episodic event detection model is as follows:

[0059] S5b-1. Encode the patient's response phenotype into a one-hot vector, and then transform it into a response phenotype embedding vector through an embedding matrix:

[0060]

[0061] in, For the patient response phenotype one-hot vector, In response to the phenotypic embedding matrix, For bias terms, For response phenotype embedding vectors; d e For the embedding dimension, such as 8, 16, or 32; C is the number of response phenotype categories;

[0062] S5b-2. Temporal features extracted from multimodal signals using an epidural event detection model. Global average pooling is performed to obtain fixed-length signal features. :

[0063]

[0064] in, , The length of the signal segment;

[0065] S5b-3. Fixed-length signal features Concatenate with the response phenotype embedding vector:

[0066]

[0067] in, .

[0068] The second objective of this invention is achieved as follows:

[0069] A seizure monitoring system based on cross-modal temporal and individual responses, comprising:

[0070] The dataset preprocessing module, connected to the event pre-filtering module, is used to preprocess the signals in the dataset, including denoising and alignment, and then send the preprocessed dataset to the event pre-filtering module.

[0071] The event pre-screening module is connected to the dataset preprocessing module and the event-level time series graph construction module, respectively. It is used to screen candidate events based on the signals of the preprocessed dataset and determine the labels of the candidate events, including epileptic seizures and non-epileptic seizures; and send the candidate events to the event-level time series graph construction module.

[0072] The event-level time series graph construction module is connected to the event pre-screening module, the patient response phenotype training module, and the epidural event detection model training module, respectively. It is used to build graphs of candidate events and send the graphs of candidate events to the patient response phenotype training module and the epidural event detection model training module.

[0073] The patient response phenotype training module is connected to the event-level time series graph construction module, the seizure event detection model training module, and the response phenotype determination module, respectively. It is used to train the unsupervised clustering model based on the graph of candidate events labeled as epileptic seizures to obtain the response phenotype of each patient in the dataset and adjust the parameters of the unsupervised clustering model. The patient's response phenotype is input into the seizure event detection model training module, and the trained unsupervised clustering model is input into the response phenotype determination module.

[0074] The seizure event detection model training module is connected to the event-level time series graph construction module, the patient response phenotype training module, and the seizure detection output module, respectively. It is used to construct the seizure event detection model, and train the seizure event detection model using the graph of candidate events and the response phenotype of the patients corresponding to the candidate events. The trained seizure event detection model is then sent to the seizure detection output module.

[0075] The real-time signal preprocessing module, connected to the real-time pre-screening module, is used to preprocess the multimodal peripheral physiological signals of the monitored subjects acquired in real time, including denoising and alignment, and send the preprocessed multimodal peripheral physiological signals to the real-time pre-screening module.

[0076] The real-time pre-screening module is connected to the real-time signal preprocessing module and the detected event map construction module, respectively. It is used to determine the detected events of the monitored subject based on the preprocessed multimodal peripheral physiological signals of the monitored subject sent by the real-time signal preprocessing module, and send them to the detected event map construction module.

[0077] The detected event graph construction module is connected to the real-time pre-screening module, the response phenotype determination module, and the seizure detection output module, respectively. It is used to build a graph of the detected events based on the detected events of the monitored subjects sent by the real-time pre-screening module, and send it to the response phenotype determination module and the seizure detection output module.

[0078] The response phenotype determination module is connected to the detected event graph construction module, the patient response phenotype training module, and the seizure detection output module, respectively. It is used to determine the response phenotype of the monitored person based on the detected event graph sent by the detected event graph construction module and the trained unsupervised clustering model sent by the patient response phenotype training module; and send the response phenotype of the monitored person to the seizure detection output module.

[0079] The seizure detection output module is connected to the response phenotype determination module, the seizure event detection model training module, and the detected event atlas construction module, respectively. It is used to input the atlas of the detected event sent by the detected event atlas construction module and the response phenotype of the monitored subject sent by the response phenotype determination module into the trained seizure event detection model sent by the seizure event detection model training module, and output the seizure probability of the detected event. When the seizure probability is greater than the preset probability threshold, the detected event is determined to be a seizure, and an early warning signal is issued.

[0080] Furthermore, the event-level time series graph construction module, when building a graph for candidate events, is specifically used for:

[0081] Calculate the maximum correlation delay corresponding to different modal signals for a single candidate event:

[0082] τ ij (w)=argmax {τ∈[-T,T]} {corr(x i (t),x j (t+τ))}

[0083] Where T is the maximum search delay, x i (t) is the signal corresponding to mode i at time t, xj (t+τ) is the signal corresponding to mode j at time t+τ;

[0084] Calculate the DTW distance between signals of different modes;

[0085] Calculate the transfer entropy between different modes :

[0086]

[0087] Where X is the source signal, Y is the target signal, and y t+1 Let Y be the state at the next moment. Let Y be the historical state composed of u points in the past. For X past The historical state consists of _ ...

[0088] Calculate the coupling strength between signals of different modes;

[0089] The multimodal signal of each candidate event is used as the node feature matrix, and the maximum correlation delay, DTW distance, propagation entropy and coupling strength of each candidate event are used as the adjacency matrix to obtain the event-level time series map of each candidate event.

[0090] Furthermore, when training the unsupervised clustering model, the patient response phenotype training module is specifically used for:

[0091] All candidate events labeled as epileptic seizures from the same patient were integrated using median and interquartile range to obtain the integrated feature P for each patient. k :

[0092] P k =(median{E k,1 E k,2 ,…,E k,n ,…,E k,N ),IQR{E k,1 E k,2 ,…,E k,n ,…,E k,N}), n=1,...,N

[0093] Among them, E k,n This is the atlas corresponding to the nth seizure event of the kth patient;

[0094] The integrated features of each patient are input into the unsupervised clustering model, and the parameters of the unsupervised clustering model are adjusted.

[0095] Furthermore, the patient response phenotype training module obtains integrated characteristics of patients whose number of seizures is less than a preset seizure frequency threshold. Specifically used for:

[0096] By using a weighted approach to fuse individual characteristics and prior group characteristics, the integrated characteristics of patients are obtained. :

[0097]

[0098] in, This is the patient-level response feature vector obtained from the existing epidemiological events of the k-th patient. To train the average response feature vector of the total number of patients in the set or patients with similar response phenotypes, , The number of seizures for the kth patient is given, and c is the smoothing coefficient.

[0099] Furthermore, in the episodic event detection model training module, the episodic event detection model, when processing the spectrum of candidate events and the response phenotypes of the patients corresponding to the candidate events, is specifically used for:

[0100] Based on the adjacency matrix of the candidate event spectrum, feature extraction is performed on the multimodal signal segments of the candidate events, and temporal features are output. Finally, the output temporal features are conditionally fused or conditionally modulated using the response phenotype and input into the classification head to obtain the epileptic seizure probability.

[0101] Furthermore, the event detection model is a CNN-LSTM or multimodal Transformer architecture;

[0102] When the seizure event detection model is a CNN-LSTM, the seizure event detection model training module, when extracting features from multimodal signal segments, is specifically used for:

[0103] The input layer of CNN-LSTM receives multimodal signal segments, with each modality using an independent channel input. After that, it passes through 3 sets of convolutional-pooling blocks, batch normalization layers, ReLU activation functions, and max pooling layers. After the first set of convolutional-pooling blocks, a channel attention mechanism is introduced. The channel attention mechanism uses the adjacency matrix of the graph to perform attention calculation on the features extracted by the first set of convolutional-pooling blocks.

[0104] When the seizure event detection model is a multimodal Transformer architecture, the seizure event detection model training module, when extracting features from multimodal signal segments, is specifically used for:

[0105] The input embedding layer of the multimodal Transformer architecture divides the multimodal signal segment into fixed-length patches. Each patch is mapped to a d-dimensional embedding vector through a linear projection layer, and position encoding and modality type embedding are added. Then, it is input to several stacked multi-head attention layers. Attention is calculated in the multi-head attention layers using the adjacency matrix of the graph.

[0106] Furthermore, in the epileptic event detection model training module, when the epileptic event detection model uses the adjacency matrix of the graph to perform attention calculation, it is specifically used for:

[0107] The adjacency matrix is ​​normalized and then weighted and fused to obtain the intermodal prior relation scores. :

[0108]

[0109] in, and For preset or learnable weights, The maximum correlation delay between mode i and mode j. The DTW distance between mode i and mode j The coupling strength between mode i and mode j is... Let i be the transfer entropy from mode i to mode j. , , and All calculations are normalized.

[0110] All intermodal prior relation scores Construct a modality-level prior adjacency matrix, and expand it into a prior matrix with the same dimension as the attention matrix based on the modality to which the input token belongs. :

[0111]

[0112] In attention calculation, the original attention score matrix is ​​added to the prior matrix, and then the attention weights are obtained by softmax normalization.

[0113]

[0114] Where α is the prior strength hyperparameter, Q represents the Query matrix, K represents the Key matrix, and V represents the Value matrix. Q, K, and V are all obtained by projecting the features from the previous layer through different weight matrices.

[0115] Furthermore, when the epileptic event detection model uses response phenotypes for conditional fusion in the epileptic event detection model training module, it is specifically used for:

[0116] The patient's response phenotype is concatenated with the temporal features output by the episodic event detection model, or channel weighting is performed by generating modal weight vectors through a fully connected layer.

[0117] In the episodic event detection model training module, when the episodic event detection model uses response phenotypes for conditional modulation, it is specifically used for:

[0118] Channel weighting is performed by generating modal weight vectors through fully connected layers.

[0119] Furthermore, when the epileptic event detection model training module concatenates the patient's response phenotype with the temporal features output by the epileptic event detection model, it is specifically used for:

[0120] The patient's response phenotype is encoded as a one-hot vector, and then transformed into a response phenotype embedding vector through an embedding matrix:

[0121]

[0122] in, For the patient response phenotype one-hot vector, In response to the phenotypic embedding matrix, For bias terms, For response phenotype embedding vectors; d e For the embedding dimension, such as 8, 16, or 32; C is the number of response phenotype categories;

[0123] Temporal features extracted from multimodal signals using an epidural event detection model. Global average pooling is performed to obtain fixed-length signal features. :

[0124]

[0125] in, , The length of the signal segment;

[0126] Fixed-length signal characteristics Concatenate with the response phenotype embedding vector:

[0127]

[0128] in, .

[0129] Furthermore, in the epileptic event detection model training module, when the epileptic event detection model generates modal weight vectors through fully connected layers for channel weighting, it is specifically used for:

[0130] Gating weights for different modalities are generated using response phenotype embedding vectors:

[0131]

[0132] in, ,and , and These are the gating parameters;

[0133] The modal features are then weighted and fused to obtain the fused features. :

[0134]

[0135] in, This represents the contribution weight of the m-th modality in the k-th patient, where m is the m-th modality. The features output by the event detection model are subjected to global average pooling to obtain fixed-length signal features.

[0136] This invention reveals for the first time the asymmetric asynchronous coupling mechanism of peripheral responses to epileptic seizures, treating multimodal responses as nonlinear, asynchronously triggered dynamic systems. By constructing cross-modal temporal response maps, it systematically quantifies the asynchronous response relationships between different peripheral physiological modalities during epileptic seizures for the first time, solving the problem of high noise sensitivity in models caused by simple synchronous fusion in existing technologies. This invention constructs a cross-individual heterogeneity modeling system based on robust response phenotype aggregation, achieving a leap from individualized fragmented data to group response phenotype patterns. Through unsupervised clustering to identify patient physiological response phenotypes, it realizes structured modeling of the "heterogeneous within the same disease" problem, providing a new methodological path for cross-individual generalization.

[0137] This invention proposes a deep learning monitoring framework based on physiological constraints and data-driven approaches, significantly improving recognition robustness and cross-individual generalization ability. By embedding cross-modal temporal and patient response phenotypic information as physiological prior constraints into the deep learning model, the model can proactively perceive evolutionary differences between different modalities and individuals, significantly improving detection robustness and cross-individual generalization ability, and enhancing the accuracy of epileptic seizure event detection. Furthermore, this invention effectively filters redundant computations through a two-stage asynchronous wake-up pre-screening mechanism based on physiological temporal features; the pre-screening step effectively filters non-seizure signal segments, significantly reducing computational power consumption and extending the battery life of wearable devices. Attached Figure Description

[0138] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0139] Figure 2 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0140] The invention will now be described in further detail with reference to the accompanying drawings.

[0141] like Figure 1 As shown, the epileptic seizure monitoring method of the present invention includes the following steps:

[0142] S1. Perform preprocessing on the signals in the dataset, including denoising and alignment.

[0143] The datasets used in this invention are the SeizeIT2 public dataset and the HebmuSeize clinical dataset. The SeizeIT2 public dataset includes 125 patients and 883 seizure events, which serve as the primary modeling and internal validation data; the HebmuSeize clinical dataset includes 30 patients and 69 seizure events, which serve as independent external validation data.

[0144] The dataset includes electrocardiogram (ECG), electromyogram (EMG), and acceleration (ACM) signals. Noise reduction was performed on the signals: an eighth-order Butterworth low-pass filter (10Hz cutoff) was applied to the ACM signal to remove high-frequency interference; baseline drift filtering and R-peak detection were performed on the ECG signal based on wavelet transform. The denoised signals were then aligned by mapping the ECG, EMG, and ACM signals to the same time grid to ensure temporal alignment. The original temporal information of each modality was preserved during preprocessing, without forcibly assuming synchronous changes across different modalities, thus providing a foundation for subsequent cross-modal temporal feature modeling.

[0145] S2. Filter candidate events based on the signals in the preprocessed dataset and determine the labels of the candidate events.

[0146] The labels include epileptic seizures and non-epileptic seizures.

[0147] The signals in the dataset include signals before, during, and after a seizure. Each patient's signals appear multiple times in the same segment, but are labeled only during seizures, with the start and end of each seizure event marked.

[0148] The processing method for a segment of signal from the same patient is as follows:

[0149] For the preprocessed acceleration signal, calculate the standard deviation (SD) within a sliding window (e.g., 10 seconds). When the SD within the sliding window exceeds a first threshold (e.g., 0.4) determined based on the statistical distribution of the training data, mark the sliding window as a pre-screening segment of the acceleration signal.

[0150] For heart rate signal processing, the increase in heart rate relative to baseline within a sliding window is calculated. When the increase exceeds a second threshold (e.g., 10%) determined based on the statistical distribution of training data, the sliding window is marked as a pre-screened segment of the heart rate signal.

[0151] The system calculates the root mean square value, integral electromyography (EMG), zero-crossing rate, and spectral energy or high-frequency energy ratio within a sliding window for the EMG signal. When the EMG energy exceeds a third threshold, it is marked as a pre-screening segment of the EMG signal.

[0152] The threshold can be the median or mean of the statistical data.

[0153] Preprocessing is performed on the pre-screening segments of different modal signals. Specifically, pre-screening segments of different modal signals with an interval smaller than a preset window interval are merged into a merged pre-screening segment. The start time of the earliest pre-screening segment to be merged is used as the start time of the merged pre-screening segment, and the end time of the latest pre-screening segment to be merged is used as the end time of the merged pre-screening segment. Both the merged pre-screening segment and the unmerged pre-screening segment are used as the initial screening segment.

[0154] When the time interval between two initial screening segments of the same person is less than a merging threshold (e.g., 100 seconds) determined based on the statistical characteristics of the duration of the attack, the two segments are merged into a single merged initial screening segment. The start time of the earliest segment to be merged is used as the start time of the merged initial screening segment, and the end time of the latest segment to be merged is used as the end time of the merged initial screening segment. This results in merged initial screening segments and unmerged initial screening segments. The start time of a merged initial screening segment is used as the start time of a candidate event, and the end time of that merged initial screening segment is used as the end time of that candidate event. Similarly, the start time of an unmerged initial screening segment is used as the start time of a candidate event, and the end time of that unmerged initial screening segment is used as the end time of that candidate event. A candidate event includes ECG, EMG, and accelerometer signals from the start time to the end time of the candidate event.

[0155] The preset window interval is usually less than the merging threshold.

[0156] The multimodal signals in the candidate events are converted into fixed-length feature representations. Signals shorter than the preset length are zero-padded, and signals longer than the preset length are truncated.

[0157] The screening method of this invention for candidate events identifies cases where patients exhibit signal fluctuations that are not epileptic seizures, such as movement. To avoid prediction errors in subsequent seizure event detection models, these cases need to be labeled. Specifically, it determines whether the candidate event contains a label indicating the onset of an epileptic seizure in the dataset. If not, the candidate event is labeled as a non-epileptic seizure event; if it does, it is labeled as an epileptic seizure event.

[0158] S3. Construct a graph of candidate events, specifically:

[0159] Cross-correlation analysis was used to determine the macroscopic response order of each modality relative to the onset point. Assume the time series representation of the multimodal peripheral physiological signals corresponding to a single candidate event is as follows:

[0160] X(t) = [x ECG (t),x EMG (t),x ACM (t)]

[0161] Where, x ECG (t) represents the ECG signal at time t, x EMG (t) represents the EMG signal at time t, x ACM (t) represents the ACM signal at time t.

[0162] S3-1. Calculate the maximum correlation delay τ of any modal i and j corresponding to a single candidate event within the time window w. ij (w):

[0163] τ ij (w)=argmax {τ∈[-T,T]} {corr(x i (t),x j (t+τ))}

[0164] Where T is the maximum search delay, x i (t) is the signal corresponding to mode i at time t, x j (t+τ) is the signal corresponding to mode j at time t+τ.

[0165] The time window w is the time period of the candidate event.

[0166] Maximum correlation delay τ ij (w) serves as a static benchmark for cross-modal temporal structures, describing the macroscopic differences in the response order of different physiological systems to the same event. Within a set delay search range, it finds the time offset corresponding to the highest correlation between the two modes.

[0167] S3-2. Calculate the DTW distance between different modal signals.

[0168] Based on the coarse-grained delay estimation, the DTW method is then used to nonlinearly align the signals of the three modes. Specifically, a cost matrix is ​​constructed, where C(i,j)=d(x i ,x j ) represents the distance metric between modal signals (i.e., DTW distance), where each element in the cost matrix C represents the difference between two modes at two time points; d represents the distance metric (i.e., cost), which can be Euclidean distance, absolute difference, or standardized feature distance.

[0169] S3-3. Causal feature analysis is also performed on the multimodal signal to calculate the transfer entropy between modes:

[0170]

[0171] Where X is the source signal, Y is the target signal, and y t+1 Let Y be the state at the next moment. Let Y be the historical state composed of u points in the past. For X past The historical state consists of _p_ points in time, where p is a probability distribution. It represents the intensity of directional information flow (i.e., transfer entropy) from the source signal X to the target signal Y.

[0172] S3-4. Next, the linear coupling strength between modal signals is quantified by calculating the correlation coefficient or performing coherence analysis. The coupling strength can be calculated using the correlation coefficient, cross-correlation peak value, coherence, or mutual information. A simple implementation can use the Pearson correlation coefficient: calculate the linear correlation between two standardized modal sequences within the same sliding window, and take the absolute value or peak value as the linear coupling strength α. ij .

[0173] S3-5. Finally, the multimodal signal of each candidate event is used as the node feature matrix, and the maximum correlation delay, DTW distance, transfer entropy and coupling strength between two modes are used as the adjacency matrix to obtain the spectrum of each candidate event.

[0174] The node feature matrix records the multimodal signals corresponding to candidate events, with one node corresponding to one modality signal. The adjacency matrix records the relationships between modalities (such as maximum correlation delay, DTW distance, propagation entropy, and coupling strength). For example, three modalities, HR, EMG, and ACM, can form a 3×3 matrix, where the i-th row and j-th column represents the relationship between modality i and modality j. If there are multiple edge attributes, a multi-relationship adjacency matrix can be formed.

[0175] S4. Train the unsupervised clustering model based on the atlas labeled with epileptic seizure events to obtain the response phenotype for each patient in the dataset, and adjust the parameters of the unsupervised clustering model as follows:

[0176] S4-1. Integrate all candidate events labeled as epileptic seizures for each patient in the dataset.

[0177] For example, N candidate events labeled as epileptic seizures were recorded for the k-th patient. These events were integrated using the median and interquartile range (IQR) to obtain the integrated feature P for the k-th patient. k :

[0178] P k =(median{E k,1 E k,2 ,…,E k,n ,…,E k,N ),IQR{E k,1 E k,2 ,…,E k,n ,…,E k,N}), n=1,...,N

[0179] Among them, E k,n This is the atlas corresponding to the nth seizure event of the kth patient.

[0180] For all candidate events labeled as epileptic seizures of the same patient, the median of the elements at the same position in the adjacency matrix is ​​taken. If there are two medians, the average of the two medians is taken to obtain the adjacency matrix composed of the medians for that patient. The same process is applied to the node feature matrix to obtain the node feature matrix composed of the medians of all candidate events labeled as epileptic seizures for that patient. k,1 E k,2 ,…,E k,n ,…,E k,N ) represents the adjacency matrix and the node feature matrix composed of the median of the patient.

[0181] For all candidate events labeled as epileptic seizures of the same patient, the interquartile range (IQR) of elements at the same position in the adjacency matrix is ​​taken, resulting in an adjacency matrix composed of IQR values. Similarly, for all candidate events labeled as epileptic seizures of the same patient, the interquartile range (IQR) of elements at the same position in the node feature matrix is ​​taken, resulting in a node feature matrix composed of IQR values ​​for that patient. k,1 E k,2 ,…,E k,n ,…,E k,N} represents the adjacency matrix and node feature matrix composed of the interquartile ranges of the patient.

[0182] For those with fewer seizures (e.g., N) kFor newly diagnosed patients (<3), the representation is modified by incorporating the population prior distribution or Bayesian estimation to enhance stability. Specifically, when a patient has few seizures, the representation does not rely entirely on individual statistics but rather on a weighted fusion of the patient's event characteristics with the average response patterns of similar patients or the overall population. The fewer the number of seizures, the greater the weight of the population prior; as the patient's seizure record increases, the weight of individual data gradually increases. This reduces the impact of random noise from single seizures on the judgment of response phenotypes.

[0183] For patients with a low frequency of seizures, the following weighted method is used to fuse individual characteristics and prior group characteristics to obtain the patient's integrated characteristics. :

[0184]

[0185] in, For the k-th patient, the patient-level response feature vector is obtained from all candidate events labeled as epileptic seizures. To obtain the average response feature vector of the total number of patients in the training set, , The number of seizures this patient has had is given, and c is the smoothing coefficient.

[0186] The average response eigenvector is the average value of the element at the same position in the matrix of all patient profiles.

[0187] The value of c can be customized. For example, c equals 10, and the weight of the patient's individual characteristics is increased when the number of candidate events for all patient-labeled seizures exceeds 10.

[0188] The fewer the number of seizures, The smaller the group size, the greater the prior weight; as the frequency of seizures increases, As the number of patients increases, the weight of individual patient characteristics gradually rises.

[0189] This invention uses unsupervised clustering models (such as K-means, Gaussian mixture model, hierarchical clustering, etc.) to identify the following three typical response phenotypes:

[0190] Autonomic nervous system-priority type: characterized by a significant shift in ECG signals in the very early stages of an attack, with heart rate changes preceding body motion signals; Motor response-dominant type: characterized by explosive and highly synchronous changes in EMG and ACM signals, with short cross-modal response delays; Hybrid coupled response type: exhibiting multi-system coupling characteristics, with relatively balanced contributions from each modality.

[0191] Clustering results are not the sole criterion for judgment; instead, they are validated in conjunction with clinical manifestations, modal contribution weights, and subsequent detection performance. Semi-supervised correction or expert review may be employed when necessary to improve the stability of response phenotype segmentation.

[0192] S4-2. By training the unsupervised clustering model with the ensemble features of each patient, the parameters of the unsupervised clustering model are optimized and adjusted, and the response phenotype of each patient in the dataset is obtained.

[0193] S5. Construct an episodic event detection model and train the model using the spectrum of candidate events and the response phenotypes of patients corresponding to the candidate events.

[0194] The event detection model uses a CNN-LSTM or multimodal Transformer architecture.

[0195] Each candidate event's graph and the patient's response phenotype corresponding to the candidate event are used as a sample. The samples in the dataset are divided into a training set and a validation set. The samples in the training set and the validation set are input into the epidural event detection model for training.

[0196] Based on the adjacency matrix of the candidate event graph, feature extraction is performed on the multimodal signal segments of the candidate events to output temporal features. Finally, the output temporal features are weighted using the response phenotype of the patients corresponding to the candidate events and input into the classification head to obtain the probability of epileptic seizures.

[0197] The node feature matrix of the sample is input into the CNN-LSTM. The CNN-LSTM processes the input data by receiving standardized multimodal signal segments (5 minutes × 20Hz = 6000 time points × 3 modalities) at the input layer, with each modality using an independent channel input. Feature extraction is then performed on each modality, specifically through three sets of convolutional-pooling blocks, batch normalization layers, ReLU activation functions, and max-pooling layers. A channel attention mechanism is introduced after the first set of convolutional-pooling blocks.

[0198] Each convolution-pooling block has a kernel size of 3-7, a stride of 1, and padding to maintain the size.

[0199] Transform the adjacency matrix in the graph of candidate events into a prior matrix M. prior Specifically, the following steps are taken: First, the latency, DTW distance, coupling strength, and propagation entropy are normalized. Then, relationships that are conducive to modal interaction are assigned higher scores, while weakly correlated relationships or those that do not conform to physiological time sequences are assigned lower scores. Finally, a prior matrix compatible with the attention matrix dimension is obtained. This matrix can be fixed during model training or set as a learnable parameter for fine-tuning. The specific method for attention calculation is as follows:

[0200] S5a-1. Based on weighting coefficients The adjacency matrices of the samples are weighted and fused to obtain the prior relation scores between modes. :

[0201]

[0202] in, and The four weights are preset or learnable weights, and their sum is 1. The maximum correlation delay between mode i and mode j. The DTW distance between mode i and mode j The coupling strength between mode i and mode j is... Let be the transfer entropy from mode i to mode j.

[0203] Through conversion functions respectively , , and Normalize the attributes of each side to a uniform value range. , , and All calculations are normalized. Used to convert the maximum relevant delay into a delay prior score. Used to convert DTW distance into response similarity score Used to convert coupling strength into correlation score Used to convert transitive entropy into directional information flow scores.

[0204] In one implementation, the attributes of each edge are weighted equally, i.e. In another implementation, the weighting coefficients are determined by a validation set search, with detection sensitivity, false alarm rate, or F1 score as the optimization objective.

[0205] S5a-2. All Construct a modality-level prior adjacency matrix, and expand it into a prior matrix with the same dimension as the attention matrix based on the modality to which the input token belongs. .

[0206]

[0207] In cross-modal attention calculation, the original attention score matrix is ​​first calculated from the Query matrix and the Key matrix. Then, the prior attention bias matrix obtained from the cross-modal time-series response map is added to the original attention score matrix according to the coefficients. Subsequently, the attention weights are obtained by softmax normalization, and the Value matrix is ​​weighted and summed.

[0208] S5a-3. In the channel attention mechanism calculation, the original attention score matrix is ​​added to the prior matrix, and then the attention weights are obtained by softmax normalization:

[0209]

[0210] Where α is the prior strength hyperparameter, Q represents the Query matrix, K represents the Key matrix, and V represents the Value matrix. Q, K, and V are all obtained by projecting the features from the previous layer through different weight matrices.

[0211] The extracted temporal feature sequences are then input into an LSTM layer to capture the time-series dependency relationships of the temporal feature sequences. A Dropout layer is then added after the LSTM layer to prevent overfitting.

[0212] The LSTM layer contains 32 hidden units, and the dropout layer has a dropout rate of 0.5.

[0213] Feature extraction can also be performed using a multimodal Transformer architecture. The input embedding layer of this architecture divides the multimodal signal into fixed-length patches (e.g., 20 sampling points per second × 20Hz). Each patch is mapped to a d-dimensional embedding vector through a linear projection layer, with positional encoding and modality type embedding added before feature extraction. The cross-modal Transformer encoder in the multimodal Transformer architecture is a multi-head self-attention layer consisting of stacked L layers (e.g., 4 layers). Each attention head contains three projection matrices: Query, Key, and Value. Cross-modal attention allows tokens from different modalities to interact with each other. The cross-modal temporal response map is then transformed into a prior attention bias matrix M. prior In self-attention computation, the following is introduced:

[0214]

[0215] Self-attention computation enables the model to prioritize cross-modal interactions that conform to physiological principles.

[0216] Furthermore, patient response phenotypes are used to conditionally fuse or conditionally modulate the temporal features extracted by the seizure event detection model. For CNN-LSTM models or multimodal Transformer architectures, the patient's response phenotype is encoded as a one-hot vector. Then, the response phenotype embedding vector is calculated based on the one-hot vector and concatenated with the temporal features output by the CNN-LSTM model or multimodal Transformer architecture (i.e., conditional fusion), or channel-weighted (conditional modulation) is performed by generating modality weight vectors through fully connected layers. Finally, global average pooling is used to compress the temporal feature sequence into a fixed-length vector, which is then processed by a classification head to output the seizure probability.

[0217] The classification head used in this invention consists of two fully connected layers.

[0218] For example, the encoded response phenotype, the autonomic nervous system-preferred type: Motion response-driven type: Hybrid Coupled Response Type: .

[0219] S5b-1 one-hot vectors are sparse and can typically be transformed into continuous response phenotype embedding vectors using an embedding matrix:

[0220]

[0221] in, For the patient response phenotype one-hot vector, In response to the phenotypic embedding matrix, For bias terms, For response phenotype embedding vectors; d e For the embedding dimension, such as 8, 16, or 32; C is the number of response phenotype categories, and k represents the kth patient.

[0222] The meaning is: to convert category labels such as "autonomic nervous system priority type, motor response dominant type, and hybrid coupling type" into continuous numerical vectors that the model can calculate.

[0223] S5b-2. Temporal features extracted from multimodal signals using CNN-LSTM or multimodal Transformer architecture. Global average pooling is performed to obtain fixed-length signal features. :

[0224]

[0225] in, , This represents the length of the signal segment.

[0226] The specific method for conditionally fusing the temporal features extracted by the episodic event detection model using the patient's response phenotype is as follows:

[0227] S5b-3. Then, the signal features are concatenated with the response phenotype embedding vector:

[0228]

[0229] in, .

[0230] And input the classification header to obtain the probability of epileptic seizures:

[0231]

[0232] in, This represents the probability of an epileptic seizure. For the Sigmoid function; and These are the classification header parameters.

[0233] The specific method for conditionally modulating the temporal features extracted by the episodic event detection model using the patient's response phenotype is as follows:

[0234] Gating weights for different modalities are generated using response phenotype embedding vectors:

[0235]

[0236] in, ,and , and These are the gating parameters.

[0237] The modal features are then weighted and fused to obtain the fused features. :

[0238]

[0239] in, This represents the contribution weight of the m-th modality in the k-th patient.

[0240] Input the classification head, and the classification head outputs the probability of epileptic seizures. :

[0241]

[0242] in, and These are the classification header parameters.

[0243] In this way, the seizure event detection model can dynamically adjust the contribution weights of cardiac-related modalities, electromyographic modalities, and acceleration modalities according to the patient's response phenotype, enabling patients with different response phenotypes to adopt differentiated modality fusion strategies, thereby achieving individualized seizure detection.

[0244] This invention uses a loss function to train the event detection model. The loss function is: classification loss (cross-entropy) + λ·mechanism consistency constraint loss.

[0245] Among them, the mechanism consistency constraint loss is: constraining the model attention matrix and the prior attention bias matrix M. prior The KL divergence between the values. Optimizer: Adam, initial learning rate 0.001, batch size 64. Early stopping strategy: based on validation set loss, patience value 10 epochs.

[0246] S6. Perform preprocessing, including denoising and alignment, on the multimodal peripheral physiological signals of the monitored subjects acquired in real time.

[0247] Multimodal peripheral physiological signals of the monitored subject are collected synchronously using wearable sensors. Wearable sensors include, but are not limited to, armband devices, wristband devices, and chest patch devices. Multimodal peripheral physiological signals include at least: (1) electrocardiogram (ECG) signals or their derived heart rate (HR) time-series data, reflecting the autonomic nervous system response of the heart. Heart rate information can be obtained through photoplethysmography (PPG) signals; (2) electromyography (EMG) signals, reflecting the muscle activation process; and (3) acceleration signals (ACM), reflecting motor behavior.

[0248] The sampling frequencies of the signals can be different. For example, the sampling frequency of an acceleration signal is 11-12 Hz, while the sampling frequency of a photoplethysmography (PPG) signal is 100 Hz. To ensure uniform processing, the raw signals are preprocessed as follows:

[0249] Noise reduction was performed on different modal signals. Specifically, a second-order Butterworth bandpass filter (0.5-3.5Hz) was applied to the PPG signal to remove noise and frequency components outside the heart rate range; an eighth-order Butterworth lowpass filter (cutoff frequency 10Hz) was applied to the acceleration signal to remove high-frequency interference; and baseline drift filtering and R-peak detection were performed on the ECG signal based on wavelet transform.

[0250] The original heart rate time series, such as PPG / ECG, electromyography, and acceleration, or their derived signals, are uniformly mapped to the same time grid. During preprocessing, the original temporal information of each modality is preserved, and the assumption of synchronous changes across different modalities is not enforced, thus providing a foundation for subsequent cross-modal temporal feature modeling.

[0251] S7. Determine the monitored events based on the pre-processed multimodal peripheral physiological signals of the monitored subjects.

[0252] In this invention, the collected signals are not input into the event detection model for prediction each time. Instead, the collected signals are filtered to determine the event to be detected, and then the event to be detected is predicted. This can reduce computational power consumption in real-time monitoring scenarios.

[0253] For the preprocessed acceleration signal, calculate the standard deviation (SD) within a sliding window (e.g., 10 seconds). When the SD within the sliding window exceeds a first threshold (e.g., 0.4) determined based on the statistical distribution of the training data, mark the sliding window as an active segment of the acceleration signal.

[0254] For heart rate signal processing, specifically, the increase in heart rate relative to baseline within a sliding window is calculated. When the increase exceeds a second threshold (e.g., 10%) determined based on the statistical distribution of training data, the sliding window is marked as an active segment of the heart rate signal.

[0255] The root mean square value, integral electromyography (EMG), zero-crossing rate, and spectral energy or high-frequency energy ratio within a sliding window are calculated for the EMG signal. When the EMG energy exceeds a third threshold, it is marked as an active segment of the EMG signal.

[0256] The first threshold, second threshold, and third threshold can be the median or mean of the statistical data.

[0257] The active segments of different modal signals are preprocessed. Specifically, active segments of different modal signals with an interval smaller than a preset window interval are merged into a merged active segment. The start time of the earliest active segment to be merged is taken as the start time of the merged active segment, and the end time of the latest active segment to be merged is taken as the end time of the merged active segment. Both the merged active segment and the active segments that have not been merged are regarded as events to be screened.

[0258] When the time interval between two screening seizure events of the same person is less than a merging threshold (e.g., 100 seconds) determined based on the statistical characteristics of seizure event duration, the two are merged into a single screening seizure event. The start time of the earliest screening seizure event to be merged is used as the start time of the merged event, and the end time of the latest screening seizure event to be merged is used as the end time of the merged event. Both merged and unmerged screening seizure events are considered as detected events.

[0259] Among them, the preset window interval is less than the merging threshold.

[0260] The multimodal signal in the detected event is converted into a fixed-length feature representation. Signals shorter than the preset length are zero-padded, and signals longer than the preset length are truncated.

[0261] S8. Establish a map of the detected events based on the detected events of the monitored subjects.

[0262] After identifying the event to be detected, the same steps are performed on the event to be detected as on the candidate events. Based on the ECG, EMG, and accelerometer signals of the event to be detected, the maximum correlation delay, DTW distance, propagation entropy, and coupling strength are extracted. Similar to the method used to build the map of the candidate events, the multimodal signal of the event to be detected is used as the node feature matrix, and the maximum correlation delay, DTW distance, propagation entropy, and coupling strength are used as the adjacency matrix to obtain the map of the event to be detected.

[0263] S9. Based on the graph of the detected events and the trained unsupervised clustering model, determine the response phenotype of the monitored individuals.

[0264] Historical epileptic seizure events of the monitored individual are obtained. When the monitored individual has more than two historical epileptic seizure events, the historical epileptic seizure events and the detected events are directly integrated based on the spectra of the monitored individual's historical epileptic seizure events and detected events to obtain the monitored individual's integrated features. The integrated features are then input into the trained unsupervised clustering model to obtain the monitored individual's response phenotype. When the monitored individual has fewer than two historical epileptic seizure events, the monitored individual's historical epileptic seizure events, the detected events, and the average response feature vector of the overall population are weighted and fused. The overall population can be the group of patients with the most response phenotypes in the training set, or the group of patients with the most response phenotypes statistically analyzed during the use of this invention. The fewer the number of seizures, the greater the group prior weight; as the number of seizure records for the patient increases, the individual data weight gradually increases. This can reduce the influence of random noise from a single seizure on the judgment of the response phenotype.

[0265] A type discrimination and matching mechanism based on cross-modal temporal features is constructed to quickly locate the corresponding temporal response type in new observation data. Specifically, the cross-modal temporal feature vector of a new patient or new event is compared with the response phenotype center or response phenotype probability distribution formed during the training phase. The distance or posterior probability between the vector and the autonomic nervous system-dominated, motor response-dominated, or mixed-coupled response phenotypes is calculated, and the vector is assigned to the closest response phenotype. Commonly used methods include nearest-center matching, Mahalanobis distance matching, Gaussian mixture model posterior probability matching, or lightweight classifier discrimination.

[0266] S10. Input the spectrum of the event to be detected and the response phenotype of the monitored person into the trained seizure event detection model, and output the epileptic seizure probability of the detected event; when the epileptic seizure probability is greater than the preset probability threshold, the detected event is determined to be an epileptic seizure and an early warning signal is issued.

[0267] The graph of the event to be detected and the response phenotype of the monitored subject are input into the trained seizure event detection model. The adjacency matrix of the graph of the event to be detected is used to extract features from the node feature matrix and output temporal features. Then, the response phenotype of the monitored subject is used to weight the features output by the detection model. Finally, the seizure probability of the event to be detected is obtained by inputting the classification head of the seizure event detection model.

[0268] The preset probability threshold can be set manually, for example, to 60%. When the probability of an epileptic seizure exceeds the preset threshold, the detected event can be identified as an epileptic seizure, and an early warning signal can be issued. This early warning signal, which includes the probability of an epileptic seizure, can be sent to the medical terminal. When the probability of an epileptic seizure is less than or equal to the preset threshold, normal monitoring is maintained, and no early warning signal is output.

[0269] like Figure 2 As shown, the epileptic seizure monitoring system of the present invention includes:

[0270] The dataset preprocessing module, connected to the event pre-filtering module, is used to preprocess the signals in the dataset, including denoising and alignment, and then send the preprocessed dataset to the event pre-filtering module.

[0271] The event pre-screening module is connected to the dataset preprocessing module and the event-level time series graph construction module, respectively. It is used to screen candidate events based on the signals of the preprocessed dataset and determine the labels of the candidate events, including epileptic seizures and non-epileptic seizures; and send the candidate events to the event-level time series graph construction module.

[0272] The event-level time series graph construction module is connected to the event pre-screening module, the patient response phenotype training module, and the epidural event detection model training module, respectively. It is used to build graphs of candidate events and send the graphs of candidate events to the patient response phenotype training module and the epidural event detection model training module.

[0273] The patient response phenotype training module is connected to the event-level time series graph construction module, the seizure event detection model training module, and the response phenotype determination module, respectively. It is used to train the unsupervised clustering model based on the graph of candidate events labeled as epileptic seizures to obtain the response phenotype of each patient in the dataset and adjust the parameters of the unsupervised clustering model. The patient's response phenotype is input into the seizure event detection model training module, and the trained unsupervised clustering model is input into the response phenotype determination module.

[0274] The seizure event detection model training module is connected to the event-level time series graph construction module, the patient response phenotype training module, and the seizure detection output module, respectively. It is used to construct the seizure event detection model, and train the seizure event detection model using the graph of candidate events and the response phenotype of the patients corresponding to the candidate events. The trained seizure event detection model is then sent to the seizure detection output module.

[0275] The real-time signal preprocessing module, connected to the real-time pre-screening module, is used to preprocess the multimodal peripheral physiological signals of the monitored subjects acquired in real time, including denoising and alignment, and send the preprocessed multimodal peripheral physiological signals to the real-time pre-screening module.

[0276] The real-time pre-screening module is connected to the real-time signal preprocessing module and the detected event map construction module, respectively. It is used to determine the detected events of the monitored subject based on the preprocessed multimodal peripheral physiological signals of the monitored subject sent by the real-time signal preprocessing module, and send them to the detected event map construction module.

[0277] The detected event graph construction module is connected to the real-time pre-screening module, the response phenotype determination module, and the seizure detection output module, respectively. It is used to build a graph of the detected events based on the detected events of the monitored subjects sent by the real-time pre-screening module, and send it to the response phenotype determination module and the seizure detection output module.

[0278] The response phenotype determination module is connected to the detected event graph construction module, the patient response phenotype training module, and the seizure detection output module, respectively. It is used to determine the response phenotype of the monitored person based on the detected event graph sent by the detected event graph construction module and the trained unsupervised clustering model sent by the patient response phenotype training module; and send the response phenotype of the monitored person to the seizure detection output module.

[0279] The seizure detection output module is connected to the response phenotype determination module, the seizure event detection model training module, and the detected event atlas construction module, respectively. It is used to input the atlas of the detected event sent by the detected event atlas construction module and the response phenotype of the monitored subject sent by the response phenotype determination module into the trained seizure event detection model sent by the seizure event detection model training module, and output the seizure probability of the detected event. When the seizure probability is greater than the preset probability threshold, the detected event is determined to be a seizure, and an early warning signal is issued.

Claims

1. A method for monitoring epileptic seizures based on cross-modal temporal and individual responses, characterized in that, include: S1. Perform preprocessing on the signals in the dataset, including denoising and alignment; S2. Filter candidate events based on the signals of the preprocessed dataset and determine the labels of the candidate events, the labels including epileptic seizures and non-epileptic seizures; S3. Build a graph of candidate events; S4. Train the unsupervised clustering model based on the map of candidate events labeled as epileptic seizures to obtain the response phenotype of each patient in the dataset, and adjust the parameters of the unsupervised clustering model. S5. Construct an episodic event detection model and train the model using the candidate event map and the response phenotypes of the patients corresponding to the candidate events; S6. Perform preprocessing, including denoising and alignment, on the multimodal peripheral physiological signals of the monitored subjects collected in real time; S7. Determine the monitored events based on the pre-processed multimodal peripheral physiological signals of the monitored subjects; S8. Establish a map of the monitored events based on the monitored events; S9. Based on the graph of the detected events and the trained unsupervised clustering model, determine the response phenotype of the monitored individuals; S10. Input the spectrum of the event to be detected and the response phenotype of the monitored person into the trained seizure event detection model, and output the epileptic seizure probability of the detected event; when the epileptic seizure probability is greater than the preset probability threshold, the detected event is determined to be an epileptic seizure and an early warning signal is issued.

2. The method for monitoring epileptic seizures according to claim 1, characterized in that, The specific method for building a graph of candidate events in step S3 is as follows: S3-1. Calculate the maximum correlation delay corresponding to different modal signals for a single candidate event: τ ij (w)=argmax {τ∈[-T,T]} {corr(x i (t),x j (t+τ))}, Where T is the maximum search delay, x i (t) is the signal corresponding to mode i at time t, x j (t+τ) is the signal corresponding to mode j at time t+τ; S3-2. Calculate the DTW distance between signals of different modes; S3-3. Calculate the transfer entropy between different modes. : , Where X is the source signal, Y is the target signal, and y t+1 Let Y be the state at the next moment. Let Y be the historical state composed of u points in the past. For X in the past The historical state consists of _ ... S3-4. Calculate the coupling strength between signals of different modes; S3-5. Use the multimodal signal of each candidate event as the node feature matrix, and the maximum correlation delay, DTW distance, propagation entropy and coupling strength of each candidate event as the adjacency matrix to obtain the event-level time series map of each candidate event.

3. The method for monitoring epileptic seizures according to claim 1, characterized in that, The specific method for training the unsupervised clustering model in step S4 is as follows: S4-1. All candidate events labeled as epileptic seizures from the same patient are integrated using median and interquartile range to obtain the integrated feature P for each patient. k : P k =(median{E k,1 ,AND k,2 ,…,AND k,n ,…,AND k,N ),IQR{E k,1 ,AND k,2 ,…,AND k,n ,…,AND k,N }),n=1,...,N, Among them, E k,n This is the atlas corresponding to the nth seizure event of the kth patient; S4-2. Input the ensemble features of each patient into the unsupervised clustering model and adjust the parameters of the unsupervised clustering model.

4. The method for monitoring epileptic seizures according to claim 3, characterized in that, For patients whose seizure frequency is less than a preset seizure frequency threshold, the patient's integrated characteristics are obtained. The specific method is: By using a weighted approach to fuse individual characteristics and prior group characteristics, the integrated characteristics of patients are obtained. : , in, This is the patient-level response feature vector obtained from the existing epidemiological events of the k-th patient. To train the average response feature vector of the total number of patients in the set or patients with similar response phenotypes, , The number of seizures for the kth patient is given, and c is the smoothing coefficient.

5. The method for monitoring epileptic seizures according to claim 1, characterized in that, The specific method by which the episodic event detection model processes the candidate event profile and the corresponding patient response phenotypes is as follows: Based on the adjacency matrix of the candidate event spectrum, feature extraction is performed on the multimodal signal segments of the candidate events, and temporal features are output. Finally, the output temporal features are conditionally fused or conditionally modulated using the response phenotype and input into the classification head to obtain the epileptic seizure probability.

6. The method for monitoring epileptic seizures according to claim 5, characterized in that, The event detection model is a CNN-LSTM or multimodal Transformer architecture; When the event detection model is a CNN-LSTM, the specific method by which the event detection model extracts features from multimodal signal segments is as follows: The input layer of CNN-LSTM receives multimodal signal segments, with each modality using an independent channel input. After that, it passes through 3 sets of convolutional-pooling blocks, batch normalization layers, ReLU activation functions, and max pooling layers. After the first set of convolutional-pooling blocks, a channel attention mechanism is introduced. The channel attention mechanism uses the adjacency matrix of the graph to perform attention calculation on the features extracted by the first set of convolutional-pooling blocks. When the seizure event detection model is a multimodal Transformer architecture, the specific method by which the seizure event detection model extracts features from the multimodal signal segments is as follows: The input embedding layer of the multimodal Transformer architecture divides the multimodal signal segment into fixed-length patches. Each patch is mapped to a d-dimensional embedding vector through a linear projection layer, and position encoding and modality type embedding are added. Then, it is input to several stacked multi-head attention layers. Attention is calculated in the multi-head attention layers using the adjacency matrix of the graph.

7. The method for monitoring epileptic seizures according to claim 6, characterized in that, The specific method by which the event detection model uses the adjacency matrix of the graph to calculate attention is as follows: S5a-1. Normalize the adjacency matrix and perform weighted fusion to obtain the prior relation scores between modes. : , in, and For preset or learnable weights, The maximum correlation delay between mode i and mode j. The DTW distance between mode i and mode j The coupling strength between mode i and mode j is... Let i be the transfer entropy from mode i to mode j. , , and All calculations are normalized. S5a-2. Score all intermodal prior relations. Construct a modality-level prior adjacency matrix, and expand it into a prior matrix with the same dimension as the attention matrix based on the modality to which the input token belongs. : , S5a-3. In attention calculation, the original attention score matrix is ​​added to the prior matrix, and then the attention weights are obtained by softmax normalization: , Where α is the prior strength hyperparameter, Q represents the Query matrix, K represents the Key matrix, and V represents the Value matrix. Q, K, and V are all obtained by projecting the features from the previous layer through different weight matrices.

8. The method for monitoring epileptic seizures according to claim 6, characterized in that, The specific method by which the event detection model utilizes response phenotypes for conditional fusion is as follows: The patient's response phenotype is concatenated with the temporal features output by the episodic event detection model; The specific method by which the episodic event detection model uses response phenotypes for conditional modulation is as follows: Channel weighting is performed by generating modal weight vectors through fully connected layers.

9. The method for monitoring epileptic seizures according to claim 8, characterized in that, The specific method by which the episodic event detection model concatenates the patient's response phenotype with the temporal features output by the model is as follows: S5b-1. Encode the patient's response phenotype into a one-hot vector, and then transform it into a response phenotype embedding vector through an embedding matrix: , in, For the patient response phenotype one-hot vector, In response to the phenotypic embedding matrix, For bias terms, For response phenotype embedding vectors; d e For the embedding dimension, such as 8, 16, or 32; C is the number of response phenotype categories; S5b-2. Temporal features extracted from multimodal signals using an epidural event detection model. Global average pooling is performed to obtain fixed-length signal features. : , in, , The length of the signal segment; S5b-3. Fixed-length signal features Concatenate with the response phenotype embedding vector: , in, .

10. A seizure monitoring system based on cross-modal temporal and individual responses, characterized in that, include: The dataset preprocessing module, connected to the event pre-filtering module, is used to preprocess the signals in the dataset, including denoising and alignment, and then send the preprocessed dataset to the event pre-filtering module. The event pre-screening module is connected to the dataset preprocessing module and the event-level time series graph construction module, respectively. It is used to screen candidate events based on the signals of the preprocessed dataset and determine the labels of the candidate events, including epileptic seizures and non-epileptic seizures; and send the candidate events to the event-level time series graph construction module. The event-level time series graph construction module is connected to the event pre-screening module, the patient response phenotype training module, and the epidural event detection model training module, respectively. It is used to build graphs of candidate events and send the graphs of candidate events to the patient response phenotype training module and the epidural event detection model training module. The patient response phenotype training module is connected to the event-level time series graph construction module, the seizure event detection model training module, and the response phenotype determination module, respectively. It is used to train the unsupervised clustering model based on the graph of candidate events labeled as epileptic seizures to obtain the response phenotype of each patient in the dataset and adjust the parameters of the unsupervised clustering model. The patient's response phenotype is input into the seizure event detection model training module, and the trained unsupervised clustering model is input into the response phenotype determination module. The seizure event detection model training module is connected to the event-level time series graph construction module, the patient response phenotype training module, and the seizure detection output module, respectively. It is used to construct the seizure event detection model, and train the seizure event detection model using the graph of candidate events and the response phenotype of the patients corresponding to the candidate events. The trained seizure event detection model is then sent to the seizure detection output module. The real-time signal preprocessing module, connected to the real-time pre-screening module, is used to preprocess the multimodal peripheral physiological signals of the monitored subjects acquired in real time, including denoising and alignment, and send the preprocessed multimodal peripheral physiological signals to the real-time pre-screening module. The real-time pre-screening module is connected to the real-time signal preprocessing module and the detected event map construction module, respectively. It is used to determine the detected events of the monitored subject based on the preprocessed multimodal peripheral physiological signals of the monitored subject sent by the real-time signal preprocessing module, and send them to the detected event map construction module. The detected event graph construction module is connected to the real-time pre-screening module, the response phenotype determination module, and the seizure detection output module, respectively. It is used to build a graph of the detected events based on the detected events of the monitored subjects sent by the real-time pre-screening module, and send it to the response phenotype determination module and the seizure detection output module. The response phenotype determination module is connected to the detected event graph construction module, the patient response phenotype training module, and the attack detection output module, respectively. It is used to determine the response phenotype of the monitored person based on the detected event graph sent by the detected event graph construction module and the trained unsupervised clustering model sent by the patient response phenotype training module. The response phenotype of the monitored individual is then sent to the seizure detection output module. The seizure detection output module is connected to the response phenotype determination module, the seizure event detection model training module, and the detected event atlas construction module, respectively. It is used to input the atlas of the detected event sent by the detected event atlas construction module and the response phenotype of the monitored subject sent by the response phenotype determination module into the trained seizure event detection model sent by the seizure event detection model training module, and output the seizure probability of the detected event. When the seizure probability is greater than the preset probability threshold, the detected event is determined to be a seizure, and an early warning signal is issued.