Newborn epilepsy electroencephalogram detection method based on multi-modal spatial-temporal feature fusion
By using multimodal spatiotemporal feature fusion and CBAM3D attention mechanism methods in neonatal epilepsy detection, the shortcomings of the existing two-dimensional CNN models in capturing transient epilepsy discharge characteristics and processing multimodal bioelectric signals are solved, and a high accuracy and high sensitivity detection of neonatal epilepsy is achieved.
Patent Information
- Application Number
- CN202510302167.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing two-dimensional CNN model has problems in the detection of neonatal epilepsy in low efficiency in capturing transient epilepsy discharge characteristics and asynchronous time-frequency scale of multimodal bioelectric signals, resulting in insufficient accuracy and sensitivity.
Using a multimodal spatiotemporal feature fusion method, the space-time-domain-frequency-domain features are extracted to achieve accurate recognition of epilepsy detection by constructing a two-dimensional dynamic topological mapping mechanism and a CBAM3D attention mechanism, combined with three-dimensional deformable convolution and wavelet decomposition technology.
The accuracy and sensitivity of neonatal epilepsy detection was significantly improved, the false detection rate was reduced, and the detection accuracy was reached 95.8% and the sensitivity was 94.7%, and the specificity and AUC also reached 98.4%.
Smart Images

Figure CN120203601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical artificial intelligence, and particularly to a neonatal epilepsy electroencephalogram detection method based on multi-modal spatio-temporal feature fusion. Background Art
[0002] Neonatal epilepsy is a common critical illness in clinical practice. When it occurs, the electroencephalogram (EEG) signal features are weak and transient, posing a severe challenge to traditional detection methods. During an epileptic seizure, abnormal discharges of brain neurons may trigger complications such as apnea and metabolic disorders. If not diagnosed and treated in a timely manner, the fatality rate of neonatal epilepsy can be as high as 30%, and sequelae such as mental retardation and paralysis may occur. Currently, clinical diagnosis mainly relies on physicians' visual analysis of electroencephalograms (EEGs). However, neonatal epileptiform discharges have characteristics such as low amplitude (attenuated by 30% - 50% compared to adults), short duration (0.5 - 2 seconds), and spatial dispersion. In the low signal-to-noise ratio NICU monitoring environment, the risk of missed diagnosis by manual interpretation is as high as 40%.
[0003] In recent years, automatic detection methods based on deep learning have been gradually applied to epilepsy recognition. For example, a convolutional neural network (CNN) can effectively extract spatio-temporal features of EEGs through local receptive fields. However, existing two-dimensional CNN models still have significant limitations in neonatal epilepsy detection: on the one hand, the fixed geometric structure of traditional convolutional kernels is difficult to adapt to the non-uniform distribution characteristics of EEG signals in the time-frequency-spatial domain, resulting in low capture efficiency for transient epileptiform discharges; on the other hand, conventional feature fusion methods cannot solve the asynchrony problem of multi-modal bioelectric signals on the time-frequency scale, and the simple splicing of frequency domain information and time domain features causes loss of effective information. According to tests on public datasets, although three-dimensional convolutional neural networks (3D-CNNs) perform well in decoding electroencephalograms (EEGs), in the neonatal epilepsy detection task, their accuracy is only 83.2%, and the sensitivity to focal seizures is less than 75%. Summary of the Invention
[0004] Object of the Invention: To address the limitations of existing two-dimensional CNN models in neonatal epilepsy detection, the present invention proposes a method for epilepsy detection of EEG signals based on multi-modal spatio-temporal feature fusion and attention mechanism. This method can better fuse the spatial information between electrode pairs, extract spatio-temporal-frequency domain features, and is particularly suitable for accurately identifying neonatal epilepsy seizures under low signal-to-noise ratio conditions, especially for accurately identifying weak epileptiform discharges in neonatal electroencephalograms.
[0005] Technical Solution: A neonatal epilepsy electroencephalogram detection method based on multi-modal spatio-temporal feature fusion, comprising the following steps:
[0006] Step 1: Obtain EEG signals through electrode pairs;
[0007] Step 2: Preprocess the acquired EEG signals to obtain preprocessed EEG signals;
[0008] Step 3: Input the preprocessed EEG signals into the trained epilepsy EEG detection model to obtain detection results;
[0009] Among them, the specific operations for preprocessing the acquired EEG signals to obtain preprocessed EEG signals include:
[0010] Perform filtering, slicing, and conversion to frequency-domain signals on the acquired EEG signals in sequence to obtain three-dimensional time-domain signals and three-dimensional frequency-domain signals;
[0011] Perform two-dimensional plane mapping on the three-dimensional time-domain signals and three-dimensional frequency-domain signals to obtain mapped one-dimensional time-domain signals and one-dimensional frequency-domain signals;
[0012] After normalizing the three-dimensional time-domain signals, three-dimensional frequency-domain signals, one-dimensional time-domain signals, and one-dimensional frequency-domain signals respectively, obtain preprocessed EEG signals;
[0013] Among them, the epilepsy EEG detection model includes: a first convolutional branch for performing convolutional operations on three-dimensional time-domain signals, a second convolutional branch for performing convolutional operations on three-dimensional frequency-domain signals, a third convolutional branch for performing convolutional operations on one-dimensional time-domain signals, a fourth convolutional branch for performing convolutional operations on one-dimensional frequency-domain signals, and a fusion decision layer;
[0014] The fusion decision layer is used to fuse the features output by the four convolutional branches and perform epilepsy detection based on the fused features, and output epilepsy detection results.
[0015] Further, the specific operation of performing two-dimensional plane mapping on the three-dimensional time-domain signals and three-dimensional frequency-domain signals is: perform two-dimensional plane mapping on the three-dimensional time-domain signals and three-dimensional frequency-domain signals according to the relative positions of the electrode pairs in the plane.
[0016] Further, the first convolutional branch for performing convolutional operations on three-dimensional time-domain signals includes at least two cascaded three-dimensional convolutional units, which sequentially include a first three-dimensional convolutional unit and a second three-dimensional convolutional unit;
[0017] In both the first three-dimensional convolutional unit and the second three-dimensional convolutional unit, three-dimensional convolutional kernels are used to perform convolutional operations in the spatial dimension, time domain dimension, and channel dimension;
[0018] In the second three-dimensional convolutional unit, the structure after performing convolutional operations using the three-dimensional convolutional kernel is input into the first three-dimensional convolutional attention module, and the first three-dimensional convolutional attention module recalibrates the features in the spatial dimension, time domain dimension, and channel dimension to obtain the features output by the first convolutional branch.
[0019] Furthermore, the first three-dimensional convolutional attention module includes a channel attention sub-module and a spatio-temporal attention sub-module;
[0020] The following calculations are performed in the channel attention sub-module:
[0021] Let the elements in the feature map input to the three-dimensional convolutional attention module be represented as: x ∈ R B×C×H×W×T / F ; where B represents the batch size, C represents the number of channels, H×W represents the spatial dimension, and T / F represents the time / frequency domain dimension;
[0022] Perform global average pooling on the input feature map, denoted as:
[0023]
[0024] In the formula, X[·] represents the input feature map, c represents the channel dimension of the input feature map, h, w represent the spatial dimensions of the input feature map, and t represents the time dimension of the input feature map;
[0025] Perform global maximum on the input feature map, denoted as:
[0026]
[0027] Calculate the weight coefficient M through a multi-layer perceptron with shared weights c , denoted as:
[0028] M c = σ(MLP(F avg ) + MLP(F max ))
[0029] In the formula, σ represents the Sigmoid activation function;
[0030] Perform an element-wise multiplication operation on the calculated weight coefficient M c with each channel of the input feature map to achieve channel recalibration, denoted as:
[0031]
[0032] In the formula, represents the per-channel product;
[0033] The following calculations are performed in the spatio-temporal attention sub-module:
[0034] Perform average pooling on the channel-recalibrated feature map along the channel dimension, denoted as:
[0035]
[0036] The feature map after channel recalibration is subjected to max pooling along the channel dimension, expressed as:
[0037]
[0038] Concatenate S avg and S max in spatio-temporal features, expressed as:
[0039] S concat = F Concat (S avg , S max )
[0040] Convolve the features after spatio-temporal feature concatenation using a 3×3×3 three-dimensional convolutional kernel, followed by Sigmoid to generate attention weights between 0 and 1, expressed as:
[0041] M s = σ(Conv3D 3*3*3 (S concat ))
[0042] where Conv3D 3*3*3 represents a 3×3×3 three-dimensional convolutional kernel;
[0043] Perform an element-wise multiplication operation on each channel of the feature map after channel recalibration with the calculated attention weights, expressed as:
[0044]
[0045] Furthermore, the second convolutional branch for performing convolution operations on three-dimensional frequency-domain signals includes at least two cascaded three-dimensional convolutional units, successively including a third three-dimensional convolutional unit and a fourth three-dimensional convolutional unit;
[0046] In both the third three-dimensional convolutional unit and the fourth three-dimensional convolutional unit, convolution operations are performed using a three-dimensional convolutional kernel in the spatial dimension, frequency-domain dimension, and channel dimension;
[0047] In the fourth three-dimensional convolutional unit, the structure after convolution using a three-dimensional convolutional kernel is input into the second three-dimensional convolutional attention module, and the second three-dimensional convolutional attention module recalibrates the features in the spatial dimension, frequency-domain dimension, and channel dimension to obtain the features output by the second convolutional branch.
[0048] Furthermore, the third convolutional branch for performing convolution operations on one-dimensional time-domain signals and the fourth convolutional branch for performing convolution operations on one-dimensional frequency-domain signals both include: a WTConv1d layer, a batch normalization layer, a Conv1d layer, a batch normalization layer, an activation layer, a max pooling layer, a one-dimensional Dropout layer, and a one-dimensional mean layer.
[0049] Furthermore, the following steps are performed in the WTConv1d layer:
[0050] Assume that the input one-dimensional time / frequency domain signal is represented as X;
[0051] Perform multi-level wavelet decomposition according to the following formula:
[0052]
[0053] where ll (k) is the K-th level low-frequency component, h (k) is the K-th level high-frequency component, and DWT(·) is the discrete wavelet transform;
[0054] Generate an adaptive convolution kernel for each level of high-frequency component according to the following formula:
[0055]
[0056] where GlobalAttn is the global attention module, is the kernel parameter generation function, is the adaptive convolution kernel;
[0057] Perform dynamic convolution processing on the high-frequency component using the adaptive convolution kernel, expressed as:
[0058]
[0059] where * represents the one-dimensional convolution operation, α (k) , β (k) are the learnable scaling bias parameters;
[0060] Reconstruct the signal level by level through the inverse wavelet transform, expressed as:
[0061]
[0062] where IDWT(·) is the inverse discrete wavelet transform, and W base is the static depthwise separable convolution kernel.
[0063] Furthermore, the specific operations of the fusion decision layer include:
[0064] Perform global average pooling operations on the features output by the four convolutional branches;
[0065] Automatically adjust the importance of each feature using a learnable parameter matrix;
[0066] Output the epilepsy detection result by the fully connected layer.
[0067] Furthermore, when training the epilepsy EEG detection model, the bilinear cross-entropy loss function is used to optimize the decision boundary of the epilepsy EEG detection model.
[0068] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0069] (1) When neonatal epilepsy occurs, the EEG signal has the characteristics of low amplitude (20 - 50 μV) and short duration (0.5 - 2 seconds). Due to the fixed receptive field, the traditional convolutional neural network is difficult to capture transient discharge characteristics, and conventional time-frequency analysis methods are easily interfered by motion artifacts. There are problems in the prior art such as the lack of three-dimensional EEG signal spatial topology modeling and the asynchronous fusion of time-frequency domain features, resulting in a clinical false detection rate as high as 40%. The method of the present invention innovatively constructs a two-dimensional dynamic topology mapping mechanism, projects 18 bipolar leads onto a 5×5 electrode grid according to anatomical positions, compensates for the potential of vacant positions through the sliding window mean, and retains the spatial diffusion characteristics of brain regions;
[0070] (2) Aiming at the problem of decoupling time-frequency domain features, the method of the present invention innovatively constructs a multi-modal spatio-temporal feature fusion architecture and designs a four-way parallel processing architecture: the three-dimensional time domain path uses 3×3×3 deformable convolution to capture spatial-temporal correlation patterns; the three-dimensional frequency domain path analyzes the spectral energy distribution through 5×5×5 convolution; the one-dimensional time domain path combines three-level db1 wavelet decomposition and dynamic convolution to enhance transient high-frequency components; the one-dimensional frequency domain path uses an improved Welch algorithm to extract multi-scale power spectrum features; a CBAM3D attention mechanism is added at the end of the three-dimensional time-frequency convolution path to perform weighted output processing on the output feature values; the classification boundary is optimized through bilinear cross-entropy loss;
[0071] (3) The method of the present invention realizes dynamic alignment in the time-frequency domain through three-dimensional deformable convolution, and combines the improved CBAM3D attention mechanism to enhance feature discriminability, which can greatly improve the feature extraction and fusion capabilities of the epilepsy detection model for the spatial-time-frequency domain, enabling the model to have higher accuracy and better robustness;
[0072] (4) While maintaining the advantage of model lightweight (the number of parameters is only 10110), the method of the present invention has been verified by more than 10,000 clinical samples, achieving a high detection accuracy of 95.8%, and improving the epilepsy detection rate by 24.6% compared with the traditional CNN model, which is significantly better than the performance standard of the algorithm recommended by the International League Against Epilepsy;
[0073] (5) Clinical verification shows that the method of the present invention has a sensitivity of 94.7%, the false detection rate is reduced to 0.5 times / hour, the epilepsy detection rate is improved by 24.6% compared with the traditional CNN model, the specificity is 95.4%, and the AUC is 98.4. Description of the Drawings
[0074] Figure 1 This is the overall flowchart of a method for epilepsy detection of electroencephalogram (EEG) signals based on multi-modal spatio-temporal feature fusion and attention mechanism proposed by the present invention;
[0075] Figure 2 This is a schematic diagram of the channel attention part in the CBAM3D attention mechanism;
[0076] Figure 3 This is a schematic diagram of the spatial attention part in the CBAM3D attention mechanism;
[0077] Figure 4 This is a schematic diagram of the wavelet convolution module;
[0078] Figure 5 This is a schematic diagram of the overall model;
[0079] Figure 6 This is a schematic diagram of designing different convolution branches for one-dimensional time-domain signals, one-dimensional frequency-domain signals, three-dimensional time-domain signals, and three-dimensional frequency-domain signals. Specific implementation mode
[0080] Now, the technical solution of this embodiment will be further elaborated in combination with the accompanying drawings and embodiments.
[0081] Embodiment 1:
[0082] This embodiment discloses a method for epilepsy detection of electroencephalogram (EEG) signals based on multi-modal spatio-temporal feature fusion and attention mechanism. As Figure 1 shown, it mainly includes the following steps;
[0083] Step 1: Convert 19 monopolar lead signals into 18 bipolar differential signals to effectively eliminate common-mode interference. The specific operation includes: obtaining the original EEG signals from 19 electrode channels. There is a reference electrode channel among the 19 electrode channels. Subtract the EEG signals collected from the other 18 electrode channels from the EEG signal of the reference electrode channel, that is, perform differential operation, so as to obtain 18 processed bipolar differential signals, expressed as:
[0084]
[0085] In the formula, represents the original electrode channel signal data, represents the signal data of the electrode pair formed by subtracting the original electrode channel signal from the reference electrode potential signal V Cz The reference electrode potential signal, V Cz represents the reference electrode potential;
[0086] This operation can effectively suppress electrocardiogram artifacts and respiratory movement interference, and improve the signal-to-noise ratio of epileptiform discharge signals.
[0087] In this embodiment, a composite filtering strategy is adopted to eliminate noise in stages, including: filtering the processed EEG signals using a Butterworth band-pass filter (0.5 - 64 Hz) to eliminate baseline drift and high-frequency noise, and cooperating with a dual notch filter to suppress power frequency interference. In this embodiment, a 50 Hz / 60 Hz comb filter is used.
[0088] According to the relevant parameters of the EEG signals, slice the filtered EEG signals to obtain EEG signal slices. For example, slicing can be performed according to a sliding window with a sampling frequency of 128 Hz for 3 s; in this embodiment, the relevant parameters of the EEG signals include but are not limited to the sample length, the slice length within the sample, and the slice window shift. The length of the sliding window is the sampling time * sampling frequency.
[0089] The EEG signal is a time-domain signal. An improved Welch method is adopted, where the Hanning window length is 129 points to perform signal windowing, the number of FFT points is 256 to obtain the fine frequency domain structure, and the overlap rate is 50% to ensure the balance of time-frequency resolution. Calculate the power spectral density characteristics of the EEG signal slices to obtain the frequency domain signals of the EEG signal slices.
[0090] Establish a triple verification mechanism, that is, multiple experts label the tags for each slice window. For the same slice window, only when the tags labeled by no less than three experts are exactly the same, will the slice window and the corresponding tag be included in the training set.
[0091] Step 2: Establish a 5×5 two-dimensional grid (row numbers 0 - 4, column numbers 0 - 4). Using channel mapping, map the EEG signals to the position coordinates within the range of 5*5 according to the relative positions on the decoupling plane for two-dimensional plane mapping, and obtain a 5×5 electrode mapping matrix, expressed as:
[0092]
[0093] See Figure 2 For the two-dimensional plane mapping rule:
[0094] Frontal region (row 0): Fp1-REF electrode: mapped to grid coordinates (0,1); Fp2-REF electrode: mapped to grid coordinates (0,3);
[0095] Frontal region (row 1): Central frontal region: Fz-REF is located on the central axis (1,2); Left frontal region: F3-REF (1,1), F7-REF (1,0); Right frontal region: F4-REF (1,3), F8-REF (1,4);
[0096] Central region (row 2): Motor cortex region: C3-REF (2,1), C4-REF (2,3); Anterior temporal region: T3-REF (2,0), T4-REF (2,4);
[0097] Parieto-occipital region (rows 3 - 4); Parietal region of row 3: P3 - REF(3,1), Pz - REF(3,2), P4 - REF(3,3);
[0098] Posterior temporal region: T5 - REF(3,0), T6 - REF(3,4);
[0099] Occipital region of row 4: O1 - REF(4,1), O2 - REF(4,3);
[0100] Where REF is the electrode channel Cz.
[0101] In the 5×5 electrode mapping matrix, M represents the places not yet mapped, which are automatically identified by traversing the difference set of grid coordinates and the coordinates of the assigned electrodes. Adopting a dynamic filling strategy, the grid coordinates not yet mapped (a total of 7 vacancies) are dynamically filled, expressed as:
[0102]
[0103] In the formula, C is the number of channel - electrode pairs, which is 18 here, and X i (t) represents the channel signal data of the electrode pair.
[0104] This step can effectively maintain the coherence of the spatial potential distribution and prevent the sharp change of features caused by zero filling.
[0105] After the above operations, the time - domain and frequency - domain data of the EEG signal after two - dimensional plane mapping and the time - domain and frequency - domain data of the EEG signal before mapping are obtained. Based on this, a multi - modal training dataset is constructed. The time - domain and frequency - domain data of the EEG signal after mapping are abbreviated as: one - dimensional time - domain signal and one - dimensional frequency - domain signal, and the time - domain and frequency - domain data of the EEG signal before mapping are abbreviated as: three - dimensional time - domain signal and three - dimensional frequency - domain signal. The shape of the three - dimensional time - domain signal is: (batch, channel = 1, length = 5, width = 5, time = 328), where time = 328 corresponds to 3 - second time - domain sampling points (128Hz×3s≈384 points after PSD calculation and dimensionality reduction); the shape of the three - dimensional frequency - domain signal is (batch, channel = 1, length = 5, width = 5, freq = 129), which is obtained through STFT transformation.
[0106] Normalize the one - dimensional time - domain signal, one - dimensional frequency - domain signal, three - dimensional time - domain signal, and three - dimensional frequency - domain signal respectively to obtain the data input to the model.
[0107] Step 3: As Figure 6 shown, different convolutional branches are designed for the one - dimensional time - domain signal, one - dimensional frequency - domain signal, three - dimensional time - domain signal, and three - dimensional frequency - domain signal.
[0108] Among them, the convolutional branch for three-dimensional time-domain signals is as follows: The input layer of this convolutional branch receives a five-dimensional time-domain tensor with a shape of (batch, 1, 5, 5, 328). This convolutional branch contains at least two cascaded three-dimensional convolutional units (Convd3D and Convd3D with CBAM) to achieve hierarchical feature abstraction.
[0109] Each three-dimensional convolutional unit performs the following operations: Convolution operations are carried out in the spatial dimension (N×N), time domain dimension (T), and channel dimension (C) using three-dimensional convolutional kernels. The specific operations include: effectively capturing local time-frequency domain patterns using a 1×1×3 convolutional kernel. Effectively capturing local spatial patterns using a 3×3×1 convolutional kernel. Effectively capturing time-frequency-spatial features using a 3×3×3 convolutional kernel. The convolutional operations of these three different convolutional kernels are equivalent to cascading, and the output of the previous one serves as the input of the next one. Figure 6 Both Convd3D and CNN with CBAM 3D Block in [[ ]] are such cascading operations, and only CNN with CBAM 3D Block adds the attention module of CBAM at the end.
[0110] The three-dimensional convolutional attention module (CBAM3D) is a multi-dimensional extension based on the traditional channel-spatial attention mechanism to form an attention mechanism with spatio-temporal joint perception ability. Therefore, the three-dimensional convolutional attention module (CBAM3D) is used to achieve three-dimensional feature recalibration of channels, space, and time domain. Its working process includes two core stages: the channel attention stage and the spatio-temporal attention stage. The channel attention stage is used to effectively compress the global features across spatio-temporal dimensions. The spatio-temporal attention stage is used to achieve 5×5 electrode spatial weight mapping.
[0111] Now, in combination with [[ ]] Figure 3 further explanation is given to the channel attention stage.
[0112] The main purpose of the channel attention stage is: to learn the importance among channels, suppress noise-related channels, and enhance epilepsy feature channels. The specific steps include:
[0113] Let the elements of the input feature map be: x∈R B×C×H×W×T / F ; where B represents the batch size, C represents the number of channels, H×W represents the spatial dimension (electrode layout), and T / F represents the time / frequency domain dimension.
[0114] The first step is to perform dual-path compression on the input feature map:
[0115] The first compression path performs global average pooling: Compression is carried out along the spatial (H, W) and time (T) dimensions, expressed as:
[0116]
[0117] The second compression path performs a global maximum to achieve capturing significant activations, expressed as:
[0118]
[0119] F avg , F max ∈R B×C
[0120] In the second step, the channel relationship is learned through a multi-layer perceptron MLP with shared weights, that is, the double-path compression result is input into an MLP (dimensionality reduction rate r = 4) with two fully connected layers. The MLP structure is: FC(C→C / r)→ReLU→FC(C / r→C), where FC(C→C / r) represents a fully connected layer, the input data shape is (B, C), C is the original input channel number, and the fully connected layer reduces the dimension to C / r channels. FC(C / r→C) represents a fully connected layer, the input data shape is (Batch, C / r), C / r is the input channel number after dimensionality reduction, and here it is upsampled to C channels again.
[0121] The output of the MLP is expressed as:
[0122] M c = σ(MLP(F avg ) + MLP(F max ))
[0123] In the formula, σ represents the Sigmoid activation function, and M c represents the weight coefficient of each channel obtained after calculation, and M c ∈R B×C×1×1×1 .
[0124] In the third step, the calculated weight coefficient is multiplied element-wise with each channel of the input feature map to achieve channel recalibration, expressed as:
[0125]
[0126] In the formula, represents the per-channel product.
[0127] Calculation steps:
[0128]
[0129] Among them, W0 ∈ R C / r×C , W1 ∈ R C×C / r are the MLP weights, and δ represents the Sigmoid activation function.
[0130] Such asFigure 4 As shown in the figure, the spatio-temporal attention stage will be further described below.
[0131] The goal of this spatio-temporal attention stage is to focus on key spatial regions (such as epileptic foci) and key temporal segments (such as seizure onset points). It includes the following steps:
[0132] The first step is to aggregate features across channel dimensions, including:
[0133] Average pooling along the channel dimension, denoted as:
[0134]
[0135] Max pooling along the channel dimension, denoted as:
[0136]
[0137] S avg ,S max ∈R B×1×H×W×T .
[0138] Spatio-temporal feature concatenation, denoted as:
[0139] S concat =F Concat (S avg ,S max )
[0140] S concat ∈R B×2×H×W×T .
[0141] The second step is to use 3D convolution to capture local spatio-temporal patterns, including:
[0142] Use a 3×3×3 convolutional kernel (spatial 3x3, temporal depth 3), concatenate dual-path features, and apply a Sigmoid function after convolution to generate attention weights between 0 and 1, denoted as:
[0143] M s =σ(Conv3D 3*3*3 (S concat ))
[0144] In the formula, Conv3D 3*3*3 represents a 3D convolutional kernel (3×3×3), and M s ∈R B×1×H×W×T .
[0145] The third step is to perform an element-wise multiplication operation between the calculated attention weights and each channel of the input feature map to achieve spatio-temporal feature enhancement, denoted as:
[0146]
[0147] In this embodiment, a 3×3×3 convolutional kernel (spatial 3x3, temporal depth 3) is adopted, where the spatial 3x3 can cover most of the electrode-dense regions (5×5 mapping) of the neonatal head; it can better capture the spatial features between electrodes; the temporal depth 3 can capture the short-time diffusion pattern of epileptic waves (0.5 - 2 seconds).
[0148] The calculation of the overall three-dimensional convolutional attention module (CBAM3D) is expressed as:
[0149]
[0150] Among them, the convolutional branch for three-dimensional frequency-domain signals is: a series of three-dimensional convolutional modules are used to process the 5×5×F frequency-domain feature matrix to extract spatial-spectral joint features; the type of this convolutional branch is the same as that of the convolutional branch for three-dimensional time-domain signals, and each three-dimensional convolutional module includes: a three-dimensional convolutional kernel and a three-dimensional convolutional attention module (CBAM3D).
[0151] Receive three-dimensional frequency-domain signals and perform convolutional operations using a three-dimensional convolutional kernel in the spatial dimension (N×N), frequency-domain dimension (F), and channel dimension (C).
[0152] In the three-dimensional convolutional attention module (CBAM3D), channel attention calculates the global statistics across the spatial-frequency domain; spatial-frequency domain attention captures the correlation between the electrode spatial distribution and band energy through three-dimensional convolution.
[0153] Among them, the convolutional branch for one-dimensional time-domain signals and the convolutional branch for one-dimensional frequency-domain signals both include: a WTConv1d layer, a batch normalization layer, a Conv1d layer, a batch normalization layer, an activation layer, a max pooling layer, a one-dimensional Dropout layer, and a one-dimensional mean layer.
[0154] Figure 5 The structure of the WTConv1d layer is shown, including the basic path ( Figure 5 the topmost path in ), the WTleval-1 path, and the WTleval-2 path. In this embodiment, only the convolutional operation and residual connection are performed on this basic path. The WTleval-1 path indicates that the number of layers for wavelet decomposition of the input signal is 1. First, the input signal is decomposed into low-frequency and high-frequency components; then, convolutional operations are respectively performed on these two parts, and then wavelet inverse decomposition is performed, and then combined into a signal with the same length as the original signal, and the values of this part are residually connected to the basic path. The WTleval-2 path indicates that the low-frequency and high-frequency parts obtained by WTleval-1 are decomposed again into low-frequency, sub-low-frequency, high-frequency, and sub-high-frequency parts, and then convolutional operations are performed, and then inverse decomposition is performed to obtain the processed low-frequency and high-frequency parts and then residually connected to the low-frequency and high-frequency parts of WTleval-1.
[0155] Specifically, the WTConv1d layer in the convolutional branch for one-dimensional time-domain signals includes: a multi-level wavelet decomposition layer, a high-frequency component non-linear processing layer, and a multi-level transposed convolution reconstruction layer. Among them, the multi-level wavelet decomposition layer is used to gradually separate the low-frequency components and high-frequency components to reveal the multi-scale characteristics of the signal. The high-frequency component non-linear processing layer is used to extract multi-resolution time-domain features using 1D convolution to enhance the feature recognition ability of the signal. The multi-level transposed convolution reconstruction layer is used to restore the signal resolution through transposed convolution to achieve the precise reconstruction of the signal. The above processing flow reflects the idea of multi-resolution analysis: Wavelet decomposition process: Use the db1 wavelet basis for three-level decomposition to effectively separate the low-frequency baseline signal and high-frequency transient components. High-frequency processing: Use deformable 1D convolution kernels to extract electrical sample discharge characteristics. Cross-scale reconstruction: Restore the signal resolution through transposed convolution and retain multi-level feature maps.
[0156] The WTConv1d layer in the convolutional branch for one-dimensional frequency-domain signals includes: a frequency band multi-scale decomposition layer, a frequency-domain feature enhancement layer, and a cross-frequency band fusion reconstruction layer. The frequency band multi-scale decomposition layer is used to gradually separate different frequency band components to analyze the frequency-domain characteristics of the signal. The frequency-domain feature enhancement layer is used to analyze multi-scale frequency band features using 1D convolution to improve the recognizability of the features. The cross-frequency band fusion reconstruction layer is used to reconstruct the frequency-domain signal through transposed convolution to ensure the integrity of the signal.
[0157] Now, the operations involved in the WTConv1d layer are described as follows:
[0158] Assume that the input one-dimensional time-domain / frequency-domain signal is represented as: x ∈ R B×C×T / F .
[0159] Perform multi-level wavelet decomposition according to the following formula:
[0160]
[0161] In the formula, is the K-level low-frequency component, is the K-level high-frequency component, and DWT(·) is the discrete wavelet transform.
[0162] Generate an adaptive convolution kernel for each level of high-frequency component according to the following formula:
[0163]
[0164] In the formula, GlobalAttn is the global attention module, is the kernel parameter generation function (implemented by MLP), is the dynamically generated convolution kernel (S is the kernel size).
[0165] Performing dynamic convolution processing on high-frequency components using an adaptive convolution kernel to achieve high-frequency feature enhancement, which is expressed as:
[0166]
[0167] In the formula, * represents a one-dimensional convolution operation, and α (k) , β (k) ∈R C are learnable scaling bias parameters.
[0168] Reconstructing the signal step by step through inverse wavelet transform, which is expressed as:
[0169]
[0170] In the formula, IDWT(·) is the inverse discrete wavelet transform, and W base is a static depthwise separable convolution kernel.
[0171] Step 4: Through a collaborative fusion mechanism, fuse the features of the four groups of signals to achieve information complementarity; that is, perform global average pooling on the features of the four groups of signals, flatten them and then input them into the fully connected layer. Among them, perform global average pooling on each output to generate a 32-dimensional feature vector; use a learnable parameter matrix to automatically adjust the importance of each feature to achieve self-adaptation; in the fully connected stage, adopt a bilinear cross-entropy loss function to optimize the decision boundary of the model. Finally, use the fused features for epilepsy detection and output the classification result.
[0172] In this embodiment, a joint spatio-temporal-frequency domain representation system is established, and a multi-scale epilepsy feature fusion is achieved by using a four-channel heterogeneous feature extraction network. The entire processing flow strictly follows the technical route of "spatial projection - feature decoupling - collaborative decision-making", and data interaction is realized between modules through a feature tensor pipeline, forming a closed-loop processing link.
[0173] After being strictly verified by 10,000 cases of clinical data, the method of the present invention shows the following excellent performance indicators: the accuracy rate reaches 95.8%, the sensitivity reaches 94.7%, the specificity reaches 95.4%, and the AUC reaches 98.4; the method of the present invention has been verified by clinical trials and has significant progressiveness and industrial application value.
Claims
1. A method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion, characterized by: The following steps are involved: Step 1: Obtain EEG signals through electrode pairs; Step 2: preprocessing the acquired EEG signal to obtain a preprocessed EEG signal; Step 3: Input the preprocessed EEG signal into the trained epilepsy EEG detection model to obtain the detection result; The step of preprocessing the acquired EEG signal to obtain the preprocessed EEG signal specifically includes: The acquired EEG signals are filtered, sliced, and converted into frequency domain signals in sequence to obtain three-dimensional time domain signals and three-dimensional frequency domain signals; Performing two-dimensional plane mapping on the three-dimensional time domain signal and the three-dimensional frequency domain signal to obtain a mapped one-dimensional time domain signal and a one-dimensional frequency domain signal; After normalizing the three-dimensional time domain signal, the three-dimensional frequency domain signal, the one-dimensional time domain signal and the one-dimensional frequency domain signal, a preprocessed EEG signal is obtained; The epilepsy EEG detection model includes: a first convolution branch for performing convolution operations on three-dimensional time domain signals, a second convolution branch for performing convolution operations on three-dimensional frequency domain signals, a third convolution branch for performing convolution operations on one-dimensional time domain signals, a fourth convolution branch for performing convolution operations on one-dimensional frequency domain signals, and a fusion decision layer; The fusion decision layer is used to fuse the features output by the four convolution branches, perform epilepsy detection based on the fused features, and output epilepsy detection results; The first convolution branch for performing a convolution operation on a three-dimensional time domain signal comprises at least two cascaded three-dimensional convolution units, which sequentially include a first three-dimensional convolution unit and a second three-dimensional convolution unit; In both the first three-dimensional convolution unit and the second three-dimensional convolution unit, a three-dimensional convolution kernel is used to perform convolution operations in the spatial dimension, the time domain dimension, and the channel dimension; In the second three-dimensional convolution unit, the result of the convolution operation using the three-dimensional convolution kernel is input into the first three-dimensional convolution attention module, and the first three-dimensional convolution attention module recalibrates the spatial dimension, time domain dimension and channel dimension features to obtain the features output by the first convolution branch; The first three-dimensional convolutional attention module includes a channel attention submodule and a spatiotemporal attention submodule; The following calculations are performed in the channel attention submodule: Assume that the element in the feature map of the input 3D convolutional attention module is represented as: x∈R B×C×H×W×T / F ; Where B represents the batch size, C represents the number of channels, H×W represents the spatial dimension, and T / F represents the time / frequency domain dimension; Perform global average pooling on the input feature map, expressed as: Where X[·] represents the input feature map, c represents the channel dimension of the input feature map, h,w represents the spatial dimension of the input feature map, and t represents the temporal dimension of the input feature map; Perform global maximum on the input feature map, expressed as: The weight coefficient M is calculated by the multi-layer perceptron with shared weights c , expressed as: In the formula, σ represents the Sigmoid activation function; The calculated weight coefficient M c Perform element-by-element multiplication with each channel of the input feature map to achieve channel recalibration, which is expressed as: In the formula, represents channel-wise product; The following computations are performed in the spatiotemporal attention submodule: The channel-recalibrated feature map is averagely pooled along the channel dimension, expressed as: The feature map after channel recalibration is pooled along the channel dimension, which is expressed as: S avg and S max Perform spatiotemporal feature concatenation, expressed as: S concat =F Concat (S avg ,S max ) The concatenated spatiotemporal features are convolved using a 3×3×3 three-dimensional convolution kernel, followed by a Sigmoid to generate a 0-1 attention weight, expressed as: M s =σ(Conv3D 3*3*3 (S concat )) In the formula, Conv3D 3*3*3 Represents a 3×3×3 three-dimensional convolution kernel; The calculated attention weight is multiplied element-by-element with each channel of the feature map after the channel is recalibrated, expressed as:
2. The method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion according to claim 1, characterized in that: The two-dimensional plane mapping of the three-dimensional time domain signal and the three-dimensional frequency domain signal is specifically: two-dimensional plane mapping of the three-dimensional time domain signal and the three-dimensional frequency domain signal according to the relative positions of the planes where the electrode pairs are located.
3. The method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion according to claim 1, characterized in that: The second convolution branch for performing a convolution operation on a three-dimensional frequency domain signal comprises at least two cascaded three-dimensional convolution units, which sequentially include a third three-dimensional convolution unit and a fourth three-dimensional convolution unit; In the third three-dimensional convolution unit and the fourth three-dimensional convolution unit, a three-dimensional convolution kernel is used to perform convolution operations in the spatial dimension, the frequency domain dimension, and the channel dimension; In the fourth three-dimensional convolution unit, the structure after the convolution operation using the three-dimensional convolution kernel is input into the second three-dimensional convolution attention module, and the second three-dimensional convolution attention module recalibrates the spatial dimension, frequency domain dimension and channel dimension features to obtain the features output by the second convolution branch.
4. The method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion according to claim 1, characterized in that: The third convolution branch for performing convolution operations on one-dimensional time domain signals and the fourth convolution branch for performing convolution operations on one-dimensional frequency domain signals both include: a WTConv1d layer, a batch normalization layer, a Conv1d layer, a batch normalization layer, an activation layer, a maximum pooling layer, a one-dimensional Dropout layer and a one-dimensional mean layer.
5. The method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion according to claim 4, characterized in that: The following steps are performed in the WTConv1d layer: Assume that the input one-dimensional time domain / frequency domain signal is represented as X; Perform multi-level wavelet decomposition according to the following formula: In the formula, ll (k) is the Kth low-frequency component, h (k) is the K-th level high frequency component, DWT(·) is discrete wavelet transform; Generate an adaptive convolution kernel for each level of high-frequency components according to the following formula: Where GlobalAttn is the global attention module, is the kernel parameter generation function, is the adaptive convolution kernel; Adaptive convolution kernel is used to perform dynamic convolution processing for high-frequency components, which can be expressed as: In the formula, * represents a one-dimensional convolution operation, α (k) ,β (k) is a learnable scaling bias parameter; The signal is reconstructed step by step through inverse wavelet transform, which is expressed as: Where IDWT(·) is the inverse discrete wavelet transform, W base is a static depth-wise separable convolution kernel.
6. The method for detecting neonatal epilepsy EEG based on multimodal spatiotemporal feature fusion according to claim 1, characterized in that: The fusion decision layer specifically operates as follows: Perform global average pooling on the features output by the four convolution branches; Automatically adjust the importance of each feature using a learnable parameter matrix; The epilepsy detection result is output by the fully connected layer.
7. The method for detecting neonatal epilepsy based on multimodal spatiotemporal feature fusion according to claim 1, characterized in that: When training the epilepsy EEG detection model, the bilinear cross entropy loss function is used to optimize the decision boundary of the epilepsy EEG detection model.
Citation Information
Patent Citations
Neonatal convulsion electroencephalogram signal classification system of time-frequency-space domain CNN-LSTM introducing attention mechanism
CN115211871A
Interseizure epilepsy sample discharge detection method and device based on double-view feature fusion framework
CN116172517A
Epilepsy prediction method and system in combination with spatial-temporal characteristics
CN118072947A
Epilepsy electroencephalogram signal classification method based on time-space-frequency attention
CN118319247A
Automated seizure detection
WO2024007036A2