Multi-scale epileptic seizure cross-object detection method based on power spectral density
Through a multi-scale epilepsy seizure detection method based on power spectral density, the MSSTDCN model is used to extract the spatiotemporal characteristics of EEG signals, solving the time-consuming and subjective impact problems of traditional detection, and improving the accuracy and efficiency of epilepsy detection.
Patent Information
- Application Number
- CN202510342363.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-21
AI Technical Summary
传统癫痫发作检测依赖人工分析脑电信号耗时且易受主观因素影响,难以在个体间实现准确检测,个体生理差异和噪声影响导致自动化检测挑战。
The multi-scale epilepsy seizure detection method based on power spectral density is used to extract spatiotemporal features using adaptive multi-band power spectral density and multi-scale spatiotemporal depth convolution network (MSSTDCN) to construct an MSSTDCN model for epilepsy detection.
It improves the accuracy and efficiency of epilepsy detection, reduces the time cost of medical staff, enhances the generalization ability of the model, is suitable for different patients, and reduces the calculation cost.
Smart Images

Figure CN120296490A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of biomedical signal processing and artificial intelligence, and particularly relates to a multi-scale cross-object seizure detection method based on power spectral density. Background Art
[0002] Epilepsy is a temporary brain dysfunction caused by sudden abnormal discharges of brain neurons. Patients may experience transient symptoms such as loss of consciousness or perception, and motor function disorders during seizures. Currently, there are approximately 50 million epilepsy patients worldwide, regardless of age. The onset of epilepsy is sudden and recurrent, bringing great physical and mental distress to patients and their families. Therefore, accurate detection of epileptic seizures is crucial for the diagnosis, treatment, and management of patients.
[0003] Traditional epileptic seizure detection mainly relies on the manual analysis of electroencephalogram (EEG) signals by neurologists. Electroencephalogram can capture the complex dynamic electrical activities of the brain, identify epileptic types, assist in the diagnosis of epileptic syndromes, and evaluate the risk of epileptic recurrence. However, this method is both time-consuming and easily affected by subjective factors, making it difficult to detect epileptic seizures in a timely and accurate manner. Therefore, automated processing of EEG signals for epileptic seizure detection is particularly important.
[0004] In recent years, with the development of deep learning technology, more and more studies have attempted to apply artificial intelligence to the prediction of epileptic seizures, mainly by analyzing electroencephalogram (EEG) signals to discover potential abnormal patterns. However, due to the huge physiological differences among individuals, EEG signals have highly personalized characteristics in the time domain, frequency domain, and spatial domain, and are easily affected by noise. These factors pose challenges to the application of automatic epilepsy detection technology in unknown patients. Summary of the Invention
[0005] In order to solve the above problems, the object of the present invention is to provide a multi-scale cross-object seizure detection method based on power spectral density, which uses adaptive multi-band power spectral density to significantly improve the accuracy of epilepsy detection; extracts spatio-temporal features of different scales through a multi-scale spatio-temporal deep convolutional network, thereby improving the flexibility of epilepsy detection; reduces the time cost of medical staff, and largely protects the safety of patients.
[0006] To achieve the above object of the invention, the present invention adopts the following technical solutions:
[0007] A multi-scale cross-object seizure detection method based on power spectral density, comprising the following steps:
[0008] S1. Obtain multi-channel EEG signal data of epilepsy patients and perform preprocessing;
[0009] S2. Regularize the EEG signal data preprocessed in step S1;
[0010] S3. Construct the adaptive power spectral density, that is, reconstruct the power spectral density features;
[0011] S4. Balance the EEG signal data using the SMOTE oversampling method according to the power spectral density features obtained in step S3 to form a balanced dataset;
[0012] S5. Build and train the MSSTDCN model
[0013] The MSSTDCN model is a multi-scale spatio-temporal deep convolutional network model, which consists of a time feature extraction network, a spatial feature extraction network, a first multi-branch attention layer, a second multi-branch attention layer, a feature adaptive fusion module, a deep convolutional network, and multiple fully connected layers; The balanced adaptive PSD features obtained in step S4 are first used to extract time and spatial features respectively through a parallel dual-branch module composed of a time feature extraction network and a spatial feature extraction network. The extracted time features are fused according to the corresponding frequency bands and then input into the first multi-branch attention layer, and the spatial features are fused according to the corresponding frequency bands and then input into the second multi-branch attention layer; The first multi-branch attention layer and the second multi-branch attention layer respectively perform feature mapping on the corresponding input features to obtain corresponding refined features; The feature adaptive fusion module fuses the output results of the first multi-branch attention layer and the second multi-branch attention layer, and the obtained fusion features are input into the deep convolutional network; The deep convolutional network extracts high-dimensional features; All features are flattened and classified through multiple fully connected layers to obtain the epilepsy detection result;
[0014] S6. Divide the data obtained in step S4 into a training set and a validation set, that is, shuffle the reconstructed PSD data of X - 1 objects as the training set, and shuffle the reconstructed PSD data of 1 object as the validation set, where X represents the total number of objects; Input the training data and the validation data into the MSSTDCN model in step S5 for training.
[0015] Further, in step S1 above, the operation is as follows: First, collect multi-channel EEG signal data of epilepsy patients at a set sampling rate, and filter the EEG signal data using a 0.5 - 50 Hz band-pass filter; Then, divide the EEG signal data into multiple segments of a set time, and divide the long-time signal into multiple small windows.
[0016] Further, in step S2 above, when using the Z-score normalization method to process the EEG signal data, the operation is as follows:
[0017]
[0018] Among them, X(t) represents the original signal; μ and σ represent the mean and standard deviation of the original signal X(t) respectively; X'(t) represents the normalized signal.
[0019] Furthermore, step S3 above includes the following sub-steps:
[0020] S3.1. Perform frequency division processing on the EEG signal processed in step S2, and divide it into a σ frequency band with a frequency of 0.5 - 4 Hz, a θ frequency band with a frequency of 4 - 8 Hz, an α frequency band with a frequency of 8 - 13 Hz, a β frequency band with a frequency of 13 - 30 Hz, and a γ frequency band with a frequency of 30 - 50 Hz;
[0021] S3.2. Use the Multitaper method to calculate the PSD values of the corresponding frequency bands, and obtain the PSD values of the σ frequency band, θ frequency band, α frequency band, β frequency band, and γ frequency band respectively;
[0022] S3.3. Average the obtained PSD values in the channel dimension;
[0023] S3.4. Concatenate the physiological characteristics of the obtained 5 frequency bands in sequence to obtain the adaptive PSD characteristics.
[0024] Even further, in step S3 above, when using the Multitaper method to calculate the PSD values of the corresponding frequency bands, the steps are as follows:
[0025] S3.11. Select orthogonal windows
[0026] Select a set of K mutually orthogonal data windows where n = 0, 1,..., M - 1, and N represents the time series length;
[0027] S3.12. Single PSD estimate
[0028] For each taper, multiply the original time series by this taper to obtain a new time series, and then perform Fourier transform on the new time series to obtain the corresponding PSD estimate value; assuming there is a time series x(n), using the kth taper V k (n), the PSD estimate is:
[0029]
[0030] where f represents frequency; j represents the imaginary unit; e -i2πfn represents the complex exponential function;
[0031] S3.13. Averaging processing
[0032] Add up the PSD estimate values generated by all K tapers to obtain the average PSD:
[0033]
[0034] Further, in the above step S5, in the MSSTDCN model, the spatial feature extraction network consists of five wavelet convolutional layers. First, the features of the σ band, θ band, α band, β band, and γ band are extracted through the five wavelet convolutional layers respectively. Second, the features of each band are concatenated. Finally, the fused features are input into the second multi-branch attention layer, and the features F extracted by the spatial convolutional network are obtained through the second multi-branch attention layer s ; specifically, for a given input EEG representation x at time sample t, the first wavelet convolutional layer performs a discrete wavelet transform, defined as:
[0035]
[0036] where u(r) and v(r) represent the approximation filter and the detail filter respectively; x A (t) and x D (t) represent the approximation coefficient and the detail coefficient respectively; r represents the filter index, r = 0, 1, 2,..., R; s represents the scaling factor; k represents the sequence number of the wavelet convolutional layer;
[0037] The approximation coefficient x A (t) is fed back layer by layer to the next wavelet convolutional layer, and the spectral analysis is iteratively performed continuously through five wavelet convolutional layers.
[0038] Further, in the above step S5, in the MSSTDCN model, the time feature extraction network consists of multiple standard two-dimensional convolutional layers with different kernel sizes, which respectively extract the time features of different scales in the signal. Finally, the fused features are input into the first multi-branch attention layer, and the features F extracted by the time convolutional network are obtained through the first multi-branch attention layer t .
[0039] Further, in the above step S5, each multi-branch attention layer in the MSSTDCN model includes three attention branches, for the input tensor X in ∈R C×H×W , where C represents the number of channels, that is, the number of input feature maps, and H and W represent the height and width of the input feature maps respectively; the operation is as follows:
[0040] S5.1. The multi-branch attention layer performs group convolution on the input time-frequency spectrum feature X in ∈R C×H×W .
[0041] S5.2. Pass the feature map to three attention branches respectively. The first attention branch is used to capture the interaction between the (H, W) dimensions and is performed on X HW ∈R C×H×W ; the second attention branch is used to permute the input tensor into the second input tensor X CW ∈R H×C×W , capture the interaction between the (C, W) dimensions, and transform the output tensor at the end into the original input shape; the third attention branch is used to permute the input tensor into the third input tensor X HC ∈R W×H×C , capture the interaction between the (H, C) dimensions, and transform the output tensor at the end into the original input shape;
[0042] S5.3. Aggregate the outputs of the three attention branches through average pooling to generate refined features.
[0043] Furthermore, the operations performed by the above-mentioned first attention branch include the following steps:
[0044] S5.21. Apply channel max pooling and average pooling to reduce the channel dimension, and then concatenate the pooled features, defined as follows:
[0045]
[0046] where X pool ∈R 2×H×W represents the pooled feature; Max Pool(·) and AvgPool(·) represent max pooling and average pooling by channel respectively; is concatenation by channel; X HW represents the first input tensor;
[0047] S5.22. The pooled feature X pool is fed into a standard convolution with a sigmoid activation layer, and this activation layer provides intermediate attention weights X w ∈R 1×H×W ;
[0048] S5.23. Apply the generated attention weights X w ∈R 1×H×W to the first input tensor X HW through element-wise multiplication.
[0049] Furthermore, in the above step S5, in the MSSTDCN model, the refined feature X in ∈R C×H×W is obtained from the input tensor X ref , and the operations of each multi-branch attention layer are represented as:
[0050]
[0051] Among them, the first input tensor X HW is equal to the input tensor X in ; the second input tensor X CW and the third input tensor X HC respectively represent the permutation results from the input tensor X in ; represents the attention weight generated from the i-th attention branch; ⊙ is the broadcast element multiplication; —— represents the permutation operation.
[0052] Furthermore, in the above step S5, in the MSSTDCN model, the feature adaptive fusion module fuses the features F t extracted by the temporal convolutional network output from the first multi-branch attention layer, s the features F t extracted by the spatial convolutional network of the second multi-branch attention layer; specifically, for the feature F t extracted by the temporal convolutional network, the adaptive temporal feature Z t is expressed by the formula:
[0053] Z t = λF t + (1 - λ)∈
[0054] Among them, ∈ represents random noise; λ represents the ratio of the control signal to the noise;
[0055] for the feature F s extracted by the spatial convolutional network, the adaptive spatial feature Z s is expressed by the formula:
[0056] Z s = λF s + (1 - λ)∈
[0057] Fuse the obtained adaptive features, where the feature fusion strategy is expressed as:
[0058] F fusion = w t Z t + w s Z s
[0059] Among them, w t and w s respectively represent the weights of the temporal feature and the spatial feature; the two are dynamically adjusted according to the model classification loss.
[0060] Further, in the above step S5, in the MSSTDCN model, the deep convolutional network includes multiple deep convolutional modules, and each deep convolutional module is composed of a convolutional layer, a batch normalization layer, and a max pooling layer connected in sequence, and is used to extract the fused features output by the feature adaptive fusion module.
[0061] Due to the adoption of the above technical solution, the present invention has the following advantages:
[0062] The multi-scale epilepsy seizure cross-object detection method based on power spectral density of the present invention captures EEG signals with non-stationary and time-varying statistical characteristics using power spectral density (PSD), and then processes the non-stationary time series data using a deep convolutional network, which can not only improve the accuracy and efficiency of epilepsy detection, but also solve the problems caused by individual differences to a certain extent, making the model more generalizable; the MSSTDCN model extracts temporal and spatial features in EEG signals through a dual-branch network, making up for the problem of information loss caused by single features; the MSSTDCN model can solve the problem of low performance in cross-object epilepsy detection, and improve the epilepsy detection performance of the model by reconstructing PSD features in different frequency bands; it can reduce the dependence on a large amount of labeled data, reduce the complexity and computational cost of the model, and improve the applicability of the model among different patients, providing a new and effective technical path for the automation of epilepsy seizure detection. Description of the Drawings
[0063] Figure 1 is a flowchart of the multi-scale epilepsy seizure cross-object detection method based on power spectral density of the present invention;
[0064] Figure 2 is Figure 1 a flowchart of the adaptive power spectral density construction step in
[0065] Figure 3 is a schematic structural diagram of an embodiment of the MSSTDCN model. Detailed Embodiments
[0066] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can fully understand and implement the present invention.
[0067] The multi-channel EEG signals of epilepsy patients in the embodiment include 940h of long-term continuous multi-channel scalp EEG records from 23 epilepsy subjects aged 1.5 to 22 years; according to the international 10 / 20 standard, the EEG signals are sampled at a frequency of 256HZ and include at least 19 EEG channels. The above records include 198 epilepsy seizures, and the onset and end are accurately annotated by clinicians with neuroscience expertise.
[0068] As Figure 1 、 2 shown in Fig. 3, a multi-scale epilepsy seizure cross-object detection method based on power spectral density includes the following steps:
[0069] S1. Obtain multi-channel electroencephalogram (EEG) signal data of epilepsy patients and perform preprocessing;
[0070] First, collect multi-channel EEG signals of epilepsy patients at a set sampling rate, filter the EEG signals using a 0.5 - 50 Hz band-pass filter to remove low-frequency and high-frequency noises as well as electrocardiogram signals, and retain the effective information related to epileptic seizures; then, divide the EEG signals into multiple segments according to a 4s time window, and divide the long-time signals into multiple small windows for easy feature extraction;
[0071] S2. Perform regularization processing on the EEG signals preprocessed in step S1 to ensure signal consistency, reduce inter-individual differences, and help the deep learning model converge more efficiently; use the Z-score normalization method to process the EEG signals, which can effectively cope with abnormal fluctuations in the data and provide a stable and reliable signal representation; the operation is as follows:
[0072]
[0073] Among them, X(t) represents the original signal; μ and σ represent the mean and standard deviation of the original signal X(t) respectively; X'(t) represents the normalized signal;
[0074] S3. Adaptive power spectral density construction, that is, reconstruct power spectral density features; the operation is as follows:
[0075] S3.1. Divide the EEG signals processed in step S2 into frequency bands, and divide them into a σ frequency band with a frequency of 0.5 - 4 Hz, a θ frequency band with a frequency of 4 - 8 Hz, an α frequency band with a frequency of 8 - 13 Hz, a β frequency band with a frequency of 13 - 30 Hz, and a γ frequency band with a frequency of 30 - 50 Hz;
[0076] S3.2. Use the Multitaper method to calculate the PSD values of the corresponding frequency bands, and obtain the PSD values of the σ frequency band, θ frequency band, α frequency band, β frequency band, and γ frequency band respectively; the steps are as follows:
[0077] S3.11. Select orthogonal windows
[0078] Select a set of K mutually orthogonal data windows where n = 0, 1,..., N - 1, and N represents the time series length;
[0079] S3.12. Single PSD estimate value
[0080] For each taper, multiply the original time series by this taper to obtain a new time series, and then perform a Fourier transform on the new time series to obtain the corresponding PSD estimate value; assume there is a time series x(n), and use the k-th taper V k (n) The PSD estimate is:
[0081]
[0082] where f represents frequency; j represents the imaginary unit; e -j2πfn represents the complex exponential function, which is used to convert the time-domain signal to the frequency domain;
[0083] S3.13. Averaging processing
[0084] Add the PSD estimate values generated by all K tapers to obtain the average PSD:
[0085]
[0086] S3.3. Average the obtained PSD values in the channel dimension;
[0087] S3.4. Concatenate the physiological features of the obtained 5 frequency bands in sequence to obtain the adaptive PSD feature;
[0088] S4. According to the power spectral density features obtained in step S3, use the SMOTE (Synthetic Minority Over-sampling Technique) oversampling method to balance the EEG signal data, forming a balanced data set, which can enhance the training of seizure data and alleviate the problem of class imbalance; the operation is as follows:
[0089] SMOTE selects samples from the minority class, calculates its M nearest neighbors, and then randomly selects a neighbor B to generate synthetic samples for the minority class; specifically, first select a sample A from the minority class, and then use the following formula to generate a new synthetic sample C:
[0090] C = A + r × (B - A)
[0091] where r is a random number between 0 and 1, representing the position along the AB direction; repeat this process until the desired class balance ratio is reached;
[0092] S5. MSSTDCN model construction and training
[0093] The MSSTDCN model is a multi-scale spatio-temporal depth convolutional network model, which consists of a time feature extraction network, a spatial feature extraction network, a first multi-branch attention layer, a second multi-branch attention layer, a feature adaptive fusion module, a depth convolutional network, and a multi-layer fully connected layer; the operation is as follows:
[0094] S5-1. The balanced adaptive PSD features obtained in step S4 are respectively used to extract time and spatial features through a parallel dual-branch module composed of a time feature extraction network and a spatial feature extraction network. Specifically, the spatial feature extraction network consists of five wavelet convolutional layers. First, the features in the σ frequency band, θ frequency band, α frequency band, β frequency band, and γ frequency band are respectively extracted through the five wavelet convolutional layers. Secondly, the features of each frequency band are connected. The fused features are input into the second multi-branch attention layer. Specifically, for a given input EEG representation x at time sample t, the first wavelet convolutional layer performs a discrete wavelet transform, which is defined as:
[0095]
[0096]
[0097] where u(r) and v(r) respectively represent the approximation filter and the detail filter; x A (t) and x D (t) respectively represent the approximation coefficient and the detail coefficient; r represents the filter index r = 0, 1, 2,..., R. Preferably, R = 5 is set; s represents the scaling factor, which is used to control the scale of the wavelet transform. Preferably, it is set to 0.6; k represents the sequence number of the wavelet convolutional layer.
[0098] The approximation coefficient x A (t) is fed back layer by layer to the next wavelet convolutional layer, and the spectral analysis is iteratively performed through five wavelet convolutional layers continuously.
[0099] The time convolutional network consists of multiple standard two-dimensional convolutional layers with different kernel sizes. Preferably, the specific kernel sizes are set to 1, 3, 5, 7, 9, and the stride is set to 2, respectively extracting time features of different scales in the signal. Finally, the fused features are input into the first multi-branch attention layer.
[0100] S5-2. In the first multi-branch attention layer and the second multi-branch attention layer, each multi-branch attention layer includes three attention branches, which can model channel attention and spatial attention at low cost and effectively without involving dimensionality reduction. For the input tensor X in ∈R C×H×W , where C represents the channel, that is, the number of input feature maps, and H and W respectively represent the height and width of the input feature map; the operation is as follows:
[0101] S5.1. Considering the heterogeneity among multi-domain feature mappings, the multi-branch attention layer performs group convolution on the input time-frequency feature X in ∈R C×H×W to reduce the computational cost and eliminate the aliasing effect while setting the parameter group to 2;
[0102] S5.2. The feature maps are respectively passed to three attention branches. The first attention branch is used to capture the interaction between (H, W) dimensions and is performed on X HW ∈R C×H×W ; The second attention branch is used to permute the input tensor into the second input tensor X CW ∈R H×C×W , capture the interaction between (C, W) dimensions, and transform the output tensor at the end into the original input shape; The third attention branch is used to permute the input tensor into the third input tensor X HC ∈R W×H×C , capture the interaction between (H, C) dimensions, and transform the output tensor at the end into the original input shape;
[0103] The operations performed by the first attention branch include the following steps:
[0104] (1). Applying channel max pooling and average pooling to reduce the channel dimension, and then concatenating the pooled features can shrink the feature map and further lighten the calculation, which is defined as follows:
[0105]
[0106] where X pool ∈R 2×H×W represents the pooled feature; Max Pool(·) and AvgPool(·) respectively represent max pooling and average pooling by channel; is concatenation by channel; X HW represents the first input tensor;
[0107] (2). The pooled feature X pool is fed into a standard convolution with a sigmoid activation layer, and this activation layer provides the intermediate attention weight X w ∈R 1×H×W ;
[0108] (3). The generated attention weight X w ∈R 1×H×W is applied to the first input tensor X HW through element-wise multiplication;
[0109] The second attention branch performs operations similar to those of the first attention branch, except that a permutation operation is added to permute the input tensor into the second input tensor X CW ∈R H×C×W , with the aim of capturing the dependencies between the (C, W) dimensions. Another permutation operation at the end of the second attention branch transforms the output tensor into the original input shape;
[0110] The third attention branch is constructed in the same way, with the first permutation operation on the third input tensor X HC ∈R W×H×C , to build the dependencies between the (H, C) dimensions. Another permutation operation at the end of the third attention branch transforms the output tensor into the original input shape;
[0111] S5.3. Aggregate the outputs of the three branches through average pooling to generate refined features;
[0112] Obtain the refined feature X in ∈R C×H×W from the input tensor X ref . The operations of each multi-branch attention layer are expressed as:
[0113]
[0114] where the first input tensor X HW is equal to the input tensor X in ; the second input tensor X CW and the third input tensor X HC respectively represent the permutation results from the input tensor X in ; represents the attention weights generated from the i-th attention branch; ⊙ is the broadcast element-wise multiplication; —— represents the permutation operation;
[0115] S5-3. The feature adaptive fusion module fuses the features F t extracted by the temporal convolutional network from the output of the first multi-branch attention layer and the features F s extracted by the spatial convolutional network from the output of the second multi-branch attention layer; specifically, for the features F t extracted by the temporal convolutional network, the adaptive temporal feature Z t is expressed by the formula:
[0116] Z t = λF t +(1 - λ)∈
[0117] where ∈ represents random noise; λ represents the ratio of the control signal to the noise;
[0118] For the features Z s, the adaptive spatial feature Z s It is expressed by the formula as:
[0119] Z s = λF s +(1 - λ)∈
[0120] Fuse the obtained adaptive features. Among them, the feature fusion strategy is expressed as:
[0121] F fusion = w t Z t + w s Z s
[0122] Among them, w t , w s respectively represent the weights of the temporal feature and the spatial feature; the two are dynamically adjusted according to the classification loss of the model;
[0123] Constructing adaptive features can help identify the features that are most important for the final classification decision, thereby improving the transparency of the model;
[0124] S5-4. The deep convolutional network includes multiple deep convolutional modules. Each deep convolutional module is composed of a convolutional layer, a batch normalization layer, and a max pooling layer connected in sequence. Preferably, the specific kernel size is set to 3, which is used to extract the fused features output by the feature adaptive fusion module; in Figure 3 , the deep convolutional network is described by taking 4 deep convolutional modules as an example;
[0125] S5-5. Flatten all features and complete the final classification of the seizure period and the interictal period through two fully connected layers;
[0126] S6. Divide the data obtained in step S4 into a training set and a validation set, that is, shuffle the PSD data combinations reconstructed from X - 1 objects as the training set, and shuffle the PSD data reconstructed from 1 object as the validation set, where X represents the total number of objects; input the training data and the validation data into the MSSTDCN model for training related tasks.
[0127] The protection scope of the present invention is not limited to the above specific embodiments. Any form of change based on the core principle of the present invention, including but not limited to equivalent replacement of technical solutions, structural improvements, etc., shall be regarded as falling within the protection scope of the claims of the present invention. Various modifications and adjustments made to the implementation solutions by those skilled in the art without departing from the design concept of the present invention also belong to the protection scope of the present invention.
Claims
1. A multi-scale cross-object detection method for epileptic seizures based on power spectral density, characterized in that: It includes the following steps: S1. Obtain multi-channel EEG signal data of epilepsy patients and perform preprocessing; S2. Perform regularization processing on the EEG signal data preprocessed in step S1; S3. Adaptive power spectral density construction, that is, reconstruct power spectral density features; S4. Balance the EEG signal data using the SMOTE oversampling method according to the power spectral density features obtained in step S3 to form a balanced data set; S5. Build and train the MSSTDCN model The MSSTDCN model is a multi-scale spatio-temporal deep convolutional network model, which consists of a time feature extraction network, a spatial feature extraction network, a first multi-branch attention layer, a second multi-branch attention layer, a feature adaptive fusion module, a deep convolutional network, and multiple fully connected layers; the balanced adaptive PSD features obtained in step S4 are first input into the parallel dual-branch module composed of the time feature extraction network and the spatial feature extraction network to extract time and spatial features respectively. The extracted time features are fused according to the corresponding frequency bands and then input into the first multi-branch attention layer, and the spatial features are fused according to the corresponding frequency bands and then input into the second multi-branch attention layer; the first multi-branch attention layer and the second multi-branch attention layer respectively perform feature mapping on the corresponding input features to obtain corresponding refined features; The feature adaptive fusion module fuses the output results of the first multi-branch attention layer and the second multi-branch attention layer, and the obtained fusion features are input into the deep convolutional network; the deep convolutional network extracts high-dimensional features; all features are flattened and classified through multiple fully connected layers to obtain epilepsy detection results; S6. Divide the data obtained in step S4 into a training set and a validation set, that is, shuffle the PSD data combinations reconstructed from X - 1 objects as the training set, and shuffle the PSD data reconstructed from 1 object as the validation set, where X represents the total number of objects; input the training data and validation data into the MSSTDCN model in step S5 for training.
2. The multi-scale seizure cross-object detection method based on power spectral density according to claim 1, wherein: The step S3 includes the following sub-steps: S3.
1. Perform frequency division processing on the EEG signal processed in step S2, and divide it into a σ frequency band with a frequency of 0.5 - 4 Hz, a θ frequency band with a frequency of 4 - 8 Hz, an α frequency band with a frequency of 8 - 13 Hz, a β frequency band with a frequency of 13 - 30 Hz, and a γ frequency band with a frequency of 30 - 50 Hz; S3.
2. Use the Multitaper method to calculate the PSD values of the corresponding frequency bands, and obtain the PSD values of the σ frequency band, θ frequency band, α frequency band, β frequency band, and γ frequency band respectively; S3.
3. Average the obtained PSD values in the channel dimension; S3.
4. Concatenate the physiological features of the obtained 5 frequency bands in sequence to obtain adaptive PSD features.
3. The multi-scale seizure cross-object detection method based on power spectral density according to claim 2, characterized in that: In the step S3, the Multitaper method is used to calculate the PSD values of the corresponding frequency bands, and the steps are as follows: S3.
11. Select orthogonal windows Select a set of K mutually orthogonal data windows where n = 0, 1, ..., N-1 and N represents the length of the time series; S3.
12. Single PSD estimate value For each taper, multiply the original time series by this taper to obtain a new time series, and then perform a Fourier transform on the new time series to obtain the corresponding PSD estimate; assuming there is a time series x(n), the PSD estimate using the k-th taper V k (n) is: where f represents frequency; j represents the imaginary unit; e -j2πfn represents the complex exponential function; S3.
13. Averaging processing Add the PSD estimates generated by all K tapers to obtain the average PSD:
4. The multi-scale seizure cross-object detection method based on power spectral density according to claim 1, characterized in that: In the step S5, the spatial feature extraction network in the MSSTDCN model consists of five wavelet convolutional layers. First, the features of the σ frequency band, θ frequency band, α frequency band, β frequency band, and γ frequency band are extracted through the five wavelet convolutional layers respectively. Second, the features of each frequency band are concatenated. Finally, the fused features are input into the second multi-branch attention layer, and the features F extracted by the spatial convolutional network are obtained through the second multi-branch attention layer. s Specifically, for a given input EEG representation x at time sample t, the first wavelet convolutional layer performs a discrete wavelet transform, which is defined as: where u(r) and v(r) represent the approximation filter and the detail filter respectively; x A (t) and x D (t) represent the approximation coefficient and the detail coefficient respectively; r represents the filter index, r = 0, 1, 2,......, R; s represents the scaling factor; k represents the sequence number of the wavelet convolution layer; Approximation coefficient x A (t) is fed back layer by layer to the next wavelet convolutional layer, and spectral analysis is iteratively performed continuously through five wavelet convolutional layers.
5. The multi-scale seizure cross-object detection method based on power spectral density according to claim 4, characterized in that: In the step S5, the time feature extraction network in the MSSTDCN model consists of multiple standard two-dimensional convolutional layers with different kernel sizes, which respectively extract time features of different scales in the signal. Finally, the fused features are input into the first multi-branch attention layer, and the features F extracted by the time convolutional network are obtained through the first multi-branch attention layer t .
6. The multi-scale seizure cross-object detection method based on power spectral density according to claim 4 or 5, characterized in that: In the step S5, each multi-branch attention layer in the MSSTDCN model includes three attention branches for the input tensor X in ∈R C×H×W , where C represents the number of channels, that is, the number of input feature maps, and H and W respectively represent the height and width of the input feature map; the operation is as follows: S5.
1. The multi-branch attention layer performs group convolution on the input time-frequency feature X in ∈R C×H×W ; S5.
2. Pass the feature map to three attention branches respectively. The first attention branch is used to capture the interaction between (H, W) dimensions and perform it on X HW ∈R C×H×W ; The second attention branch is used to permute the input tensor into the second input tensor X CW ∈R H×C×W , capture the interaction between (C, W) dimensions, and transform the output tensor at the end into the original input shape; The third attention branch is used to permute the input tensor into the third input tensor X HC ∈R W×H×C , capture the interaction between (H, C) dimensions, and transform the output tensor at the end into the original input shape; S5.
3. Aggregate the outputs of the three attention branches through average pooling to generate refined features.
7. The multi-scale seizure cross-object detection method based on power spectral density according to claim 6, characterized in that: The operations performed by the first attention branch include the following steps: S5.
21. Apply channel max pooling and average pooling to reduce the channel dimension, and then concatenate the pooled features, defined as follows: where X pool ∈R 2×G×W represents the pooled feature; Max Pool(·) and AvgPool(·) respectively represent the per-channel maximum pooling and average pooling; is the per-channel concatenation; X HW represents the first input tensor; S5.22, Pooled Feature X pool is fed into a standard convolution with a sigmoid activation layer, which provides intermediate attention weights X w ∈R 1×H×W ; S5.
23. Apply the generated attention weights X w ∈R 1×H×W to the first input tensor X HW .
8. The multi-scale epileptic seizure cross-object detection method based on power spectral density according to claim 6, wherein: In the step S5, in the MSSTDCN model, from the input tensor X in ∈R C×H×W obtain the refined feature X ref , and the operation of each multi-branch attention layer is expressed as: Among them, the first input tensor X HW is equal to the input tensor X in ; the second input tensor X CW and the third input tensor X HC respectively represent the permutation results from the input tensor X in ; represents the attention weight generated from the i-th attention branch; ⊙ is the broadcast element multiplication; —— represents the permutation operation.
9. The method for cross-object detection of multi-scale epileptic seizures based on power spectral density according to claim 1 or 2, characterized in that: In the step S5, in the MSSTDCN model, the feature adaptive fusion module fuses the feature F extracted by the temporal convolutional network output by the first multi-branch attention layer t and the feature F extracted by the spatial convolutional network of the second multi-branch attention layer s Specifically, for the feature F extracted by the temporal convolutional network t , the adaptive temporal feature Z t is expressed by the formula as follows: Z t = λF t + (1 - λ)∈ where ∈ represents random noise; λ represents the ratio of the control signal to the noise; For the feature F extracted by the spatial convolutional network s , the adaptive spatial feature Z s is expressed by the formula as follows: Z s = λF s + (1 - λ)∈ Fuse the obtained adaptive features, where the feature fusion strategy is expressed as: F fusion = w t Z t + w s Z s where, w t and w s represent the weights of the time feature and the space feature respectively; the two are dynamically adjusted according to the classification loss of the model.
10. The multi-scale seizure cross-object detection method based on power spectral density according to claim 9, characterized in that: In step S5, the deep convolutional network in the MSSTDCN model includes multiple deep convolutional modules, each of which consists of a convolutional layer, a batch normalization layer, and a max pooling layer connected in sequence, and is used to extract the fused features output by the feature adaptive fusion module.
Citation Information
Patent Citations
Children epilepsy syndrome auxiliary analysis method based on double-flow 3D deep neural network
CN114224363A
Electroencephalogram signal classification method and Parkinson's disease detection device
CN116226710A
Multi-task epilepsy electroencephalogram automatic detection and prediction system
CN116304575A
Epilepsy classification method and system based on electroencephalogram feature fusion deep learning model
CN116898454A
Epileptic seizure prediction system and method based on space-time multi-scale attention mechanism
CN117122281A
Cited By
Cross-modal sEEG-to-iEEG generation method based on multi-scale condition regularization stream
CN121144700A