Spiking wave detection method based on multi-channel time-frequency diagram and biomimetic pac-gated spike superposition
Patent Information
- Application Number
- CN202611018796.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-09
AI Technical Summary
[0007]本发明针对现有深度学习方法在RonS检测中忽视神经生理耦合机制、抗噪能力差以及样本不平衡等问题,提供一种基于多通道时频图与仿生PAC门控的棘波叠加涟波检测方法
[0069]本发明针对现有深度学习方法在RonS检测中忽视神经生理耦合机制、低信噪比环境下抗噪能力差以及样本不平衡的问题,提出一种基于多通道时频图与仿生PAC门控的深度神经网络检测方法。通过构建三通道输入与PAC门控机制,结合频域注意力骨干网络,显著提升了在复杂脑电背景下对微弱病理性高频振荡的检测精度与模型泛化能力。其创新性主要体现在:采用仿生PAC门控模块建模低频棘波相位对高频涟波幅值的调制作用,通过时序同步投影与残差进行物理调制,增强RonS信号并抑制非耦合的高频噪声,降低肌电伪迹导致的误报率;利用基于全局物理阈值的归一化策略保留信号绝对强度信息,结合FCA的ConvNeXt骨干网络,通过离散余弦变换筛选频域关键特征通道,解决了传统空间注意力忽略频域相关性的问题;引入针对样本不平衡和物理约束的复合损失函数,通过Focal Loss解决正负样本间的分布差异,同时利用门控约束损失强制网络学习符合神经生理学规律的稀疏门控行为,有效提升了模型在不平衡数据集下的训练稳定性与结果可解释性。
Smart Images

Figure CN122527484B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical signal processing and artificial intelligence-assisted diagnosis technology, specifically involving a spike-and-ripple (RonS) detection method based on scalp electroencephalography. By using a multi-channel time-frequency graph input and a biomimetic phase-amplitude coupling (PAC) gating network, it solves the problems of robust RonS detection and sample imbalance in low signal-to-noise ratio environments. Background Technology
[0002] Epilepsy is one of the most common chronic neurological disorders worldwide, and accurate localization of the epileptogenic zone is crucial in preoperative evaluation for epilepsy surgery. Traditionally, interictal epileptiform discharges (IEDs) have been the primary localization biomarkers. However, recent studies have shown that high-frequency brain oscillations (HFOs) have higher specificity and spatial accuracy than IEDs in identifying epileptogenic tissue. RonS has demonstrated high pathological specificity in various HFO events. Studies have shown that RonS has higher specificity and spatial accuracy in localizing the epileptogenic zone than spikes or ripples alone, and is strongly correlated with a favorable postoperative prognosis. Existing detection methods mainly focus on detecting spikes or ripples alone, with less attention paid to RonS. Compared to individual waveforms, the high correlation between complex waves and the seizure zone (SOZ) makes RonS more accurate in predicting epileptic seizures.
[0003] Currently, there are few automatic detection methods for RonS. Although some studies have proposed RonS detection methods based on time-frequency analysis or deep learning, the following technical problems still exist:
[0004] 1. Most existing detection methods perform continuous wavelet transform on high-frequency signals to obtain a single-channel time-frequency plot as model input. This method removes the multi-band signal from the EEG, failing to obtain low-frequency components and global contextual information. This results in the loss of correlations between different frequency components in the model input, making it difficult for the model to establish effective distinguishing criteria at the feature level.
[0005] 2. Existing deep learning models often neglect the neurophysiological coupling mechanism in RonS. Most current image detection frameworks directly treat time-frequency images as general natural images for processing, ignoring the influence of low-frequency spike phase modulation on the amplitude of high-frequency ripples. This easily leads to misclassification of high-frequency noise that is not superimposed on spikes as ripples, resulting in a high false detection rate.
[0006] To address the above problems, this invention proposes a deep neural network detection model based on multi-channel time-frequency maps and a biomimetic PAC gating mechanism. This model achieves feature enhancement and model constraint through neurophysiological coupling mechanisms, effectively reducing broadband noise interference and resolving data imbalance issues, thereby improving the detection accuracy of RonS. Summary of the Invention
[0007] This invention addresses the shortcomings of existing deep learning methods in RonS detection, such as neglecting neurophysiological coupling mechanisms, poor noise resistance, and sample imbalance. It provides a spike-ripple detection method based on multi-channel time-frequency maps and biomimetic PAC gating. First, based on neurophysiological mechanisms, the single-channel EEG signal is decomposed into a three-channel time-frequency map signal, and global physical normalization is used to preserve energy information. Second, a biomimetic PAC gating module is designed to enhance high-frequency ripple features using low-frequency spike generation gating, while suppressing high-frequency noise interference not superimposed on the spikes. Finally, the ConvNeXt backbone network is combined with Fourier Channel Attention (FCA), using a joint loss function that addresses data imbalance and gating constraints to achieve high-precision end-to-end detection, improving the model's detection performance and generalization ability.
[0008] To achieve the above objectives, the following technical solution is adopted:
[0009] In a first aspect, embodiments of this application provide a method for detecting spike-and-ripple superposition based on multi-channel time-frequency maps and biomimetic PAC gating, comprising the following steps:
[0010] Step 1: Obtain complete single-channel EEG data, preprocess the input EEG signals, slice and expand them.
[0011] Step 2: Process the extended signal obtained in Step 1 through the R, G, and B channels to obtain a three-channel signal; perform wavelet transform and normalization on the three-channel signal to obtain the corresponding three-channel time-frequency diagram.
[0012] Step 3: Process the three-channel time-frequency diagram using the constructed biomimetic PAC gating module to obtain the enhanced multi-channel time-frequency diagram.
[0013] Step 4: Construct a frequency domain attention-based (FCA-ConvNeXt) backbone network.
[0014] Based on the ConvNeXt backbone network, a frequency attention module (FCA) is embedded in Stages 1-3 to obtain the FCA-ConvNeXt backbone network. The enhanced multi-channel time-frequency map is used as input to output the classification result.
[0015] Step 5: Design a composite loss function that combines classification loss and gating constraint loss, and train and optimize a hybrid neural network that combines a biomimetic PAC gating module and an FCA-ConvNeXt backbone network.
[0016] Step 6: Input the test set data into the trained hybrid neural network model to achieve automatic detection of spikes and ripples (RonS).
[0017] In one possible implementation, the preprocessing includes filtering; the slicing involves using a sliding window technique to temporally segment the EEG signal into fixed-length sliding windows, and then expanding these fixed-length sliding windows horizontally to ensure the integrity of the ripples within the windows, resulting in expanded signals. All expanded signals are then divided into datasets, and the training set is downsampled.
[0018] In one possible implementation, step 1 is specifically implemented as follows:
[0019] Step 1-1: Obtain complete EEG signal data from multiple patients as raw signals to construct a dataset.
[0020] Steps 1-2: Remove the DC component from the original signal, remove the baseline drift of extremely low frequencies by high-pass filtering, and remove the power line noise of the data by notch filtering to obtain the processed EEG signal.
[0021] Steps 1-3: Window the processed EEG signal. Use the sliding window method to perform time-series segmentation of the EEG signal. Set a rectangular window with a duration of 200 milliseconds and slide it continuously along the time axis with a 50% overlap rate (corresponding to a step size of 100 milliseconds) to form a 200ms sliding window.
[0022] Steps 1-4: Define windows with an overlap rate greater than 70% between the sliding window and the annotation window as RonS windows, and define the rest as background windows. Extend all windows left and right by 50ms each to form a complete 300ms sliding window.
[0023] Steps 1-5: Divide all generated 300ms sliding window samples, with the RonS window as positive samples and the background window as negative samples. Divide all samples into training set, validation set, and test set.
[0024] Steps 1-6: Downsample the ratio of RonS window to background window in the training set, and preserve the original distribution of the validation set and test set to simulate the real scene.
[0025] In one possible implementation, step 2 is specifically implemented as follows:
[0026] Step 2-1: Filter the complete sliding window obtained in Step 1 using filters of different frequency bands: Perform a bandpass filter of 80-250Hz on the R channel to obtain a high-frequency ripple signal, which serves as the modulated amplitude signal in the PAC coupling mechanism; perform a bandpass filter of 1-70Hz on the B channel to extract low-frequency spikes and capture their morphology and phase information as the modulation signal; leave the G channel unprocessed to retain the extended signal from Step 1 and provide global background information.
[0027] Step 2-2: Perform continuous wavelet transform (CWT) on the three processed channels using complex Morlet wavelet to obtain the time-frequency amplitude matrix.
[0028] Steps 2-3: Calculate the maximum absolute amplitude of all samples within the window of different channels, and then calculate the 95th percentile of the maximum absolute amplitude of all samples, which will be used as the global amplitude benchmark value for the corresponding channel.
[0029] A channel scaling factor is introduced to fine-tune the global amplitude reference value. The time-frequency amplitude matrix is divided by the global amplitude reference value adjusted by the channel scaling factor, while a small constant is introduced to prevent the denominator from being zero. Then, a truncation function is used to constrain the calculation results, truncating high amplitude noise that significantly exceeds the threshold to 1.0, and constraining the input features to be uniformly mapped to the [0, 1] interval, resulting in the normalized time-frequency graph.
[0030] Finally, the normalized time-frequency image is bilinearly interpolated to obtain a two-dimensional time-frequency image for each channel. The R, G, and B channel images are stacked according to the channel dimension to form the input tensor.
[0031] In one possible implementation, step 3 is specifically implemented as follows:
[0032] Step 3-1: Construct a spike encoder using the PAC coupling mechanism. The encoder input is a normalized B-channel time-frequency map tensor. Encode the intermediate feature map using a two-dimensional convolutional layer with a large receptive field. Apply the Sigmoid activation function to map the feature values to the (0,1) interval, generating the original two-dimensional gated map.
[0033] Step 3-2: Generate a time-series projection gating graph using the original two-dimensional gating graph.
[0034] First, global max pooling is performed along the frequency axis on the original two-dimensional gating graph to extract the strongest activation value at each time step and generate a one-dimensional temporal gating vector.
[0035] When a spike feature is detected on any frequency component at any time, that time is marked as the gating open state. Then, using a broadcast mechanism, the one-dimensional temporal gating vector is copied H times along the frequency axis, expanding it into a two-dimensional projective gating graph with the same dimension as the original input.
[0036] Step 3-3: Construct a ripple encoder. The input to the encoder is a normalized R-channel time-frequency tensor for feature extraction.
[0037] High-frequency features are extracted using 3×3 convolution kernels, and then... The activation function enhances the network's expressive power, resulting in a ripple feature map. A two-dimensional projection-gated map is used as the gain, and its modulation term is calculated by element-wise multiplication with the ripple feature map.
[0038] Steps 3-4: Introduce learnable scalar parameters and adopt... The function transforms the scalar parameter to obtain the actual scaling factor.
[0039] Finally, the weighted modulation terms are superimposed back onto the original ripple features using residual connections to generate an enhanced multi-channel time-frequency map.
[0040] In one possible implementation, the frequency attention module is specifically implemented as follows:
[0041] First, global average pooling is performed along the spatial dimension on the enhanced multi-channel time-frequency map to generate channel feature vectors.
[0042] Performing a one-dimensional discrete cosine transform on the vector maps the spatial channel features to the frequency domain, yielding a spectral vector.
[0043] Set the dimensionality reduction ratio and perform a truncation operation on the spectrum vector, retaining only the first k low-frequency coefficients of the spectrum vector to remove high-frequency noise.
[0044] The truncated vector is input into a multilayer perceptron and activated by a sigmoid function to generate channel weights, which are then used to calibrate the original feature map by adding weights to each channel.
[0045] In one possible implementation, the ConvNeXt network architecture uses the ConvNeXt-Tiny variant as its basic backbone and includes four feature extraction stages, where Stage 0 retains the standard ConvNeXt module and is responsible for extracting shallow texture features.
[0046] FCA-ConvNeXt embeds the FCA module in Stages 1-3 of ConvNeXt. Specifically, the embedding location is after depthwise convolution and layer normalization, and before the first pointwise convolution. The specific process is as follows:
[0047] The input features are processed by depthwise convolution and layer normalization to obtain intermediate features; the intermediate features are then fed into the FCA module for channel attention calibration to generate weighted features.
[0048] Finally, the channel dimensions of the weighted features are enlarged by the first pointwise convolution, then activated by the GELU function, and finally restored by the second pointwise convolution. The processed features are scaled by the layers and then residually connected to the original input to obtain the final output.
[0049] In one possible implementation, step 5 is specifically implemented as follows:
[0050] Constructing a composite loss function Classification loss and gating constraints Weighted composition.
[0051] Among them, the category of loss item The definition is as follows:
[0052]
[0053] in Predict the probability of positive samples for the model.
[0054] Gating constraints The definition is as follows:
[0055] When the input contains RonS, the gating mechanism is activated within the time window corresponding to the spike, while remaining silent in the non-spike region to reflect sparsity. When the input is background, there are no spikes that can generate ripples, and the gating remains closed throughout the time period, blocking noise from passing through the R channel. By minimizing the gating mean, the network is forced to output all-zero gating when there is no spike input.
[0056] The total loss is calculated based on the constructed composite loss function. The AdamW optimizer is used to iteratively adjust the weight coefficients and bias terms in the hybrid neural network through the gradient backpropagation algorithm until the preset optimization termination condition is met. Finally, the model weights with the best generalization performance are saved.
[0057] Secondly, embodiments of this application provide a spike-ripple detection system based on multi-channel time-frequency diagrams and biomimetic PAC gating, comprising the following modules:
[0058] Data processing module: Acquires complete single-channel EEG data, preprocesses the input EEG signals, slices and expands them.
[0059] Time-frequency graph acquisition module: The extended signal obtained from the data processing module is processed through the R, G, and B channels to obtain a three-channel signal; the three-channel signal is then subjected to wavelet transform and normalization to obtain the corresponding three-channel time-frequency graph.
[0060] Enhancement Module: The three-channel time-frequency diagram is processed by the constructed biomimetic PAC gating module to obtain the enhanced multi-channel time-frequency diagram.
[0061] Detection module: Based on the ConvNeXt backbone network, a frequency attention module (FCA) is embedded in Stage 1-3 to obtain the FCA-ConvNeXt backbone network. The enhanced multi-channel time-frequency map is used as input, and the classification result is output.
[0062] Training module: Design a composite loss function that combines classification loss and gating constraint loss to train and optimize a hybrid neural network that combines a biomimetic PAC gating module with an FCA-ConvNeXt backbone network.
[0063] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory.
[0064] The memory is used to store computer programs.
[0065] When the processor executes the program stored in the memory, it implements any of the spike-and-ripple detection methods described in this application.
[0066] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the spike-and-ripple detection methods described in this application.
[0067] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the spike superimposed ripple detection methods described in this application.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] This invention addresses the shortcomings of existing deep learning methods in RonS detection, such as neglecting neurophysiological coupling mechanisms, poor noise resistance in low signal-to-noise ratio environments, and sample imbalance. It proposes a deep neural network detection method based on multi-channel time-frequency maps and biomimetic PAC gating. By constructing a three-channel input and PAC gating mechanism, combined with a frequency-domain attention backbone network, the detection accuracy and model generalization ability for weak pathological high-frequency oscillations are significantly improved in complex EEG backgrounds. Its innovations are mainly reflected in the following aspects: First, it employs a biomimetic PAC gating module to model the modulation effect of low-frequency spike phase on high-frequency ripple amplitude. Physical modulation is achieved through temporal synchronous projection and residuals, enhancing the RonS signal and suppressing uncoupled high-frequency noise, thus reducing the false alarm rate caused by EMG artifacts. Second, it utilizes a normalization strategy based on a global physical threshold to preserve absolute signal strength information. Combined with the ConvNeXt backbone network of FCA, it filters key frequency domain feature channels through discrete cosine transform, solving the problem of traditional spatial attention ignoring frequency domain correlation. Third, it introduces a composite loss function targeting sample imbalance and physical constraints. Focal Loss addresses the distribution differences between positive and negative samples, while gating constraint loss forces the network to learn sparse gating behavior that conforms to neurophysiological laws, effectively improving the training stability and interpretability of the model under imbalanced datasets. Attached Figure Description
[0070] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.
[0071] Figure 2 This is a diagram illustrating the detection effect of a specific application of the present invention on complete electroencephalogram (EEG) signals. Detailed Implementation
[0072] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0073] like Figure 1 As shown, this application provides a method for detecting spike-and-ripple superposition based on multi-channel time-frequency maps and biomimetic PAC gating, including the following steps:
[0074] Step 1: Acquire complete single-channel EEG data. Preprocess the input EEG signal, then slice and expand it. The preprocessing includes filtering. Slicing involves using a sliding window technique to divide the EEG signal into 200ms sliding windows, extending each window by 50ms to the left and right to ensure the integrity of the ripples within the window, thus obtaining the expanded signal. All expanded signals are then divided into datasets, and the training set is downsampled.
[0075] Step 2: Process the extended signal obtained in Step 1 through the R, G, and B channels to obtain a three-channel signal. Perform wavelet transform and normalization on the three-channel signal to obtain the corresponding three-channel time-frequency diagram.
[0076] Step 3: Process the three-channel time-frequency map using the constructed biomimetic PAC gating module; extract the features of the B channel through the PAC coupling mechanism to generate a time-series projection gating, enhance the features of the R channel, add the enhanced result to the original data to form a new R channel, and obtain the enhanced multi-channel time-frequency map.
[0077] Step 4: Construct a frequency-attention-based (FCA-ConvNeXt) backbone network. Based on the complete ConvNeXt backbone network, embed frequency attention modules (FCA) in Stages 1-3; the FCA-ConvNeXt backbone network takes the enhanced multi-channel time-frequency map as input and outputs the classification result.
[0078] Step 5: Design a composite loss function that combines classification loss and gating constraint loss to address sample imbalance while forcing the network to learn gating behavior that conforms to neurophysiological laws; train and optimize a hybrid neural network that combines a biomimetic PAC gating module with an FCA-ConvNeXt backbone network.
[0079] Step 6: Input the test set data into the trained hybrid neural network model to achieve automatic detection of spikes and ripples (RonS).
[0080] The specific implementation of step 1 is as follows:
[0081] Step 1-1: Obtain complete EEG signal data from 12 patients as raw signals to construct a dataset. The data is a 21-channel EEG conforming to the international 10-20 system, with a sampling frequency of 1000Hz.
[0082] Steps 1-2: Remove the DC component from the original signal, remove the baseline drift of extremely low frequencies by using a 0.5Hz high-pass filter, and remove the power line noise of the data by using a 50Hz notch filter to obtain the processed EEG signal.
[0083] Steps 1-3: Window the processed EEG signal. Use the sliding window method to perform time-series segmentation of the EEG signal. Set a rectangular window with a duration of 200 milliseconds and slide it continuously along the time axis with a 50% overlap rate (corresponding to a step size of 100 milliseconds) to form a 200ms sliding window.
[0084] Steps 1-4: To ensure the integrity of the ripples, windows with an overlap rate greater than 70% between the sliding window and the annotation window are defined as RonS windows, and the rest are defined as background windows. Simultaneously, to prevent edge effects caused by the filter and to ensure that the window contains the complete period of both the spikes and ripples, all windows are extended 50ms to the left and right, forming a complete 300ms sliding window.
[0085] Steps 1-5: Divide all 300ms sliding window samples generated in Steps 1-4, with the RonS window as positive samples and the background window as negative samples. Divide all samples into training set, validation set, and test set in a 6:2:2 ratio. The training set is used for model parameter optimization, the validation set is used to monitor the training process and implement early stopping strategy, and the test set is used for final performance evaluation.
[0086] Steps 1-6: Downsample the ratio of RonS window to background window in the training set to 1:5, and preserve the original distribution of the validation set and test set to simulate the real scene.
[0087] The specific implementation of step 2 is as follows:
[0088] Step 2-1: Filter the complete sliding window obtained in Step 1 using filters of different frequency bands: Perform a bandpass filter of 80-250Hz on the R channel to obtain a high-frequency ripple signal, which serves as the modulated amplitude signal in the PAC coupling mechanism; perform a bandpass filter of 1-70Hz on the B channel to extract low-frequency spikes and capture their morphology and phase information as the modulation signal; leave the G channel unprocessed to retain the extended signal from Step 1 and provide global background information.
[0089] Step 2-2: Perform continuous wavelet transform (CWT) on the three processed channels using complex Morlet wavelet:
[0090]
[0091] in, Indicates the channel index. For scale parameters, For translation parameters, This is the complex conjugate of the wavelet function. 224 scale parameters are selected, corresponding to a frequency range of 1-250Hz. Through each scale parameter... and each translation parameter Calculate wavelet coefficients The time-frequency amplitude matrix formed by wavelet coefficients is defined as .
[0092] Steps 2-3: To preserve the absolute signal strength information and suppress background noise, instead of using single-sample Min-Max normalization, we employ statistical global threshold normalization. The specific process is as follows:
[0093] First, calculate the maximum absolute amplitude of all samples within the window of different channels. Then, calculate the 95th percentile of the maximum absolute amplitude of all samples, and use this as the global amplitude benchmark value for the corresponding channel.
[0094]
[0095] in, The total number of training samples, This is the global amplitude reference value, representing The 95th percentile of the largest amplitude across all samples in the channel.
[0096] Introducing channel scaling factor The global amplitude reference value is fine-tuned, with the R channel set to 1.0 to maintain sensitivity; the B and G channels are set to 3.0 to reduce relative brightness and prevent image overexposure, thereby highlighting the main role of the ripples.
[0097] The formula for calculating the normalized time-frequency graph is:
[0098]
[0099] in, It is the numerical stability constant. The function forcibly truncates outliers that significantly exceed the physical threshold to 1.0, constraining the input features to the [0,1] interval to prevent gradient explosion caused by high-amplitude noise. The time-frequency plot is only activated when the absolute value of the signal energy significantly exceeds the background noise benchmark, effectively suppressing artifacts and noise in the background.
[0100] Finally, the normalized time-frequency graph is uniformly scaled to a scaled value using bilinear interpolation. The resolution is determined to obtain a two-dimensional time-frequency image for each channel. The R, G, and B channel images are stacked according to the channel dimension to form an input tensor of (3, 224, 224) (including the R, G, and B channel time-frequency image tensors).
[0101] The specific implementation of step 3 is as follows:
[0102] Step 3-1: Construct a spike encoder using the PAC coupling mechanism. The encoder input is a normalized B-channel time-frequency tensor. To capture the morphological features of spikes on the time-frequency plot, a two-dimensional convolutional layer with a large receptive field was used for encoding to obtain intermediate feature maps. The kernel size was set to 13×13, padding to 6, and stride to 1. Intermediate feature map. The calculation formula is as follows:
[0103]
[0104] in, For the training convolutional kernel weights, For bias terms, These represent the time step and frequency index, respectively. Then, the Sigmoid activation function is applied to map the feature values to the (0,1) interval, generating the original two-dimensional gated graph. :
[0105]
[0106] Step 3-2: Utilize the original two-dimensional gating graph Generate a time-series projection gating graph.
[0107] First, the original two-dimensional gated graph... Perform global max pooling along the frequency axis to extract the strongest activation value at each time step and generate a one-dimensional temporal gating vector. :
[0108]
[0109] At any time A spike feature is detected on any frequency component, and the moment at which the gate is opened is marked.
[0110] Then, by utilizing the broadcast mechanism, the one-dimensional time-gated vector The graph is copied H times along the frequency axis and expanded into a two-dimensional projective gated graph with the same dimensions as the original input. :
[0111]
[0112] in This represents a vector of all 1s along the frequency axis. At this point, It appears as a vertical stripe mask on the time-frequency graph, ensuring that the same gain is applied to all frequency components at the same time.
[0113] Step 3-3: Construct a ripple encoder and process the time-frequency tensor of the R channel. Feature extraction is performed.
[0114] High-frequency features are extracted using 3×3 convolution kernels, and the following is introduced: Activation functions enhance the expressive power of networks:
[0115]
[0116] in This represents the extracted ripple feature map.
[0117] Two-dimensional projection gated image As a gain, with ripple feature map Perform element-wise multiplication to calculate the modulation term. For final feature overlay:
[0118]
[0119] in The time axis coordinates of the time-frequency plot are represented. The frequency axis coordinates of the time-frequency plot, That is, at a specific time and specific frequencies On the coordinates The specific value.
[0120] This achieves cross-frequency coupling, that is, only when Ripple characteristics Only those features are preserved; otherwise, the features are suppressed.
[0121] Steps 3-4: To control the modulation intensity and ensure the stability of the gradient flow, learnable scalar parameters are introduced. (Initialized to 2.0). To ensure the gain coefficient is always positive, the following is used: The function transforms the scalar parameters to obtain the actual scaling factor. :
[0122]
[0123] Finally, the weighted modulation terms are superimposed back onto the original ripple features using residual connections to generate an enhanced multi-channel time-frequency plot. :
[0124]
[0125] The specific implementation of step 4 is as follows:
[0126] Step 4-1: Transform the spatial domain features to the frequency domain using the Frequency Attention (FCA) module, and then use the principle of spectral sparsity to filter key components. The specific process is as follows:
[0127] First, analyze the enhanced multi-channel time-frequency diagram. Global average pooling is performed along the spatial dimension to generate channel feature vectors. :
[0128]
[0129] For vectors Performing a one-dimensional discrete cosine transform maps the spatial channel features to the frequency domain, yielding the spectrum vector. :
[0130]
[0131] Set the dimensionality reduction ratio For the spectrum vector Perform a truncation operation, that is, retain only the spectral vector. The former A few low-frequency coefficients are used to remove high-frequency noise. The resulting vector after truncation is... :
[0132]
[0133] Will Inputting the data into a multilayer perceptron and activating it via a sigmoid function generates channel weights. For the original feature map Perform channel-by-channel weighted calibration:
[0134]
[0135] in , These are the parameters for the dimensionality reduction and dimensionality expansion layers of the MLP, respectively.
[0136] Step 4-2: The ConvNeXt network architecture uses the ConvNeXt-Tiny variant as the basic backbone and includes 4 feature extraction stages. Stage 0 retains the standard ConvNeXt module, which is responsible for extracting shallow texture features.
[0137] FCA-ConvNeXt embeds the FCA module in Stages 1-3 of ConvNeXt. Specifically, the embedding location is after depthwise convolution and layer normalization, and before the first pointwise convolution. The specific process is as follows:
[0138] Input features After depthwise convolution and layer normalization, intermediate features are obtained. :
[0139]
[0140] Then, The signal is fed into the FCA module for channel attention calibration, generating weighted features. :
[0141]
[0142] Finally, through the first pointwise convolution... The channel dimensions are enlarged, then activated by the GELU function, and finally restored by a second pointwise convolution. The processed features are then scaled by layers and compared to the original input. Perform residual connections to obtain the final output. :
[0143]
[0144] Introducing FCA here can more effectively distinguish signals from artifacts based on spectral characteristics while maintaining computational efficiency.
[0145] The specific implementation of step 5 is as follows:
[0146] Step 5-1: Construct a composite loss function Classification loss and gating constraints Weighted composition. Its mathematical expression is:
[0147]
[0148] in, This is a hyperparameter used to balance the weights between classification accuracy and physical consistency. In this embodiment, it will... The value is set to 0.7 to ensure that the model pursues high classification accuracy without sacrificing the interpretability of the gating mechanism.
[0149] Step 5-2, where the classification loss term The definition is as follows:
[0150]
[0151] in This determines the probability of a positive sample for the model. The focus parameter in this example... Set to 2.0, balance parameter Set to 0.5.
[0152] To ensure the gating vector generated in step 3 It has clear neurophysiological significance, namely, it is turned on only when spikes occur and turned off at other times, and is a gating constraint term. The definition is as follows:
[0153] When the input contains RonS, the gating mechanism is activated within the time window corresponding to the spike. Simultaneously, it remains silent in the non-spike region to reflect sparsity. At this point, the label y=1, and the loss function is defined as:
[0154]
[0155] in This represents the maximum activation value of the gate on the time axis. This indicates the average activation level of the gating. (First item) The gating system must be fully open at least at some point; the second item To constrain the overall sparsity of the gate and prevent it from being open at all times, the sparsity coefficient is... Set it to 0.2.
[0156] When the input is background, there are no spikes that can generate ripples, and the gating remains closed at all times. Noise is blocked from passing through the R channel. At this point, the label y=0, and the loss function is defined as:
[0157]
[0158] By minimizing the gating mean, the network is forced to output all-zero gating when there is no spike input.
[0159] The final gate constraint loss is as follows:
[0160]
[0161] Step 5-3: Calculate the total loss based on the constructed composite loss function. Using the AdamW optimizer, iteratively adjust the weight coefficients and bias terms in the hybrid neural network through the gradient backpropagation algorithm until the preset optimization termination condition is met. Finally, save the model weights with the best generalization performance (i.e., the highest F1 score on the validation set).
[0162] The specific implementation of step 6 is as follows:
[0163] Step 6-1: Input the test set data into the trained hybrid neural network model and perform forward propagation to obtain the classification and recognition results of EEG signal segments. The recognition results include RonS (positive class) and background (negative class) signal segments.
[0164] Step 6-2: Compare the identification results with the true annotations and calculate the four basic confusion matrix parameters: true positive, true negative, false positive, and false negative. Finally, calculate the precision, recall, and F1 score of the test set based on the four parameters to evaluate the model's detection performance.
[0165] The implementation method for optimizing the termination condition described in this invention adopts the following approach: 1. Fixed-period termination strategy: Set a maximum training period threshold of 30 training periods. When the cumulative number of training iterations reaches the threshold, the optimization process is automatically terminated. 2. Early stopping strategy: Set a patience value window P (P=10), using the F1 score of the validation set as the core monitoring indicator. When the F1 score of the validation set fails to break the historical high within P consecutive training periods, the early stopping mechanism is triggered. Simultaneously configure the fixed-period threshold and the dynamic early stopping mechanism, prioritizing the execution of the dynamic early stopping condition. If the early stopping condition is not triggered, the process is forcibly terminated after 30 periods.
[0166] Figure 2 This is a diagram illustrating the RonS recognition effect of the present invention on a complete signal.
[0167] In steps 1-3, a sample length of 0.2 s (200 ms) is chosen because the discharge duration of spikes is typically 0.02~0.07 s, while the accompanying ripple oscillation duration is typically between 0.08~0.25 s. Setting a 200 ms sliding window ensures complete capture of a typical spike-ripple superimposed event and its contextual information. To ensure similarity in sample spatial distribution between training and test data, the test data uses the same sliding window cutting method as the training phase. To prevent signal truncation and improve the temporal resolution of detection, a 50% sample overlap rate is set, meaning the starting point of each new sample is 0.1 s (100 ms) from the previous sample. According to the RonS point marking rules, if the starting or ending interval of a sample contains more than 70% of the marked RonS points, then that sample and its adjacent overlapping samples are classified as RonS samples, while other samples that do not contain RonS points and have little or no overlap with the marked intervals are classified as normal samples.
[0168] To further demonstrate the effectiveness of this invention in detecting RonS signals from electroencephalograms (EEGs), it was tested on real EEG data from the Children's Hospital Affiliated to Zhejiang University School of Medicine, and compared with several current mainstream detection algorithms. The specific experimental results are as follows:
[0169] The experimental data was sampled at a frequency of 1000Hz and divided into 12 different datasets of diseased individuals, with an average data length of 10 minutes. The detection algorithm based on 1D+2DCNN proposed by Li et al. in 2024 achieved an average precision of 82.01% on the same dataset, but the average F1 score was only 72.00%, indicating its difficulty in achieving both high recall and high accuracy. Nadalin et al.'s 2022 CNN model based on transfer learning achieved an average precision of 77.69% and an average F1 score of 79.60%. This invention achieves an average precision of 87.69%, a recall of 88.57%, and an average F1 score of 88.12% on the dataset. This invention proposes a novel RonS detection algorithm, which, unlike traditional deep learning that relies solely on data statistics for feature learning, explicitly models the PAC mechanism. By constructing a biomimetic PAC gating module, the gating is generated using the phase features of low-frequency spikes, enhancing the feature weights of the high-frequency ripple channel and significantly improving the model's ability to capture weak high-frequency oscillations. Experiments show that after changing the input from a single-channel grayscale image to a three-channel physical time-frequency image, the F1 score improved from 79.08% to 84.22%. Furthermore, the introduction of FCA and PAC gating further improved the model performance, verifying the crucial role of frequency domain feature selection and the PAC coupling mechanism in detection performance. In addition, to address the extreme positive-negative sample imbalance problem in RonS detection, this invention designs a composite loss function combining Focal Loss and gating constraints. This not only effectively mitigates the impact of sample imbalance but also forces the network to learn sparse gating behavior that conforms to neurophysiological principles. This invention demonstrates superior robustness and generalization performance on real clinical data compared to existing models.
[0170] This application also provides a spike-and-ripple detection system based on multi-channel time-frequency maps and biomimetic PAC gating, including the following modules:
[0171] Data processing module: Acquires complete single-channel EEG data, preprocesses the input EEG signals, slices and expands them.
[0172] Time-frequency graph acquisition module: The extended signal obtained from the data processing module is processed through the R, G, and B channels to obtain a three-channel signal; the three-channel signal is then subjected to wavelet transform and normalization to obtain the corresponding three-channel time-frequency graph.
[0173] Enhancement Module: The three-channel time-frequency diagram is processed by the constructed biomimetic PAC gating module to obtain the enhanced multi-channel time-frequency diagram.
[0174] Detection module: Based on the ConvNeXt backbone network, a frequency attention module (FCA) is embedded in Stage 1-3 to obtain the FCA-ConvNeXt backbone network. The enhanced multi-channel time-frequency map is used as input, and the classification result is output.
[0175] Training module: Design a composite loss function that combines classification loss and gating constraint loss to train and optimize a hybrid neural network that combines a biomimetic PAC gating module with an FCA-ConvNeXt backbone network.
[0176] In one possible implementation, the preprocessing includes filtering; the slicing involves using a sliding window technique to temporally segment the EEG signal into fixed-length sliding windows, and then expanding these fixed-length sliding windows horizontally to ensure the integrity of the ripples within the windows, resulting in expanded signals. All expanded signals are then divided into datasets, and the training set is downsampled.
[0177] In one possible implementation, the data processing module is specifically implemented as follows:
[0178] Step 1-1: Obtain complete EEG signal data from multiple patients as raw signals to construct a dataset.
[0179] Steps 1-2: Remove the DC component from the original signal, remove the baseline drift of extremely low frequencies by high-pass filtering, and remove the power line noise of the data by notch filtering to obtain the processed EEG signal.
[0180] Steps 1-3: Window the processed EEG signal. Use the sliding window method to perform time-series segmentation of the EEG signal. Set a rectangular window with a duration of 200 milliseconds and slide it continuously along the time axis with a 50% overlap rate (corresponding to a step size of 100 milliseconds) to form a 200ms sliding window.
[0181] Steps 1-4: Define windows with an overlap rate greater than 70% between the sliding window and the annotation window as RonS windows, and define the rest as background windows. Extend all windows left and right by 50ms each to form a complete 300ms sliding window.
[0182] Steps 1-5: Divide all generated 300ms sliding window samples, with the RonS window as positive samples and the background window as negative samples. Divide all samples into training set, validation set, and test set.
[0183] Steps 1-6: Downsample the ratio of RonS window to background window in the training set, and preserve the original distribution of the validation set and test set to simulate the real scene.
[0184] In one possible implementation, the time-frequency diagram acquisition module is specifically implemented as follows:
[0185] Step 2-1: Filter the complete sliding window obtained in Step 1 using filters of different frequency bands: Perform a bandpass filter of 80-250Hz on the R channel to obtain a high-frequency ripple signal, which serves as the modulated amplitude signal in the PAC coupling mechanism; perform a bandpass filter of 1-70Hz on the B channel to extract low-frequency spikes and capture their morphology and phase information as the modulation signal; leave the G channel unprocessed to retain the extended signal from Step 1 and provide global background information.
[0186] Step 2-2: Perform continuous wavelet transform (CWT) on the three processed channels using complex Morlet wavelet to obtain the time-frequency amplitude matrix.
[0187] Steps 2-3: Calculate the maximum absolute amplitude of all samples within the window of different channels, and then calculate the 95th percentile of the maximum absolute amplitude of all samples, which will be used as the global amplitude benchmark value for the corresponding channel.
[0188] A channel scaling factor is introduced to fine-tune the global amplitude reference value. The time-frequency amplitude matrix is divided by the global amplitude reference value adjusted by the channel scaling factor, while a small constant is introduced to prevent the denominator from being zero. Then, a truncation function is used to constrain the calculation results, truncating high amplitude noise that significantly exceeds the threshold to 1.0, and constraining the input features to be uniformly mapped to the [0, 1] interval, resulting in the normalized time-frequency graph.
[0189] Finally, the normalized time-frequency image is bilinearly interpolated to obtain a two-dimensional time-frequency image for each channel. The R, G, and B channel images are stacked according to the channel dimension to form the input tensor.
[0190] In one possible implementation, the enhancement module is specifically implemented as follows:
[0191] Step 3-1: Construct a spike encoder using the PAC coupling mechanism. The encoder input is a normalized B-channel time-frequency map tensor. Encode the intermediate feature map using a two-dimensional convolutional layer with a large receptive field. Apply the Sigmoid activation function to map the feature values to the (0,1) interval, generating the original two-dimensional gated map.
[0192] Step 3-2: Generate a time-series projection gating graph using the original two-dimensional gating graph.
[0193] First, global max pooling is performed along the frequency axis on the original two-dimensional gating graph to extract the strongest activation value at each time step and generate a one-dimensional temporal gating vector.
[0194] When a spike feature is detected on any frequency component at any time, that time is marked as the gating open state. Then, using a broadcast mechanism, the one-dimensional temporal gating vector is copied H times along the frequency axis, expanding it into a two-dimensional projective gating graph with the same dimension as the original input.
[0195] Step 3-3: Construct a ripple encoder. The input to the encoder is a normalized R-channel time-frequency tensor for feature extraction.
[0196] High-frequency features are extracted using 3×3 convolution kernels, and then... The activation function enhances the network's expressive power, resulting in a ripple feature map. A two-dimensional projection-gated map is used as the gain, and its modulation term is calculated by element-wise multiplication with the ripple feature map.
[0197] Steps 3-4: Introduce learnable scalar parameters and adopt... The function transforms the scalar parameter to obtain the actual scaling factor.
[0198] Finally, the weighted modulation terms are superimposed back onto the original ripple features using residual connections to generate an enhanced multi-channel time-frequency map.
[0199] In one possible implementation, the frequency attention module is specifically implemented as follows:
[0200] First, global average pooling is performed along the spatial dimension on the enhanced multi-channel time-frequency map to generate channel feature vectors.
[0201] Performing a one-dimensional discrete cosine transform on the vector maps the spatial channel features to the frequency domain, yielding a spectral vector.
[0202] Set the dimensionality reduction ratio and perform a truncation operation on the spectrum vector, retaining only the first k low-frequency coefficients of the spectrum vector to remove high-frequency noise.
[0203] The truncated vector is input into a multilayer perceptron and activated by a sigmoid function to generate channel weights, which are then used to calibrate the original feature map by adding weights to each channel.
[0204] In one possible implementation, the ConvNeXt network architecture uses the ConvNeXt-Tiny variant as its basic backbone and includes four feature extraction stages, where Stage 0 retains the standard ConvNeXt module and is responsible for extracting shallow texture features.
[0205] FCA-ConvNeXt embeds the FCA module in Stages 1-3 of ConvNeXt. Specifically, the embedding location is after depthwise convolution and layer normalization, and before the first pointwise convolution. The specific process is as follows:
[0206] The input features are processed by depthwise convolution and layer normalization to obtain intermediate features; the intermediate features are then fed into the FCA module for channel attention calibration to generate weighted features.
[0207] Finally, the channel dimensions of the weighted features are amplified by the first pointwise convolution, then activated by the GELU function, and finally restored by the second pointwise convolution. The processed features are then scaled and residually concatenated with the original input to obtain the final output.
[0208] In one possible implementation, the training module is specifically implemented as follows:
[0209] Constructing a composite loss function Classification loss and gating constraints Weighted composition. Its mathematical expression is:
[0210]
[0211] in, This is a hyperparameter used to balance the weight between classification accuracy and physical consistency.
[0212] Among them, the category loss item The definition is as follows:
[0213]
[0214] in Predict the probability of positive samples for the model.
[0215] Gating constraints The definition is as follows:
[0216] When the input contains RonS, the gating mechanism is activated within the time window corresponding to the spike, while remaining silent in the non-spike region to reflect sparsity.
[0217] When the input is background, there are no spikes that can generate ripples, and the gating remains closed at all times, blocking noise from passing through the R channel.
[0218] By minimizing the gating mean, the network is forced to output all-zero gating when there is no spike input.
[0219] The total loss is calculated based on the constructed composite loss function. The AdamW optimizer is used to iteratively adjust the weight coefficients and bias terms in the hybrid neural network through the gradient backpropagation algorithm until the preset optimization termination condition is met. Finally, the model weights with the best generalization performance are saved.
[0220] This application also provides an electronic device, including a processor and a memory.
[0221] The memory is used to store computer programs.
[0222] When the processor executes a program stored in the memory, it implements any of the methods described in this application.
[0223] In one possible implementation, the electronic device of this application embodiment further includes a communication interface and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0224] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0225] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0226] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0227] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0228] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the methods described in this application.
[0229] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the methods described in this application.
[0230] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0231] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0232] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.
[0233] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A method for detecting spike-and-ripple superposition based on multi-channel time-frequency plots and biomimetic PAC gating, characterized in that, Includes the following steps: Step 1: Obtain complete single-channel EEG data, preprocess the input EEG signals, slice and expand them; Step 2: Process the extended signal obtained in Step 1 through the R, G, and B channels to obtain a three-channel signal; perform wavelet transform and normalization on the three-channel signal to obtain the corresponding three-channel time-frequency diagram; Step 3: Process the three-channel time-frequency diagram using the constructed biomimetic PAC gating module to obtain the enhanced multi-channel time-frequency diagram; The specific implementation of step 3 is as follows: Step 3-1: Construct a spike encoder using the PAC coupling mechanism. The encoder input is a normalized B-channel time-frequency map tensor. Use a two-dimensional convolutional layer with a large receptive field to encode the intermediate feature map. Apply the Sigmoid activation function to map the feature values to the (0,1) interval to generate the original two-dimensional gated map. Step 3-2: Generate a time-series projection gating graph using the original two-dimensional gating graph; First, global max pooling is performed along the frequency axis on the original two-dimensional gating graph to extract the strongest activation value at each time step and generate a one-dimensional temporal gating vector. When a spike feature is detected on any frequency component at any time, that time is marked as the gate open state. Then, by using the broadcast mechanism, the one-dimensional time-series gating vector is copied H times along the frequency axis and expanded into a two-dimensional projection gating graph with the same dimension as the original input. Step 3-3: Construct a ripple encoder. The input to the encoder is a normalized R-channel time-frequency tensor for feature extraction. High-frequency features are extracted using 3×3 convolution kernels, and then... The activation function enhances the network's expressive power to obtain the ripple feature map; the two-dimensional projection gating map is used as the gain and multiplied element-wise with the ripple feature map to calculate the modulation term; Steps 3-4: Introduce learnable scalar parameters and adopt... The function transforms the scalar parameters to obtain the actual scaling factor; Finally, the weighted modulation terms are superimposed back onto the original ripple features using residual connections to generate an enhanced multi-channel time-frequency plot. Step 4: Construct a backbone network based on frequency domain attention; Based on the ConvNeXt backbone network, a frequency attention module is embedded in Stages 1-3 to obtain the FCA-ConvNeXt backbone network. The enhanced multi-channel time-frequency map is used as input to output the classification result. The ConvNeXt network architecture uses the ConvNeXt-Tiny variant as its basic backbone and includes four feature extraction stages. Stage 0 retains the standard ConvNeXt module and is responsible for extracting shallow texture features. FCA-ConvNeXt embeds the FCA module in Stages 1-3 of ConvNeXt. Specifically, the embedding location is after depthwise convolution and layer normalization, and before the first pointwise convolution. The specific process is as follows: The input features are processed by depthwise convolution and layer normalization to obtain intermediate features; the intermediate features are then fed into the FCA module for channel attention calibration to generate weighted features. Finally, the channel dimension of the weighted features is enlarged by the first pointwise convolution, and after passing through the GELU activation function, the dimension is restored by the second pointwise convolution. The processed features are scaled by the layers and then residually connected to the original input to obtain the final output. Step 5: Design a composite loss function, combining classification loss and gating constraint loss, and train and optimize a hybrid neural network that combines the biomimetic PAC gating module and the FCA-ConvNeXt backbone network; Step 6: Input the test set data into the trained hybrid neural network model to achieve automatic detection of spikes and ripples (RonS).
2. The spike-ripple detection method based on multi-channel time-frequency diagrams and biomimetic PAC gating according to claim 1, characterized in that, The preprocessing includes filtering; the slicing involves using a sliding window technique to temporally segment the EEG signal into fixed-length sliding windows, and then expanding the fixed-length sliding windows left and right to ensure the integrity of the ripples within the window to obtain expanded signals; all expanded signals are divided into datasets, and the training set is downsampled.
3. The method for detecting spike-ripple superposition based on multi-channel time-frequency diagrams and biomimetic PAC gating according to claim 2, characterized in that, The specific implementation of step 1 is as follows: Step 1-1: Obtain complete EEG signal data from multiple patients as raw signals to construct a dataset; Steps 1-2: Remove the DC component from the original signal, remove the baseline drift of extremely low frequencies by high-pass filtering, and remove the power line noise of the data by notch filtering to obtain the processed EEG signal. Steps 1-3: Window the processed EEG signal. Use the sliding window method to segment the EEG signal temporally. Set a rectangular window with a duration of 200 milliseconds and slide it continuously along the time axis with a 50% overlap rate to form a 200ms sliding window. Steps 1-4: Define windows with an overlap rate greater than 70% between the sliding window and the annotation window as RonS windows, and define the rest as background windows; extend all windows left and right by 50ms each to form a complete sliding window of 300ms. Steps 1-5: Divide all the generated 300ms sliding window samples into two groups, with the RonS window as the positive sample and the background window as the negative sample. All samples are divided into training set, validation set and test set; Steps 1-6: Downsample the ratio of RonS window to background window in the training set, and preserve the original distribution of the validation set and test set to simulate the real scene.
4. The method for detecting spike-ripple superposition based on multi-channel time-frequency diagrams and biomimetic PAC gating according to claim 1, characterized in that, The specific implementation of step 2 is as follows: Step 2-1: Filter the complete sliding window obtained in Step 1 using filters of different frequency bands: Perform a bandpass filter of 80-250Hz on the R channel to obtain a high-frequency ripple signal, which serves as the modulated amplitude signal in the PAC coupling mechanism; perform a bandpass filter of 1-70Hz on the B channel to extract low-frequency spikes, capturing their morphology and phase information as the modulation signal; leave the G channel unprocessed to retain the extended signal from Step 1, providing global background information. Step 2-2: Perform continuous wavelet transform on the three processed channels using complex Morlet wavelets to obtain the time-frequency amplitude matrix; Steps 2-3: Calculate the maximum absolute amplitude of all samples within the window of different channels, and then calculate the 95th percentile of the maximum absolute amplitude of all samples, which will be used as the global amplitude benchmark value for the corresponding channel. A channel scaling factor is introduced to fine-tune the global amplitude reference value. The time-frequency amplitude matrix is divided by the global amplitude reference value adjusted by the channel scaling factor, and a small constant is introduced to prevent the denominator from being zero. Then, the calculation results are constrained by a truncation function to cut off high amplitude noise that significantly exceeds the threshold to 1.
0. The input features are constrained to be uniformly mapped to the [0,1] interval to obtain the normalized time-frequency graph. Finally, the normalized time-frequency image is bilinearly interpolated to obtain a two-dimensional time-frequency image for each channel. The R, G, and B channel images are stacked according to the channel dimension to form the input tensor.
5. The spike-ripple detection method based on multi-channel time-frequency diagrams and biomimetic PAC gating according to claim 1, characterized in that, The specific implementation of the frequency attention module is as follows: First, global average pooling is performed along the spatial dimension on the enhanced multi-channel time-frequency plot to generate channel feature vectors; Performing a one-dimensional discrete cosine transform on the vector maps the spatial channel features to the frequency domain, yielding a spectrum vector; Set the dimensionality reduction ratio and perform a truncation operation on the spectrum vector, retaining only the first k low-frequency coefficients of the spectrum vector to remove high-frequency noise; The truncated vector is input into a multilayer perceptron and activated by a sigmoid function to generate channel weights, which are then used to calibrate the original feature map by adding weights to each channel.
6. The method for detecting spike-ripple superposition based on multi-channel time-frequency diagrams and biomimetic PAC gating according to claim 1, characterized in that, The specific implementation of step 5 is as follows: Constructing a composite loss function Classification loss and gating constraints Weighted composition; Classification loss item The definition is as follows: in To predict the probability of positive samples for the model, To focus parameters, For balance parameters; Gating constraints The definition is as follows: When the input contains RonS, the gating mechanism is activated within the time window corresponding to the spike. Meanwhile, it remains silent in the non-spike region to reflect sparsity; at this time, the label y=1, and the loss function is defined as: in This represents the maximum activation value of the gate on the time axis. Indicates the average activation level of the gating; the first item The gating system must be fully open at least at some point; the second item Constrain the overall sparsity of gating to prevent gating from being open at all times. The sparsity coefficient; When the input is background, there are no spikes that can generate ripples, and the gating remains closed at all times. Noise is blocked from passing through the R channel; at this point, the label y=0, and the loss function is defined as: By minimizing the gating mean, the network is forced to output all-zero gating when there is no spike input; The final gate constraint loss is as follows: The total loss is calculated based on the constructed composite loss function. The AdamW optimizer is used to iteratively adjust the weight coefficients and bias terms in the hybrid neural network through the gradient backpropagation algorithm until the preset optimization termination condition is met. Finally, the model weights with the best generalization performance are saved.
7. A spike-and-ripple detection system based on multi-channel time-frequency plots and biomimetic PAC gating, used to implement the method described in any one of claims 1-6, characterized in that, Includes the following modules: Data processing module: Acquires complete single-channel EEG data, preprocesses the input EEG signals, slices and expands them; Time-frequency graph acquisition module: The extended signal obtained from the data processing module is processed through the R, G, and B channels to obtain a three-channel signal; the three-channel signal is then subjected to wavelet transform and normalization to obtain the corresponding three-channel time-frequency graph; Enhancement module: The time-frequency diagram of the three channels is processed by the constructed biomimetic PAC gating module to obtain the enhanced multi-channel time-frequency diagram; Detection module: Based on the ConvNeXt backbone network, a frequency attention module is embedded in Stages 1-3 to obtain the FCA-ConvNeXt backbone network. The enhanced multi-channel time-frequency map is used as input, and the classification result is output. Training module: Design a composite loss function that combines classification loss and gating constraint loss to train and optimize a hybrid neural network that combines a biomimetic PAC gating module with an FCA-ConvNeXt backbone network.
8. An electronic device, characterized in that, Including processor and memory; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the spike-and-ripple detection method according to any one of claims 1-6.
Citation Information
Patent Citations
Phase detecting device and method of alternative-current generator
CN102841253A
Electroencephalogram signal abnormal state recognition method based on theta-gamma cross frequency coupling characteristics
CN121943352A