A system and method for epileptic signal recognition based on multi-view image modal reconstruction
Patent Information
- Application Number
- CN202511122853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-08-12
AI Technical Summary
[0005]然而,当前大多数研究多集中于单一视角脑电图数据,数据利用效率差、无法充分贴合模型训练,难以充分训练达到最好的识别效果
[0015] In summary, compared with existing technologies, the technical solution conceived in this invention overcomes the limitations of traditional methods that rely on manual feature extraction and single-view data modeling by introducing a deep learning-driven multi-view EEG signal recognition method. This method, without requiring manual feature design, automatically learns discriminative features from raw EEG signals using deep neural networks. Combined with representations such as spectrograms, graph structures, or attention mechanisms, it effectively improves the accuracy and robustness of epilepsy signal classification. Especially under multi-view image modal conditions, by constructing a deep neural network model, it extracts spatiotemporal feature information from different perspectives, achieving more accurate modeling and recognition of epileptic seizure-related signals. While improving the model's generalization ability and recognition accuracy, this method effectively solves the problems of data redundancy and insufficient representation in traditional methods, providing a reliable new path for efficient recognition and intelligent diagnosis of epilepsy signals.
Smart Images

Figure CN121167377B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of epilepsy signal recognition, and more specifically, relates to an epilepsy signal recognition system and method based on multi-view image modal reconstruction. Background Technology
[0002] Epilepsy EEG signal recognition tasks take many forms, with different types serving different diagnostic goals and research needs. For example, some studies focus on identifying signals during the pre-ictal and interictal periods to predict seizures in advance, helping patients take timely intervention measures and reduce the risks associated with sudden seizures. Other studies aim to differentiate EEG signals between epileptogenic and non-epileptic areas to assist doctors in accurately locating the epileptogenic zone and providing preoperative support information for surgery. Still other studies focus on identifying signals during and outside ictal periods to optimize the diagnostic and treatment process by improving the accuracy of seizure detection. Some studies also attempt to improve the diagnostic efficiency and accuracy of initial epilepsy screening by differentiating EEG signals between healthy individuals and epilepsy patients.
[0003] Traditional methods for epilepsy signal recognition primarily rely on two key technical steps: manually designed feature extraction and classifier construction. Before entering the feature extraction module, signals typically undergo preprocessing operations such as denoising, filtering, and segmentation to improve data quality and reduce artifact interference. In feature extraction, researchers often extract key feature information from the time domain, frequency domain, time-frequency domain, and nonlinear dynamics, and employ techniques such as Principal Component Analysis (PCA) or Linear Discriminant Analysis (LDA) for feature dimensionality reduction to reduce computational complexity and improve model generalization ability. Subsequently, classic classifiers such as Support Vector Machines, Decision Trees, and Artificial Neural Networks are used to construct recognition models to classify and judge signals from different epileptic states.
[0004] In recent years, deep learning has made significant progress in the application of epilepsy signal recognition. Compared with traditional methods, deep neural networks have the ability to automatically learn feature representations, directly extracting discriminative information from raw signals, reducing human intervention and improving model performance. For example, some studies have converted EEG signals into spectrograms or graph-structured data and applied convolutional neural networks (CNNs), graph convolutional networks (GCNs), or attention mechanism networks for modeling, achieving high-accuracy classification results on multiple publicly available epilepsy databases. These advances indicate that deep learning-based automatic epilepsy signal recognition methods have become one of the research hotspots in this field.
[0005] However, most current research focuses on single-view EEG data, which has poor data utilization efficiency, cannot fully fit the model training, and is difficult to train to achieve the best recognition results. Summary of the Invention
[0006] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides an epilepsy signal recognition system and method based on multi-view image modal reconstruction, which has high recognition accuracy and robustness.
[0007] To achieve the above objectives, according to one aspect of the present invention, an epilepsy signal recognition system based on multi-view image modal reconstruction is provided, comprising: The EEG preprocessing module is used to preprocess the EEG sampled signals; The signal-to-image conversion module is used to convert the preprocessed EEG sampling signals into a time-frequency graph generated based on continuous wavelet transform, a recursive graph generated based on phase space reconstruction, and a relative position matrix graph generated based on position coding technology, respectively. The feature extraction module is used to extract depth feature vectors for three image modalities from the time-frequency map, recursive map, and relative position matrix map, respectively. The dimensionality reduction and feature fusion module is used to fuse the depth feature vectors of three image modalities and reduce the dimensionality of the fused multimodal image features. The classification module is used to output epilepsy recognition results based on the features of the dimensionality-reduced multimodal image.
[0008] Preferably, the EEG preprocessing module includes: The baseline correction module is used to remove the sampling mean from the EEG sampling signal; The filtering module is used to filter the output signal of the baseline correction module, retaining only the signal in the preset frequency band.
[0009] Preferably, when generating the time-frequency diagram based on the continuous wavelet transform, the frequency integration range of the continuous wavelet transform is set to 0.5Hz to 80Hz.
[0010] Preferably, generating a recursive graph based on phase space reconstruction includes the following steps: The preprocessed EEG sampling signal is denoted as Reconstruct using phase space. in, It is the embedded dimension, used to define the reconstructed spatial dimension. It is the delay time, used to control the sampling interval of data points in a time series. , Represents the reconstructed trajectory in phase space, with the midpoint of phase space as the midpoint. and points The formula for calculating the distance between them is as follows: Distance information between phase points The image is converted into a two-dimensional grayscale texture image with different levels of brightness and darkness, and then the two-dimensional grayscale texture image is mapped to the RGB color space to convert it into a time-frequency graph in color image format.
[0011] Preferably, generating a relative position matrix based on position encoding technology includes the following steps: right To standardize, the calculation formula is as follows: in, yes The mean, yes From the standard deviation, the standard normal distribution can be obtained. Using a piecewise aggregation approximation method, the following is applied: The dimension is determined by Reduce to ', Select dimensionality reduction factor To generate a new smooth time series The calculation formula is as follows: Indicates rounding up. Indicates rounding down; Create a size of × matrix , each timestamp The value in the matrix Each row serves as a reference point to calculate the relative position between two timestamps, and the preprocessed EEG sampling signal is then used. Convert to a two-dimensional matrix The conversion formula is as follows: Using the Min-Max normalization method to normalize the matrix Convert to a grayscale matrix and obtain the final relative position matrix. The calculation formula is: in, for The minimum value, for The maximum value; Will Mapping to the RGB color space yields a relative position matrix in color format.
[0012] Preferably, the feature extraction module includes: The first SE-ResNet18 network is used to extract the deep feature vector of the time-frequency map; The second SE-ResNet18 network is used to extract the deep feature vectors of the recurrent graph; The third SE-ResNet18 network is used to extract the deep feature vectors of the relative position matrix graph.
[0013] Preferably, the first SE-ResNet18 network, the second SE-ResNet18 network, the third SE-ResNet18 network, and the network of the classification module are trained independently.
[0014] According to another aspect of the present invention, a method for epilepsy signal recognition based on multi-view image modal reconstruction is provided, comprising the steps of: Preprocessing of EEG sampling signals; The preprocessed EEG sampling signals were converted into a time-frequency graph based on continuous wavelet transform, a recursive graph based on phase space reconstruction, and a relative position matrix graph based on position coding technology, respectively. Depth feature vectors for three image modalities were extracted from the time-frequency graph, recursive graph, and relative position matrix graph, respectively. The depth feature vectors of the three image modalities are fused, and the dimensionality of the fused multimodal image features is reduced. The epilepsy identification result is output based on the features of the dimensionality-reduced multimodal image.
[0015] In summary, compared with existing technologies, the technical solution conceived in this invention overcomes the limitations of traditional methods that rely on manual feature extraction and single-view data modeling by introducing a deep learning-driven multi-view EEG signal recognition method. This method, without requiring manual feature design, automatically learns discriminative features from raw EEG signals using deep neural networks. Combined with representations such as spectrograms, graph structures, or attention mechanisms, it effectively improves the accuracy and robustness of epilepsy signal classification. Especially under multi-view image modal conditions, by constructing a deep neural network model, it extracts spatiotemporal feature information from different perspectives, achieving more accurate modeling and recognition of epileptic seizure-related signals. While improving the model's generalization ability and recognition accuracy, this method effectively solves the problems of data redundancy and insufficient representation in traditional methods, providing a reliable new path for efficient recognition and intelligent diagnosis of epilepsy signals. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the composition of the epilepsy signal recognition system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the operation of the epilepsy signal recognition system according to an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0018] In the description of the embodiments of this application, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, apparatus, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product or device.
[0019] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0020] This invention provides an epilepsy signal recognition system and method based on multi-view image modal reconstruction, which will be described below.
[0021] like Figure 1 As shown, an epilepsy signal recognition system based on multi-view image modal reconstruction includes: The EEG preprocessing module is used to preprocess the EEG sampled signals; The signal-to-image conversion module is used to convert the preprocessed EEG sampling signals into a time-frequency graph generated based on continuous wavelet transform, a recursive graph generated based on phase space reconstruction, and a relative position matrix graph generated based on position coding technology, respectively. The feature extraction module is used to extract depth feature vectors for three image modalities from the time-frequency map, recursive map, and relative position matrix map, respectively. The dimensionality reduction and feature fusion module is used to fuse the depth feature vectors of three image modalities and reduce the dimensionality of the fused multimodal image features. The classification module is used to output epilepsy recognition results based on the features of the dimensionality-reduced multimodal image.
[0022] This system can be applied to the extraction of signal features through deep learning and the classification of epilepsy signals through machine learning in multi-view image modalities, but it is not only applicable to multi-view image modalities or epilepsy.
[0023] The following describes the preferred implementation method of the EEG preprocessing module.
[0024] The EEG preprocessing module includes a baseline correction module and a filtering module. It preprocesses the input data, including removing baseline shifts and filtering. Input EEG sampling data... T is the number of original signal sampling points. Here, M represents the number of channels at the collection points. Preliminary screening of the input data is necessary to ensure its validity and integrity, including data format conversion and parameter verification, as well as sampling data information.
[0025] The formula for removing the sampling offset is as follows: (1) Filtering: Filtering the acquired data is to eliminate frequency bands that are irrelevant to brain activity, thereby processing and improving the signal. The main method used is time-frequency filtering based on Fourier or wavelet transform, which retains the data of the frequency band of interest. The preprocessed data can ensure the accuracy of subsequent epileptic wave labeling and localization.
[0026] The signal-to-image conversion module is used to convert the raw EEG signals into three different image modalities.
[0027] Preferably, the three types of image modal information are a time-frequency graph generated based on continuous wavelet transform, a recursive graph generated based on phase space reconstruction, and a relative position matrix generated based on position coding technology, respectively.
[0028] Continuous wavelet transform generates time-frequency plots: Time-frequency plots are a visualization tool for time-series signals, used to show the distribution characteristics of signals in both time and frequency dimensions.
[0029] The time-frequency plot is a two-dimensional matrix with dimensions of frequency × time, where: frequency is the number of scale parameters (corresponding to frequency resolution); time is the number of displacement parameters (corresponding to time resolution).
[0030] Time-frequency plots are typically represented as two-dimensional images. The horizontal axis represents time, and the vertical axis represents frequency. Color changes at different locations reflect dynamic changes in signal strength and energy. When using continuous wavelet transform, the frequency range is set from 0.5 to 80 Hz to cover the main frequency band of the epilepsy signal while eliminating high-frequency and low-frequency noise. The number of wavelets within each octave is set to 12, meaning 12 wavelets are evenly distributed within any octave interval to analyze the signal characteristics within that frequency band. This ensures that the time-frequency plot captures the high-frequency components of the signal without being overly coarse in the low-frequency range.
[0031] Phase space reconstruction generates recursive graphs: Recursive graphs are tools for analyzing the intrinsic characteristics of time series data. They can explore complex dynamic behaviors such as periodicity, chaos, and non-stationarity in time series and are widely used in the research of nonlinear signals. Recursive graphs reconstruct the phase space, mapping the state of a dynamic system at a certain moment to a two-dimensional image. During the reconstruction process, an appropriate threshold ε is selected, and the similarity between phase points is judged based on their distance. When the distance between two phase points is less than the pre-selected threshold, the two points are considered similar; otherwise, they are considered not similar. The recursive graph is displayed as a binary image, where each pixel represents the similarity between two data points in the time series. When a pixel is black, it indicates that there is similarity between the two data points; when a pixel is white, it indicates that there is no similarity. The calculation formula for the recursive graph is as follows: For a given length of One-dimensional time series Reconstruct using phase space: in, It is the embedded dimension, used to define the reconstructed spatial dimension. It is the delay time, used to control the sampling interval of data points in a time series. . This represents a series of reconstructed trajectories in phase space, the form of which will change with the embedding dimension. and delay time It changes accordingly depending on the different values it takes.
[0032] midpoint of phase space and points The formula for calculating the distance between them is as follows: Among them, symbols Operators that calculate Euclidean norms.
[0033] When calculating the threshold recursion graph, the recursive value also needs to be calculated at the end: in, For a pre-set distance threshold, and This represents the Heaviside function.
[0034] When applying a thresholded recursive graph, due to the distance threshold parameter... Due to the influence of thresholded recursive graphs, the recursive graph matrix only contains 0 and 1 values. Binarization of this matrix may cause the recursive graph to obscure or lose some feature information. To overcome the potential problems of thresholded recursive graphs, this study skips the thresholding step and adopts a thresholdless recursive graph construction method. Unlike thresholded recursive graphs, the two-dimensional image generated by thresholdless recursive graphs is no longer a simple binarized image, but directly converts the distance information between phase points into two-dimensional grayscale texture images with different brightness levels. This method avoids information loss caused by forced binarization, better preserves the similarity information between time series data points, and provides a richer and more accurate information foundation for subsequent analysis. Furthermore, to match the input of subsequent deep learning models, the recursive graph matrix is mapped to the RGB color space, thus converting the original grayscale texture image into a color image. This conversion not only enhances the richness of the image but also allows the color image to carry more information, thereby more comprehensively expressing the differences and correlations between data.
[0035] Position encoding generates relative position matrices: Relative position matrices are advanced mathematical tools specifically designed to capture and quantify spatial relationships between objects. This matrix, through sophisticated mathematical modeling, transforms complex relative position information into an intuitive data format. The core principle of relative position matrices lies in their ability to convert object position information into a matrix form, where each element represents the relative position between two objects. In natural language processing, relative position matrices often refer to relative position encoding matrices, a mechanism that enhances the performance of Transformer models. This encoding allows the model to consider the relative distances between words when processing sequential data, thereby better understanding sentence structure and semantics. The formula for calculating a relative position matrix is as follows: Given a sequence of length... Time series are denoted as ,right Perform Z-Score standardization: in, yes The mean, yes From the standard deviation, the standard normal distribution can be obtained. A piecewise aggregation approximation method is adopted to... The dimension is determined by Reduce to Choose a suitable dimensionality reduction factor. To generate a new smooth time series The calculation formula is as follows: Specifically, by calculating the average of the piecewise constants, the standardized time series data is transformed from... Dimensions reduced to This preserves the approximate trend of the original sequence while maintaining its dimension. Then, a size of [dimensionality missing] is created. × matrix , each timestamp The value in the matrix Each row serves as a reference point to calculate the relative position between two timestamps, thus processing the preprocessed time series. Convert to a two-dimensional matrix The conversion formula is as follows: Among them, matrix Each row and column is based on a reference to a specific timestamp, encompassing information from the entire time series. The matrix is normalized using the Min-Max method. Convert to a grayscale matrix and obtain the final relative position matrix. : To match the input of subsequent deep learning models and enhance visual representation, the relative position matrix is mapped to the RGB color space, thereby converting the original grayscale texture image into a color image.
[0036] The feature extraction module is used to extract depth features from the multimodal images by inputting the three image modalities into the SE-ResNet18 network structure respectively.
[0037] This patent utilizes the deep network designer in MATLAB R2023b, employing transfer learning to import a pre-trained ResNet18 model. An SE module is embedded in the residual block of the ResNet18 model to introduce an SE-ResNet18 model. The SE-ResNet18 model is then used to train three image modalities of different types of EEG signals on the training set. After training, the SE-ResNet18 model is used as a feature extractor, extracting the global average pooling layer features before classification as depth features to obtain depth features for images from each viewpoint.
[0038] Three image modalities , Each image modal has dimensions H×W×C (height, width, number of channels).
[0039] Define the SE-ResNet18 model as a function 。 This function maps the input image to a high-dimensional feature vector.
[0040] Branch 1: Input time-frequency diagram Features are extracted from the first SE-ResNet18 model. The model first captures the local texture patterns of the time-frequency map through convolutional layers, then deepens the feature hierarchy through residual blocks, and finally dynamically weights the channel features by the SE module to enhance the representation of key frequency band information.
[0041] Branch 2: Input recursion graph The model then proceeds to the second SE-ResNet18 processing. Convolutional layers identify periodic and chaotic structures in the recurrence graph, residual connections alleviate gradient vanishing, and the SE module highlights significant pattern channels related to the dynamic characteristics of epilepsy.
[0042] Branch 3: Input relative position matrix Features are extracted in the third SE-ResNet18 module. The convolutional layer focuses on spatial location correlation, the residual block learns deep location dependence, and the SE module recalibrates the channel weights to enhance the feature contribution of key electrode locations.
[0043] in These are the depth feature vectors of the time-frequency graph, the recursive graph, and the relative position matrix, respectively.
[0044] Different channels contribute differently to features; some channels may be rich in valuable information, while others may be filled with noise or redundant data. Channel attention mechanisms analyze the global information of each layer's feature map and assign appropriate weights to each channel, thereby enhancing the saliency of important features while suppressing unimportant features, thus optimizing the network's feature extraction capabilities. The SE module is a channel attention mechanism that enhances the network's representational ability by recalibrating the feature responses of channels. The core idea of the SE module is to dynamically recalibrate the feature map of each channel, enabling the network to focus more on more useful features while suppressing irrelevant features.
[0045] The main operations implemented by the SE module are compression, activation, and weighting. In the compression stage, global average pooling is performed on the input feature layer, compressing the spatial dimension of each channel into a real value to generate global features that capture global spatial information. In the activation stage, the compressed features are processed through two fully connected layers, with a non-linear activation function used in between to enhance the model's expressive power, and the weights are constrained between 0 and 1 using the sigmoid function.
[0046] in These are the original weight values output by the fully connected layer. It is the output weight. The range is [0, 1], representing the relative importance of each channel.
[0047] During the weighting stage, the obtained weights are multiplied by each channel of the original input feature layer to recalibrate the features, emphasizing important channels and suppressing unimportant channels, thereby improving model performance.
[0048] The dimensionality reduction and feature fusion module is used to fuse the depth features of three image modalities by stitching them together. Principal component analysis (PCA) is then performed on the fused multimodal image features to reduce dimensionality and achieve complementarity of multi-view image information.
[0049] After extracting depth features from three perspective images using the SE-ResNet18 model, the image features from each perspective are concatenated and fused. Then, PCA is used to reduce the dimensionality of the fused features in the training set, retaining 95% of the cumulative variance contribution rate.
[0050] in,' 'Represents a splicing operation, where each spliced feature has a dimension of '. ,so The dimension is .
[0051] The classification module is used to input the dimensionality-reduced features into the classifier for training and inference, thereby enabling automatic recognition of epilepsy signals.
[0052] This patent utilizes the classification learner in MATLAB R2023b, setting up three types of classifiers, as detailed below: Support Vector Machines (SVMs): When processing nonlinear data, SVMs introduce kernel functions to map the original data to a higher-dimensional feature space, thereby making originally nonlinearly separable data linearly separable. Common kernel function types include linear kernels, polynomial kernels, and Gaussian kernels.
[0053] Different kernel function scales significantly influence the shape of the classification decision boundary, determining its smoothness and adaptability. This study employs a refined Gaussian support vector machine based on the kernel function and kernel scale. This patent selects a "one-to-one" strategy.
[0054] Before training the model, the data is preprocessed to have zero mean and unit variance, which improves the training effect and convergence efficiency of the model.
[0055] Decision Trees: In decision tree algorithms, maximum depth is a core parameter controlling model complexity. If the maximum depth is set too small, the decision tree may become too simple, leading to underfitting and failing to capture complex patterns in the data; if the maximum depth is too large, the decision tree may become too complex, easily overfitting the training data and thus reducing the model's generalization ability. Considering all factors, we set a medium-depth tree of "20" as the classifier. The classification criterion is the standard for evaluating the optimal split point selected by the decision tree at each split, affecting the tree's structure and performance. This patent uses the Gini diversity index as the classification criterion. The Gini index assesses the impurity of the dataset by measuring the dispersion of sample class distributions.
[0056] in, , For all quantities of samples, denoted as the number of samples in the k-th class.
[0057] The Gini index is: Naive Bayes: This study uses Gaussian Naive Bayes, which assumes that the conditional probability of each feature follows a Gaussian distribution. When the data features exhibit obvious Gaussian distribution characteristics, Gaussian Naive Bayes can effectively capture these characteristics.
[0058] in Represents the i-th feature dimension. and Distribution is represented in categories Lower features The corresponding standard deviation and expected value. After calculating the conditional probability for each feature dimension, the posterior probability is then maximized. The dimensionality-reduced features are input into the classifier for training. After training, the classifier is tested on the test set to achieve the recognition of epileptic EEG signals.
[0059] The three SE-ResNet18 networks and the classifier network are trained independently. The input to the SE-ResNet18 network is three signal images, and its training objective is binary classification (labeled "0" or "1"). The function serves as the output layer. The input to the classifier network is a feature representation obtained by concatenating or fusing the feature vectors output by the three SE-ResNet18 networks (i.e., the output results of the layer before each fully connected layer). Its training objective is also binary classification (label "0" or "1").
[0060] After successfully obtaining the epilepsy wave classification results, the results need to be displayed. Furthermore, the displayed results are based on existing classification model parameters, such as waveform features and classifier parameter information. The results are as follows: Figure 2 As shown.
[0061] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An epilepsy signal recognition system based on multi-view image modal reconstruction, characterized in that, include: The EEG preprocessing module is used to preprocess the EEG sampled signals; The signal-to-image conversion module is used to convert the preprocessed EEG sampling signals into a time-frequency graph generated based on continuous wavelet transform, a recursive graph generated based on phase space reconstruction, and a relative position matrix graph generated based on position coding technology, respectively. The feature extraction module is used to extract depth feature vectors for three image modalities from the time-frequency map, recursive map, and relative position matrix map, respectively. The dimensionality reduction and feature fusion module is used to fuse the depth feature vectors of three image modalities and reduce the dimensionality of the fused multimodal image features. The classification module is used to output epilepsy recognition results based on the features of the dimensionality-reduced multimodal image; Generating a recursive graph based on phase space reconstruction includes the following steps: The preprocessed EEG sampling signal is denoted as Reconstruct using phase space. in, It is the embedded dimension, used to define the reconstructed spatial dimension. It is the delay time, used to control the sampling interval of data points in a time series. , Represents the reconstructed trajectory in phase space, with the midpoint of phase space as the midpoint. and points The formula for calculating the distance between them is as follows: Distance information between phase points The image is converted into a two-dimensional grayscale texture image with different levels of brightness and darkness, and then the two-dimensional grayscale texture image is mapped to the RGB color space to convert it into a time-frequency graph in color image format. Generating a relative position matrix based on position encoding technology includes the following steps: right To standardize, the calculation formula is as follows: in, yes The mean, yes From the standard deviation, the standard normal distribution can be obtained. Using a piecewise aggregation approximation method, the following is applied: The dimension is determined by Reduce to ', Select dimensionality reduction factor To generate a new smooth time series The calculation formula is as follows: This indicates rounding up. Indicates rounding down; Create a size of × matrix , each timestamp The value in the matrix Each row serves as a reference point to calculate the relative position between two timestamps, and the preprocessed EEG sampling signal is then used. Convert to a two-dimensional matrix The conversion formula is as follows: Using the Min-Max normalization method to normalize the matrix Convert to a grayscale matrix and obtain the final relative position matrix. The calculation formula is: in, for The minimum value, for The maximum value; Will Mapping to the RGB color space yields a relative position matrix in color format.
2. The epilepsy signal recognition system based on multi-view image modal reconstruction as described in claim 1, characterized in that, The EEG preprocessing module includes: The baseline correction module is used to remove the sampling mean from the EEG sampling signal; The filtering module is used to filter the output signal of the baseline correction module, retaining only the signal in the preset frequency band.
3. The epilepsy signal recognition system based on multi-view image modal reconstruction as described in claim 1, characterized in that, When generating time-frequency graphs based on continuous wavelet transform, the frequency integration range of the continuous wavelet transform is set to 0.5Hz to 80Hz.
4. The epilepsy signal recognition system based on multi-view image modal reconstruction as described in claim 1, characterized in that, The feature extraction module includes: The first SE-ResNet18 network is used to extract the deep feature vector of the time-frequency map; The second SE-ResNet18 network is used to extract the deep feature vectors of the recurrent graph; The third SE-ResNet18 network is used to extract the deep feature vectors of the relative position matrix graph.
5. The epilepsy signal recognition system based on multi-view image modal reconstruction as described in claim 4, characterized in that, The first SE-ResNet18 network, the second SE-ResNet18 network, the third SE-ResNet18 network, and the network of the classification module are trained independently.
6. A method for epilepsy signal recognition based on multi-view image modal reconstruction, characterized in that, Including the following steps: Preprocessing of EEG sampling signals; The preprocessed EEG sampling signals were converted into a time-frequency graph based on continuous wavelet transform, a recursive graph based on phase space reconstruction, and a relative position matrix graph based on position coding technology, respectively. Depth feature vectors for three image modalities were extracted from the time-frequency graph, recursive graph, and relative position matrix graph, respectively. The depth feature vectors of the three image modalities are fused, and the dimensionality of the fused multimodal image features is reduced. The epilepsy identification result is output based on the features of the dimensionality-reduced multimodal image; Generating a recursive graph based on phase space reconstruction includes the following steps: The preprocessed EEG sampling signal is denoted as Reconstruct using phase space. in, It is the embedded dimension, used to define the reconstructed spatial dimension. It is the delay time, used to control the sampling interval of data points in a time series. , Represents the reconstructed trajectory in phase space, with the midpoint of phase space as the midpoint. and points The formula for calculating the distance between them is as follows: Distance information between phase points The image is converted into a two-dimensional grayscale texture image with different levels of brightness and darkness, and then the two-dimensional grayscale texture image is mapped to the RGB color space to convert it into a time-frequency graph in color image format. Generating a relative position matrix based on position encoding technology includes the following steps: right To standardize, the calculation formula is as follows: in, yes The mean, yes From the standard deviation, the standard normal distribution can be obtained. Using a piecewise aggregation approximation method, the following is applied: The dimension is determined by Reduce to ', Select dimensionality reduction factor To generate a new smooth time series The calculation formula is as follows: This indicates rounding up. Indicates rounding down; Create a size of × matrix , each timestamp The value in the matrix Each row serves as a reference point to calculate the relative position between two timestamps, and the preprocessed EEG sampling signal is then used. Convert to a two-dimensional matrix The conversion formula is as follows: Using the Min-Max normalization method to normalize the matrix Convert to a grayscale matrix and obtain the final relative position matrix. The calculation formula is: in, for The minimum value, for The maximum value; Will Mapping to the RGB color space yields a relative position matrix in color format.
Citation Information
Patent Citations
Epilepsy multi-class classification method and system based on deep learning
CN116049655A